NVIDIA Nemotron 3.5 ASR
Nemotron 3.5 ASR is a streaming speech recognition model built for high-quality transcription in both low-latency real-time and batch workloads.
Model details
View repository
Latency vs concurrency — how median latency holds up as concurrent streams scale, by GPU.This deployment is a Baseten docker_server model: the Baseten WebSockettransport proxies your connection straight through to the NIM's native Riva Realtime WebSocket API. There is no custom Baseten schema, and you speak Riva's OpenAI-Realtime-style JSON event protocol directly.
Two kinds of transcription messages
As audio streams in, the realtime API returns two kinds of transcription events:
Partial results (
...transcription.delta) — live, in-progress hypotheses that update continuously as you speak. Ideal for low-latency captions that refine in real time. Each event is an incremental delta, not the full hypothesis — concatenate the deltas within a segment to build the running caption.Final results (
...transcription.completed) — stable, punctuated transcripts for each completed segment. These won't change and represent the authoritative output.
Prerequisites
pip install websockets
export BASETEN_API_KEY=...Input must be a mono, 16-bit, 16 kHz PCM WAV. The server treats the bytes as raw pcm16 at the declared sample rate, so stereo or off-rate input transcribes to garbage (or nothing) — convert it before streaming (see the note under load_pcm).
Running it
python transcribe.py audio.wav1import asyncio, base64, json, os, sys, wave
2import websockets
3
4URL = "wss://model-<id>.api.baseten.co/environments/production/websocket?intent=transcription"
5MODEL = "cache-aware-parakeet-rnnt-multi-asr-streaming-sortformer"
6
7async def stream_wav(ws, path):
8 with wave.open(path, "rb") as wf:
9 while chunk := wf.readframes(4000): # 250 ms chunks
10 audio = base64.b64encode(chunk).decode()
11 await ws.send(json.dumps({"type": "input_audio_buffer.append", "audio": audio}))
12 await ws.send(json.dumps({"type": "input_audio_buffer.commit"}))
13 await asyncio.sleep(0.25) # stream at real-time speed
14 await ws.send(json.dumps({"type": "input_audio_buffer.done"}))
15
16async def main(path):
17 headers = {"Authorization": f"Api-Key {os.environ['BASETEN_API_KEY']}"}
18 async with websockets.connect(URL, additional_headers=headers) as ws:
19 await ws.send(json.dumps({"type": "transcription_session.update", "session": {
20 "input_audio_format": "pcm16",
21 "input_audio_transcription": {"model": MODEL, "language": "auto"},
22 "input_audio_params": {"sample_rate_hz": 16000, "num_channels": 1},
23 "recognition_config": {"enable_automatic_punctuation": True},
24 }}))
25 sender = asyncio.create_task(stream_wav(ws, path))
26 async for raw in ws:
27 event = json.loads(raw)
28 if event.get("type", "").endswith("transcription.completed"):
29 print(event["transcript"])
30 if event.get("is_last_result"):
31 break
32
33asyncio.run(main(sys.argv[1]))