Use your OpenAI commitments on Baseten open models. Learn more
Transcription

NVIDIA logoNVIDIA Nemotron 3.5 ASR

Nemotron 3.5 ASR is a streaming speech recognition model built for high-quality transcription in both low-latency real-time and batch workloads.

Model details

View repository
Latency vs concurrency — how median latency holds up as concurrent streams scale, by GPU.Latency vs concurrency — how median latency holds up as concurrent streams scale, by GPU.

This deployment is a Baseten docker_server model: the Baseten WebSockettransport proxies your connection straight through to the NIM's native Riva Realtime WebSocket API. There is no custom Baseten schema, and you speak Riva's OpenAI-Realtime-style JSON event protocol directly.

Two kinds of transcription messages

As audio streams in, the realtime API returns two kinds of transcription events:

  • Partial results (...transcription.delta) — live, in-progress hypotheses that update continuously as you speak. Ideal for low-latency captions that refine in real time. Each event is an incremental delta, not the full hypothesis — concatenate the deltas within a segment to build the running caption.

  • Final results (...transcription.completed) — stable, punctuated transcripts for each completed segment. These won't change and represent the authoritative output.

Prerequisites

pip install websockets
export BASETEN_API_KEY=...

Input must be a mono, 16-bit, 16 kHz PCM WAV. The server treats the bytes as raw pcm16 at the declared sample rate, so stereo or off-rate input transcribes to garbage (or nothing) — convert it before streaming (see the note under load_pcm).

Running it

python transcribe.py audio.wav
Input
1import asyncio, base64, json, os, sys, wave
2import websockets
3
4URL = "wss://model-<id>.api.baseten.co/environments/production/websocket?intent=transcription"
5MODEL = "cache-aware-parakeet-rnnt-multi-asr-streaming-sortformer"
6
7async def stream_wav(ws, path):
8    with wave.open(path, "rb") as wf:
9        while chunk := wf.readframes(4000):  # 250 ms chunks
10            audio = base64.b64encode(chunk).decode()
11            await ws.send(json.dumps({"type": "input_audio_buffer.append", "audio": audio}))
12            await ws.send(json.dumps({"type": "input_audio_buffer.commit"}))
13            await asyncio.sleep(0.25)  # stream at real-time speed
14    await ws.send(json.dumps({"type": "input_audio_buffer.done"}))
15
16async def main(path):
17    headers = {"Authorization": f"Api-Key {os.environ['BASETEN_API_KEY']}"}
18    async with websockets.connect(URL, additional_headers=headers) as ws:
19        await ws.send(json.dumps({"type": "transcription_session.update", "session": {
20            "input_audio_format": "pcm16",
21            "input_audio_transcription": {"model": MODEL, "language": "auto"},
22            "input_audio_params": {"sample_rate_hz": 16000, "num_channels": 1},
23            "recognition_config": {"enable_automatic_punctuation": True},
24        }}))
25        sender = asyncio.create_task(stream_wav(ws, path))
26        async for raw in ws:
27            event = json.loads(raw)
28            if event.get("type", "").endswith("transcription.completed"):
29                print(event["transcript"])
30                if event.get("is_last_result"):
31                    break
32
33asyncio.run(main(sys.argv[1]))

🔥 Trending models