Consider playing a fast-paced game of table tennis. If your opponent hits a ball across the net, you do not write a formal letter describing the ball's trajectory, mail it to a post office, wait for a written reply, and then swing your paddle three seconds after the ball has already bounced off the floor. Human conversation is a high-speed, full-duplex acoustic dance that demands responses in under 300 milliseconds.
The Stitched Pipeline Latency Barrier
Traditional voice assistants relied on a cascaded three-tier pipeline:
- Speech-to-Text (STT): Whisper waits for the user to finish speaking, buffers audio, and transcribes it to text (~800ms).
- Large Language Model (LLM): The model processes text and generates the first sentence of a response (~600ms).
- Text-to-Speech (TTS): A neural synthesizer converts text back into audio waveform chunks (~800ms).
Including network transport and audio buffer handshakes, total turn-taking latency hovered around 2.5 to 3.5 seconds. This created an awkward, stilted experience where natural interruptions and emotional timing were physically impossible.
[The Cascaded Voice Pipeline (2.5s Latency Wall)]
User Audio ──► [Whisper STT (800ms)] ──► [LLM Decode (600ms)] ──► [Neural TTS (800ms)] ──► Audio Out
(High latency, complete loss of emotional tone, impossible to interrupt!)
[End-to-End Full-Duplex Audio Architecture (Sub-300ms)]
Continuous Audio Stream ──► [WebRTC DataChannel / Opus Packets]
│
▼ (Direct Neural Audio Codec: SNAC / Mimi)
[Native Audio Transformer Tokens]
│ (Direct Streamed Audio Latents)
▼
[Instant Audio Packet Synthesis (250ms TTFT)]
(Supports mid-sentence barge-in and emotional modulation!)
The Native Audio Token Revolution
Modern real-time voice architectures (such as GPT-4o Realtime and open architectures like Mini-Omni and Moshi) eliminate the intermediate text step entirely:
- Neural Audio Codecs (e.g. SNAC, EnCodec, Mimi): Compress raw 24kHz audio waveforms into discrete acoustic and semantic tokens at high compression ratios.
- Unified Sequence-to-Sequence Modeling: The foundation model consumes audio tokens directly as input and autoregressively generates output audio tokens in parallel with text tokens.
- Full-Duplex WebRTC Transport: Using WebRTC over UDP ensures zero packet buffering lag. If the user begins speaking while the model is outputting voice, an acoustic VAD (Voice Activity Detector) halts generation instantly—achieving seamless human 'barge-in' interruptions.
Engineering Takeaway
Conversational presence is a function of latency, not parameter size. By adopting native end-to-end audio tokenization over WebRTC, you transition from robotic turn-based chatbots to fluid, human-like voice intelligence.