How We Cut Voice AI Latency Down to 200 Milliseconds
The Monday Morning Triage Problem
At 8:03 on a Monday morning, a busy urban outpatient clinic receives roughly forty inbound calls per minute. An anxious father dials in trying to secure a same-day appointment for an ear infection. An elderly patient calls to reschedule an oncology consult before public transit shuts down. A third caller needs to verify whether a surgical referral went through.
When healthcare facilities deploy voice agents to absorb this administrative crush, the bottleneck rarely centers on whether the language model knows how to query the electronic health record or read the scheduling calendar. The failure happens in the silence between breaths. In ordinary human interactions, conversational pauses feel natural when they mirror human biology. The moment that pause stretches beyond a fraction of a second, the caller assumes the call dropped, says "Hello? Are you still there?", talks over the system, and hangs up in frustration.
For decades, enterprise telephony has treated interactive voice response systems as blunt IVR trees: slow, deterministic, and tolerated only because patients had no choice. Modern generative voice systems promised to eradicate those frustrating phone trees. Yet the first wave of automated telephone agents introduced an equally maddening flaw known as conversational drag. Callers were met with heavy, robotic pauses lasting two to three seconds while cloud servers transcribed, analyzed, and synthesized responses.
Human speech is radically fast. Linguistic research published by the Proceedings of the National Academy of Sciences shows that across cultures and languages, natural turn-taking gaps average between 200 and 250 milliseconds. Bridging the chasm between a typical 2,000-millisecond machine lag and the biological 200-millisecond threshold requires rethinking the entire conversational stack, from the transport protocols carrying audio packets to the neural networks generating words.
The Anatomy of Conversational Drag
To understand why conversational software feels sluggish, one must examine the legacy cascading stack. Historically, voice agents operated as a relay race of three independent, sequential processes: Automatic Speech Recognition (ASR), followed by a Large Language Model (LLM), followed by Text-to-Speech (TTS).
In this sequential model, the patient finishes speaking. The ASR engine buffers the audio, waits for a prolonged silence to verify the speaker has finished, and returns a block of transcribed text. That text is shipped over an HTTP connection to an LLM. The model processes the prompt, generates a complete response, and delivers the full text payload to a TTS engine. Finally, the TTS engine converts the entire string into an audio file, buffers it, and streams it back through a telephone interface. By the time the caller hears the first phoneme, two full seconds have elapsed.
In a healthcare scheduling scenario, that two-second pause is disastrous. Telephony networks operate in full duplex, meaning audio flows both ways simultaneously. When an engine waits two seconds to speak, the patient almost invariably fills the dead air with an interjection. That sudden interjection confuses the transcription engine, forces the model to regenerate its response, and traps both parties in an awkward loop of mutual interruption.
| Pipeline Architecture | VAD / Turn Detection | ASR Processing | LLM TTFT | TTS Audio Synthesis | Network Transport | Total Round-Trip Latency |
|---|---|---|---|---|---|---|
| Legacy Cascaded Stack | 500ms to 700ms | 350ms to 500ms | 600ms to 1,000ms | 400ms to 600ms | 150ms to 250ms (HTTP/WS) | 2,000ms to 3,050ms |
| Early Streaming Pipelines | 250ms to 400ms | 150ms to 200ms | 250ms to 400ms | 150ms to 200ms | 80ms to 120ms (WebSockets) | 880ms to 1,320ms |
| Sub-200ms Target Architecture | 20ms to 40ms (Edge VAD) | Chunked (0ms perceived) | 80ms to 110ms (vLLM/TensorRT) | 30ms to 60ms (TTFA) | 15ms to 30ms (Edge WebRTC) | 180ms to 240ms |
Eradicating Silence: Fast Turn Detection at the Edge
The first place where latency piles up is Voice Activity Detection (VAD). Conventional telephony stacks rely on crude volume thresholding. When a patient speaks, the system registers sound. When the patient stops, the system holds its breath, waiting anywhere from 500 to 700 milliseconds of absolute silence to guarantee the caller has completed their sentence.
That waiting period alone consumes more than double our total latency budget. If a patient takes a brief, natural breath between clauses ("I need an appointment with Dr. Miller... for my hip"), an aggressive VAD cuts them off. If the VAD is overly conservative, the user waits three-quarters of a second before the machine even starts working.
Slashing this threshold to under 50 milliseconds requires replacing blunt energy heuristics with lightweight, neural VAD models running directly on edge servers. Models like Silero VAD evaluate tiny audio chunks (often 30 milliseconds long) to analyze voice characteristics rather than simple volume decibels. By coupling neural voice activity detection with semantic turn-taking predictors, the system determines not just whether sound has ceased, but whether the syntactic rhythm of the utterance suggests a completed thought.
If a caller says, "Can I schedule a checkup for next Tuesday at four," the intonation drops and grammatical completion signals an end-of-turn. The engine closes the audio capture window within 40 milliseconds, shaving more than half a second off the conversational loop before text transcription even reaches completion.
Zero-Perceived Transcription via Chunked Streaming ASR
Legacy speech recognition processed complete audio files. Modern real-time voice AI architecture treats speech as a continuous, unidirectional stream of tiny audio frames. Instead of waiting for an entire phrase, the system streams raw PCM audio directly over low-overhead transport layers to modern streaming ASR engines, such as Deepgram Nova-2, which achieve word-level transcription latencies well below 150 milliseconds.
Crucially, transcription latency must not run sequentially against language model latency. By the time the patient finishes pronouncing the final syllable of their request, ninety-five percent of their sentence has already been transcribed, normalized, and tokenized in the background. The speech recognition pipeline overlaps almost entirely with the patient's vocalization window. The actual latency contribution of the transcription layer shrinks to the time it takes to finalize the very last word (roughly 30 to 50 milliseconds).
Accelerating the Inference Engine: Chasing Time to First Token
Once text tokens stream out of speech recognition, they hit the central reasoning engine. In medical appointment scheduling, the language model must perform entity recognition, confirm clinical specialties, consult slot availability, and honor cancellation policies. Doing this with an unoptimized 70-billion-parameter model hosted on standard inference servers guarantees a 600-millisecond delay before the first token appears.
To break the 200-millisecond barrier, two architectural shifts are mandatory: model rightsizing and inference acceleration.
Handling hospital switchboards and scheduling queues does not require frontier models designed for advanced legal reasoning or open-ended creative prose. Compact models, specifically fine-tuned 8-billion-parameter variants, deliver exceptional domain accuracy when constrained to clear operational workflows. These smaller models fit entirely within high-bandwidth memory caches on modern GPUs, drastically reducing memory bandwidth bottlenecks.
Furthermore, standard Hugging Face or unoptimized PyTorch runtime environments introduce unacceptable overhead. Deploying high-throughput inference engines like vLLM or NVIDIA TensorRT-LLM allows for continuous batching, fused kernel operations, and pre-allocated Key-Value (KV) caching. KV caching avoids recomputing the mathematical attention vectors for static prompt components, such as hospital routing rules, system prompts, and calendar formats.
With pre-warmed KV caches, specialized CUDA kernels, and 8B parameter models, the Time to First Token (TTFT) drops from 800 milliseconds down to an astonishing 80 to 110 milliseconds.
The secret to sub-200ms conversational performance lies in overlapping every layer of the stack. True real-time orchestration never waits for a full sentence when a single token can trigger the next stage.
Token-Level Streaming Text-to-Speech
The next historic latency sink sits inside audio synthesis. Traditional Text-to-Speech engines accept a paragraph, spend hundreds of milliseconds computing acoustic representations, and generate a wave file. If your language model produces a thirty-word confirmation message, waiting for the complete sentence before generating sound guarantees immediate conversational failure.
The solution is sub-sentence streaming TTS. Modern non-autoregressive and State Space Model (SSM) architectures, such as Cartesia Sonic, ElevenLabs Flash, and Deepgram Aura, generate audio incrementally. Instead of waiting for punctuation or sentence completions, these systems trigger audio synthesis on the first three to five generated tokens.
The primary metric here is Time to First Audio (TTFA). State Space Models can consume five words ("I can help with that...") and output audible, high-fidelity PCM audio buffers within 40 to 90 milliseconds. As the user listens to those opening syllables, the remaining words stream through the LLM and the voice generator in parallel. Perceived wait time drops by up to 65 percent compared to full-sentence inference batching.
- The caller completes their request.
- The streaming ASR finalizes the trailing phoneme within 40 milliseconds.
- The lightweight LLM emits its initial three tokens within 80 milliseconds.
- The streaming TTS engine converts those three tokens into playable audio frames within 50 milliseconds.
- The first sound wave hits the caller's telephone speaker at 170 to 190 milliseconds.
Transport Layers: Why WebSockets Fail and WebRTC Wins
Even an infinitely fast machine stack will crawl if audio transport chokes on network hops. For years, developers built real-time voice applications using standard WebSockets. While WebSockets improve on basic HTTP polling, they rely on TCP. Under adverse network conditions, TCP mandates in-order packet delivery. If a single audio packet drops over an uneven cellular connection, TCP halts transmission while waiting for a retransmission, triggering audible stutter and conversational freezes.
Modern low-latency architectures abandon WebSockets in favor of WebRTC (Web Real-Time Communication) media and data channels, coordinated by orchestrators such as LiveKit. WebRTC runs over UDP, prioritizing speed and real-time delivery over complete packet retransmission. When an audio packet drops over a cellular tower, the jitter buffer conceals the drop seamlessly instead of pausing the entire conversational thread.
Network latency is also a function of physical geography. Sending a patient's voice from a phone in Chicago to an inference data center in Oregon adds 70 milliseconds of pure round-trip transit time. Overcoming this reality demands distributing telephony edge ingress and inference nodes across localized regional points of presence. Bridging Session Initiation Protocol (SIP) trunks directly to edge-deployed WebRTC media servers keeps network transit latency strictly below 30 milliseconds.
Full-Duplex Interruption Handling: The Barge-In Challenge
Speed creates a downstream engineering problem: if an automated voice agent responds within 200 milliseconds, it will inevitably clash with callers who change their mind mid-sentence or clarify their requests. If a patient says, "Actually, wait, let me look at my calendar," while the agent is reciting available morning slots, a system lacking true full-duplex capabilities will bulldoze through the patient's voice.
Nothing breaks the conversational illusion faster than a machine that talks over a human. Solving this requires zero-latency Acoustic Echo Cancellation (AEC) and instant audio buffer flushing.
Because the agent's outgoing audio is being played through the telephone network, some of that sound bounces back into the input stream. High-grade echo cancellation filters out the agent's own voice in real time, preventing the system from interrupting itself. Simultaneously, the moment the edge VAD detects genuine caller audio while the TTS engine is outputting speech, the orchestrator issues a hardware-level interrupt. The audio playback buffer flushes in zero milliseconds, the generation pipeline terminates current tasks, and the system pivots immediately to listening mode.
Beyond Cascades: The Dawn of Native Speech-to-Speech
The 200-millisecond threshold represents the absolute physical limit of cascaded pipelines. Even with world-class streaming ASR, aggressive token pipelining, and sub-100ms TTS engines, developers are constantly shaving fractions of milliseconds across four distinct network hops.
The industry's ultimate destination lies in Native Speech-to-Speech (S2S) foundation models. Architectures like Kyutai's Moshi and recent multimodal models bypass intermediate text transcription entirely. They ingest audio tokens directly and emit audio tokens natively, preserving tone, pacing, hesitation, and emotional inflection while dropping end-to-end processing latencies to 160 milliseconds.
In high-volume hospital call centers and busy neighborhood clinics, these milliseconds translate directly to human relief. When front-desk systems answer on the first ring, comprehend urgent inquiries instantly, and reply with the brisk, natural cadence of an experienced receptionist, administrative burnout plummets. Patients do not want to wait on hold for twenty minutes, nor do they want to navigate clunky voice bots that pause like broken satellites. By conquering the 200-millisecond threshold, healthcare communications finally moves past the era of digital friction, delivering voice interactions that feel as swift, capable, and responsive as human care itself.