How to Build Sub-300ms Barge-In for Real-Time Voice Agents
Picture a parent calling an outpatient pediatric clinic at seven in the morning. Their toddler has a spiking fever, and the parent needs an appointment before the morning clinic schedule locks up. Instead of reaching a frazzled front-desk receptionist juggling three ringing lines, the call connects instantly to an automated voice agent. The agent begins its greeting: "Thank you for calling City Pediatrics. Our clinic opens at eight, but I can help you book an urgent appointment or leave a message for the triage nurse right now. If this is a medical emergency, please hang up and call..."
"Wait, I just need the first available opening with Dr. Chen," the parent interrupts.
In standard interactive voice response setups, the machine plows ahead, oblivious to the human voice. It finishes its scripted legal disclaimer, recites clinic hours, and forces the caller to wait eight seconds before processing the request. In that tiny window of deaf, robotic persistence, trust evaporates. The caller grows agitated, punches zero to bypass the machine, or hangs up, compounding administrative overhead when that caller rings back twenty minutes later.
Human speech is not a series of rigid half-duplex transmissions. It is an intricate, overlapping dance. When voice automation is deployed across hospital switchboards and clinic scheduling lines, interrupting the machine naturally, known technically as barge-in, is the defining benchmark of usability. Achieving real-time conversational AI interruption handling under 300 milliseconds requires re-engineering the voice pipeline from the hardware transport layer up to the neural audio synthesis engine.
The Biological Imperative of the 300ms Threshold
Human conversational dynamics operate on surprisingly narrow tolerances. Extensive research from the Max Planck Institute for Psycholinguistics, led by Stephen Levinson and Francisco Torreira, revealed that the gap between turns in natural human conversation sits consistently between 200 and 250 milliseconds across diverse languages and cultures. When two people speak, the listener does not wait for the speaker to fall silent before formulating a response. They predict the end of the sentence, parse the prosody, and prepare to seize the floor.
Natural human-to-human conversational turn-taking latency averages between 200ms and 250ms, making sub-300ms total interrupt and response the threshold for natural conversational flow.
If an enterprise voice agent takes 700 milliseconds to cut its audio playback after a patient speaks, that delay registers psychologically as disrespect, confusion, or machine failure. The patient repeats themselves, creating double-talk collisions that confuse speech-to-text models. To establish authentic, full-duplex conversational AI for clinic call centers, the total round-trip barge-in time, from the instant acoustic sound waves leave the caller's mouth to the moment the speaker falls dead silent, must drop below 300 milliseconds.
Ditching WebSockets: The WebRTC Voice AI Architecture
The quest for sub-300ms barge-in begins with the wire protocol. For years, voice AI developers relied on bidirectional WebSockets to stream audio bytes between client devices and backend cloud providers. While WebSockets are straightforward to deploy, they run atop TCP, an inherently flawed choice for real-time speech.
TCP guarantees packet delivery via retransmissions. When a patient speaks over a cellular connection with slight packet loss, TCP halts transmission to re-request missing bytes, a phenomenon known as head-of-line blocking. The audio stream stalls, jitter buffers swell, and latency spikes unpredictably into hundreds of milliseconds. In conversational telephony, an audio packet that arrives 300 milliseconds late is useless; it belongs in the garbage, not in a retransmission queue.
Modern low-latency architectures abandon WebSockets in favor of WebRTC running over UDP. WebRTC standardizes on the Opus audio codec, transmitting 20ms audio frames with adaptive bitrate scaling. Under poor network conditions, WebRTC drops individual packets without pausing the media stream, preserving temporal continuity.
WebRTC media streaming reduces end-to-end packet transport latency by 40% to 65% compared to standard TCP WebSocket transport over mobile networks experiencing 2% packet loss.
By routing inbound telephony streams through edge-deployed Selective Forwarding Units (SFUs) and media gateways, engineers keep regional network hops under 20 milliseconds. This transport layer foundation frees up the latency budget for the complex computation that follows.
Dual-Stage Voice Activity Detection
Once audio packets transit the network without queuing delay, the system must determine whether the caller made a meaningful sound. Traditional server-side transcription pipelines fail here because running continuous automatic speech recognition (ASR) to detect interruptions introduces massive computational overhead and latency.
The solution is a dual-stage Voice Activity Detection (VAD) architecture:
- Stage One (Edge Mute): An ultra-lightweight client-side or media-server VAD operates on 10ms to 30ms audio windows. Running optimized engines like Silero VAD v5 via WebAssembly or ONNX Runtime Web directly on the media gateway requires merely 1.5 to 3 milliseconds of compute time per frame. The instant this stage detects acoustic energy resembling human vocal tract resonances, it triggers an immediate local mute command to stop speaker playback within 30 milliseconds.
- Stage Two (Semantic Validation): While the local speaker is silenced, the audio chunk streams concurrently to a server-side deep learning VAD or micro-classifier. This model evaluates whether the utterance constitutes an actual conversational interruption or accidental background noise, such as a car door closing or a cough. If validated, the pipeline proceeds to conversational generation; if invalidated, the agent gracefully resumes its previous output.
Decoupling the physical audio mute from conversational intent verification allows the agent to shut up instantly without waiting for a full natural language understanding cycle.
Acoustic Echo Cancellation: Preventing Self-Interruption
A persistent hazard in conversational voice systems is echo bleed. When a voice agent speaks through a device speaker or over a telephone hybrid, its own synthesized output bleeds back into the microphone input. Without rigorous acoustic isolation, the agent's VAD hears its own synthetic voice, misinterprets the sound as the caller interrupting, and abruptly silences itself mid-sentence.
Engineering a reliable acoustic echo cancellation (AEC) pipeline requires software like the WebRTC Audio Processing Module (APM) or SpeexDSP. The system maintains a continuous reference buffer of the synthesized audio being pushed to the telephone line. The AEC algorithm computes an adaptive digital filter that models the acoustic coupling between speaker and microphone, subtracting the reference signal from the incoming audio capture stream in real time.
Only the differential signal, representing the caller's actual voice, passes forward to the VAD. Without robust AEC, sub-300ms barge-in turns into a hair-trigger liability, causing the bot to stutter uncontrollably whenever it attempts to speak.
The Mechanics of the Kill Switch: Immediate Buffer Flushing
Detecting that a caller wants to interrupt is only half the battle; stopping the sound from reaching their ears is where many voice platforms stumble. When a speech synthesis service generates an audio response, it buffers several seconds of raw PCM audio data across internal queues, network sockets, and client-side playback devices.
A naive implementation sends a cancellation signal to the server and waits for the pipeline to drain gracefully. This approach wastes hundreds of milliseconds while buffered audio continues to play out of the recipient's earpiece. Achieving sub-300ms barge-in requires an immediate hardware-level flush:
- The edge media worker issues an asynchronous abort signal, severing downstream audio packet dispatch within 10 to 20 milliseconds.
- Client-side audio sinks discard their internal ring buffers instantly rather than playing out remaining frames.
- Upstream Large Language Model (LLM) and Text-to-Speech (TTS) inference tasks are canceled via async controllers, preventing wasted compute cycles and preserving API token budgets.
Crucially, this interruption creates a conversational state divergence. If the agent had generated a sixty-word response outlining appointment prep instructions, but the caller interrupted after word ten, the agent cannot record the full sixty words in its context history. Context preservation engines capture the exact millisecond mark of the audio playback cut-off, translate that back into words, and inject a truncated transcript marker, such as "[Interrupted while saying: ...]", into the conversation context. When the agent speaks next, it possesses full awareness of what the patient actually heard versus what was discarded.
Semantic vs. Acoustic Barge-In Filtering
Not every vocalization from a patient is an interruption. In clinical phone calls, patients constantly emit backchannel feedback: "uh-huh", "yeah", "okay", or quiet sighs while listening to available time slots. If an overly aggressive acoustic VAD cuts the agent's playback every time an elderly patient says "right", the interaction feels broken and hyperactive.
To eliminate this friction, advanced speech-to-speech streaming pipelines employ secondary semantic classification. Low-latency micro-classifiers analyze the initial 150 milliseconds of recognized phonemes. If the system classifies the incoming sound as non-interruptive backchannel agreement, playback volume drops subtly, an effect called audio ducking, without terminating the generation thread. If the caller follows the backchannel with sustained speech, the kill switch engages fully.
The Latency Budget for Sub-300ms Performance
Hitting sub-300ms barge-in requires strict enforcement of a real-time voice agent latency budget. Every stage of the system must operate predictably under hardware constraints.
| Pipeline Component | Legacy WebSocket Architecture | Optimized WebRTC Architecture | Latency Contribution |
|---|---|---|---|
| Audio Packetization | 50ms - 100ms frames | Opus 20ms frames | 20ms |
| Network Transport | TCP (80ms - 150ms + jitter) | UDP / WebRTC (Edge SFU) | 25ms - 40ms |
| Acoustic Echo Cancellation | Server-side roundtrip (80ms) | Local / Gateway APM | 5ms - 10ms |
| Voice Activity Detection | Cloud ASR endpointing (300ms - 500ms) | Silero VAD v5 via WASM / Edge | 15ms - 30ms |
| Buffer Flush Execution | Queue drain (150ms - 300ms) | Async AbortController | 10ms - 20ms |
| Total Barge-In Mute | 660ms - 1130ms | Sub-300ms Pipeline | 75ms - 120ms |
The total latency to mute the outbound voice sits well under 120 milliseconds. The remaining 180 milliseconds of the 300ms budget are allocated to receiving the caller's new intent, routing it through an ultra-low latency model, and beginning the new synthetic speech playback stream.
The Shift to Native Speech-to-Speech
The industry is transitioning away from the traditional cascaded pipeline, which strings together separate automatic speech recognition, large language model generation, and text-to-speech modules. Cascaded pipelines incur latency penalties at every serialization boundary. Even with aggressive speculative execution, passing text tokens between three distinct engines introduces compounding delay.
The emerging frontier centers on native speech-to-speech (S2S) foundation models. Architectures like Kyutai Moshi, OpenAI Realtime API, and Google Gemini Multimodal Live process continuous audio streams directly into neural networks without intermediate text conversion. These models are inherently full-duplex. They listen and output audio simultaneously across parallel vector channels, handling barge-in as an internal neural state shift rather than an orchestration of separate software services.
In high-volume hospital operations, where patient access centers handle tens of thousands of scheduling and prescription inquiries weekly, these architectural refinements directly impact operational efficiency. Front-desk staff spend less time dealing with dropped calls and misrouted inquiries, while patients get their needs resolved through conversations that feel responsive, natural, and genuinely attentive.