Full Duplex Bench

Full Duplex Bench

Full Duplex Bench is a testing procedure that researchers use to measure how well a spoken AI system conducts a real conversation — that is, listening and speaking at the same time. What is primarily examined is timing: when the AI interrupts, when it stays silent, and when it reacts too late.

A real conversation between two people doesn’t run like a radio transmission, taking turns. Both listen while the other speaks. One tosses in an “mhm”, one interrupts, one senses that the other is finished even before the sentence ends. Spoken AI systems long couldn’t do this: they waited until the human stopped, and answered afterward. Newer systems are supposed to do better, and that’s exactly what Full Duplex Bench exists for. It is a collection of standardized tests used to measure how natural the timing of such a speech system is. “Full Duplex” comes from communications engineering and means: in both directions at the same time, as with a telephone and unlike a walkie-talkie.

Why conversational timing determines the impression

Anyone talking to a voice assistant notices delays immediately. A pause of one second already feels uncomfortably long. Conversely, a system that butts in mid-sentence is annoying. The content can be perfect — if the timing is off, the conversation still feels wrong. Such qualities were long hardly comparable, because every company used its own criteria.

That is precisely the purpose of a benchmark. A benchmark is a fixed test course that different systems run through under the same conditions. The results turn into numbers, and numbers can be placed side by side. Without such a shared standard, marketing claims like “especially natural conversations” remain unverifiable.

An important distinction: Full Duplex Bench does not test whether the AI answers cleverly or tells the truth. Other tests exist for that. Here, it is exclusively about behavior in the temporal flow of a conversation.

The four conversational situations tested

The test plays recorded conversation excerpts to the model and cuts them off at a certain point. Then it observes what the system does. Measurement typically takes place across four situations. First, the normal turn-taking: the human has finished their sentence — how quickly does the AI take over? Second, feedback given while listening, known in the jargon as backchannel — that is, short interjections like “yeah” or “I see” that signal one is still following along.

Third, interruption by the user. If the human talks over the AI, it should fall silent within fractions of a second and respond to the new input. Fourth, the pause mid-sentence: if someone briefly pauses to think, the system must not immediately seize the floor. Evaluation is done in milliseconds and via hit rates, for instance how often an interruption was correctly recognized.

The evaluation is carried out partly by software, partly by a larger language model that judges the recordings. This is practical, but it has a weakness: an automatic evaluator has its own blind spots. Good benchmark scores therefore do not automatically mean that real users find the conversation pleasant.

Language models you can interrupt

One mainly encounters the term in connection with the voice features of major AI providers. Systems such as ChatGPT’s voice mode, Google's Gemini Live, or assistants in cars and smart speakers advertise that you’re allowed to cut in on them. When a company wants to prove such progress, it turns to metrics from benchmarks of this kind.

In tech news, Full Duplex Bench therefore usually appears in tables comparing several models by latency values. Latency refers to the delay between input and response. Anyone reading such tables should know that they show only a partial picture. A model can lead in timing and be weaker in content.

This becomes practically relevant with phone hotlines and voice control in cars. There, fluent speech is not a luxury but determines whether people use the technology at all. Benchmarks like this one are an attempt to make that impression measurable instead of merely claiming it.

Related Products

Latest News

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.