FULL-DUPLEX SPEECH · MULTI-PARTY INTERACTION

Duplex-MPE.

Benchmarking Multi-Party Interaction
in Full-Duplex Dialogue

Knowing the answer is only part of the conversation.
Knowing when to speak, stay silent, and stop matters too.

* Equal contribution   † Corresponding author

2,000paired scenarios
3–4 + 1humans + assistant
61.10 hdistinct speech clips
5speech systems evaluated

01 / HEAR THE DIFFERENCE

A conversation has more
than one possible next move.

Listen to recorded model outputs.
Switch models on the same conversation.

Same scene · different model
THE SHARED AUDIO TIMELINE
Human
Model

Headphones recommended. Human speech is on the left; the model is on the right. Download clip ↗

IN THIS EXCERPT

Read the preceding conversation

Earlier turns heard by the model; outside the audio excerpt above.

Selected examples from four systems. Names are shown only when an example passes the displayed metric; other examples are anonymous. Anonymous labels are local to each metric. Aggregate results below cover all five evaluated systems.

02 / THE BENCHMARK

One assistant.
Several possible addressees.

Each scenario places Aria among three or four people. Models hear spoken role instructions followed by continuous conversation audio, without transcripts or supplied turn boundaries.

Duplex-MPE framework: requests to Aria, speech requiring silence, resolved requests, continuous audio input, and separate response metrics.

Same task. Two ways to address the assistant.

Each scenario has an explicit version that names Aria and an implicit version that identifies it through conversational cues. The task and intended answer stay the same. Local wording and audio can differ between the complete paired scenarios.

Four capabilities, measured separately.

Speaking often is not the same as participating well.

01

Fresh-onset
trigger recall

Does the model start a new response within 5 seconds after a request ends, with no speech active at that boundary?

02

Conditional
answer accuracy

When the model produces a response to the request, is its answer correct?

03

Silence
preservation

Does it stay silent when not addressed, allowing only brief acknowledgements that do not take the floor?

04

Answering-window
yield

When it starts after the question and is still speaking as a human resolves it, does it stop by the deadline and remain silent?

Explore the response-timing definitions

Response presence measures speech on T requests, including speech already underway. Window response measures speech during an N4 question or its following 3-second gap. These are coverage diagnostics that help interpret the capability scores.

Four metric timelines showing response presence, fresh response initiation, N4 window response, and stopping eligibility and deadlines.

For eligible N4 events, the model must be silent two seconds after the human answer begins and remain silent through the rest of that human turn’s observation window, including its following 0.5 seconds. A shorter human turn does not shorten the two-second deadline.

03 / THE RESULTS

Speaking more
does not mean doing better.

Five speech systems show different combinations of missed requests, inaccurate answers, unwanted speech, and difficulty stopping.

Speech-system results

Percentages · higher is better for scored capabilities

Speech system results. The first four metrics score capabilities; the final two show coverage.
SCORED CAPABILITIES ↑COVERAGE DIAGNOSTICS
ModelFresh-onset
trigger recall
Conditional
answer accuracy
Silence
preservation
Answering-window
yield
Response
presence
Window
response

Bold marks the highest reportable value for each scored capability. Yield is conditional on each model’s eligible events; sample sizes appear below its values. Freeze-Omni has only one eligible event in each condition, so its yield rate is not reported. Answer accuracy is conditional on requests with spoken responses.

95.00%

Fresh responses, not just speech

MiniCPM-o 4.5 starts a fresh response to 95.00% of explicit requests. It leads on fresh-onset recall, answer accuracy and silence preservation under both conditions.

99.55%

High presence can be misleading

Under explicit addressing, Freeze-Omni produces speech after almost every request, yet preserves silence in only 5 of 20,360 silence windows.

+64.3pp

Addressing matters in text

Gemini 3.1 Pro responds more often to explicit than implicit requests when given transcripts. Paired tests detect no significant difference for the speech systems.

04 / DATASET CONSTRUCTION

From conversation scripts
to continuous audio.

Scenarios vary setting, activity, participant relationships, nearby devices, and speaking style. Turn-wise speech synthesis provides the constructed time boundaries.

Dataset construction pipeline: scene attributes, script generation, paired addressing versions, turn-wise speech synthesis, and continuous audio construction.

BUILD ON THIS WORK

Cite Duplex-MPE.

Questions about the benchmark?
Get in touch ↗

BIBTEX
@misc{ma2026duplexmpe,
  title = {Duplex-MPE: Benchmarking Multi-Party
           Interaction in Full-Duplex Dialogue},
  author = {Ma, Chengqian and Feng, Wenhao and
            Jin, Weixuan and Dai, Gaole and
            Xie, Tianyu and Ma, Yuexiao and
            Kang, Zhaolu and Zhao, Xiangyu and
            Zheng, Xiawu and Chao, Fei},
  year = {2026},
  url = {https://github.com/step-out/MPEval}
}