DiffuPlex: Accelerating Full-Duplex Spoken Dialog Models via Rolling Masked Diffusion

Anonymous Submission

Abstract

Recent full-duplex spoken dialog models enable simultaneous listening and speaking, but fine-grained models still advance their backbone autoregressively at every interaction frame. We introduce DiffuPlex, a rolling masked diffusion framework that reduces this sequential computation by predicting multiple future user and assistant frames in a single backbone wake. DiffuPlex consumes only a confident prefix of each predicted future while interaction continues at the original frame rate. As user speech arrives, it checks the corresponding user predictions and, when the interaction diverges, preserves already played assistant content while revising only the unplayed future. We consider two inference policies over the same predictor: DiffuPlex-Listen consumes multiple future frames when they predict assistant silence, whereas DiffuPlex-Speak can also consume predicted assistant speech. Across full-duplex interaction and spoken-language evaluations, DiffuPlex substantially reduces sequential backbone computation while largely preserving interaction behavior and general capability. DiffuPlex-Listen and DiffuPlex-Speak achieve 1.46× and 1.59× deployment-path wall-clock speedups and 1.61× and 1.80× Core LM speedups, with all measured backbone invocations completing within the 80ms interaction interval. Human evaluation shows that Listen preserves speech naturalness and conversational quality, while Speak retains conversational quality with some degradation in speech naturalness.

DiffuPlex overview
A conventional full-duplex model wakes its backbone at every 80 ms frame. DiffuPlex drafts several future frames in a single wake, plays a confident prefix, and revises only the unplayed remainder when the user diverges from what was predicted.

Samples

How to read the strips

A dot is one backbone call. The bar under it is what that call produced. Here it produced a single 80 ms frame, which is the autoregressive case.
A longer bar means one call covered more frames. Everything dark blue was spoken to the listener.
Light blue is drafted but not yet spoken. Nothing tells the model which of those frames it will get to use. If the user diverges from the prediction, the unspoken remainder turns coral at that moment and a new call begins.
Amber marks the user. The dashed band is the moment they cut in. Coral is only ever discarded output. Above each row, the close-set grey ticks are where a conventional model would have to call its backbone, once for every frame.