A voice assistant usually gets two bad options: think hard and leave you sitting in silence, or answer right away and answer shallowly. This technical report refuses that trade. It builds one audio-language model that hears both sides of the conversation at once, thinks in private while it is already talking, and keeps chatting while its tools run in the background.
You are walking to the train and you ask your phone a question out loud. It is not a hard question for a person with a pen, but it has several moving parts: a departure time, a travel time, a time-zone shift, and a buffer at the other end. You want to know one thing: will you make dinner?
The assistant goes quiet. One second. Two. Three. You glance at the screen to check whether it heard you. You say "hello?" And now the assistant has a new problem, because your "hello?" just arrived in the middle of its thinking, and it has to decide what that sound means.
That silence is not a bug in one product. It is a tax that almost every voice system pays, and it comes from a very simple fact: good answers to tricky questions need deliberation (working through the problem step by step before committing), and deliberation takes time. In a text chat, the time is invisible: a spinner turns, you read something else. In a voice conversation, the time is audible. Silence is a message. It says "I am broken" or "I did not hear you".
So builders pick one of two bad options. Option one: think first, then speak. The answer is careful, and the user sits through dead air. Option two: speak immediately with no deliberation. The conversation feels alive, and the answer is shallow, sometimes confidently wrong.
Every number, table value and architecture detail in this lesson comes from the StepAudio 3 Realtime Technical Report (arXiv:2609.14005, StepFun-Audio Team). When we need made-up numbers to make a mechanism visible, the prose says so in plain words: illustrative. When the paper is silent on something, we say that too, instead of filling the gap.
The introduction of the paper lists the difficulties in a single paragraph, and each one becomes a chapter of this lesson. Read them slowly, because each describes a moment that a normal chatbot pipeline gets wrong.
Notice what the four problems share. In each one, two things are happening at the same time, and the system has to keep both alive: your speech and its speech (problems 1 and 2), its thinking and its speaking (problem 3), its tool work and its talking (problem 4). A system built around strict turns, where one thing finishes before the next begins, cannot represent any of them.
Think about two radios. A walkie-talkie is half-duplex: only one side can transmit at a time. You press the button, speak, say "over", and release. The other side cannot cut in, and you cannot hear them while you talk. A telephone is full-duplex: both sides can speak and hear at the same moment, which is why you can say "mm-hm" while a friend tells a story, and why they can stop mid-word when you gasp.
Most classic voice assistants are walkie-talkies wearing a telephone costume. A voice activity detector (a small model that decides whether a stretch of audio contains speech) waits for you to stop, a speech recognizer turns your words into text, a language model writes a reply, and a speech synthesizer reads it out. The "over" button is just hidden inside a silence timer. The turn-taking Gleam walks through why that timer is so fragile.
It helps to see exactly where a turn-based pipeline breaks against each of the four problems. The table below walks one hypothetical cascade (voice activity detector, recognizer, language model, synthesizer) through them. None of these failures is exotic; each is the default behavior of a design that only allows one thing to happen at a time.
| Moment | What the cascade does | Why it goes wrong |
|---|---|---|
| User pauses mid-request | Silence timer fires, turn ends, reply starts | The timer sees only silence, not whether the sentence is finished |
| User says "right" during the reply | Detector hears speech and stops playback | Every sound looks like an interruption; an acknowledgment kills the answer |
| Question needs reasoning | Language model thinks, then synthesizer speaks | Thinking time turns straight into dead air |
| A tool call takes a while | The turn blocks until the tool returns | The user cannot ask "is it done yet?" or add a requirement |
| TV talking in the background | Detector hears speech and starts a turn | Nothing asks whether the speech was addressed to the assistant |
The last row is a fifth problem the paper handles, background speech rejection: deciding that a voice in the room is not talking to you. Chapter 3 shows how the model uses dialogue history as evidence for that decision.
StepAudio 3 Realtime is built as a telephone from the inside. The paper calls its input path "full-duplex": the model hears the user's audio stream and its own outgoing audio stream, continuously, and it makes its turn-taking decisions every 320 milliseconds (Chapter 3 unpacks that number).
The paper organizes the whole system as a single cycle that never stops turning: listen, converse, think, act. Each verb is backed by a named capability. Here they are, with the paper's names in bold and a plain-words definition beside each.
| Loop verb | Capability (paper's name) | What it means in plain words | Chapter |
|---|---|---|---|
| Listen | Deep Perception | Hear not only the words but how they are said: tone, emotion, speaker traits, background sounds, timing. | 2 |
| Converse | Seamless Duplex | Use both audio streams to decide, moment by moment, whether to keep listening, start talking, keep talking, or stop. | 3 |
| Think | Think-While-Speaking (+ Adaptive Thinking, MTP) | Reason privately in one process while another process is already speaking; think only when it helps; decode the thinking faster. | 6, 7 |
| Act | Voice Agent | Call tools and backends, keep the conversation going while they run, and fold their results back into speech. | 8 |
Two more pieces glue the four together. Model merging (Chapter 5) combines several specialist checkpoints into one set of weights, so dialogue skill, audio understanding and text reasoning live in a single model. And a benchmark the team built, StepAudioChat (Chapter 4), measures the conversational intelligence that the thinking machinery is supposed to protect.
The paper's Figure 2 draws the loop as a ring. On the left, two input waves feed in: a user stream labelled "lexical · emotion · acoustic context", and a model stream labelled "speech · overlap · turn state". Around the ring sit the four capabilities: Deep Perception (ASR, audio understanding), Seamless Duplex (listen, speak, yield), Think While Speaking (adaptive thinking, Medusa MTP) and Streaming Action (intent, tool call, feedback). In the middle sits a single disc labelled "conversational state". On the right, one arrow leaves the ring: streaming speech.
That middle disc is the most important object in the paper, and it deserves a name of its own.
Picture a whiteboard in the middle of a small team. One person writes down what the customer just said and how they sounded. Another notes whose turn it is. A third scribbles half-finished calculations. A fourth pins up a sticky note: "booking lookup: still running." Everyone reads the same board, and everyone writes to it.
The paper's version of the whiteboard is the shared conversational context. Section 2.1 lists exactly what is on it: "acoustic and linguistic evidence, dialogue history, the current speaking turn, reasoning progress, and tool-execution status."
| Ingredient on the board | The question it answers | Who mainly writes it |
|---|---|---|
| Acoustic + linguistic evidence | What did the user say, and how did they say it? | Deep Perception |
| Dialogue history | What has been agreed, asked, or promised so far? | Every turn |
| Current speaking turn | Who holds the floor right now, and is anyone overlapping? | Seamless Duplex |
| Reasoning progress | How far has private thinking got, and is it finished? | Think-While-Speaking |
| Tool-execution status | Is an external task pending, running, or done, and what did it return? | Voice Agent |
The paper says this context "informs whether to continue listening or speaking, whether to reason further, and whether a request is ready for external action." And crucially: "Newly observed speech and returned tool results can change these decisions as the conversation proceeds." The board is never frozen. Every 320 ms, something new may be written to it.
One detail on the board is easy to miss and matters a lot: "Model-side speech provides additional context for interpreting user utterances that overlap with a response." The model does not only remember what it planned to say; it tracks what the user has actually heard so far. If you interrupt at word twelve of a forty-word answer, the useful fact is that you heard twelve words, not forty.
The simulation below puts the three strategies side by side on one timeline. The top lane is private reasoning. The bottom lane is speech the user hears. Red marks dead air before the first word. Drag the slider to make the question harder (more reasoning time) and watch what each strategy does with that time.
All times and the quality meter are illustrative, chosen to make the shape of the trade-off visible; the paper reports no per-question latency figures. "Think then speak" waits for all reasoning. "Speak now" never reasons. "Think-While-Speaking" starts talking at once, lets each spoken segment use whatever reasoning exists at that moment, and can append a final continuation once reasoning completes (the paper's Speak-First default).
Drag the slider all the way right in "Think then speak" mode, and the red bar grows linearly with difficulty: the user pays for every second of thought in silence. Switch to "Speak now" and the red bar vanishes, but the quality meter collapses, because nothing was thought through. Switch to "Think-While-Speaking" and something new happens: speech starts immediately, and each segment uses whatever reasoning exists when it is released. In this simulation we give the speaker a simple, illustrative scheduling policy (open with framing, state substance once reasoning is available); the paper does not state such a rule, and it allows an answer that began on incomplete reasoning to be corrected at the end.
The simulation hides one honest cost that the paper does not hide. Speech that begins before reasoning ends can commit to something the reasoning later contradicts. The paper handles that with "a final continuation [that] can supplement or correct an answer that began from incomplete reasoning", and it warns in Section 6.3.2 that keeping strict verification on the spoken output "does not eliminate errors arising from incomplete private reasoning." We return to this in Chapter 6.
A natural objection: just make reasoning faster. The paper does accelerate reasoning (Chapter 7), and its best measured wall-clock speedup is 2.05×, for three prediction heads with Medusa-style acceptance (Table 6). Let us see what that buys in a think-then-speak design, with an illustrative question that needs 400 reasoning tokens at an illustrative 50 tokens per second.
The lesson of the arithmetic: acceleration shrinks the tax proportionally, but concurrency changes its shape. Speed divides the silence; running thinking alongside speaking removes the dependency between silence and thinking time altogether. That is why the paper uses both: concurrency for the user experience, and multi-token prediction so that the private thinking finishes sooner and more of the answer is spoken from complete reasoning.
Put the four capabilities and the shared context together and you get the paper's operating picture. It is worth stating as a list of clocks, because each clock ticks at its own rate and none waits for the others.
| Clock | What ticks | Rate, as the paper describes it |
|---|---|---|
| Listening | User audio arrives and is encoded | Continuously, in 320 ms blocks |
| Floor control | A state or text token is emitted | After every 320 ms block |
| Speaking | Response segments are released | According to playback progress of the output audio |
| Thinking | Private reasoning tokens are decoded | As fast as decoding allows, sped up by MTP |
| Acting | A backend task runs | Asynchronously; results arrive whenever they are ready |
In a pipeline, these clocks are chained, so the slowest one sets the pace for everything. In StepAudio 3 Realtime they are decoupled, and the shared context is how they stay coherent: the speaking clock reads whatever the thinking clock has written so far, the floor clock reads what the speaking clock has actually played, and the thinking clock reads what the tool clock has returned. The paper sums it up in the introduction: "These functions operate concurrently as needed, with new user input shaping the ongoing interaction."
A useful way to test your understanding of the rest of this lesson is to ask, for each mechanism, which two clocks it decouples. Seamless Duplex decouples listening from speaking. Think-While-Speaking decouples thinking from speaking. The Voice Agent decouples acting from speaking. Adaptive Thinking and MTP do not decouple anything; they make the thinking clock tick less often and faster.
The report actually describes two models, and it is worth separating them now to avoid confusion later.
StepAudio 3 Realtime is the conversational system: the full-duplex, thinking, tool-using assistant. StepAudio 3 ASR Max is a transcription specialist. The paper states they "share the same pretraining and midtraining stages" and "diverge only during supervised fine-tuning, where the ASR branch is specialized for transcription and the realtime branch is tuned for spoken interaction." So the ASR numbers in Chapter 2 describe the specialist, and the paper is explicit that they do not describe how the realtime model transcribes.
These are the headline numbers the paper reports. Each will be derived, compared and questioned in its own chapter; for now, just notice how many different kinds of skill are on one page.
| Capability | Benchmark | StepAudio 3 | Best reported baseline |
|---|---|---|---|
| Speech recognition (ASR Max) | LibriSpeech test-clean WER | 1.18 | 1.38 (HY3.0 ASR Preview) |
| Audio understanding | MMSU | 90.6 | 83.6 (Gemini 3.1 Pro) |
| Audio understanding | 8-benchmark macro average | 81.3 | 81.8 (Gemini 3.1 Pro) |
| Dialogue, reasoning mode | StepAudioChat macro | 73.0 | 77.1 (Kimi K3) |
| Dialogue, realtime mode | StepAudioChat macro | 70.4 | 77.1 (Kimi K3, reasoning mode) |
| Full-duplex control | AA Full-Duplex Bench Overall | 98.9 | 98.4 (Qwen Audio 3.0 Realtime Plus) |
| Voice agent | τ-Voice macro task success | 56.0% | 56.5% (Grok Voice Think Fast 2.0 High) |
Read the table the way the paper asks you to: not as "wins everywhere". In realtime mode the dialogue score is comparable to reasoning-mode Doubao 2.0 Lite (70.5) and DeepSeek-V4-Flash (71.4), which is the comparison the paper draws, while Kimi K3 sits well above at 77.1. The model leads on full-duplex control and on several audio benchmarks, sits close behind on the audio macro average and on τ-Voice, and trails a dedicated reasoning model on dialogue. The paper names its own gaps: "multi-turn constraint following and retail tool-use tasks." Chapter 9 puts all of this on one scoreboard.
Before any clever behavior, there is plumbing. Here is the job, stated as a data problem. Two waveforms arrive continuously: what the user is saying, and what the assistant itself is saying. A text history sits alongside them: the system prompt, earlier turns, tool results. Out of all that, several times a second, the system must produce the next thing to do: stay quiet, say a word, think a thought, call a tool.
What machine can do that? This chapter builds it piece by piece from the paper's Section 3 and Figure 3, then follows the recipe the team used to teach it: three stages of pretraining and a midtraining stage.
The paper's architecture diagram has five boxes and two waves. Read it left to right, like a conveyor belt.
Now each box in plain words.
An audio encoder is a network that listens. Raw audio is a long list of air-pressure samples, far too many and too low-level for a language model to reason over. The encoder compresses short windows of sound into vectors that capture what matters: phonemes, pitch, loudness, texture. StepAudio 3 Realtime does not train its own encoder from scratch; it uses the Audio Transformer (AuT) encoder from the Qwen3-Omni technical report. Figure 5 of the paper labels this component a "streaming audio encoder", which tells you it processes audio as it arrives rather than waiting for a whole file.
An adapter is a translator between two vocabularies of vectors. The encoder was trained to speak "audio vector"; the language model was trained to read "token embedding". The two spaces have different sizes and different geometry. The adapter is a learned mapping that turns the first into something the second can read as if it were a sequence of token embeddings.
The LLM decoder is the brain: an autoregressive language model that reads one mixed sequence (audio-derived vectors plus ordinary text tokens) and predicts what comes next. What "comes next" can be a text token, an interaction-state token (Chapter 3), a private thinking token (Chapter 6) or a tool call (Chapter 8).
The generator is the mouth. It turns the decoder's output into streaming audio "with context-appropriate tone and rhythm." The paper adds that "natural delivery includes expressive cues such as pauses and hesitation, connecting the content of a response with its communicative intent." The report does not describe the generator's internal design (codec, vocoder, or token rate), so we will not invent one.
The concept-plus-realization rule says you should be able to follow the data. The report does not publish hidden sizes, encoder frame rates or token rates, so the table below uses symbols where the paper gives no number, and the paper's numbers where it does.
| Wire | Contents | Shape / type | Source |
|---|---|---|---|
| User audio in | Microphone samples | a stream of samples, cut into 320 ms blocks | §5.1 |
| Model audio in | The assistant's own generated speech | a parallel stream, same 320 ms blocking | Fig. 3, Fig. 5 |
| Encoder out | Acoustic feature vectors | Tf × denc (sizes not reported) | AuT, Qwen3-Omni |
| Adapter out | Audio vectors in LLM space | Tf × dmodel (sizes not reported) | §3.1 |
| Text in | System prompt, history, tool results | token ids → L × dmodel | §3.1 |
| Decoder out | Next-token distribution | state / text / think / response tokens | Fig. 5, Fig. 7 |
| Generator out | Streaming speech | audio, fed back as "model audio" | Fig. 3 |
Figure 5(A) adds a finer picture of how the two streams become one sequence. Each 320 ms slot in the drawing holds a few user-audio cells and a few model-audio cells side by side; a dashed line collects them into the encoder; the decoder row is labelled "Serialized S2T input" and "Audio LLM and full-duplex interaction states"; and above each slot sits a small circle, a state token. Two of those state outputs connect to a box labelled "Step-Audio model: streaming speech output", whose output feeds back into the model audio stream. Chapter 3 is entirely about those circles.
Imagine a hospital with a hundred specialists. Each patient is seen by a triage nurse, who sends them to the two or three specialists most relevant to their symptoms. The hospital as a whole knows a vast amount, but any single visit costs only a few doctors' time.
A mixture-of-experts (MoE) layer works the same way. Instead of one large feed-forward block that every token passes through, the layer holds many smaller feed-forward "experts" and a small router that sends each token to a few of them. The model's total knowledge scales with the number of experts; the compute per token scales only with the experts actually used. The Mixture of Experts Gleam builds routing from zero.
The paper states: "StepAudio 3 Realtime uses a mixture-of-experts architecture with approximately 196 billion total parameters and 11 billion active parameters per token. Its language backbone is based on Step 3.7 Flash." It does not report the number of experts, how many are selected per token, or the layer count. What the two numbers do tell us is the sparsity, and sparsity is what makes a realtime model of this size plausible.
Why does step 5 matter for a voice product? Because the model must emit something every 320 ms and, in Think-While-Speaking, runs two concurrent calls to the same model (Chapter 6). A dense model of 196B parameters would pay roughly 17.8 times more compute per token than this MoE. Sparse activation is part of what makes "think and speak at the same time" affordable.
Press Step to push a 320 ms block of audio through each box. The right panel shows a toy router picking experts for the current token: the expert count and the top-k are illustrative (the paper does not report them), but the readout uses the paper's 11B / 196B. Toggle ASR Max fine-tuning to see which parts the paper says are frozen and which are updated when the transcription specialist is trained.
Stages light up in order. The feedback arc from the generator back to the input is the model hearing its own voice. In ASR Max fine-tuning mode, a lock marks the frozen audio encoder (paper, §4.1.1); the adapter and decoder glow as trainable. The paper does not specify frozen modules for the realtime branch, so that mode shows no locks.
Two things to take from the simulation. First, the text path joins after the adapter: audio and text become one sequence only at the decoder, which is why the adapter's job ("map into the representation space of the language model") is so central. Second, the feedback arc means every block the model speaks becomes input a few hundred milliseconds later, so the decoder's context always contains both sides of the overlap.
Here is the architecture as pseudo-PyTorch. Module names follow the paper; internal sizes are placeholders because the report does not give them. The structure, not the numbers, is the point.
pseudo-pytorchclass StepAudio3Realtime(nn.Module): def __init__(self, aut_encoder, moe_decoder, generator, d_enc, d_model): self.encoder = aut_encoder # AuT encoder from Qwen3-Omni (streaming) self.adapter = nn.Sequential( # maps encoder space -> LLM space nn.Linear(d_enc, d_model), nn.GELU(), nn.Linear(d_model, d_model)) self.decoder = moe_decoder # ~196B total / 11B active, Step 3.7 Flash based self.generator = generator # streaming speech output def step(self, user_blk, model_blk, text_ids, cache): # user_blk, model_blk: one 320 ms block from each stream a_user = self.encoder(user_blk) # (T_f, d_enc) a_model = self.encoder(model_blk) # (T_f, d_enc) the model hears itself audio = self.adapter(torch.cat([a_user, a_model])) # (2*T_f, d_model) text = self.decoder.embed(text_ids) # (L, d_model) separate text path h, cache = self.decoder(torch.cat([audio, text]), cache) nxt = h[-1].argmax() # a state, text, think, or tool token speech = self.generator(h, nxt) if is_response(nxt) else None return nxt, speech, cache # speech re-enters as next model_blk
Two honest caveats about this sketch. The adapter here is a small MLP because that is the most common choice; the paper says only that "an adapter maps the encoder outputs into the representation space of the language model." And the paper's Figure 5 draws user and model cells side by side within each slot, but does not specify whether they are encoded jointly or separately; the sketch encodes them separately for clarity.
A model this size is only as good as the audio it has heard. Section 3.2 describes an automated curation pipeline (inherited from StepAudio 2.5) that turns raw recordings into training samples. Walk through it like a factory line.
| Stage | What it does | Why it matters |
|---|---|---|
| Sound event detection | Flags what kinds of sound are present | Keeps music, noise, and speech from being mislabelled as each other |
| Voice activity detection | Finds where speech starts and stops | Lets the pipeline cut at natural boundaries |
| Merge + resegment | Rejoins and re-cuts audio "into samples of suitable duration that preserve semantic completeness" | A sample should not end mid-sentence; half-thoughts teach bad turn-ending cues |
| Audio-level metadata | Quality, synthetic-speech likelihood, speaker count | Enables filtering TTS-generated audio and routing multi-speaker clips |
| Multi-system transcription + language ID | Several recognizers transcribe; outputs are cross-checked | Agreement is evidence of a correct transcript (the same idea powers ROVER in Chapter 2) |
| Grading | Samples graded by acoustic and semantic quality | Supports "quality-aware sampling across training stages" |
For this model, the paper says, the pipeline "is extended to broaden language coverage and support the sustained perception and interaction demands of realtime dialogue." It does not give hours, languages, or grade thresholds, so neither will we.
Think of teaching a translator who already speaks one language fluently (text) and is learning to understand a second medium (sound). First you teach the alphabet of the new medium. Then you practise mixed conversations at scale. Finally, you polish with the best material you have. The paper's three stages follow that arc exactly.
Across all three, the sequence length is fixed at 32K tokens and the model processes 1.2T training tokens in total. A cooldown stage, in current practice, is the last stretch of pretraining where the data mixture shifts toward the cleanest sources (and often the learning rate decays); the paper describes only the data side.
One mixture decision stands out: "The pretraining mixture increases the proportion of pure text to preserve the general capabilities of the base language model and support subsequent reasoning and agent training." Why would an audio model want more text? Because of catastrophic forgetting: when a pretrained network is trained hard on a new distribution, it tends to lose skills it is no longer practising. The backbone arrives already good at reasoning and knowledge; pure text keeps those muscles exercised while audio is being learned. The general-text results in Chapter 9 (86.8 on HMMT February 2026) show the text skills survived, though the report does not isolate how much the text share contributed.
Midtraining is a stage between broad pretraining and task-specific fine-tuning, where the data mixture is steered toward the capabilities the final model needs. Section 3.3 makes two changes.
Context extension. The context length grows from 32K to 128K tokens "to accommodate longer dialogue histories, earlier user requirements, and intermediate tool results." In a voice conversation, the requirement that matters might have been spoken twenty minutes ago; a tool result might be long; both must still be in view.
Mixture shift. Midtraining uses "perception, synthetic conversational, and voice-agent data" and "substantially increases the share of audio-understanding and agent-interaction data." The first broadens "speech, music, environmental sound, and audio-grounded reasoning"; the second "trains the model to carry user intent through planning, tool use, and spoken follow-up." Chapter 3 adds that duplex midtraining also includes streaming ASR, voice activity detection and utterance-completeness prediction over more than 10,000 hours of synthetic full-duplex data.
Whatever the real token rate, the method is the lesson: context length in a voice model is a time budget, and every private thought or tool payload spends some of it. That is part of why Adaptive Thinking (Chapter 7) matters beyond latency.
For the ASR specialist, Section 4.1.1 is explicit: supervised fine-tuning keeps "the audio encoder frozen" while "updating the audio-language adapter and language decoder." For the realtime branch and for the pretraining stages, the report does not say which modules are updated. We will not guess.
Why freeze an encoder at all? Our reading (not the paper's stated reason): the encoder already produces strong acoustic features from its own large-scale training, and it is shared infrastructure that other capabilities depend on. Updating it on a transcription-only objective risks narrowing its features toward words and away from tone, speaker traits and ambient sound. The adapter and decoder are where transcription-specific behavior (normalization, use of context, rare terms) naturally lives.
A system diagram looks tidy until the inputs get messy. Walk each wire and ask what a bad input does downstream. The paper does not run these stress tests box by box, so the right-hand column says which later chapter or benchmark touches each failure.
| Degraded input | First box that suffers | Downstream effect | Where the paper addresses it |
|---|---|---|---|
| Noisy or reverberant user audio | Encoder | Weaker acoustic evidence for words and for tone | SpecAugment-style masking in ASR SFT (Ch. 2); τ-Voice includes "diverse forms of background noise" (Ch. 8) |
| A second talker in the room | Encoder + decoder | Speech that is not addressed to the assistant enters the context | Background speech rejection using dialogue history (Ch. 3) |
| User speaks over the model | Full-duplex input | Two voices in one moment; ambiguity about intent | Dual-stream input + model speech as context (Ch. 3) |
| Rare names, product codes | Decoder | Homophone substitution ("sounds like a common word") | Long-tail terminology augmentation (Ch. 2) |
| Very long session | Decoder context | Early requirements scroll out of view | 128K midtraining context (this chapter); multi-turn constraint following remains a stated gap (Ch. 9) |
| Slow or failing tool | Text path (tool results) | Claims about work that has not finished | Evidence-grounded tool dialogues and negative examples (Ch. 8) |
Notice that most defenses are not extra modules. They are data: augmentations, synthetic dialogues, negative examples. The architecture stays five boxes; robustness is taught, not bolted on. That is a recurring theme of the report, and Chapter 2's "less is more" ablation puts a number on how much data quality matters.
The training story is spread across five sections of the report. Here it is gathered into one map, so later chapters can point back to it. The report does not state the exact order of every post-training step relative to teacher merging and Adaptive Thinking training, so the rows are grouped by purpose, not by a timeline.
| Stage | Used by | Data (as described) | Context | What it teaches |
|---|---|---|---|---|
| Pretraining 1: modality alignment §3.2 | Both models | Curated audio from the automated pipeline, plus text | 32K | The interface between acoustic representations and the language model |
| Pretraining 2: multimodal mixed §3.2 | Both models | Audio-text at scale, with a raised share of pure text | 32K | Joint audio-text modeling without losing text skills |
| Pretraining 3: cooldown §3.2 | Both models | Greater weight on high-quality data | 32K | A refined foundation (1.2T tokens across all three stages) |
| Midtraining §3.3, §5.3 | Both models | Perception, synthetic conversational and voice-agent data; over 10,000 hours of synthetic full-duplex data with streaming ASR, VAD and completeness supervision; text | 128K | Long histories, the time-interleaved duplex format, carrying intent through tool use |
| SFT, ASR branch §4.1.1 | ASR Max | Short labelled utterances, ROVER-fused long-form pseudo-labels, long-tail synthetic terms, optional context | packed to 32K | Normalized, context-aware transcription (encoder frozen) |
| SFT and post-training, realtime branch §4.2, §5.3, §6.2, §7.3 | Realtime | Quality-controlled audio QA (final set size not stated; the ~100K vs ~2M comparison in Chapter 2 is an ablation); self-play multi-turn dialogue; duplex interaction data; voice-agent dialogues and real trajectories | not stated | Audio understanding, conversation, floor control, tool use |
| Reasoning selection §6.3.1 | Realtime | Turns relabelled by the no-think probe and blind judge, under per-domain budgets | not stated | When to think (Adaptive Thinking) |
| Teachers and merge §6.4 | Realtime | Four data compositions from one base; 3:1:1:1 average | n/a | One model holding dialogue, audio and text strengths |
| Inference-time system §6.3 | Realtime | none (runtime) | n/a | Think-While-Speaking, adaptive routing, MTP3 acceleration |
Two patterns jump out of the map. First, almost every row's "data" column is a pipeline, not a dataset: curation, fusion, synthesis, self-play, relabelling. Second, the pure-text thread never disappears; it appears in pretraining, in duplex midtraining, and in the text-reasoning teacher, which fits an audio model staying competitive on HMMT (the report does not isolate how much the text share contributed).
A customer on a support line says a product code out loud. It is a made-up-looking string, the kind of name no dictionary contains. The recognizer hears it and writes down the nearest common words, which sound almost the same. The agent then searches for a product that does not exist.
A minute later the same customer says "fine." Once brightly, meaning "great, go ahead". Once flatly, after a long sigh, meaning "this is not fine at all". The transcript is identical: one word, one period.
Those two moments are the two halves of this chapter. The first is a lexical failure: getting the words wrong. The second is a nonverbal failure: getting the words right and the meaning wrong. The paper's Section 4 opens by naming both: "Perception combines lexical understanding with cues about the speaker, vocal delivery, acoustic events, and temporal structure."
The paper splits the work across two models. StepAudio 3 ASR Max specializes in transcription. StepAudio 3 Realtime is trained "for broader audio understanding and spoken interaction." We take them in that order.
Automatic speech recognition (ASR) turns speech into text. The standard score is the word error rate (WER): align the system's output (the hypothesis) with the correct transcript (the reference), count the edits needed to turn one into the other, and divide by the reference length.
Here S is the number of substituted words (a wrong word in place of the right one), D is the number of deleted words (a reference word with nothing in the hypothesis), I is the number of inserted words (a hypothesis word with nothing in the reference), and N is the number of words in the reference. Because insertions count, WER can exceed 100%. Lower is better.
Mandarin has no spaces between words, and deciding where one word ends is itself ambiguous. So Mandarin results use the character error rate (CER): the same formula, counted over characters instead of words. The paper follows this convention: "English results use word error rate (WER), while Mandarin results use character error rate (CER)."
Recall from Chapter 1: ASR Max and Realtime share pretraining and midtraining and split at supervised fine-tuning (SFT), the stage where the model learns from input-output pairs that show exactly the desired behavior. Section 4.1.1 lists four design decisions for the ASR branch.
Packing. Examples are "packed into sequences of up to 32K tokens." Packing means concatenating several short training examples into one long sequence so that the accelerator is not wasting compute on padding. A two-second command and a thirty-second voicemail can share one 32K row.
Augmentation. The paper applies "time-frequency masking following the augmentation principle of SpecAugment." SpecAugment (Park et al. 2019) blanks out random stretches of time and random bands of frequency in the input spectrogram. It is like training a reader on pages with coffee stains: the model learns not to depend on any single moment or pitch band, which helps when real audio has a cough, a dropout, or a missing frequency range.
What trains. The audio encoder is frozen; the adapter and the language decoder are updated "to produce normalized transcripts". A normalized transcript follows fixed formatting conventions so that the same spoken content always maps to the same text; the paper does not list its normalization rules.
Context-aware recognition. "An example may additionally provide dialogue history, a preceding model response, a scenario description, or task-specific terminology as optional evidence." Then comes the crucial guard: "The target transcript remains grounded in the input waveform, allowing the model to use relevant context without simply copying unrelated terms."
Why is that guard necessary? Give a recognizer a list of product names and it can start "hearing" them everywhere, turning ordinary words into the listed terms. The training target is always what was actually said, even when the context suggests otherwise. So the model learns that context is a hint for ambiguous sounds, not a script to recite.
The ASR mixture "combines short labeled utterances with long pseudo-labeled recordings." A pseudo-label is a transcript produced by machines rather than people. Long recordings are cheap to collect and expensive to hand-transcribe, so the question is how to make machine transcripts trustworthy.
The paper's answer is a jury. "Multiple recognition systems transcribe segmented audio, and their hypotheses are aligned and fused with Recognizer Output Voting Error Reduction (ROVER)." ROVER (Fiscus, 1997) aligns several transcripts of the same audio into one grid of word slots, then picks the most-voted word in each slot. Independent systems tend to make different mistakes, so the majority is often right where each individual is sometimes wrong.
Then two more steps. "Agreement-based filtering selects reliable segments for recomposition into longer sessions." If the jury was split on a segment, that segment is not trusted and is dropped. The surviving segments are stitched back into long sessions, and an LLM "restores punctuation and improves consistency across each session."
Each row is one recognizer's aligned hypothesis for the same audio segment; each column is a word slot (∅ marks a deletion). The fused row takes the most-voted word per slot. The agreement score is the mean winning-vote share across slots; segments below the threshold are dropped. The three segments and five recognizers are illustrative. Try the "rare term" segment: the jury agrees, and is wrong.
Push the threshold up and the noisy segment is thrown out: that is agreement-based filtering doing its job, trading quantity for reliability. Now look at the rare-term segment. Most recognizers share the same blind spot: they have never seen the product name, so they all fall back on the same common-word spelling. The jury is confident and wrong, and no threshold catches it.
"Rare names and technical terms are often confused with common words that sound similar. We therefore build targeted synthetic training examples for these cases." The pipeline, step by step, as Section 4.1.1 describes it:
Step 5 matters more than it looks. Speech synthesizers also mispronounce rare terms. A synthetic clip that says the common word while its label says the rare term would teach the model to write the rare term whenever it hears the common word, which is exactly the "unrelated lexical substitution" the paper wants to prevent. The report does not say how pronunciation consistency is checked, so we leave that detail open.
Step 6 closes the loop with context-aware recognition: for acoustically confusable terms, the example shows the model the hint and a waveform-grounded target, teaching it "to use relevant context while avoiding unrelated lexical substitutions."
ASR Max is evaluated on five standard sets (LibriSpeech test-clean and test-other, AISHELL-1, WenetSpeech test-net and test-meeting) and on ContextASR-Bench, "long-form, multi-domain, entity-rich speech in English and Mandarin." ContextASR-Bench is run in its Contextless setting: no domain labels, no entity lists, no hotword injection. That makes it a test of what the model knows, not of what it is told. All baselines "are rerun under the same evaluation setup... using the same test audio and scoring procedure."
| Test set (lower is better) | ASR Max | Doubao 2.0 ASR | Seed 2.0 Lite | HY3.0 ASR Preview |
|---|---|---|---|---|
| LibriSpeech test-clean (WER) | 1.18 | 2.94 | 1.47 | 1.38 |
| LibriSpeech test-other (WER) | 2.28 | 5.98 | 2.67 | 2.80 |
| AISHELL-1 (CER) | 0.49 | 2.07 | 1.66 | 1.22 |
| WenetSpeech test-net (CER) | 3.99 | 4.03 | 4.71 | 3.71 |
| WenetSpeech test-meeting (CER) | 4.35 | 5.09 | 4.80 | 4.12 |
| ContextASR-Speech-EN (WER) | 7.91 | 12.04 | 9.48 | 8.53 |
| ContextASR-Dialogue-EN (WER) | 3.43 | 9.09 | 3.65 | 4.66 |
| ContextASR-Speech-ZH (CER) | 1.43 | 2.80 | 2.15 | 1.74 |
| ContextASR-Dialogue-ZH (CER) | 1.02 | 10.47 | 4.15 | 1.63 |
ASR Max is best on seven of nine rows. On both WenetSpeech subsets it beats Doubao 2.0 ASR and Seed 2.0 Lite but trails HY3.0 ASR Preview "only by a small margin." On all four ContextASR-Bench subsets it is best.
Relative reductions are the right lens for low error rates. A drop from 1.22 to 0.49 is "only" 0.73 points, but it removes about three of every five remaining errors. One caution the paper itself states plainly: "these results characterize the ASR-specialized model, not the transcription behavior of the realtime model."
Now the "fine." problem. Section 4.2 builds training data from a hierarchical taxonomy (a tree of capability categories and sub-categories) covering six areas.
| Taxonomy branch | An example question it licenses (ours, illustrative) |
|---|---|
| Lexical content | What did the speaker ask for? |
| Paralinguistics (how something is said: pitch, pace, emotion, emphasis) | Does the speaker sound frustrated? |
| Acoustic events | Is there a siren in the background? |
| Speaker and temporal structure | How many people speak, and who speaks first? |
| Music | Is the accompaniment major or minor? |
| Audio-grounded reasoning | Given the sounds, where was this likely recorded? |
Figure 4 of the paper draws the construction pipeline as five numbered boxes, guided by two banners: "Sampling controls: what audio is allowed in" and "Self-defined sub-capability space: what may be asked about it."
Step 3 prevents a common failure in synthetic audio QA: asking a question the clip cannot answer. A silent room recording should never become a training example about the speaker's emotion. Tagging first, then asking only about supported aspects, keeps questions answerable.
Here is the hard part of audio data quality. To check whether an answer is grounded (actually supported by the sound), a checker must hear the sound. But the strongest, cheapest judges are text-only LLMs. The paper resolves this with a division of labor, in three layers.
Layer 1, deterministic checks. Remove "empty, truncated, malformed, or severely repetitive outputs."
Layer 2, text-only judges. They "assess query and response quality and assign a case-value score", which "jointly considers the query, the response, the amount of useful information available in the audio as represented by its annotations, and the training value of the question." The paper is explicit: "These judges do not directly evaluate audio grounding."
Layer 3, cross-model consistency. "Grounding reliability is estimated from the consistency of responses independently produced by multiple models for the same audio–question pair." If several models that did hear the clip give the same answer, the answer is probably in the audio. If they disagree, something is off: the question is ambiguous, the audio is unclear, or one model hallucinated.
python (sketch)def route_candidate(clip, question, answers, judge): # answers: independent responses from several audio models for the same pair if is_broken(answers): # empty, truncated, malformed, repetitive return "reject" quality = judge.quality(question, answers) # text-only LLM judge case_value = judge.case_value(question, answers, clip.annotations) agreement = consistency(answers) # proxy for audio grounding if quality >= Q_HI and case_value >= V_HI and agreement >= A_HI: return "sft_candidate" # the best of the best if broadly_useful(question, answers): return "midtraining" # useful, not SFT-grade return "relabel_or_review" # disagreements, correctable cases # Thresholds are placeholders: the paper describes the routing, not the numbers.
The routing is the realization of one sentence in the paper: "Only candidates with high quality, high case value, and strong cross-model consistency are retained as SFT candidates; broadly useful examples may enter midtraining, while disagreements and correctable cases are routed to relabeling or further review." Nothing is simply thrown away if it can still teach something at a lower tier.
| Benchmark (0–100, higher better) | StepAudio 3 Realtime | Doubao 2.0 Lite | Gemini 3 Flash | Gemini 3.1 Pro |
|---|---|---|---|---|
| Big Bench Audio | 98.1 | 98.8 | 99.4 | 99.6 |
| AudioMultiChallenge | 49.3 | 48.5 | 56.6 | 67.0 |
| MMSU | 90.6 | 80.0 | 77.0 | 83.6 |
| MMAU | 79.0 | 77.5 | 77.6 | 80.5 |
| WildSpeech | 77.1 | 73.9 | 74.4 | 77.7 |
| MMAR | 86.5 | 75.9 | 75.4 | 81.7 |
| Step-Caption | 78.2 | 76.8 | 67.8 | 74.8 |
| MTalk-Bench | 91.7 | 89.9 | 88.5 | 89.1 |
| Macro average | 81.3 | 77.7 | 77.1 | 81.8 |
The paper's reading: StepAudio 3 Realtime "leads the reported baselines on four of the eight benchmarks," with the largest margins on MMSU (90.6 vs 83.6, +7.0) and MMAR (86.5 vs 81.7, +4.8). It "trails Gemini 3.1 Pro by 17.7 points on AudioMultiChallenge" (67.0 − 49.3 = 17.7), and Big Bench Audio "is nearly saturated for all systems." The paper draws the conclusion itself: "maintaining and revising constraints over natural multi-turn audio remaining a clear area for improvement."
Two footnotes from the protocol section (8.3) matter for reading those rows. Step-Caption "uses judge-based scoring against annotated speaker attributes." And the MTalk-Bench entry "includes only the Paralinguistic Information and Ambient Sound components, reported as one aggregate", so it is not the full benchmark.
Section 4.2 ends with the chapter's most transferable result. The team ran an ablation where "the SFT data are the only changed factor": roughly two million randomly sampled examples versus about 100K high-quality examples retained after the quality control described above.
| Metric | ~2M random | ~100K quality-controlled | Change |
|---|---|---|---|
| MMSU | 78.78 | 89.70 | +10.92 |
| MMAR | 74.70 | 84.50 | +9.80 |
| WildSpeech | 74.20 | 77.11 | +2.91 |
| MTalk-Bench (ambient, paralinguistic, semantic subsets) | 88.83 | 90.84 | +2.01 |
Check the arithmetic: 89.70 − 78.78 = 10.92; 84.50 − 74.70 = 9.80; 77.11 − 74.20 = 2.91; 90.84 − 88.83 = 2.01. And the size ratio: 2,000,000 ÷ 100,000 = 20, the paper's "roughly one twentieth as many examples."
You are explaining a recipe to a friend over the phone. Halfway through, they say "right." Do you keep going?
It depends. If "right" came in the gentle rhythm of someone following along, you keep going. If it came sharply, followed by a breath, it is the start of "right, but you said two eggs earlier", and you should stop. The paper uses this exact word as its example: "'right' may function as a backchannel acknowledging the model's explanation or as a preface to a correction."
People resolve that ambiguity dozens of times a minute without noticing. A voice model has to do it explicitly, from audio, in real time. This chapter is about how StepAudio 3 Realtime does it.
Picture a talking stick passed around a circle: whoever holds it speaks. Linguists call the right to speak the conversational floor. The paper frames the whole problem around it: "A central challenge in full-duplex dialogue is determining when to take, retain, or yield the conversational floor."
Three verbs, and each has a matching failure. Take the floor too early, and you interrupt someone who was only pausing. Retain it too stubbornly, and you talk over a correction. Yield it too easily, and every "mm-hm" kills your answer.
Two vocabulary words before we go on. A backchannel is a short listener signal ("mm-hm", "yeah", "right") that shows engagement without asking for the floor. An interruption (in voice products, often called barge-in) is a substantive attempt to take the floor while someone else is speaking.
The paper names the two distinctions that matter: "distinguishing pauses within an unfinished utterance from turn completion, and brief acknowledgments from attempts to interrupt." And it says what resolving them requires: "acoustic evidence interpreted in the context of the unfolding dialogue." Sound alone is not enough. Words alone are not enough. The model needs both, plus history.
Figure 5(A)'s caption names three inputs that jointly inform one decision: user speech, model-side speech, and dialogue history (the drawing itself shows the two audio streams feeding the streaming encoder). Each carries evidence the others lack.
| Input | Evidence it carries | Decision it helps most |
|---|---|---|
| User speech | Is there voice? How loud, how long, what words, what prosody? | Pause vs end; backchannel vs interruption |
| Model speech | What has the user already heard at this instant? | Is the overlap a reaction to what was just said? |
| Dialogue history | What was asked, what is pending, who has been addressed | Is this speech even directed at the assistant? |
Section 5.1 gives the most concrete detail of the duplex design: "Audio is organized into 320 ms blocks, each followed by a state or text token." Every block of listening ends with the model saying, in its own token vocabulary, what it is doing now.
Figure 5(B) draws four of these blocks on a timeline from 0.00 s to 1.28 s. Each block is drawn as a start marker S, four audio cells a, an end marker E, and then a state token. The drawn state tokens, in order, are:
Read that line as a tiny story. In the first 320 ms, nobody is talking: the model notes "no voice" and keeps listening. In the next two blocks, the user speaks: "user voice", still listening. In the fourth block, the model decides the user is done and emits a "speaking transition": it takes the floor. The caption above the timeline reads: "One state or text token follows every 320 ms audio block."
Here is the same drawing as a table, with a middle column of our own reading: the kind of evidence each decision would need. The paper does not annotate the figure with evidence values; the first and last columns are taken directly from it.
| Block (Fig. 5B) | Contents | Evidence a correct decision needs (our reading) | State token (Fig. 5B) |
|---|---|---|---|
| 0.00–0.32 s | S a a a a E | No voice in the user cells | listening: no_voice |
| 0.32–0.64 s | S a a a a E | User voice present; an utterance has started | listening: user_voice |
| 0.64–0.96 s | S a a a a E | User voice continues; the utterance is not yet complete | listening: user_voice |
| 0.96–1.28 s | S a a a a E | The utterance is judged complete and the floor is free | speaking transition |
Notice what the fourth row requires. Nothing in a single block's audio says "the user is done". The decision depends on what was said in the previous blocks, which is why the state token is predicted by a decoder that sees the whole history, not by a classifier that sees one block.
Section 5.1 names the full set of decisions these tokens encode: "continue listening, initiate a response, continue speaking, or yield the conversational floor." The figure spells out three token names; the report does not list the exact spelling of the others, so in this lesson we call them by the paper's verbs.
Section 5.1 adds: "The model also uses its own ongoing speech to interpret overlapping user utterances in the context of what the user is currently hearing."
Suppose the model is reading out three restaurant options. The user says "that one!" during option two. If the model only knew what it planned to say, "that one" would be ambiguous. Because the model audio stream is part of its input (the feedback wire from Chapter 1), the decoder knows that the user had just heard the name of option two when they spoke. The overlap is interpreted against the words that were actually in the air.
Section 5.2 walks through three decisions, each with its own mix of evidence.
Pauses and turn endings. "The system combines acoustic timing with semantic completeness to distinguish within-turn pauses from turn endings." Semantic completeness asks whether the words so far form a finished request. "Book me a table for" followed by silence is acoustically a pause and semantically incomplete: keep listening. "Book me a table for two at seven" followed by the same silence is complete: respond. The silence is identical; the words decide.
Backchannels and interruptions. "During model speech, user backchannels are interpreted in relation to the ongoing response. Brief acknowledgments can signal continued engagement without requesting a floor transfer, whereas a substantive request or correction may signal an intent to interrupt." The deciding question is whether the overlap asks for something. "Mm-hm" asks for nothing: keep speaking. "Wait, I meant Tuesday" asks for a change: yield.
Background speech rejection. "Dialogue history provides contextual evidence for assessing whether incoming speech is directed at the assistant." A television announcing weather in another city, or a colleague talking to someone else, is speech, but not for the assistant. If the conversation so far is about a dinner booking, a voice saying "and now, sports" is almost certainly background. The paper's verb is careful: the assessment "informs whether the speech should be incorporated into the active exchange or treated as unrelated background conversation."
Inject events and watch two policies react block by block. The top decision row is a context-aware policy in the spirit of Section 5.2: it combines voice, semantic completeness, whether speech is addressed to the assistant, and whether an overlap is substantive. The bottom row is a plain silence timer that takes the floor after two silent blocks and yields to any voice. A red outline marks a decision that disagrees with the scenario's intended behavior.
Evidence values (completeness, addressed, substantive) are illustrative: the real model infers them implicitly from both audio streams and dialogue history. Playback is slowed so you can read it; in the model, each column is 320 ms. Chip codes: L·nv = listening: no_voice, L·uv = listening: user_voice, →S = speaking transition (these three are spelled in the paper's Figure 5); S = continue speaking, Y = yield, L·bg = keep listening, speech not addressed to the assistant (paper's verbs; token spellings not reported).
Run each scenario and compare the rows. The silence timer interrupts the paused user, is slower than necessary at genuine turn ends, abandons its answer at the first "right", and happily answers the television. It gets exactly one case right: the real interruption. The context-aware row gets all five, because each decision uses a different piece of evidence: completeness for pauses, substantiveness for overlaps, addressedness for background speech.
That pairing is the one the paper highlights in its evaluation: "respecting within-turn pauses while responding at turn completion, and accommodating user interruptions while continuing through backchannels." Each pair is a tension. A policy that is good at one side of a pair by being trigger-happy or stubborn fails the other side.
How do you turn two continuous audio streams and a decision into training data for an autoregressive decoder? The sketch below builds the time-interleaved sequence of Figure 5(B). Block markers and the three state names follow the figure; the remaining token names and the target labels are ours.
python (sketch)BLOCK_MS = 320 # paper, section 5.1 def serialize_duplex(user_audio, model_audio, floor_labels): """user_audio, model_audio: aligned streams (same clock). floor_labels[k]: the decision after block k, e.g. 'listening: no_voice', 'listening: user_voice', 'speaking transition', or a text token.""" seq, loss_mask = [], [] for k, label in enumerate(floor_labels): t0, t1 = k * BLOCK_MS, (k + 1) * BLOCK_MS blk = ["<S>"] blk += audio_slots(user_audio[t0:t1]) # user cells blk += audio_slots(model_audio[t0:t1]) # model cells: what the user is hearing blk += ["<E>"] seq += blk + [label] loss_mask += [0] * len(blk) + [1] # learn the decision, not the audio return seq, loss_mask # Training objective: ordinary next-token cross-entropy on the masked positions. # The "decision" is just the next token after <E>; no separate classifier head.
The loss mask is the realization of "turn-taking as next-token prediction." Audio positions provide context; decision positions provide supervision. The paper does not publish its exact masking; this is the standard way to express "predict a state or text token after each block."
Section 5.3 describes two phases.
Midtraining "adapts the model to the time-interleaved representation used for full-duplex interaction." It "combines supervision for streaming ASR, voice activity detection (VAD), and streaming prediction of utterance completeness." The mixture "includes over 10,000 hours of synthetic full-duplex interaction data," and "text data are also incorporated to help retain general language and reasoning capabilities."
Look at how neatly those three supervision signals line up with the decisions.
| Midtraining signal | What it teaches | Decision it feeds |
|---|---|---|
| Streaming ASR | Which words have been said so far, block by block | Semantic completeness; substantive vs brief overlap |
| VAD | Whether a block contains voice | "listening: no_voice" vs "listening: user_voice" (Figure 5's own labels) |
| Streaming utterance completeness | Whether the utterance so far is finished | Pause vs turn ending; when to emit "speaking transition" |
Why synthetic duplex data? Real two-channel conversational recordings with clean per-speaker separation and labelled floor events are scarce. Synthesis lets you control exactly where pauses, overlaps and background voices occur, so every block can carry a correct label. The report does not describe how its 10,000+ hours were synthesized.
Post-training "further refines conversational behavior using high-quality interaction data covering turn taking, user backchannel handling, interruption handling, and background speech rejection." Those four behaviors map one-to-one onto the scenarios in Sim 3, and three of them onto the benchmark below.
Each floor decision leans on a particular piece of evidence, so each has a particular way to fail when that evidence degrades. The paper does not report per-condition duplex ablations; the table below is our analysis of which decision is exposed to which degradation, using only the mechanisms the paper describes.
| Degradation | Evidence it corrupts | Likely wrong decision | Which input can rescue it |
|---|---|---|---|
| Slow, hesitant speaker with long pauses | Acoustic timing | Taking the floor mid-request | Semantic completeness from the words so far |
| Clipped, complete-sounding fragment ("two.") | Semantic completeness | Responding before the user adds "…and a high chair" | Dialogue history (was a list being built?) and prosody |
| Loud room, TV in the background | Voice presence | Answering speech not meant for the assistant | Dialogue history: is this on-topic and addressed? |
| User says "right" in a flat, ambiguous tone | Prosodic cue for intent | Yielding to an acknowledgment, or ignoring a correction | The words that follow in the next block, and the model's own speech |
| Model's own voice leaking into the user mic | User stream purity | Treating its own echo as an interruption | The model audio stream: the model knows what it just said |
Two points stand out. First, every row has a rescue that comes from a different input than the one that was corrupted. That redundancy is the practical argument for feeding the decoder all three inputs rather than a single voice-activity signal.
Second, the 320 ms cadence makes waiting cheap. When evidence is ambiguous, "keep listening for one more block" costs only 320 ms, and the next block usually disambiguates. A silence timer has no such option: it either fires or it does not. The echo row is our own inference from the dual-stream design; the paper does not discuss acoustic echo explicitly.
The paper evaluates on the Artificial Analysis (AA) subset of Full Duplex Bench v1 and v1.5, which scores four aspects: pause handling, turn taking, user interruption handling, and backchannel handling. Category scores "measure the percentage of samples satisfying the corresponding interaction criterion."
| Capability (0–100) | StepAudio 3 Realtime | GPT-realtime-2 (High) | Qwen Audio 3.0 Realtime Plus | Grok Voice Think Fast 2.0 High |
|---|---|---|---|---|
| Pause handling | 98.9 | 99.3 | 98.0 | 98.0 |
| Turn taking | 100.0 | 100.0 | 98.0 | 91.0 |
| User interruption handling | 99.0 | 95.0 | 98.0 | 97.0 |
| Backchannel handling | 98.0 | 86.7 | 100.0 | 95.0 |
| Overall | 98.9 | 95.3 | 98.4 | 95.1 |
StepAudio 3 Realtime ranks first overall at 98.9, ahead of Qwen Audio 3.0 Realtime Plus at 98.4. It is not best in every row: GPT-realtime-2 is slightly better at pauses (99.3), and Qwen is perfect on backchannels (100.0). The paper's claim is about balance: "Strong performance across both pairs indicates balanced conversational control over when to listen, speak, and yield."
Look at GPT-realtime-2's profile: 100.0 on turn taking and 86.7 on backchannels. Our reading (the paper does not analyze baseline behavior): this profile is consistent with a policy that responds promptly but also treats some acknowledgments as interruptions, exactly the tension from Sim 3.
Early in a long call, you mention that you do not eat meat. Twenty minutes and a dozen topics later, you ask for a quick dinner idea. The assistant cheerfully suggests a steak.
Its timing was perfect. It did not interrupt you. It did not go silent. It was still a bad conversation partner, because it forgot a constraint you gave it.
The paper draws this line sharply at the start of Section 6: "Seamless Duplex determines when the model should respond. Conversational intelligence determines how it should engage with the user and how much reasoning the response requires." Chapter 3 was about when. This chapter is about what, and about how the team measured and trained it.
The paper's description of the target behavior is worth reading as a checklist: "A natural voice assistant should follow intent across turns, clarify underspecified goals, and move the conversation toward a useful outcome. Routine turns should avoid unnecessary deliberation, while complex requests should retain the reasoning needed for a reliable answer."
To improve a skill, you first need a ruler. The team built its own: StepAudioChat, "a closed, text-based benchmark for foundational conversational intelligence." Two words in that sentence are design decisions.
Text-based. "Its scope isolates text-level response quality from prosody, turn timing, interruption handling, and other properties of the speech interface." If a benchmark played audio in and scored audio out, a low score could mean bad reasoning, bad synthesis, or bad timing, and you could not tell which. Scoring text separates the question "was the content right?" from "was the delivery right?" Delivery is measured elsewhere (Chapter 3).
Closed. The items are newly constructed and not published, "to reduce reliance on public test questions." Public test sets leak into training data over time, a problem called contamination: a model can score well because it has seen the questions, not because it has the skill. A closed set cannot be memorized from the web. The cost, which we should say plainly, is that outsiders cannot reproduce StepAudioChat numbers.
Section 6.1.1: "We organize conversational abilities into a hierarchy whose leaves target observable behaviors with defined evaluation boundaries." A leaf is not "is helpful"; it is something a checker can observe, with a clear line between pass and fail.
How do you know the taxonomy covers the right ground? The team mapped "tasks and representative examples from public benchmarks" onto it as "a coverage check, without reusing their test questions." Where public examples mapped to several leaves, the overlap was used to sharpen definitions; where examples mapped to nothing, the team had found a gap.
Within each capability family, items vary "the source of a constraint, its form of expression, and its interaction with other conditions." That produces the hard cases: "nested constraints and requirements that must remain consistent across turns." Every item "targets a primary capability," and "only capabilities supported by validated items enter the evaluation suite."
The benchmark reports eight dimensions. The paper names them but does not define each one in prose, so the right-hand column below is our short gloss of what the name suggests, not a quotation.
| StepAudioChat dimension | Our gloss (illustrative) |
|---|---|
| Instruction Following | Honors explicit constraints: length, format, what to include or avoid |
| Faithfulness | Stays true to given facts and context; does not invent |
| Reasoning | Works multi-step problems correctly inside dialogue |
| Memory | Uses information from earlier turns ("contextual recall") |
| Knowledge | Knows facts about the world |
| Safety & Reliability | Handles risky requests and uncertainty responsibly |
| Conversational Pragmatics | Reads intent, clarifies ambiguity, fits the social moment |
| Persona & Role Consistency | Keeps an assigned role, voice and style over the conversation |
Section 6.1.2: "Each item contains a dialogue prompt, independently checkable criteria, a valid reference response, and a deliberately flawed response."
The last two are the clever part. In a lab experiment, a positive control is a sample you know should test positive, and a negative control is one you know should test negative. If your test fails either control, you do not trust any of its results. The paper uses the valid and flawed responses exactly that way: they "provide positive and negative controls for the judging criteria while allowing multiple valid phrasings."
Here is an illustrative item in that shape (ours, not from the closed set):
| Part | Content |
|---|---|
| Dialogue prompt | Turn 1: "i'm vegetarian btw." … Turn 6: "ok what's a quick dinner, like 20 min" |
| Criterion A | The suggestion contains no meat or fish. |
| Criterion B | The suggestion is plausibly doable in about 20 minutes. |
| Criterion C | The reply is short enough to speak comfortably. |
| Valid reference | "A chickpea and spinach stir-fry: about fifteen minutes with canned chickpeas." |
| Flawed response | "A quick garlic shrimp pasta takes about twenty minutes." (violates A) |
Notice the prompt's style: lower case, fragmentary, casual. That is deliberate. "De-identified utterances from real interactions inform prompt phrasing and local context, preserving brevity, fragmentation, colloquial wording, and transcription noise." People do not speak to assistants in tidy paragraphs, and a benchmark written in tidy paragraphs flatters models.
Two privacy and scope rules follow. "Personal entities are replaced with typed placeholders" (a name becomes something like a PERSON slot). The real utterances "do not supply expected answers or grading criteria"; they only shape how prompts sound. And "except for intrinsically domain-specific capabilities, scenarios are recast across everyday settings to broaden coverage beyond individual applications."
Section 6.1.3 stacks the checks.
Difficulty calibration deserves a picture. Run a weaker and a stronger reference system on every item and sort by outcome. The paper names three kinds of items; the mapping from outcomes to kinds below is our reading of those names.
| Weaker system | Stronger system | Item kind (our mapping) | What it measures |
|---|---|---|---|
| pass | pass | Baseline | Floor competence; catches regressions |
| fail | pass | Discriminative | Separates stronger from weaker models |
| fail | fail | Difficult | Headroom above today's systems |
| pass | fail | Suspicious | Often a sign of ambiguous wording or inconsistent grading, worth auditing |
The paper's own summary of why all this matters: "These checks help distinguish capability demands from ambiguous wording or inconsistent grading." A benchmark item that fails a model for the wrong reason is noise in the score.
Section 6.2 turns from measuring to training. The goals: "follow user intent and constraints across turns, clarify underspecified requests, and adapt responses to the conversational context." And one sentence that sets up Chapter 7: "Joint training on dialogue and reasoning examples supports deliberation on difficult requests and concise responses to routine turns."
Figure 6 of the paper draws the construction as five numbered boxes with two side inputs.
The resulting training example, in the figure's words: "the first N−1 turns as context, the final query, and the labelled response with its reasoning."
Self-play means two models converse to generate data. It is cheap and unbounded, and it has one weakness: the synthetic assistant turns may be mediocre. The pipeline handles that by using the self-played turns only as context. The only thing the student model is trained to produce is the carefully labelled final answer. That is why "no answer is produced" at step 4: the final answer comes from a separate, higher-quality labelling step.
python (sketch)def build_dialogue_example(case_plan, user_model, asst_model, labeler, N): history = [] for t in range(N - 1): # self-play: N-1 turns of context history.append(("user", user_model.speak(case_plan, history))) history.append(("assistant", asst_model.reply(history))) query = user_model.final_query(case_plan, history, # capability injection is optional inject=case_plan.capabilities) reasoning, answer = labeler.label(history, query) # the ONLY target tokens, mask = [], [] for role, text in history + [("user", query)]: ids = tok(role, text); tokens += ids; mask += [0] * len(ids) # context: no loss ids = tok("assistant", think(reasoning) + answer) tokens += ids; mask += [1] * len(ids) # loss on reasoning + answer return tokens, mask
The mask line is the realization of "this one answer is the only target." Whether the paper trains on self-played assistant turns at all is not stated beyond that figure label; the sketch follows the figure.
Generation "follows four axes: topics and their concrete discussion points; participant personas that specify roles, backgrounds, and interaction styles; turn depth; and target capabilities such as contextual recall, logical reasoning, and instruction use."
If you sample each axis independently from its overall popularity, the common combinations crowd out the rare ones. Popular personas meet popular capabilities over and over, and some combinations never appear. The paper's fix: "A stratified rotation schedule coordinates these choices within topics, extending coverage beyond their global marginal distributions."
A marginal distribution is the frequency of one axis on its own (how often each persona appears, ignoring everything else). Matching the marginals does not guarantee good coverage of combinations. A stratified rotation deliberately cycles through combinations within each topic instead of sampling each axis independently. In its idealized form (our illustration, used in the simulation below), every persona meets every capability at every depth before any combination repeats. The paper states only that "a stratified rotation schedule coordinates these choices within topics, extending coverage beyond their global marginal distributions"; it gives no stronger guarantee.
One topic. 4 personas × 4 capabilities × 3 turn depths = 48 combinations (sizes and the skewed marginals are illustrative). Each cell's shade is how many dialogues landed on that combination. The curve panel tracks coverage for both strategies as you draw more dialogues.
Draw 48 dialogues in each mode. Rotation covers all 48 combinations exactly once. Independent sampling, with the same budget, leaves a sizeable share of cells empty and piles many dialogues onto the popular corner. The model trained on the second set would rarely practise, say, a terse expert persona asking for deep contextual recall on turn twelve.
"Turn depth is an explicit construction variable. Longer dialogues progress through deeper engagement with discussion points, allowing later turns to refine constraints, resolve ambiguity, revisit evidence, or change direction while remaining consistent with the history."
Why make depth explicit rather than letting it vary naturally? Because the behaviors that break in long conversations (forgotten constraints, contradicted earlier answers) only appear at depth. If most synthetic dialogues are three turns long, the model rarely practises turn fifteen. The paper later admits that "multi-turn constraint following" remains a gap; explicit depth is the data-side lever aimed at it.
"Quality control separates three concerns into independent stages."
| Stage | What it checks |
|---|---|
| 1 · Context–query review | Depth and informativeness of the history; whether the opening is self-contained; whether the final query is substantive and connected to what came before |
| 2 · Capability-specific review | When capabilities were injected: does the query actually instantiate each one, and does the response actually satisfy it? |
| 3 · Response review | Answer quality, persona and style consistency, and "naturalness as spoken dialogue" |
On top of these, "deterministic checks handle defects such as empty output, malformed tokens, severe repetition, and role confusion." Role confusion is when a self-play model forgets which side it is playing, a common failure of two-model generation.
The retention rule is strict: "Only examples that pass these checks and receive high scores in all three quality-control stages are retained for supervised fine-tuning; all other examples are rejected." Compare that with the audio data in Chapter 2, where near-misses were routed to midtraining or relabelling. For dialogue, a near-miss is simply discarded. Chapter 2's "less is more" result is a good reason to believe that strictness pays.
Before adding any realtime machinery, the team evaluated StepAudio 3 "in reasoning mode", with explicit thinking always available and no speaking deadline. "This evaluation isolates the model's conversational and reasoning capability." The realtime system, with Think-While-Speaking, Adaptive Thinking and MTP, is evaluated separately (Chapter 6 and Chapter 9).
| Dimension | StepAudio 3 (Reasoning) | Doubao 2.0 Lite | DeepSeek-V4-Flash | Kimi K3 |
|---|---|---|---|---|
| Instruction Following | 66.3 | 72.9 | 71.4 | 68.9 |
| Faithfulness | 72.4 | 67.5 | 75.3 | 78.4 |
| Reasoning | 73.0 | 72.7 | 64.8 | 81.9 |
| Memory | 72.0 | 71.3 | 71.5 | 77.6 |
| Knowledge | 73.1 | 59.9 | 71.6 | 78.6 |
| Safety & Reliability | 79.0 | 75.9 | 79.9 | 84.8 |
| Conversational Pragmatics | 67.2 | 61.5 | 62.9 | 70.3 |
| Persona & Role Consistency | 80.9 | 82.6 | 73.6 | 76.5 |
| Macro Average | 73.0 | 70.5 | 71.4 | 77.1 |
The paper's reading: StepAudio 3 "ranks second on reasoning, memory, knowledge, conversational pragmatics, and persona and role consistency." Kimi K3 "leads six of the eight dimensions," while Doubao 2.0 Lite "leads instruction following and persona and role consistency." The weakest cell for StepAudio 3 is instruction following at 66.3, the lowest of the four systems. The paper's lesson: "strong aggregate reasoning does not imply uniformly stronger instruction following or role consistency."
You fine-tune your model on a big batch of conversation data. Dialogue scores go up. Then you check math, and math has gone down. You add more math data and retrain; now audio understanding slips. Every capability you push on seems to pull another one back.
The textbook fix is to train one model on the union of every dataset, with carefully tuned proportions. But every time one team improves its data, the whole mixture has to be rebalanced and the whole model retrained. With a model of roughly 196 billion parameters, that is an expensive habit.
Section 6.4 of the paper takes a different path, and it is disarmingly simple: train several specialists separately from the same starting point, then average their weights.
Picture four apprentice chefs trained in the same kitchen by the same master. One spends a year on sauces, one on pastry, one on grilling, one on a bit of everything. Because they learned in the same kitchen, their stations are laid out identically: the same knife is in the same drawer for all four. You could reasonably blend their notebooks page by page, because page 40 means the same thing in each.
The paper's version: "We train multiple compatible teacher checkpoints from a common base model, using a different data composition for each teacher." The mixtures "emphasize complementary capabilities, including multi-turn dialogue, audio understanding, general text reasoning and knowledge, and targeted mixed-domain behavior." The result: "teachers that are individually strong in different regions of the capability space while preserving parameter alignment for merging."
Two terms need unpacking. A checkpoint is a saved copy of all of a model's weights at some point in training. Parameter alignment means that the same position in each weight tensor plays the same role in every teacher, like the same knife in the same drawer.
Why does a common base give alignment? This is background knowledge rather than a claim in the report. Neural networks have permutation symmetry: you can shuffle the order of neurons inside a layer (and shuffle the connected weights to match) without changing what the network computes. Two networks trained from different random starts may end up with the "same" features in different positions, so averaging them position by position mixes unrelated features. Fine-tuning from a shared base changes weights only modestly, so each neuron keeps its role, and position-wise averaging stays meaningful.
How do you combine four sets of weights into one? The paper uses "directly averaging their parameters" with weights:
Here θi is the full parameter vector of teacher i (every weight in the network, laid end to end), αi is how much teacher i contributes, and θmerge is the resulting model. The two constraints make this a convex combination: a weighted average with non-negative weights that sum to one, so the merged point lies "between" the teachers and never outside them. The operation is applied tensor by tensor, and within each tensor, element by element.
A useful rewriting shows what merging really combines. Write each teacher as the base plus a change: θi = θbase + Δi, where Δi is what fine-tuning taught teacher i. Substitute, one step at a time:
So the merged model is the shared base plus a weighted blend of what each specialist learned. If the specialists' updates mostly touch different directions in weight space, the blend keeps much of each. If they conflict, the blend dilutes them. This reading is standard in the model-merging literature; the report itself states only the first equation.
"The reported model uses four teachers with a normalized 3:1:1:1 weighting." The coefficients "are selected against held-out evaluations spanning dialogue, audio understanding, and general text capabilities." The report does not say which of the four teachers receives the weight of 3, so we will not assume it.
One more property the paper emphasizes: "Because integration occurs in parameter space, it introduces neither additional model components nor inference-time routing." Contrast this with an ensemble, which runs several models and combines their outputs (paying for all of them at inference), or with routing a request to the right specialist (paying for a router and keeping all specialists deployed). A merged model costs exactly one model at inference.
pytorchdef merge_teachers(state_dicts, ratios): """theta_merge = sum_i alpha_i * theta_i (paper, section 6.4) state_dicts: teacher checkpoints fine-tuned from ONE common base. ratios: e.g. [3, 1, 1, 1] -> normalized to alpha = [0.5, 1/6, 1/6, 1/6].""" assert all(r >= 0 for r in ratios) alphas = [r / sum(ratios) for r in ratios] # sum to 1 keys = state_dicts[0].keys() assert all(sd.keys() == keys for sd in state_dicts) # same architecture merged = {} for k in keys: # tensor by tensor merged[k] = sum(a * sd[k].float() for a, sd in zip(alphas, state_dicts)) return merged # one model: no router, no extra modules # Coefficient search (the paper selects alphas on held-out dialogue, audio and text evals): # for ratios in candidate_grid: score(merge_teachers(teachers, ratios)) -> keep best balance
The picture below is a two-dimensional cartoon of weight space (real weight space has about 196 billion dimensions). The shaded region is a toy low-loss basin around the shared base. Four teachers sit in different directions from the base, each pushed toward its own capability. The sliders set the merge ratios; the merged point is their convex combination. Flip to "misaligned teachers" to see what happens when the teachers do not share a base.
The landscape, teacher positions and loss readout are illustrative; only the 3:1:1:1 preset comes from the paper. In the misaligned view, two teachers came from a different starting point, so their neurons sit in permuted positions: the same skills live in a mirrored basin, and averaging across basins lands on a ridge.
Three experiments to run in the simulation. Press "Only teacher A": the merge sits exactly on teacher A, because α = (1, 0, 0, 0) is still a valid convex combination. Press "Equal 1:1:1:1": the merge moves to the centroid of the four teachers. Now switch to the misaligned view and press "Equal" again: the centroid of two basins is the ridge between them, and the toy loss jumps above every teacher's.
The two-teacher case makes the geometry exact. With only teachers 1 and 2 and weights (1 − λ, λ), the merge is θ(λ) = (1 − λ)θ1 + λθ2, the straight segment between them as λ runs from 0 to 1. Whether a merge works is therefore the question of whether the loss stays low along that segment, which is what a shared base makes likely and separate starts make unlikely.
The teacher labels in the simulation (dialogue-leaning and so on) are ours; the paper lists the capability emphases but not which numbered teacher got which mixture, nor which got the weight of 3. What the cartoon should leave you with: convex weights can only move the merged point inside the teachers' hull, and that hull is only a safe place to be if the teachers live in the same basin.
The paper compares the four teachers with their merge across three domains. Dialogue is "measured with the model in reasoning mode."
| Benchmark | Merged | T1 | T2 | T3 | T4 |
|---|---|---|---|---|---|
| BigBench Audio | 98.1 | 96.1 | 98.1 | 98.4 | 98.5 |
| AudioMultiChallenge | 49.3 | 49.1 | 49.8 | 50.2 | 47.6 |
| MMSU | 90.6 | 85.5 | 85.5 | 90.4 | 90.9 |
| MMAU | 79.0 | 78.5 | 78.8 | 77.8 | 77.7 |
| WildSpeech | 77.1 | 76.5 | 76.4 | 75.9 | 76.1 |
| MMAR | 86.5 | 85.4 | 86.4 | 87.3 | 86.5 |
| Step-Caption | 78.2 | 75.8 | 78.3 | 79.6 | 79.4 |
| MTalk-Bench | 91.7 | 91.8 | 90.3 | 90.9 | 90.7 |
| Audio macro | 81.3 | 79.8 | 80.5 | 81.3 | 80.2 |
| HMMT 2026 Feb | 86.8 | 44.0 | 82.2 | 81.3 | 79.8 |
| GPQA Diamond | 83.0 | 73.1 | 81.9 | 80.1 | 80.1 |
| MultiChallenge | 59.7 | 54.6 | 51.3 | 50.6 | 59.7 |
| General text macro | 76.5 | 57.2 | 71.8 | 70.7 | 73.2 |
| Instruction Following | 66.3 | 64.2 | 64.5 | 69.3 | 64.2 |
| Faithfulness | 72.4 | 74.0 | 65.6 | 67.1 | 73.6 |
| Reasoning | 73.0 | 75.2 | 67.0 | 67.2 | 67.8 |
| Memory | 72.0 | 73.4 | 70.0 | 68.3 | 71.5 |
| Knowledge | 73.1 | 75.1 | 63.7 | 62.0 | 75.6 |
| Safety & Reliability | 79.0 | 79.2 | 75.7 | 75.6 | 78.0 |
| Conversational Pragmatics | 67.2 | 67.8 | 62.4 | 63.4 | 64.5 |
| Persona & Role | 80.9 | 84.7 | 79.5 | 76.9 | 83.0 |
| Dialogue macro | 73.0 | 74.2 | 68.6 | 68.7 | 72.3 |
The paper's summary, domain by domain: on audio, the merge "reaches a macro average of 81.3, tying the best teacher macro average while leading on MMAU and WildSpeech"; on general text, it "achieves the highest macro average of 76.5, with the best HMMT 2026 Feb and GPQA Diamond scores and a tie on MultiChallenge"; on dialogue, it "has a macro average of 73.0, exceeding Teachers 2, 3, and 4, while remaining below the strongest dialogue teacher at 74.2."
Read the teacher profiles, which the table reveals even though the paper does not name each teacher's mixture. Teacher 1 is the best conversationalist (dialogue macro 74.2) and by far the weakest mathematician (HMMT 44.0). Teacher 3 ties the best audio macro (81.3). Teacher 4 is the strongest text teacher (73.2) and nearly the best at dialogue (72.3). No teacher is best at everything, which is exactly the premise of merging.
The punchline: the merged model is not a mixture of teacher behaviors in any simple sense. On math and science it beats every teacher it was built from. A plausible explanation (ours; the paper does not analyze it) is that the averaged fine-tuning updates act partly as a regularizer, cancelling idiosyncratic noise in each specialist while keeping directions they share.
The paper calls the merge "a balanced trade-off." We can make that word precise with Table 8 by ranking the merged model against the four teachers on every row. Rank 1 means the merge is best; rank 5 means it is worst.
The same exercise on the eight audio rows gives a wider spread: rank 1 on MMAU and WildSpeech, rank 2 on MMSU and MTalk-Bench, tied for second on MMAR (86.5, level with Teacher 4), tied for third on BigBench Audio (98.1, level with Teacher 2), rank 3 on AudioMultiChallenge, and rank 4 on Step-Caption (78.2, ahead only of Teacher 1's 75.8). On the three text rows it is first on two and tied first on the third.
That is what "balanced" looks like in numbers: a model that is rarely the single best specialist, almost never near the bottom, and never pays the kind of 19-point penalty a specialist pays outside its domain. For a product that must do all three jobs in one conversation, the worst-case row matters more than the best-case row.
Merging is one of several ways to get several specialists' skills into one product. The comparison below is general engineering background, with the merge column filled from the paper.
| Strategy | Training cost when one data mix changes | Inference cost | What the paper says about it |
|---|---|---|---|
| Train one model on the union of all data | Retrain the whole model | One model | Merging "allows teacher data mixtures to be developed independently and then recombined without retraining a single model on the full union of data." |
| Keep specialists and route or ensemble | Retrain one specialist | Several models plus a router, or several forward passes | Merging "introduces neither additional model components nor inference-time routing." |
| Weighted parameter merge (this paper) | Retrain one teacher, then re-average | One model | Coefficients selected on held-out dialogue, audio and text evaluations; four teachers at 3:1:1:1 |
For a 196B-total-parameter realtime model that already runs two concurrent calls per deliberate turn (Chapter 6), the inference column is decisive: a router or an ensemble would multiply the serving cost of every turn.
Merging is not free. Scan the rows where the merge falls below the best teacher: Persona & Role 80.9 versus 84.7 (T1), a 3.8-point gap; Instruction Following 66.3 versus 69.3 (T3), a 3.0-point gap; Knowledge 73.1 versus 75.6 (T4); Reasoning (dialogue) 73.0 versus 75.2 (T1); Step-Caption 78.2 versus 79.6 (T3).
The paper says this directly: the results "support model merging as a low-cost mechanism for combining complementary capabilities, while showing that it provides a balanced trade-off rather than uniform improvement over every specialized teacher." And its stated design goal is modest on purpose: merging "is intended to retain complementary strengths rather than make the merged model identical to the best teacher on every metric."
The practical win is organizational as much as numerical. "This also allows teacher data mixtures to be developed independently and then recombined without retraining a single model on the full union of data." A dialogue team and an audio team can iterate on their own teachers, and the merge step recombines their progress.
The report shows a merge that works. It is still worth knowing the ways a weighted average can fail, because each one explains a design choice the paper made. The failure modes below are general background from the model-merging literature, not experiments in this report.
| Failure mode | What happens | The paper's matching choice |
|---|---|---|
| Teachers from different starting points | Permutation mismatch: averaging mixes unrelated neurons, and the merge lands off the low-loss region (Sim 5, misaligned view) | All teachers are trained "from a common base model," "preserving parameter alignment" |
| Conflicting updates | Two teachers push the same weights in opposite directions; averaging cancels both skills | Teacher mixes "emphasize complementary capabilities," so updates overlap less |
| Badly chosen weights | One skill dominates and another fades | Coefficients "selected against held-out evaluations spanning dialogue, audio understanding, and general text" |
| Hidden regressions | A merged model looks fine on averages and fails a specific skill | The paper reports per-benchmark rows and says the merge is a trade-off, not a uniform win |
The last row is the one to carry into your own work. A macro average of 73.0 hides a persona score 3.8 points below the best teacher. If persona consistency matters most for your product, the 3:1:1:1 weights might not be your weights. The paper's framing, "balanced trade-off," is exactly the right warning label.
Back to the question from Chapter 0: departure time, flight time, a time-zone shift, a buffer at arrivals, and a dinner reservation. The merged model from Chapter 5 can reason its way to a good answer. In reasoning mode, it scores 73.0 on StepAudioChat. But reasoning mode, used naively in a voice product, means the user waits in silence while the thinking happens.
This chapter is about the mechanism that removes the wait: Think-While-Speaking. The paper describes it in Section 6.3, and it is the heart of the report's claim to resolve "the tension between deep deliberation and latency."
Picture a lecturer on stage with a research assistant in the wings. A student asks a hard question. The lecturer does not freeze; they start talking: restating the question, laying out what matters. Meanwhile the assistant works the problem on paper and slides notes onto the lectern, one finding at a time. The lecturer speaks from whatever notes have arrived so far. When the assistant finishes, the lecturer wraps up with the full picture, and if an early remark turned out to be off, corrects it at the end.
The paper's design follows that shape. "It builds on the two-process design of Mind-Paced Speaking. Two concurrent calls to the same audio model act as a Formulation Brain and an Articulation Brain."
The Formulation Brain is the assistant in the wings: it "generates a private reasoning trace." Private means the trace is never spoken; the user never hears it.
The Articulation Brain is the lecturer: it "produces short response segments conditioned on the reasoning available so far and on the response already spoken." Those two conditioning sources are the whole trick. "The reasoning available so far" keeps what is said consistent with what has been worked out. "The response already spoken" keeps the answer coherent as it grows segment by segment.
Figure 7(B) of the paper, "Think-While-Speaking with MTP", draws both brains at a single "time step i" with a colour legend: input tokens, think tokens, and response tokens. Here is the figure's wiring as a table.
| Brain | Reads (per Figure 7) | Writes | Who hears it |
|---|---|---|---|
| Formulation | User context (input tokens from the audio encoder and adapter) + its own thought stream so far | The "current think" tokens | Nobody (private) |
| Articulation | User context + "previous think" + "current think" + "previous response" + "current response" | Response tokens | The user, as streaming output audio |
An orange arrow in the figure carries the Formulation Brain's current think tokens into the Articulation Brain's input. A green arrow carries the Articulation Brain's response tokens down to "Streaming Output Audio". A dashed box beside both decoders, labelled "MTP Decoding Acceleration", connects to each of them (Chapter 7). And the input side shows the same audio path as Chapter 1: input audio → audio encoder → audio adapter → input tokens.
In data-flow terms, one step of the system looks like this:
If the Articulation Brain can speak at any time, when should it produce the next segment? The paper's answer: "Playback-aware scheduling releases response segments according to the progress of the streaming output audio while formulation continues in parallel."
In other words, the speaker's own audio playback is the clock. Our reading of that sentence: a new segment is produced as the audio already queued runs down, not all at once at the start. Why is that the right clock? Here is our reasoning, based on what the paper describes.
First, every second of delay before generating a segment is a second more of reasoning it can condition on. Generating the whole answer up front would throw away all the thinking that happens during playback.
Second, speech that has been generated but not yet played is a liability. If the user interrupts (Chapter 3) or new reasoning changes the answer, un-played text is either wasted or wrong. Keeping the buffer short keeps the answer revisable.
Third, the user experiences audio, not tokens. Tying generation to playback keeps the audible stream continuous without racing ahead of the reasoning.
Three more rules complete the mechanism.
After formulation ends. "Once formulation finishes, the remaining response can use the complete reasoning state." Segments released after that point are as well-informed as a think-then-speak answer.
Final continuation. "A final continuation can supplement or correct an answer that began from incomplete reasoning." If an early segment was framed on partial reasoning and the full trace changes the picture, the model says so at the end.
Speak-First vs Think-First. "The system uses Speak-First by default, starting the response without waiting for an initial reasoning prefix. Think-First waits for a short reasoning prefix before beginning the response." Speak-First minimizes silence. Think-First trades a short pause for an opening sentence that is already informed by some reasoning. The report does not quantify the prefix length.
The simulation lays out one deliberate turn on a timeline. The top lane is the Formulation Brain; three findings (F1, F2, F3) appear as the trace grows, and F3 is the final answer. The middle lane shows each response segment at the moment it is generated. The bottom lane is what the user hears. To make the effect of timing visible, the simulation uses an illustrative policy of our own, not a rule from the paper: segments 3, 4 and 5 wait for a finding before they are stated, and segments 1, 2 and 6 are framing that needs none. The paper itself says only that segments are "conditioned on the reasoning available so far" and that a final continuation "can supplement or correct an answer that began from incomplete reasoning."
All durations, the six-segment answer and the finding positions are illustrative. The speed buttons apply the paper's measured wall-clock ratios from Table 6 (1.76× for MTP3 strict, 2.05× for MTP3 with Medusa-style acceptance) to the thinking lane only. With "eager" scheduling, all segments are generated at time zero from no reasoning, so substantive claims are unsupported (red) and a final correction is needed.
Three experiments worth running. Stretch the reasoning length and watch the simulated policy insert more framing segments before the substantive ones, while the first audio still starts at time zero. Switch to Think-First and a short silence appears, but the opening is better informed. Switch to eager scheduling and the same audio timeline now contains unsupported claims, followed by a correction once F3 arrives. The faster thinking lanes shorten the framing phase: the paper's two mechanisms, concurrency and acceleration, compound.
To make the two brains concrete, here is the dinner question from Chapter 0 traced step by step. Everything in this table is illustrative: the wording, the timing and the reasoning content are ours. The paper's stated mechanisms (Speak-First, conditioning on the reasoning so far, the complete reasoning after formulation ends) are labelled as such; the rest of the "rule" column is an illustrative scheduling policy, not something the report specifies.
| Time (illustrative) | Formulation Brain (private) | Articulation Brain (spoken) | Rule at work |
|---|---|---|---|
| 0.0 s | starts: "departure 18:00, flight 2 h 15 min…" | "Let's check the timing for your dinner." | Paper: Speak-First starts without a reasoning prefix |
| 1.2 s | "…lands 20:15 origin time; +1 h zone shift → 21:15 local…" | "Your flight leaves at six and takes a bit over two hours." | Illustrative policy: restate facts already in the context |
| 2.4 s | "…+40 min at arrivals → 21:55…" | "With the time difference, you land at quarter past nine local time." | Illustrative policy: state F1 after the trace contains it |
| 3.6 s | "…21:55 > 21:30, so no" (trace complete) | "After arrivals, you'd be out around five to ten." | Paper: conditioned on the reasoning so far (here, F2) |
| 4.8 s | (finished) | "So a 9:30 table is too tight; ask for 10:15 instead." | Paper: after formulation ends, the rest uses the complete reasoning |
Two details of the trace are worth pointing out. The Articulation Brain's second line is "safe": it contains only facts the user supplied, so it can be spoken before any reasoning exists. And in this illustrative run, the reasoning happened to arrive in time for every substantive line. The paper does not promise that: a segment can be spoken from incomplete reasoning, and the final continuation is the mechanism that repairs it (Section 6.3.2 warns that strict verification of the spoken output "does not eliminate errors arising from incomplete private reasoning").
The paper makes Speak-First the default and offers Think-First as an alternative, without reporting a comparison. The trade-off follows directly from the definitions:
| Property | Speak-First (default) | Think-First |
|---|---|---|
| Silence before the first word | None from reasoning | The length of a "short reasoning prefix" |
| First sentence | Framing, conditioned on little or no reasoning | Already informed by the start of the trace |
| Risk of an early misstep | Higher; the final continuation is the safety net | Lower for the opening |
| Where it fits (our reading) | Casual conversation, where dead air is the worst outcome | Turns where the opening sentence must already commit to a direction |
One more interaction deserves mention, because it links this chapter to Chapter 3. If the user interrupts while a deliberate answer is playing, the model audio stream tells the decoder how much of the answer was actually heard. Combined with playback-aware scheduling, which (on our reading) keeps the unplayed buffer short, that means little generated-but-unheard speech has to be thrown away. This is our inference from the two mechanisms together; the paper does not describe interruptions during Think-While-Speaking explicitly.
A concurrency sketch in Python's asyncio. The two coroutines share one model and one context object; the playback buffer is the scheduler's clock. Function names are ours.
python (asyncio sketch)async def formulation(model, ctx): while not ctx.think_done: toks = await model.decode_think(ctx.user, ctx.think) # MTP-accelerated, private ctx.think += toks ctx.think_done = ends_think(toks) async def articulation(model, ctx, player, speak_first=True): if not speak_first: await ctx.wait_for_think_prefix() # Think-First while not ctx.response_done: await player.wait_until_buffer_low() # playback-aware release (our reading) seg = await model.decode_segment(ctx.user, think=list(ctx.think), # reasoning available SO FAR spoken=ctx.spoken) # response already spoken player.enqueue(seg) # strict verification for speech ctx.spoken += seg ctx.response_done = ctx.think_done and answer_complete(seg) if needs_revision(ctx.spoken, ctx.think): # final continuation player.enqueue(await model.decode_segment(ctx.user, ctx.think, ctx.spoken)) async def deliberate_turn(model, ctx, player): await asyncio.gather(formulation(model, ctx), articulation(model, ctx, player))
Two details in the sketch come straight from the paper.
The think stream is decoded with MTP and permissive "typical acceptance", while spoken segments keep "strict verification" (Section 6.3.2, next chapter).
And the final continuation exists because "a final continuation can supplement or correct an answer that began from incomplete reasoning."
How the real system decides that a revision is needed is not described, so the needs_revision check above is a placeholder.
Now the measurement. The overall evaluation (Table 10) scores StepAudio 3 Realtime on StepAudioChat in "realtime mode", with Think-While-Speaking, Adaptive Thinking and MTP all active. The baselines are in reasoning mode. Set that row beside the reasoning-mode row from Chapter 4.
| Dimension | Reasoning mode (Table 4) | Realtime mode (Table 10) | Change |
|---|---|---|---|
| Instruction Following | 66.3 | 54.1 | −12.2 |
| Faithfulness | 72.4 | 71.9 | −0.5 |
| Reasoning | 73.0 | 73.6 | +0.6 |
| Memory | 72.0 | 71.8 | −0.2 |
| Knowledge | 73.1 | 70.4 | −2.7 |
| Safety & Reliability | 79.0 | 75.1 | −3.9 |
| Conversational Pragmatics | 67.2 | 68.2 | +1.0 |
| Persona & Role Consistency | 80.9 | 78.3 | −2.6 |
| Macro Average | 73.0 | 70.4 | −2.6 |
The picture is lopsided. Reasoning itself holds up (73.0 to 73.6), and conversational pragmatics even improves (67.2 to 68.2). Most of the realtime cost concentrates in one dimension: instruction following falls from 66.3 to 54.1.
What causes that? The paper does not say, and it explicitly does not attribute scores to components: "component-ablation scores are not used to fill these entries." A hypothesis consistent with the mechanism (ours): constraints like length, format and "do not mention X" must shape the answer from its first words, and in Speak-First the first segments are generated before the reasoning has fully processed those constraints. The paper's conclusion does name "multi-turn constraint handling" as a remaining area for improvement.
The paper's positive reading is also justified by the numbers: the realtime macro of 70.4 is "comparable to Doubao 2.0 Lite at 70.5 and DeepSeek-V4-Flash at 71.4", which are measured in reasoning mode, with no speaking deadline. The claim is that the model "can deliberate while producing speech and retain dialogue and reasoning performance comparable to dedicated reasoning models, rather than providing only low-latency surface responses."
It helps to list where each point of the realtime cost could come from, as candidate hypotheses to test rather than conclusions. The report combines all three mechanisms in its realtime row, so none of these can be settled from its tables.
| Candidate source of the realtime cost | Mechanism involved | Evidence available in the paper |
|---|---|---|
| Early segments committed before constraints are processed | Think-While-Speaking (Speak-First) | None isolated; consistent with the Instruction Following drop |
| Reasoning skipped on turns that needed it | Adaptive Thinking | Table 5: Reasoning 71.89 → 66.80 for the separately trained Adaptive Thinking model, yet realtime Reasoning (73.6) did not fall, so this source does not obviously dominate |
| Distribution shift in the private trace | MTP typical acceptance | Table 6: Instruction Following falls under every MTP setting (64.15 → 61.21 to 62.73) |
"What time is it in Tokyo?" "Thanks, that's all." "Actually, make it four people." Most turns in a real conversation are like these: short, routine, and answerable at once. Running a private reasoning trace on them wastes compute and, in a system with a silence budget, wastes time.
Other turns are the dinner-timing puzzle from Chapter 0, where skipping the reasoning produces a confident wrong answer.
The paper's motivation for this chapter's first mechanism is one sentence: "Different conversational turns warrant different deliberation budgets." And its second mechanism answers the follow-up question: when a turn does need thinking, how do you make that thinking cheaper? Two tools: Adaptive Thinking decides whether to think, and multi-token prediction (MTP) makes the thinking that remains faster.
Picture a teacher grading homework who wants to know which problems students should show their work on. For each problem, she compares two attempts: one with worked steps, one written straight down. If both answers are equally good, the steps did not matter for that problem. If the worked version is clearly better, the steps are worth requiring.
Section 6.3.1 builds exactly that comparison, at the level of individual assistant turns. "For each turn, we collect the dialogue context, target answer, original reasoning trace, and task-type features."
Three design choices in that pipeline are worth pausing on.
A fixed probe, trained without the transformation. If the probe had itself been trained on adaptive-thinking data, its no-think answers would already reflect the policy being built, and the measurement would chase its own tail. A fixed probe gives a stable yardstick.
Blind, paired judging. The judge sees two answers to the same turn without knowing which one had reasoning. Blinding removes the bias of preferring the answer that "looks" more thoughtful. Pairing compares like with like: same context, same target, only the reasoning differs. The paper calls this paired judgment "the primary evidence."
Budgets and caps. Even a good label is noisy, so the paper does not simply strip reasoning from every turn labelled unnecessary. "A per-domain budget on the no-think rate decides how many are taken while the remaining turns keep their original reasoning." And: "Rather than spending that budget uniformly, we stratify retention by fine-grained topic and cap the drop rate within each capability, so that reasoning-intensive capabilities are not disproportionately stripped of supervision."
What does "replacement" produce? The paper calls the reasoning-unnecessary turns "candidates for replacement." Figure 7(A) shows the resulting behavior: the direct route gives an "immediate response" with an "empty think block." The natural reading is that replaced turns teach the model to emit an empty think block and answer directly; the report does not spell out whether the replaced answer is the original target or the probe's regeneration.
python (sketch)def label_turn(turn, probe, judge, n_repeats=3): no_think = probe.generate(turn.context, think="") # empty think block verdicts = [] for _ in range(n_repeats): # consistency across repeats a, b, flipped = shuffle_pair(turn.answer, no_think) # blind: judge can't tell which verdicts.append(judge.compare(a, b, turn.target, flipped)) unnecessary = all(v == "same_quality" for v in verdicts) confidence = combine(verdicts, coherence(turn.reasoning, turn.answer), task_prior(turn.task_type)) return unnecessary, confidence def apply_budget(domain_turns, no_think_budget, cap_per_capability): cands = [t for t in domain_turns if t.unnecessary] quota = int(no_think_budget * len(domain_turns)) # per-domain no-think rate picked = stratified_by_topic(cands, quota, # spread over fine-grained topics cap=cap_per_capability) # protect reasoning-heavy skills for t in picked: t.reasoning = "" # becomes a direct-route example return domain_turns # n_repeats, the "all same" rule and the budget values are placeholders, not the paper's.
The paper compares three configurations on StepAudioChat, "covering 46 benchmark members in eight capability categories." Direct SFT and forced no-think "share baseline weights, with explicit reasoning enabled or disabled at inference time." Adaptive Thinking "is trained separately with reasoning-selection supervision, so its comparison also includes the effect of training." Scores and think rates are unweighted means over benchmark members; evaluation uses temperature zero and no system prompt.
| Category | Direct SFT think | Direct SFT score | Forced no-think score | Adaptive think | Adaptive score |
|---|---|---|---|---|---|
| Instruction Following | 100.0 | 64.15 | 61.98 | 60.0 | 62.12 |
| Faithfulness | 100.0 | 72.35 | 70.63 | 79.2 | 72.42 |
| Reasoning | 100.0 | 71.89 | 60.52 | 59.5 | 66.80 |
| Memory | 100.0 | 67.99 | 63.86 | 51.5 | 65.99 |
| Knowledge | 100.0 | 74.59 | 68.65 | 80.4 | 71.55 |
| Safety and Reliability | 100.0 | 78.46 | 74.95 | 56.9 | 77.90 |
| Dialogue Pragmatics | 100.0 | 63.59 | 62.41 | 58.9 | 65.87 |
| Persona and Role Consistency | 100.0 | 77.17 | 68.80 | 82.0 | 74.62 |
The forced no-think column always has a think rate of 0.0. Use the explorer below to see each category's three scores side by side, together with two derived numbers: the full-thinking gain (Direct SFT minus forced no-think) and the share of that gain Adaptive Thinking recovers.
All scores and think rates are the paper's (Table 5). The "gain" and "recovered" figures are our arithmetic on those numbers. A recovered share above 100% means Adaptive Thinking scored higher than always-thinking in that category.
That last step lines up with the paper's own, franker reading. "Reasoning has a think rate of 59.5% despite benefiting most from full thinking. Faithfulness has a higher rate of 79.2%, but gains only 1.72 points from full thinking. Lower thinking frequency alone therefore does not demonstrate that reasoning is allocated to the turns that benefit most." It adds: "These aggregate comparisons also do not establish the optimal decision for individual turns."
The paper's summary of the gains: full thinking helps most in Reasoning (11.37 points), then Persona and Role Consistency (8.37) and Knowledge (5.94). Relative to Direct SFT, Adaptive Thinking "improves Dialogue Pragmatics from 63.59 to 65.87, but reduces Reasoning from 71.89 to 66.80." The conclusion of the report keeps the same tone: "Adaptive Thinking reduces the frequency of explicit reasoning, with uneven effects on answer quality."
Adaptive Thinking reduces how often the model thinks. The paper then attacks the cost of the thinking that remains: "we use MTP3 with three prediction heads, drafting up to three future tokens at each target-model step."
Start with the bottleneck. An autoregressive model produces one token per forward pass. A 400-token reasoning trace means 400 passes through an 11B-active-parameter network, one after another.
Now the idea, with an analogy. A chess player thinking about a line can pre-move: "if they play this, I will play that, then this." If the opponent's actual moves match, several moves happen in the time of one decision. If not, the pre-moves are discarded, and nothing is lost except a little effort.
Multi-token prediction gives the model extra small prediction heads that, from the same hidden state, guess the token after next, the one after that, and so on. Those guesses are drafts. In the next forward pass, the main model (the target model) checks all drafts at once, in parallel, and keeps the longest prefix it agrees with. This draft-then-verify loop is speculative decoding; the paper cites the classic speculative-sampling papers, Medusa, and multi-token prediction as its lineage. The Speculative Decoding Gleam and the Leviathan et al. Veanor build it from zero.
One piece of standard accounting (background, not stated in the report): each verification pass also yields one token of the target model's own, so the tokens produced per target step are one plus the number of accepted drafts.
How does the target decide whether to keep a draft? The paper uses two rules.
Strict verification "follows the target model's standard speculative-decoding rule." In that rule (from the speculative-sampling papers), a draft is accepted exactly when doing so leaves the output distribution unchanged; with greedy decoding, that means the draft must equal the target's own top choice. Speed is gained without changing what the model says.
Typical acceptance, from Medusa, "uses an entropy-adaptive confidence threshold to accept plausible draft tokens that strict verification may reject." The Medusa paper's form of this rule accepts a draft token x when the target's probability for it clears a bar that drops as the target becomes less certain:
Here p is the target model's next-token distribution, p(x) is the probability it gives the draft token, H(p) is the entropy of that distribution (a measure of how spread out, how uncertain, it is), and ε and δ are thresholds. When the target is confident (low H), the bar is high and only strong drafts pass. When the target is unsure (high H), many continuations are reasonable, the bar drops, and a plausible draft is kept. The StepAudio report states the principle but not its threshold values, so ε and δ here are Medusa's notation, not reported constants.
The paper is clear about the cost: typical acceptance "increases acceptance while allowing the generated distribution to change." So it adds a guard: "A repetition penalty of 1.05 discourages repeated tokens and reduces the risk of repetition loops under permissive acceptance." A repetition penalty scales down the scores of tokens that have already appeared (a common implementation divides a positive logit by the penalty; the report does not specify its form).
python (sketch)def verify(target_probs, drafts, mode, eps=0.09, delta=0.3): """target_probs[k]: target distribution at draft position k (one forward pass) drafts[k]: token proposed by MTP head k+1. Returns number accepted.""" accepted = 0 for p, x in zip(target_probs, drafts): if mode == "strict": # spoken response (greedy form) ok = (x == p.argmax()) else: # "typical": private reasoning only H = -(p * p.clamp_min(1e-9).log()).sum() ok = p[x] > min(eps, delta * torch.exp(-H)) if not ok: break # later drafts depend on this one accepted += 1 return accepted # tokens this step = accepted + 1 (the target's own token) # eps=0.09, delta=0.3 follow Medusa's public defaults, NOT StepAudio's (the report gives no values). # The cap min(eps, .) binds on confident steps; delta*exp(-H) lowers the bar when entropy H is high. # The repetition penalty of 1.05 (the paper's value) is applied to the logits first.
The break is why acceptance falls with depth: head 3's draft only counts if heads 1 and 2 were both accepted.
That is what Table 7 calls marginal acceptance: the probability that a head's draft is accepted as part of the kept prefix.
| Domain (StepAudioChat) | Baseline | MTP3 | MTP3 (Medusa) | MTP5 | MTP5 (Medusa) |
|---|---|---|---|---|---|
| Instruction Following | 64.15 | 61.21 | 62.71 | 61.25 | 62.73 |
| Faithfulness | 72.35 | 72.29 | 70.46 | 72.49 | 72.03 |
| Reasoning | 70.76 | 74.30 | 73.74 | 73.14 | 72.73 |
| Memory | 67.99 | 69.29 | 68.97 | 70.85 | 71.36 |
| Knowledge | 74.59 | 75.07 | 75.34 | 75.31 | 74.33 |
| Safety & Reliability | 81.41 | 81.07 | 81.59 | 81.23 | 81.17 |
| Conversational Pragmatics | 63.59 | 66.04 | 63.17 | 66.22 | 66.15 |
| Persona & Role | 77.17 | 77.99 | 75.92 | 77.19 | 77.56 |
| Accepted drafts / step | 0 | 1.231 | 1.801 | 1.353 | 2.153 |
| Wall-clock speedup | 1.00× | 1.76× | 2.05× | 1.49× | 1.72× |
| Depth | Verification | Head 1 | Head 2 | Head 3 | Head 4 | Head 5 |
|---|---|---|---|---|---|---|
| MTP3 | Strict | 65.0% | 37.2% | 20.8% | N/A | N/A |
| MTP3 | Typical | 82.4% | 58.2% | 39.4% | N/A | N/A |
| MTP5 | Strict | 63.7% | 35.5% | 19.6% | 10.7% | 5.7% |
| MTP5 | Typical | 80.6% | 55.7% | 37.6% | 24.9% | 16.4% |
All configurations use the repetition penalty of 1.05, and the paper warns that "quality, acceptance, and wall-clock measurements use their respective evaluation settings." One visible consequence: six rows of Table 6's Baseline column match Table 5's Direct SFT column, but Reasoning (70.76 vs 71.89) and Safety (81.41 vs 78.46) differ; the paper notes only that each table keeps its own evaluation settings, and it does not say whether the two baselines are the same checkpoint.
The paper's three findings, with the numbers behind them. First, "the tested MTP configurations improve Reasoning and Memory over the baseline, while Instruction Following declines" (Reasoning 70.76 rises to between 72.73 and 74.30; Instruction Following 64.15 falls to between 61.21 and 62.73). Second, deeper drafting shows "diminishing gains in accepted drafts per step": heads 1 to 3 behave similarly at both depths, and the strict marginal rates of heads 4 and 5 fall to 10.7% and 5.7%. Third, "acceptance alone does not determine net efficiency, which also depends on reasoning length and decoding costs," so the wall-clock ratios "should not be interpreted as a controlled comparison of draft depth." That is why MTP5 strict (1.353 accepted) shows a lower speedup (1.49×) than MTP3 strict (1.231 accepted, 1.76×).
Each step, the heads propose drafts; the verifier walks them left to right and stops at the first rejection. The acceptance probabilities come straight from Table 7 (each head's conditional chance is its marginal rate divided by the previous head's). Run many steps and watch the running average converge to the sum in worked example 8b.
Green = accepted draft, red = first rejection, grey = discarded because an earlier draft was rejected, teal = the target model's own token for this step. Acceptance rates are the paper's (Table 7); the random draws are the simulation's.
"Can you book me a car to the airport tomorrow?" Simple to say, not simple to do. The assistant needs your flight time, which lives in your private calendar. It needs a pickup time, which depends on the flight. It needs to call a booking service, which may take a while to respond.
While it works, you keep living. "Oh, and it's for two people." "Is it done yet?" "What's the weather going to be there?" A walkie-talkie agent would make you wait for the booking to finish before any of those could be heard.
The paper's introduction names this as the fourth problem: "Tool use extends this challenge because an external task may outlast the spoken exchange that initiated it." Section 7 is the answer: a full-duplex voice agent whose tool work runs "alongside the conversation, enabling the user to ask about progress or provide additional requirements while work is underway."
Think of a hotel concierge. Some questions they answer from memory ("breakfast is until ten"). Some need a quick look at a screen ("rain tomorrow, bring an umbrella"). Some they hand off to someone in the back office and promise to follow up ("I'll arrange the car and call your room"). A good concierge picks the right route without being told.
Section 7 gives the model the same three routes: "The model selects among direct responses, lightweight tool calls, and asynchronous backend execution based on the requirements of each request."
| Route | When the paper uses it | Examples from the paper | Timing |
|---|---|---|---|
| Direct response | "Routine conversation and questions about stable knowledge" | chit-chat, well-known facts | Immediate |
| Lightweight tool | "Requests for up-to-date public information" | "weather lookup or web search" | Short |
| Asynchronous backend | "Requests involving private context, multi-step processing, or work extending beyond the current conversational turn" | tasks needing the user's data or several steps | Runs alongside the conversation |
The paper places this routing in the tradition of tool-augmented language models, citing Toolformer and Gorilla. The Toolformer Veanor shows how a model can learn when a tool call is worth making.
Routing also interacts with Adaptive Thinking from Chapter 7. A direct response on a routine turn can take the empty-think route. A backend task that needs planning is where "reasoning supports constraint resolution and task planning" (Section 7.2).
Routing can fail in both directions, and each direction has a different cost in a voice product. The paper's training explicitly targets both: routing examples teach when a tool is needed, and negative examples "discourage unnecessary tool invocation."
| Routing mistake | Example (illustrative) | Cost to the user |
|---|---|---|
| Tool when a direct answer would do | A web search for "what does a backchannel mean?" | Extra delay on a turn that should feel instant |
| Direct answer when fresh data is needed | Answering "will it rain tomorrow?" from memory | A confident, stale, possibly wrong answer |
| Lightweight tool when private context is needed | Web-searching for "my" booking | The task cannot succeed; the agent may bluff |
| Backend for a one-line fact | Launching a multi-step job to tell the time | Wasted compute and a needlessly long exchange |
Here is the most important idea in Section 7, and it is subtle enough to read twice.
"A grammatically complete phrase may still leave a spoken request incomplete or open to revision. Spoken responses can be extended or explicitly corrected in subsequent segments, whereas tool execution requires sufficiently specified intent and arguments."
Recall Chapter 6. Think-While-Speaking is comfortable starting to talk before reasoning is complete, because speech can be amended: "a final continuation can supplement or correct." A booking cannot be amended by a later sentence. Once the car is booked for the wrong time, saying "sorry, I meant 7:30" does not change the database.
So the paper makes the agent more conservative exactly where the cost of being wrong is higher: "Before committing to an external action, the model is therefore trained to gather missing information through clarification with the user or context retrieval using an appropriate tool, and to obtain any required confirmation."
Three ways to fill a gap before acting, all named in the paper:
| Gap | How to close it | Example (ours, illustrative) |
|---|---|---|
| Missing argument the user knows | Clarification with the user | "What time is your flight?" |
| Missing argument in private context | "Context retrieval using an appropriate tool" | Look up the flight in the user's calendar |
| Consequential action | "Obtain any required confirmation" | "Pick-up at 7:30 for two, shall I book it?" |
Section 7.2: "Backend tasks execute asynchronously while the conversation continues. During execution, the user may request progress updates, provide additional requirements, or shift to another topic."
Three kinds of user input can arrive while a task is running, and each needs a different treatment. A progress query must be answered from the task's actual status, not from optimism. An additional requirement must be attached to the right task. An unrelated topic must be handled on its own without disturbing the task. The paper states the training goal: "The model is trained to associate task-related user input with the ongoing task while distinguishing it from unrelated dialogue."
When results arrive: "As results become available, they are incorporated into the conversational context to inform subsequent spoken responses." That is the tool-execution status slot on the shared whiteboard from Chapter 0, being written by the backend and read by the speaker.
The division of labor between thinking and acting follows ReAct, the pattern of interleaving reasoning steps with actions and observations (see the ReAct Veanor): "For complex requests, reasoning supports constraint resolution and task planning, while tool and backend outputs provide evidence of what has actually been completed. This evidence guides subsequent reasoning and action."
And the paper ties the two concurrency mechanisms together in one sentence: "Think-While-Speaking supports spoken responses during deliberation, while asynchronous execution allows the conversation to continue during external task execution." Chapter 6 decoupled thinking from speaking. This chapter decouples acting from speaking.
Run the car-booking scenario. Once the backend task is running, use the interjection buttons: ask for progress, add a requirement, or change the topic. The right-hand board is the shared conversational context: watch which slot each event writes to. Then switch to a blocking agent that waits for the tool before it can talk, or skip clarification, and compare.
The scenario, durations, flight time and weather reply are illustrative; the routing, clarify-and-confirm-before-commit rule, task association and evidence-grounded reporting are the behaviors the paper describes (§7.1–7.3). In blocking mode, interjections queue until the task returns, so a late requirement misses the booking.
In the async run, every interjection is handled within seconds while the backend bar keeps growing, and the requirement "two people" is attached to the task before it completes. In the blocking run, the same interjections pile up as dead air and the requirement arrives after the booking has already been made. With clarification switched off, the task commits with a guessed pickup time, and the final report has to disclose a wrong result, which is exactly what the paper's negative examples train against claiming as success.
Here is one full exchange, written out turn by turn, with the route each turn takes and the slot of the shared context it writes. The dialogue is illustrative; the annotations follow the behaviors listed in Sections 7.1 to 7.3.
| Speaker | Words (illustrative) | Route / behavior | Context slot written |
|---|---|---|---|
| User | "Book me a car to the airport tomorrow." | Backend candidate: private context, work beyond this turn | Request |
| Model | (private) flight time unknown; check the calendar | Context retrieval with a tool before committing | Arguments: pending |
| Model | "Your calendar shows a 9:40 flight. Pick-up at 7:30?" | Clarification plus confirmation of a consequential action | Arguments: 9:40 → 7:30 |
| User | "Yes. Oh, and what's the weather there?" | Confirmation; then an unrelated question | Confirmation: yes |
| Model | "Booking now. Tomorrow looks sunny there." | Backend starts asynchronously; weather goes to a lightweight tool | Task: running; Side topic: answered |
| User | "Make it for two people." | Task-related input, associated with the running task | Requirements: 2 passengers |
| User | "Is it done?" | Progress query, answered from real status | (read) Task: 60% |
| Model | "Not yet, it's still confirming." | No unsupported success claim | none |
| Backend | (returns: booked, 7:30, 2 passengers) | Evidence enters the context | Result |
| Model | "Done: a car at 7:30 for two." | Result reporting grounded in the returned evidence | none |
Every row maps onto an item in the paper's list of targeted training behaviors: routing, clarification, private-context retrieval, confirmation before consequential actions, execution-time updates, progress queries and result reporting. The weather lookup shows the other half of routing: a lightweight tool, handled inside the same conversation while the backend keeps working.
python (asyncio sketch)class VoiceAgent: def __init__(self, model, tools, backend): self.model, self.tools, self.backend = model, tools, backend self.ctx = SharedContext() # evidence, history, turn, reasoning, TOOL STATUS async def on_user_utterance(self, utt): route = self.model.route(self.ctx, utt) # direct | tool | backend task = self.ctx.related_task(utt) # associate with an ongoing task? if task and is_progress_query(utt): return await self.say(status_report(task)) # from real status only if task: task.add_requirement(utt) # extra constraint, same task return await self.say("Added to the booking.") if route == "direct": return await self.say(self.model.answer(self.ctx, utt)) if route == "tool": # weather, web search obs = await self.tools.call(self.model.tool_call(self.ctx, utt)) return await self.say(self.model.answer(self.ctx, utt, evidence=obs)) # backend: never commit on an under-specified request args = await self.fill_arguments(utt) # clarify with user / retrieve context if not await self.confirm(args): # consequential action return job = asyncio.create_task(self.backend.run(args)) # does NOT block the conversation self.ctx.track(job, args) job.add_done_callback(lambda j: asyncio.create_task(self.report(j))) await self.say("On it. I'll let you know when it's done.") async def report(self, job): result = job.result() self.ctx.write_tool_result(result) # evidence enters the context await self.say(self.model.summarize(self.ctx, evidence=result)) # claim only what it returned
In the real system, none of those branches is hand-written; the model makes each decision by predicting tokens, and the orchestration details are not published. The sketch shows the behaviors the training data targets, in the order the paper lists them.
Section 7.3 combines two data sources.
Targeted voice-agent dialogues cover "request routing, clarification, private-context retrieval, confirmation before consequential actions, execution-time updates, progress queries, and result reporting." They "train the model to ground claims about private information and completed work in user-provided context or evidence returned by tools." And, importantly, "negative examples discourage unnecessary tool invocation and unsupported claims of successful execution."
Two failure modes are singled out by those negative examples. Unnecessary tool invocation wastes time and money, and in voice, adds delay to turns that could have been answered directly. Unsupported success claims ("Done, your car is booked!" before any tool has returned) are the voice-agent version of hallucination, and they are worse than silence because the user acts on them.
Real multi-step agent trajectories complement the dialogues "by exposing the model to longer sequences of reasoning and tool use." The team filters and normalizes them, "focusing on tool-call structure, argument consistency, evidence grounding, and suitability for spoken interaction." That last criterion matters: a trajectory that ends with a 40-row table is fine for a text agent and useless to read aloud.
Recall from Chapter 1 that midtraining already "substantially increases the share of... agent-interaction data," training the model "to carry user intent through planning, tool use, and spoken follow-up." The voice-agent training here builds on that foundation.
| Behavior the data targets | Failure it prevents |
|---|---|
| Request routing | Calling a backend for "hello"; answering "what's my balance" from memory |
| Clarification | Committing with a guessed argument |
| Private-context retrieval | Asking the user for something already in their data |
| Confirmation before consequential actions | Irreversible actions the user did not approve |
| Execution-time updates, progress queries | Dead air; fabricated progress |
| Result reporting grounded in evidence | Claiming success the tool did not return |
τ-Voice deliberately injects the conditions that break voice agents. The table connects each condition the benchmark lists to the mechanism in this report that is meant to absorb it. The mapping is ours; the paper reports only the aggregate task-success rates.
| τ-Voice condition | What can go wrong | Mechanism aimed at it |
|---|---|---|
| "Interruptions in which users revise their requests" | The agent commits to the old request | Duplex yielding (Ch. 3) + clarify and confirm before commit (§7.1) + task association (§7.2) |
| Backchannels | The agent stops or restarts on every "mm-hm" | Backchannel handling (Ch. 3) |
| "Diverse forms of background noise" | Misheard names and numbers become wrong tool arguments | Robust perception and data augmentation (Ch. 2); confirmation before consequential actions |
| Multi-step customer-service tasks | A dropped constraint leaves the database in the wrong state | Reasoning for constraint resolution; real multi-step trajectories in training (§7.3) |
The paper evaluates with the Artificial Analysis implementation of τ-Voice, which "assesses tool-grounded task completion in full-duplex spoken interaction under challenging conversational and acoustic conditions, including interruptions in which users revise their requests, backchannels, and diverse forms of background noise." Tasks are customer-service scenarios in three domains: airline, retail and telecom.
The success criterion is strict and objective: a task succeeds "when the final database state matches its target." Saying the right words is not enough; the backend state must actually be right. "Domain task-success rates average three trials where available, and the reported macro average gives equal weight to Airline, Retail, and Telecom."
| Domain (task success, %) | StepAudio 3 Realtime | Grok Voice Think Fast 2.0 High | Qwen Audio 3.0 Realtime Plus | GPT-Realtime-2.1 High |
|---|---|---|---|---|
| Airline | 60.0 | 56.0 | 61.3 | 62.0 |
| Retail | 37.7 | 49.7 | 49.0 | 45.6 |
| Telecom | 70.2 | 63.7 | 53.5 | 29.4 |
| Macro Average | 56.0 | 56.5 | 54.6 | 45.7 |
The paper's reading: StepAudio 3 Realtime reaches "a macro task-success rate of 56.0%, close to Grok's 56.5% and above Qwen's 54.6% and GPT's 45.7%." It has "the highest telecom score among the evaluated models at 70.2%," its airline score "is within 2.0 percentage points of the highest reported score of 62.0%," and "retail performance leaves room for further improvement."
The domain profiles are strikingly different across systems. GPT-Realtime-2.1 High is best at airline (62.0) and far behind at telecom (29.4). StepAudio 3 is best at telecom and weakest at retail. A macro average hides this; the paper is right to report domains separately. The report does not analyze why retail is hard for this model, so we will not guess at a cause.
You now know every mechanism in the report. This last chapter does three things: puts all the evaluations on one board, lists what the paper admits it has not solved (and what it does not report), and points you to what to read next.
Section 8 frames the evaluation as a staircase: six capability domains that "measure the progression from recognizing an utterance to understanding its context, managing the conversational floor, and completing an external task." Each step of that staircase matches a verb of the loop from Chapter 0.
| Domain | Benchmarks | Protocol notes (§8.3) | Chapter |
|---|---|---|---|
| Speech recognition | LibriSpeech, AISHELL-1, WenetSpeech, ContextASR-Bench | WER (English) / CER (Mandarin); Contextless setting; baselines rerun on the same audio | 2 |
| Audio understanding | Big Bench Audio, MMSU, MMAU, MMAR, WildSpeech-Bench, AudioMultiChallenge, Step-Caption, MTalk-Bench | Each benchmark's reported accuracy or normalized score; Step-Caption judge-scored; MTalk-Bench = paralinguistic + ambient components only | 2 |
| Dialogue and reasoning | StepAudioChat (8 dimensions) | Each dimension = unweighted mean of validated capability lines; reasoning-mode and realtime results both reported; no ablation scores used to fill entries | 4, 6 |
| Full-duplex interaction | AA subset of Full Duplex Bench v1 and v1.5 | Category = % of samples meeting the criterion; source-reported Overall, not a mean of categories | 3 |
| Agentic task completion | AA implementation of τ-Voice | Success = final database state matches target; up to three trials averaged; equal-weight macro over three domains | 8 |
| General text | HMMT February 2026, GPQA Diamond, MultiChallenge | Recorded accuracy; MultiChallenge uses instance-specific rubrics; no average taken across the three | 1, 5 |
The baselines change by domain (Section 8.2), which is worth knowing before you read any comparison. ASR: Doubao 2.0 ASR, Seed 2.0 Lite, HY3.0 ASR Preview. Audio understanding: Doubao 2.0 Lite, Gemini 3 Flash, Gemini 3.1 Pro. Dialogue: Doubao 2.0 Lite, DeepSeek-V4-Flash, Kimi K3. General text: Doubao 2.0 Lite and Gemini 3 Flash. Full-duplex: GPT-realtime-2 (High), Qwen Audio 3.0 Realtime Plus, Grok Voice Think Fast 2.0 High. Agentic: the same Qwen and Grok variants, with GPT-Realtime-2.1 High replacing GPT-realtime-2. "Model versions and effort labels follow the corresponding evaluation records."
Pick a domain to see StepAudio 3 Realtime beside the baselines the paper uses for that domain. Bars use a 0 to 100 scale, except where a domain starts its bars higher so small gaps stay visible (noted under the bars); the best score in each row is marked.
Every number is from the paper: Table 10, plus the Table 3 duplex categories and the Table 9 τ-Voice domains. For Dialogue, StepAudio 3 is in realtime ("interactive") mode while the baselines are in reasoning mode, as the paper notes.
Here is the full table in static form, for reference.
| Audio understanding | StepAudio 3 | Doubao 2.0 Lite | Gemini 3 Flash | Gemini 3.1 Pro |
|---|---|---|---|---|
| Big Bench Audio | 98.1 | 98.8 | 99.4 | 99.6 |
| AudioMultiChallenge | 49.3 | 48.5 | 56.6 | 67.0 |
| MMSU | 90.6 | 80.0 | 77.0 | 83.6 |
| MMAU | 79.0 | 77.5 | 77.6 | 80.5 |
| WildSpeech | 77.1 | 73.9 | 74.4 | 77.7 |
| MMAR | 86.5 | 75.9 | 75.4 | 81.7 |
| Step-Caption | 78.2 | 76.8 | 67.8 | 74.8 |
| MTalk-Bench | 91.7 | 89.9 | 88.5 | 89.1 |
| General text (accuracy) | StepAudio 3 | Doubao 2.0 Lite | Gemini 3 Flash |
|---|---|---|---|
| HMMT 2026 Feb | 86.8 | 73.9 | 85.9 |
| GPQA Diamond | 83.0 | 82.4 | 90.3 |
| MultiChallenge | 59.7 | 60.8 | 68.1 |
The dialogue rows (70.4 macro in realtime mode) are in Chapter 6, the full-duplex Overall (98.9) in Chapter 3, and τ-Voice (56.0) in Chapter 8.
Put in one line: the model is first on full-duplex control, leads 4 of 8 audio benchmarks and sits 0.5 below the best audio macro (with a large gap on AudioMultiChallenge), close on agentic success, competitive on dialogue given a realtime handicap, and behind a strong general model on two of three text benchmarks. The paper's own synthesis says much the same: it "combines broad audio understanding with strong dialogue, reasoning, and interaction control," and its balanced duplex profile "suggests that it can preserve conversational flow without treating every user sound as an interruption."
Notice one more pattern across the text rows. Of the three, the model leads on HMMT, the benchmark most purely about reasoning, and trails most on MultiChallenge, the one about "multi-turn conversational reliability." That echoes AudioMultiChallenge (multi-turn audio) and the realtime instruction-following drop. Three different benchmarks point at the same soft spot: holding and revising constraints across a long conversation.
The report is candid about its gaps. Collected in one place, each with where it shows up:
| Admitted limitation | Paper's words (abridged) | Evidence |
|---|---|---|
| Multi-turn constraint handling | "Multi-turn constraint handling and retail task completion remain areas for improvement." | AudioMultiChallenge 49.3 vs 67.0 (the paper's link); realtime Instruction Following 54.1 and text MultiChallenge 59.7 vs 68.1 (our link) |
| Retail tool use | "Retail performance leaves room for further improvement." | τ-Voice retail 37.7 vs 49.7 |
| Adaptive Thinking allocation | "Adaptive Thinking reduces the frequency of explicit reasoning, with uneven effects on answer quality." | Reasoning think rate 59.5% despite the largest gain (+11.37); Reasoning 71.89 → 66.80 |
| Per-turn optimality unproven | "These aggregate comparisons also do not establish the optimal decision for individual turns." | Table 5 is category-level only |
| Incomplete-reasoning errors | "Keeping strict verification for spoken output does not eliminate errors arising from incomplete private reasoning." | Think-While-Speaking design |
| MTP timing is not a controlled comparison | Wall-clock ratios "should not be interpreted as a controlled comparison of draft depth." | MTP5 strict 1.49× vs MTP3 strict 1.76× |
| Merging is a trade-off | "A balanced trade-off rather than uniform improvement over every specialized teacher." | Dialogue 73.0 vs best teacher 74.2 |
| ASR numbers are the specialist's | "These results characterize the ASR-specialized model, not the transcription behavior of the realtime model." | Table 1 |
The conclusion turns those into a research agenda: the findings "motivate more effective allocation of reasoning effort and more reliable task execution over extended conversations."
Separate from the admitted limitations, a careful reader should note what a technical report of this kind leaves out. These are observations about the document, not criticisms of the model.
No end-to-end latency measurements. For a report titled "Realtime", the only timing constant given for the interaction is the 320 ms block; there are no first-response latency or overlap-reaction figures. The MTP section reports relative wall-clock speedups, not absolute times.
Architecture internals. Expert count, experts per token, layer count, hidden sizes, the adapter's design, the generator's design and token rates are not given; only ~196B total and 11B active parameters, the Step 3.7 Flash backbone and the Qwen3-Omni AuT encoder.
Data scale per stage. The 1.2T pretraining tokens, 32K and 128K context lengths, "over 10,000 hours" of synthetic duplex data, and the ~2M vs ~100K ablation are reported; mixture proportions, hours per language and teacher data mixes are not.
Reproducibility of the dialogue score. StepAudioChat is closed by design, which protects it from contamination and also means outsiders cannot rerun it.
Component attribution in realtime mode. The realtime StepAudioChat row combines Think-While-Speaking, Adaptive Thinking and MTP; the paper deliberately does not fill it from ablations, so the 2.6-point realtime cost cannot be split among them from the report alone.
If you remember nothing else, remember these, each with the idea it stands for.
| Number | What it is | The idea behind it |
|---|---|---|
| 320 ms | Duplex block length | Floor decisions are tokens emitted 3.125 times a second |
| 11B / 196B | Active vs total parameters | Sparse experts make two concurrent calls affordable |
| 128K | Midtraining context | Long conversations and tool results stay in view |
| 100K vs 2M | Curated vs random SFT examples | Quality beat volume on every reported audio metric |
| 98.9 | AA Full-Duplex Overall | Balanced pauses, turns, interruptions and backchannels |
| 3:1:1:1 | Merge weights | Specialize, then average from a common base |
| 73.0 → 70.4 | StepAudioChat, reasoning vs realtime | Realtime mode (Think-While-Speaking + Adaptive Thinking + MTP) scores 2.6 below reasoning mode, mostly on instruction following |
| 59.5% | Think rate on Reasoning | Table 5 does not show that Adaptive Thinking spends reasoning where it helps most |
| 1.231 / 1.801 | Accepted drafts per step, MTP3 strict / typical | Sum of per-head marginal rates; typical acceptance used only in private |
| 2.05× | Best wall-clock speedup | Acceleration divides silence; concurrency removes it |
| 56.0 vs 56.5 | τ-Voice macro vs best | Strong telecom, weak retail; database state is the judge |
| Mechanism | Problem it solves | Key number(s) from the paper | Ch. |
|---|---|---|---|
| Shared conversational context | Keeps decoupled clocks coherent | Five ingredients (evidence, history, turn, reasoning, tool status) | 0 |
| MoE audio-language model | Capacity without per-token cost | ~196B total, 11B active; 32K → 128K context; 1.2T pretraining tokens | 1 |
| ASR Max specialization | Accurate, context-aware transcripts | LibriSpeech clean 1.18; AISHELL-1 0.49; ContextASR macro 5.67 / 1.23 | 2 |
| Curated audio QA | Grounded audio understanding | ~100K beats ~2M (MMSU 78.78 → 89.70); MMSU 90.6; macro 81.3 | 2 |
| 320 ms blocks + state tokens | Take, retain or yield the floor | AA Full-Duplex 98.9; turn taking 100.0; interruptions 99.0 | 3 |
| StepAudioChat + self-play data | Measure and teach conversation | Reasoning mode 73.0 (Kimi K3 77.1) | 4 |
| Weighted merging | One model, several specialties | 3:1:1:1; HMMT 86.8 above every teacher | 5 |
| Think-While-Speaking | Deliberate without dead air | Realtime 70.4 vs reasoning 73.0 | 6 |
| Adaptive Thinking | Skip reasoning when it does not help | Think rates 51.5–82.0%; Reasoning 71.89 → 66.80 | 7 |
| MTP3 + typical acceptance | Faster private reasoning | 1.801 accepted/step; 2.05× wall-clock; repetition penalty 1.05 | 7 |
| Asynchronous voice agent | Keep talking while tools run | τ-Voice 56.0 (telecom 70.2, retail 37.7) | 8 |
Strip away the specific model, and the report leaves six engineering ideas that transfer to other systems.
1. Make control decisions tokens. Floor management is a token predicted after every 320 ms block, conditioned on everything the decoder sees. Any decision that needs context (when to speak, when to call a tool, when to think) benefits from living inside the sequence rather than in a separate rule.
2. Be permissive where it is private and exact where it is public. Typical acceptance and a repetition penalty speed up the hidden reasoning; the spoken answer keeps strict verification. The same split applies to any system with an internal scratchpad and an external output.
3. Treat irreversible actions differently from revisable words. Speech can start early and be corrected; tool execution waits for specified arguments and confirmation.
4. Curate before you scale. One twentieth of the SFT data, chosen by quality, case value and cross-model agreement, won on every reported audio metric.
5. Build benchmarks with controls. A valid and a deliberately flawed response per item turn the judge itself into something you can test.
6. Specialize, then merge. Independent teacher mixes, recombined by a convex average, let separate teams improve separate skills without retraining on the union.
Several full-duplex or streaming designs have lessons on this site. At a high level, they place the boundary between thinking and speaking in different spots. The descriptions of the other systems are brief summaries of their own lessons, not claims from this report.
| Design | How listening and speaking overlap | Where reasoning lives |
|---|---|---|
| Cascade (VAD → ASR → LLM → TTS) | They do not; a silence timer passes the turn | In the LLM, before any speech |
| Moshi | User and model audio modelled as parallel streams | An inner text stream aligned with the model's own speech |
| Qwen2.5-Omni | Streaming input and output | A Thinker that produces text, with a Talker that turns it into speech |
| StepAudio 3 Realtime | Dual audio streams, 320 ms blocks, state tokens | A private Formulation Brain running alongside a speaking Articulation Brain, both calls to the same model, with adaptive routing and MTP |
Each admitted limitation suggests a concrete next experiment. These are our suggestions, built on the paper's own diagnosis.
Allocate reasoning by benefit, not by frequency. Table 5's category think rates do not demonstrate allocation by benefit (the paper's own caveat), and aggregate numbers cannot settle per-turn decisions. A policy trained on the size of the paired-judgment gain, rather than on a same/different label under a budget, is a natural next step.
Protect constraints in Speak-First. Instruction following carries most of the realtime drop. Testing Think-First, or a constraint-extraction step before the first segment, on exactly that dimension would show whether early commitment is the cause.
Explain retail. A per-task error analysis of the 37.7% retail result (argument errors, missed confirmations, noise-induced mishearings) would say which of this report's mechanisms to strengthen.
Measure the clock. Publishing first-audio latency, reaction time to interruptions, and the fraction of answers needing a final continuation would let others compare realtime systems on the property the title promises.
Every link below points to a lesson that exists on Engineermaxxing today.
| If you want to go deeper on… | Read | Why |
|---|---|---|
| How audio becomes LLM input | Audio LLMs, Qwen2-Audio | Encoder + adapter + decoder, the Chapter 1 pattern from zero |
| Turn-taking and barge-in | Turn-Taking, Endpointing & Barge-In | The cascade failures Chapter 3 replaces |
| Streaming recognition and synthesis | Streaming Speech | Why block-based streaming works and what it costs |
| Full-duplex predecessors | Moshi, DuplexSLA, PersonaPlex | Parallel-stream dialogue, synchronized speech/language/action, and voice/role control in full-duplex models |
| Thinking while talking, another design | Qwen2.5-Omni | Thinker-Talker: a different split between reasoning and speaking, from the Qwen omni line whose later Qwen3-Omni supplies this model's AuT encoder |
| Speech recognition foundations | Whisper (paper), Whisper (Gleam), Self-Supervised Speech | The paper's first citation for LLM-era ASR; WER in practice; learned audio features |
| How speech is generated | TTS Architectures, Neural Audio Codecs, VALL-E, AudioLM | What a "generator" can be; AudioLM is cited by the report |
| Draft-and-verify decoding | Speculative Decoding (Gleam), Speculative Decoding (Leviathan et al.) | The strict-verification rule of Chapter 7 |
| Multi-token prediction in a frontier LLM | DeepSeek-V3 | MTP heads at scale, in a different model |
| Sparse experts | Mixture of Experts, MoE: Sparse Computation | What "11B active of ~196B" buys |
| Reasoning and acting | ReAct, Toolformer, Agents & Tool Use | The interleaving pattern and tool-use lineage the agent chapter cites |
| Why multi-turn constraints are hard | LLMs Get Lost in Multi-Turn Conversation | Independent evidence for the gap this report names |
Beyond the site, the report's own references are the natural next step: Mind-Paced Speaking (arXiv:2510.09592) for the dual-brain design, Medusa (arXiv:2401.10774) and Gloeckle et al. (arXiv:2404.19737) for multi-token decoding, and the Artificial Analysis speech-to-speech methodology for the full-duplex and τ-Voice protocols.
A good test of understanding is whether you can regenerate the report's derived figures from its raw tables. Every one of these was worked in this lesson:
Without scrolling up: (1) draw Figure 3's five boxes and say why the model audio stream feeds back into the input; (2) write the Figure 5 block sequence and name the four floor decisions, with the evidence each one uses; (3) explain the Formulation and Articulation Brains, playback-aware scheduling, and the Speak-First default; (4) reproduce 1.231 accepted drafts per step from Table 7 and say why typical acceptance is used only for private reasoning; (5) name the two admitted open problems and the numbers that reveal them.
For years, voice assistants were built like relay races: one runner listens, hands the baton to one who thinks, who hands it to one who speaks, and nobody moves until the baton arrives. The quiet idea in this report is that a conversation is not a relay; it is a band. The listener, the thinker, the speaker and the doer all play at once, at their own tempo, reading from the same score. The hard engineering is not any single instrument. It is keeping them in time with each other while the audience keeps interrupting. The next time a voice system makes you wait, ask which two clocks it has chained together that did not need to be.