StepFun-Audio Team (Lin et al. 2026, contributors listed alphabetically) — arXiv:2609.14005, September 2026

StepAudio 3 Realtime

A voice assistant usually gets two bad options: think hard and leave you sitting in silence, or answer right away and answer shallowly. This technical report refuses that trade. It builds one audio-language model that hears both sides of the conversation at once, thinks in private while it is already talking, and keeps chatting while its tools run in the background.

2026 · arXiv v1, 12 Sep StepFun-Audio Team cs.SD · speech & audio ~196B MoE · 11B active Full-duplex voice agent
Prerequisites: what an audio LLM is (Audio LLMs) + what turn-taking and barge-in mean (Turn-Taking). Helpful: Streaming Speech, Speculative Decoding, Mixture of Experts. Everything else is built from zero.
10
Chapters
9
Interactive Sims
98.9
AA Full-Duplex
90.6
MMSU

Chapter 0: The Silence Tax

You are walking to the train and you ask your phone a question out loud. It is not a hard question for a person with a pen, but it has several moving parts: a departure time, a travel time, a time-zone shift, and a buffer at the other end. You want to know one thing: will you make dinner?

The assistant goes quiet. One second. Two. Three. You glance at the screen to check whether it heard you. You say "hello?" And now the assistant has a new problem, because your "hello?" just arrived in the middle of its thinking, and it has to decide what that sound means.

That silence is not a bug in one product. It is a tax that almost every voice system pays, and it comes from a very simple fact: good answers to tricky questions need deliberation (working through the problem step by step before committing), and deliberation takes time. In a text chat, the time is invisible: a spinner turns, you read something else. In a voice conversation, the time is audible. Silence is a message. It says "I am broken" or "I did not hear you".

So builders pick one of two bad options. Option one: think first, then speak. The answer is careful, and the user sits through dead air. Option two: speak immediately with no deliberation. The conversation feels alive, and the answer is shallow, sometimes confidently wrong.

The trade-off this paper attacks, in its own words. The abstract opens: "Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking." Three demands, and the first two pull in opposite directions. The paper's central claim is that you do not have to choose: it "resolve[s] the tension between deep deliberation and latency via Think-While-Speaking, executing private reasoning in parallel with spoken delivery."

Every number, table value and architecture detail in this lesson comes from the StepAudio 3 Realtime Technical Report (arXiv:2609.14005, StepFun-Audio Team). When we need made-up numbers to make a mechanism visible, the prose says so in plain words: illustrative. When the paper is silent on something, we say that too, instead of filling the gap.

Silence is only the first of four problems

The introduction of the paper lists the difficulties in a single paragraph, and each one becomes a chapter of this lesson. Read them slowly, because each describes a moment that a normal chatbot pipeline gets wrong.

Problem 1: the pause that is not an ending
"A pause may occur before a request is complete." You say "book me a table for…" and stop to think. A system that answers every pause interrupts you mid-sentence.
↓
Problem 2: the sound that is not an interruption
"An utterance during model speech may be an acknowledgment or a substantive interruption." You say "right" while it talks. Did you mean "go on" or "right, but that's wrong"?
↓
Problem 3: the thought that takes time
"The model must reason carefully while keeping the conversation responsive." This is the silence tax above.
↓
Problem 4: the task that outlives the sentence
"An external task may outlast the spoken exchange that initiated it." You ask it to check a booking; the lookup takes a while; you keep talking.

Notice what the four problems share. In each one, two things are happening at the same time, and the system has to keep both alive: your speech and its speech (problems 1 and 2), its thinking and its speaking (problem 3), its tool work and its talking (problem 4). A system built around strict turns, where one thing finishes before the next begins, cannot represent any of them.

Walkie-talkie versus telephone

Think about two radios. A walkie-talkie is half-duplex: only one side can transmit at a time. You press the button, speak, say "over", and release. The other side cannot cut in, and you cannot hear them while you talk. A telephone is full-duplex: both sides can speak and hear at the same moment, which is why you can say "mm-hm" while a friend tells a story, and why they can stop mid-word when you gasp.

Most classic voice assistants are walkie-talkies wearing a telephone costume. A voice activity detector (a small model that decides whether a stretch of audio contains speech) waits for you to stop, a speech recognizer turns your words into text, a language model writes a reply, and a speech synthesizer reads it out. The "over" button is just hidden inside a silence timer. The turn-taking Gleam walks through why that timer is so fragile.

It helps to see exactly where a turn-based pipeline breaks against each of the four problems. The table below walks one hypothetical cascade (voice activity detector, recognizer, language model, synthesizer) through them. None of these failures is exotic; each is the default behavior of a design that only allows one thing to happen at a time.

MomentWhat the cascade doesWhy it goes wrong
User pauses mid-requestSilence timer fires, turn ends, reply startsThe timer sees only silence, not whether the sentence is finished
User says "right" during the replyDetector hears speech and stops playbackEvery sound looks like an interruption; an acknowledgment kills the answer
Question needs reasoningLanguage model thinks, then synthesizer speaksThinking time turns straight into dead air
A tool call takes a whileThe turn blocks until the tool returnsThe user cannot ask "is it done yet?" or add a requirement
TV talking in the backgroundDetector hears speech and starts a turnNothing asks whether the speech was addressed to the assistant

The last row is a fifth problem the paper handles, background speech rejection: deciding that a voice in the room is not talking to you. Chapter 3 shows how the model uses dialogue history as evidence for that decision.

Inline check: before reading on. A cascade stops talking whenever the user makes any sound. Which of the paper's problems does that design get wrong, and which does it get accidentally right? Answer: it gets backchannels wrong (a "mm-hm" kills a good answer) and gets substantive interruptions accidentally right (a correction does stop the model). The paper's point is that the same sound, "right", can be either, so the decision needs context, not a volume threshold.

StepAudio 3 Realtime is built as a telephone from the inside. The paper calls its input path "full-duplex": the model hears the user's audio stream and its own outgoing audio stream, continuously, and it makes its turn-taking decisions every 320 milliseconds (Chapter 3 unpacks that number).

The paper's answer: one loop, four capabilities

The paper organizes the whole system as a single cycle that never stops turning: listen, converse, think, act. Each verb is backed by a named capability. Here they are, with the paper's names in bold and a plain-words definition beside each.

Loop verbCapability (paper's name)What it means in plain wordsChapter
ListenDeep PerceptionHear not only the words but how they are said: tone, emotion, speaker traits, background sounds, timing.2
ConverseSeamless DuplexUse both audio streams to decide, moment by moment, whether to keep listening, start talking, keep talking, or stop.3
ThinkThink-While-Speaking (+ Adaptive Thinking, MTP)Reason privately in one process while another process is already speaking; think only when it helps; decode the thinking faster.6, 7
ActVoice AgentCall tools and backends, keep the conversation going while they run, and fold their results back into speech.8

Two more pieces glue the four together. Model merging (Chapter 5) combines several specialist checkpoints into one set of weights, so dialogue skill, audio understanding and text reasoning live in a single model. And a benchmark the team built, StepAudioChat (Chapter 4), measures the conversational intelligence that the thinking machinery is supposed to protect.

The paper's Figure 2 draws the loop as a ring. On the left, two input waves feed in: a user stream labelled "lexical · emotion · acoustic context", and a model stream labelled "speech · overlap · turn state". Around the ring sit the four capabilities: Deep Perception (ASR, audio understanding), Seamless Duplex (listen, speak, yield), Think While Speaking (adaptive thinking, Medusa MTP) and Streaming Action (intent, tool call, feedback). In the middle sits a single disc labelled "conversational state". On the right, one arrow leaves the ring: streaming speech.

That middle disc is the most important object in the paper, and it deserves a name of its own.

The shared conversational context

Picture a whiteboard in the middle of a small team. One person writes down what the customer just said and how they sounded. Another notes whose turn it is. A third scribbles half-finished calculations. A fourth pins up a sticky note: "booking lookup: still running." Everyone reads the same board, and everyone writes to it.

The paper's version of the whiteboard is the shared conversational context. Section 2.1 lists exactly what is on it: "acoustic and linguistic evidence, dialogue history, the current speaking turn, reasoning progress, and tool-execution status."

Ingredient on the boardThe question it answersWho mainly writes it
Acoustic + linguistic evidenceWhat did the user say, and how did they say it?Deep Perception
Dialogue historyWhat has been agreed, asked, or promised so far?Every turn
Current speaking turnWho holds the floor right now, and is anyone overlapping?Seamless Duplex
Reasoning progressHow far has private thinking got, and is it finished?Think-While-Speaking
Tool-execution statusIs an external task pending, running, or done, and what did it return?Voice Agent

The paper says this context "informs whether to continue listening or speaking, whether to reason further, and whether a request is ready for external action." And crucially: "Newly observed speech and returned tool results can change these decisions as the conversation proceeds." The board is never frozen. Every 320 ms, something new may be written to it.

One detail on the board is easy to miss and matters a lot: "Model-side speech provides additional context for interpreting user utterances that overlap with a response." The model does not only remember what it planned to say; it tracks what the user has actually heard so far. If you interrupt at word twelve of a forty-word answer, the useful fact is that you heard twelve words, not forty.

The one sentence to carry through every chapter. Section 2.2 opens: "Conversational timing and reasoning progress need not advance at the same pace." Classic pipelines lock them together: think, then speak, then listen. Every mechanism in this paper is a way of letting those clocks run separately while keeping them on the same board.

Feel the tax before we remove it

The simulation below puts the three strategies side by side on one timeline. The top lane is private reasoning. The bottom lane is speech the user hears. Red marks dead air before the first word. Drag the slider to make the question harder (more reasoning time) and watch what each strategy does with that time.

Sim 0 — the silence tax, three ways

All times and the quality meter are illustrative, chosen to make the shape of the trade-off visible; the paper reports no per-question latency figures. "Think then speak" waits for all reasoning. "Speak now" never reasons. "Think-While-Speaking" starts talking at once, lets each spoken segment use whatever reasoning exists at that moment, and can append a final continuation once reasoning completes (the paper's Speak-First default).

Reasoning needed

Drag the slider all the way right in "Think then speak" mode, and the red bar grows linearly with difficulty: the user pays for every second of thought in silence. Switch to "Speak now" and the red bar vanishes, but the quality meter collapses, because nothing was thought through. Switch to "Think-While-Speaking" and something new happens: speech starts immediately, and each segment uses whatever reasoning exists when it is released. In this simulation we give the speaker a simple, illustrative scheduling policy (open with framing, state substance once reasoning is available); the paper does not state such a rule, and it allows an answer that began on incomplete reasoning to be corrected at the end.

The simulation hides one honest cost that the paper does not hide. Speech that begins before reasoning ends can commit to something the reasoning later contradicts. The paper handles that with "a final continuation [that] can supplement or correct an answer that began from incomplete reasoning", and it warns in Section 6.3.2 that keeping strict verification on the spoken output "does not eliminate errors arising from incomplete private reasoning." We return to this in Chapter 6.

Worked example: why a faster thinker alone does not fix silence

A natural objection: just make reasoning faster. The paper does accelerate reasoning (Chapter 7), and its best measured wall-clock speedup is 2.05×, for three prediction heads with Medusa-style acceptance (Table 6). Let us see what that buys in a think-then-speak design, with an illustrative question that needs 400 reasoning tokens at an illustrative 50 tokens per second.

Worked example 0 · illustrative reasoning length, paper speedup Step 1. Reasoning time without acceleration = tokens ÷ rate = 400 ÷ 50 = 8.0 s of silence.
Step 2. Apply the paper's best wall-clock ratio from Table 6: 8.0 ÷ 2.05.
Step 3. 2.05 × 3.9 = 7.995, so 8.0 ÷ 2.05 ≈ 3.90 s of silence.
Step 4. Silence removed = 8.0 − 3.90 = 4.10 s, or 4.10 ÷ 8.0 = 51% of the wait.
Step 5. Silence remaining = 3.90 s: still far longer than a comfortable conversational gap.
Step 6. Think-While-Speaking with Speak-First starts the reply without waiting for any reasoning prefix, so by definition no silence is spent waiting for reasoning, whatever the reasoning length (generating the first segment still takes some time, and the report gives no measured first-response latency).

The lesson of the arithmetic: acceleration shrinks the tax proportionally, but concurrency changes its shape. Speed divides the silence; running thinking alongside speaking removes the dependency between silence and thinking time altogether. That is why the paper uses both: concurrency for the user experience, and multi-token prediction so that the private thinking finishes sooner and more of the answer is spoken from complete reasoning.

Where this sits in the Step-Audio line. The report says StepAudio 3 Realtime "builds on the Step-Audio series' shared audio-language foundation", citing Step-Audio, Step-Audio-R1, Step-Audio-R1.5 and StepAudio 2.5. Think-While-Speaking builds on the team's earlier Mind-Paced Speaking dual-brain design (arXiv:2510.09592), and the duplex work sits next to their DuplexSLA paper, which has its own Veanor walkthrough. This report is about coordination: "Its focus is the coordination of perception, reasoning, and action as a conversation unfolds."

One loop, many clocks

Put the four capabilities and the shared context together and you get the paper's operating picture. It is worth stating as a list of clocks, because each clock ticks at its own rate and none waits for the others.

ClockWhat ticksRate, as the paper describes it
ListeningUser audio arrives and is encodedContinuously, in 320 ms blocks
Floor controlA state or text token is emittedAfter every 320 ms block
SpeakingResponse segments are releasedAccording to playback progress of the output audio
ThinkingPrivate reasoning tokens are decodedAs fast as decoding allows, sped up by MTP
ActingA backend task runsAsynchronously; results arrive whenever they are ready

In a pipeline, these clocks are chained, so the slowest one sets the pace for everything. In StepAudio 3 Realtime they are decoupled, and the shared context is how they stay coherent: the speaking clock reads whatever the thinking clock has written so far, the floor clock reads what the speaking clock has actually played, and the thinking clock reads what the tool clock has returned. The paper sums it up in the introduction: "These functions operate concurrently as needed, with new user input shaping the ongoing interaction."

A useful way to test your understanding of the rest of this lesson is to ask, for each mechanism, which two clocks it decouples. Seamless Duplex decouples listening from speaking. Think-While-Speaking decouples thinking from speaking. The Voice Agent decouples acting from speaking. Adaptive Thinking and MTP do not decouple anything; they make the thinking clock tick less often and faster.

Two models share one upbringing

The report actually describes two models, and it is worth separating them now to avoid confusion later.

StepAudio 3 Realtime is the conversational system: the full-duplex, thinking, tool-using assistant. StepAudio 3 ASR Max is a transcription specialist. The paper states they "share the same pretraining and midtraining stages" and "diverge only during supervised fine-tuning, where the ASR branch is specialized for transcription and the realtime branch is tuned for spoken interaction." So the ASR numbers in Chapter 2 describe the specialist, and the paper is explicit that they do not describe how the realtime model transcribes.

The results, previewed, so every chapter has a destination

These are the headline numbers the paper reports. Each will be derived, compared and questioned in its own chapter; for now, just notice how many different kinds of skill are on one page.

CapabilityBenchmarkStepAudio 3Best reported baseline
Speech recognition (ASR Max)LibriSpeech test-clean WER1.181.38 (HY3.0 ASR Preview)
Audio understandingMMSU90.683.6 (Gemini 3.1 Pro)
Audio understanding8-benchmark macro average81.381.8 (Gemini 3.1 Pro)
Dialogue, reasoning modeStepAudioChat macro73.077.1 (Kimi K3)
Dialogue, realtime modeStepAudioChat macro70.477.1 (Kimi K3, reasoning mode)
Full-duplex controlAA Full-Duplex Bench Overall98.998.4 (Qwen Audio 3.0 Realtime Plus)
Voice agentτ-Voice macro task success56.0%56.5% (Grok Voice Think Fast 2.0 High)

Read the table the way the paper asks you to: not as "wins everywhere". In realtime mode the dialogue score is comparable to reasoning-mode Doubao 2.0 Lite (70.5) and DeepSeek-V4-Flash (71.4), which is the comparison the paper draws, while Kimi K3 sits well above at 77.1. The model leads on full-duplex control and on several audio benchmarks, sits close behind on the audio macro average and on τ-Voice, and trails a dedicated reasoning model on dialogue. The paper names its own gaps: "multi-turn constraint following and retail tool-use tasks." Chapter 9 puts all of this on one scoreboard.

Chapter 0 recap. (1) Voice makes thinking time audible, so deliberation and responsiveness pull against each other. (2) The paper names four simultaneous-activity problems: pauses, overlapping acknowledgments, slow reasoning, long-running tools. (3) Its answer is one listen-converse-think-act loop, backed by Deep Perception, Seamless Duplex, Think-While-Speaking and the Voice Agent, all reading and writing one shared conversational context. (4) Acceleration divides silence; concurrency removes its dependence on thinking time. (5) ASR Max and Realtime share pretraining and midtraining and split only at fine-tuning.
You already do this every day
When a friend asks you something tricky, you rarely go silent for five seconds. You say "hmm, good question, so the flight leaves at six…" and keep working it out while you talk, sometimes ending with "wait, no, I forgot the time difference". People think while speaking and correct themselves out loud. The paper's Formulation Brain and Articulation Brain are an engineered version of that habit, with one addition people lack: the thinking half is completely private, and the talking half is conditioned on whatever reasoning has arrived so far, with a final continuation that can correct an early remark.
According to the paper, what is the core design move that resolves the tension between deep deliberation and low latency?

Chapter 1: The Machine

Before any clever behavior, there is plumbing. Here is the job, stated as a data problem. Two waveforms arrive continuously: what the user is saying, and what the assistant itself is saying. A text history sits alongside them: the system prompt, earlier turns, tool results. Out of all that, several times a second, the system must produce the next thing to do: stay quiet, say a word, think a thought, call a tool.

What machine can do that? This chapter builds it piece by piece from the paper's Section 3 and Figure 3, then follows the recipe the team used to teach it: three stages of pretraining and a midtraining stage.

The five boxes of Figure 3

The paper's architecture diagram has five boxes and two waves. Read it left to right, like a conveyor belt.

1 · Full-duplex input path
Takes two audio streams: the user audio stream and the model audio stream.
↓
2 · Audio encoder
The Audio Transformer (AuT) encoder from Qwen3-Omni. Turns sound into a sequence of vectors.
↓
3 · Adapter
"Maps the encoder outputs into the representation space of the language model."
↓  +  a separate Text input path joins here
4 · LLM decoder
A mixture-of-experts language model, based on Step 3.7 Flash, that "jointly condition[s] on acoustic information and textual context".
↓
5 · Generator
"Produces streaming model audio, which returns to the model audio stream for subsequent interaction."

Now each box in plain words.

An audio encoder is a network that listens. Raw audio is a long list of air-pressure samples, far too many and too low-level for a language model to reason over. The encoder compresses short windows of sound into vectors that capture what matters: phonemes, pitch, loudness, texture. StepAudio 3 Realtime does not train its own encoder from scratch; it uses the Audio Transformer (AuT) encoder from the Qwen3-Omni technical report. Figure 5 of the paper labels this component a "streaming audio encoder", which tells you it processes audio as it arrives rather than waiting for a whole file.

An adapter is a translator between two vocabularies of vectors. The encoder was trained to speak "audio vector"; the language model was trained to read "token embedding". The two spaces have different sizes and different geometry. The adapter is a learned mapping that turns the first into something the second can read as if it were a sequence of token embeddings.

The LLM decoder is the brain: an autoregressive language model that reads one mixed sequence (audio-derived vectors plus ordinary text tokens) and predicts what comes next. What "comes next" can be a text token, an interaction-state token (Chapter 3), a private thinking token (Chapter 6) or a tool call (Chapter 8).

The generator is the mouth. It turns the decoder's output into streaming audio "with context-appropriate tone and rhythm." The paper adds that "natural delivery includes expressive cues such as pauses and hesitation, connecting the content of a response with its communicative intent." The report does not describe the generator's internal design (codec, vocoder, or token rate), so we will not invent one.

The loop in the wiring. The generator's output does not only go to the speaker. It "returns to the model audio stream for subsequent interaction", which re-enters box 1. The model literally hears itself. That is how Section 5.1 can say the model "uses its own ongoing speech to interpret overlapping user utterances in the context of what the user is currently hearing." Without this feedback wire, an interruption at word twelve would be indistinguishable from an interruption at word forty.

What flows through the wires

The concept-plus-realization rule says you should be able to follow the data. The report does not publish hidden sizes, encoder frame rates or token rates, so the table below uses symbols where the paper gives no number, and the paper's numbers where it does.

WireContentsShape / typeSource
User audio inMicrophone samplesa stream of samples, cut into 320 ms blocks§5.1
Model audio inThe assistant's own generated speecha parallel stream, same 320 ms blockingFig. 3, Fig. 5
Encoder outAcoustic feature vectorsTf × denc (sizes not reported)AuT, Qwen3-Omni
Adapter outAudio vectors in LLM spaceTf × dmodel (sizes not reported)§3.1
Text inSystem prompt, history, tool resultstoken ids → L × dmodel§3.1
Decoder outNext-token distributionstate / text / think / response tokensFig. 5, Fig. 7
Generator outStreaming speechaudio, fed back as "model audio"Fig. 3

Figure 5(A) adds a finer picture of how the two streams become one sequence. Each 320 ms slot in the drawing holds a few user-audio cells and a few model-audio cells side by side; a dashed line collects them into the encoder; the decoder row is labelled "Serialized S2T input" and "Audio LLM and full-duplex interaction states"; and above each slot sits a small circle, a state token. Two of those state outputs connect to a box labelled "Step-Audio model: streaming speech output", whose output feeds back into the model audio stream. Chapter 3 is entirely about those circles.

Mixture of experts: 196 billion parameters, 11 billion at a time

Imagine a hospital with a hundred specialists. Each patient is seen by a triage nurse, who sends them to the two or three specialists most relevant to their symptoms. The hospital as a whole knows a vast amount, but any single visit costs only a few doctors' time.

A mixture-of-experts (MoE) layer works the same way. Instead of one large feed-forward block that every token passes through, the layer holds many smaller feed-forward "experts" and a small router that sends each token to a few of them. The model's total knowledge scales with the number of experts; the compute per token scales only with the experts actually used. The Mixture of Experts Gleam builds routing from zero.

The paper states: "StepAudio 3 Realtime uses a mixture-of-experts architecture with approximately 196 billion total parameters and 11 billion active parameters per token. Its language backbone is based on Step 3.7 Flash." It does not report the number of experts, how many are selected per token, or the layer count. What the two numbers do tell us is the sparsity, and sparsity is what makes a realtime model of this size plausible.

Worked example 1 · what 196B / 11B means (paper numbers; standard rules of thumb labelled) Step 1. Active fraction = active ÷ total = 11 ÷ 196.
Step 2. 196 × 0.05 = 9.8, and 196 × 0.006 = 1.176, so 196 × 0.056 = 10.976 ≈ 11. Active fraction ≈ 5.6%.
Step 3. Ratio the other way: 196 ÷ 11 ≈ 17.8. A dense model with the same per-token compute would hold about 1/17.8 of the parameters.
Step 4. Memory to store the weights must cover all 196B. At an assumed 2 bytes per parameter (16-bit; the paper does not state serving precision): 196 × 109 × 2 = 392 × 109 bytes = 392 GB.
Step 5. Compute per generated token uses only the 11B active parameters. The common rule of thumb (not from the paper) is about 2 FLOPs per active parameter for a forward pass: 2 × 11 × 109 = 2.2 × 1010 FLOPs per token.
Step 6. The paper's pretraining processes 1.2T tokens. Using the rule-of-thumb 6 FLOPs per active parameter per training token (again not a paper figure, and ignoring attention cost): 6 × 11 × 109 × 1.2 × 1012 = 6 × 13.2 × 1021 = 7.92 × 1022 FLOPs as a rough order-of-magnitude estimate, if every stage updated the whole decoder (it ignores attention, encoder and adapter compute, and it overstates the backward cost if some modules were frozen).
Step 7. Number of 32K-token sequences in 1.2T tokens: 1.2 × 1012 ÷ 3.2 × 104 = 0.375 × 108 = 37.5 million sequences.

Why does step 5 matter for a voice product? Because the model must emit something every 320 ms and, in Think-While-Speaking, runs two concurrent calls to the same model (Chapter 6). A dense model of 196B parameters would pay roughly 17.8 times more compute per token than this MoE. Sparse activation is part of what makes "think and speak at the same time" affordable.

Sim 1: send one block through the machine

Press Step to push a 320 ms block of audio through each box. The right panel shows a toy router picking experts for the current token: the expert count and the top-k are illustrative (the paper does not report them), but the readout uses the paper's 11B / 196B. Toggle ASR Max fine-tuning to see which parts the paper says are frozen and which are updated when the transcription specialist is trained.

Sim 1 — trace a 320 ms block through Figure 3

Stages light up in order. The feedback arc from the generator back to the input is the model hearing its own voice. In ASR Max fine-tuning mode, a lock marks the frozen audio encoder (paper, §4.1.1); the adapter and decoder glow as trainable. The paper does not specify frozen modules for the realtime branch, so that mode shows no locks.

Two things to take from the simulation. First, the text path joins after the adapter: audio and text become one sequence only at the decoder, which is why the adapter's job ("map into the representation space of the language model") is so central. Second, the feedback arc means every block the model speaks becomes input a few hundred milliseconds later, so the decoder's context always contains both sides of the overlap.

A forward step, written out

Here is the architecture as pseudo-PyTorch. Module names follow the paper; internal sizes are placeholders because the report does not give them. The structure, not the numbers, is the point.

pseudo-pytorchclass StepAudio3Realtime(nn.Module):
    def __init__(self, aut_encoder, moe_decoder, generator, d_enc, d_model):
        self.encoder   = aut_encoder                 # AuT encoder from Qwen3-Omni (streaming)
        self.adapter   = nn.Sequential(              # maps encoder space -> LLM space
            nn.Linear(d_enc, d_model), nn.GELU(), nn.Linear(d_model, d_model))
        self.decoder   = moe_decoder                 # ~196B total / 11B active, Step 3.7 Flash based
        self.generator = generator                   # streaming speech output

    def step(self, user_blk, model_blk, text_ids, cache):
        # user_blk, model_blk: one 320 ms block from each stream
        a_user  = self.encoder(user_blk)              # (T_f, d_enc)
        a_model = self.encoder(model_blk)             # (T_f, d_enc)  the model hears itself
        audio   = self.adapter(torch.cat([a_user, a_model]))   # (2*T_f, d_model)
        text    = self.decoder.embed(text_ids)       # (L, d_model)  separate text path
        h, cache = self.decoder(torch.cat([audio, text]), cache)
        nxt = h[-1].argmax()                        # a state, text, think, or tool token
        speech = self.generator(h, nxt) if is_response(nxt) else None
        return nxt, speech, cache                  # speech re-enters as next model_blk

Two honest caveats about this sketch. The adapter here is a small MLP because that is the most common choice; the paper says only that "an adapter maps the encoder outputs into the representation space of the language model." And the paper's Figure 5 draws user and model cells side by side within each slot, but does not specify whether they are encoded jointly or separately; the sketch encodes them separately for clarity.

Schooling, part 1: the data pipeline

A model this size is only as good as the audio it has heard. Section 3.2 describes an automated curation pipeline (inherited from StepAudio 2.5) that turns raw recordings into training samples. Walk through it like a factory line.

StageWhat it doesWhy it matters
Sound event detectionFlags what kinds of sound are presentKeeps music, noise, and speech from being mislabelled as each other
Voice activity detectionFinds where speech starts and stopsLets the pipeline cut at natural boundaries
Merge + resegmentRejoins and re-cuts audio "into samples of suitable duration that preserve semantic completeness"A sample should not end mid-sentence; half-thoughts teach bad turn-ending cues
Audio-level metadataQuality, synthetic-speech likelihood, speaker countEnables filtering TTS-generated audio and routing multi-speaker clips
Multi-system transcription + language IDSeveral recognizers transcribe; outputs are cross-checkedAgreement is evidence of a correct transcript (the same idea powers ROVER in Chapter 2)
GradingSamples graded by acoustic and semantic qualitySupports "quality-aware sampling across training stages"

For this model, the paper says, the pipeline "is extended to broaden language coverage and support the sustained perception and interaction demands of realtime dialogue." It does not give hours, languages, or grade thresholds, so neither will we.

Schooling, part 2: three stages of pretraining

Think of teaching a translator who already speaks one language fluently (text) and is learning to understand a second medium (sound). First you teach the alphabet of the new medium. Then you practise mixed conversations at scale. Finally, you polish with the best material you have. The paper's three stages follow that arc exactly.

Stage 1 · Modality alignment
"Establishes the interface between acoustic representations and the language model."
↓
Stage 2 · Multimodal mixed training
"Develops joint audio-text modeling at scale."
↓
Stage 3 · Cooldown
"Places greater weight on high-quality data to refine the resulting foundation."

Across all three, the sequence length is fixed at 32K tokens and the model processes 1.2T training tokens in total. A cooldown stage, in current practice, is the last stretch of pretraining where the data mixture shifts toward the cleanest sources (and often the learning rate decays); the paper describes only the data side.

One mixture decision stands out: "The pretraining mixture increases the proportion of pure text to preserve the general capabilities of the base language model and support subsequent reasoning and agent training." Why would an audio model want more text? Because of catastrophic forgetting: when a pretrained network is trained hard on a new distribution, it tends to lose skills it is no longer practising. The backbone arrives already good at reasoning and knowledge; pure text keeps those muscles exercised while audio is being learned. The general-text results in Chapter 9 (86.8 on HMMT February 2026) show the text skills survived, though the report does not isolate how much the text share contributed.

Schooling, part 3: midtraining for realtime interaction

Midtraining is a stage between broad pretraining and task-specific fine-tuning, where the data mixture is steered toward the capabilities the final model needs. Section 3.3 makes two changes.

Context extension. The context length grows from 32K to 128K tokens "to accommodate longer dialogue histories, earlier user requirements, and intermediate tool results." In a voice conversation, the requirement that matters might have been spoken twenty minutes ago; a tool result might be long; both must still be in view.

Mixture shift. Midtraining uses "perception, synthetic conversational, and voice-agent data" and "substantially increases the share of audio-understanding and agent-interaction data." The first broadens "speech, music, environmental sound, and audio-grounded reasoning"; the second "trains the model to carry user intent through planning, tool use, and spoken follow-up." Chapter 3 adds that duplex midtraining also includes streaming ASR, voice activity detection and utterance-completeness prediction over more than 10,000 hours of synthetic full-duplex data.

Worked example 2 · how much conversation fits in 128K? (derived from a schematic; illustrative) The paper's Figure 5(B) draws each 320 ms block as S, four audio cells, E, then one state token: 7 tokens per block in the drawing. The paper calls the waveforms schematic and never states a token rate, so treat this as an illustration of the method, not a spec.
Step 1. Blocks per second = 1000 ms ÷ 320 ms = 3.125.
Step 2. Tokens per second = 7 × 3.125 = 21.875.
Step 3. Seconds in 32K = 32,000 ÷ 21.875 = 1,462.9 s; ÷ 60 = 24.4 minutes.
Step 4. Seconds in 128K = 128,000 ÷ 21.875 = 5,851.4 s; ÷ 60 = 97.5 minutes.
Step 5. Ratio = 128 ÷ 32 = 4× more conversation in view, before subtracting text turns, private reasoning and tool results, which share the same budget.

Whatever the real token rate, the method is the lesson: context length in a voice model is a time budget, and every private thought or tool payload spends some of it. That is part of why Adaptive Thinking (Chapter 7) matters beyond latency.

Frozen or trained? What the paper does and does not say

For the ASR specialist, Section 4.1.1 is explicit: supervised fine-tuning keeps "the audio encoder frozen" while "updating the audio-language adapter and language decoder." For the realtime branch and for the pretraining stages, the report does not say which modules are updated. We will not guess.

Why freeze an encoder at all? Our reading (not the paper's stated reason): the encoder already produces strong acoustic features from its own large-scale training, and it is shared infrastructure that other capabilities depend on. Updating it on a transcription-only objective risks narrowing its features toward words and away from tone, speaker traits and ambient sound. The adapter and decoder are where transcription-specific behavior (normalization, use of context, rare terms) naturally lives.

What happens when the inputs degrade

A system diagram looks tidy until the inputs get messy. Walk each wire and ask what a bad input does downstream. The paper does not run these stress tests box by box, so the right-hand column says which later chapter or benchmark touches each failure.

Degraded inputFirst box that suffersDownstream effectWhere the paper addresses it
Noisy or reverberant user audioEncoderWeaker acoustic evidence for words and for toneSpecAugment-style masking in ASR SFT (Ch. 2); τ-Voice includes "diverse forms of background noise" (Ch. 8)
A second talker in the roomEncoder + decoderSpeech that is not addressed to the assistant enters the contextBackground speech rejection using dialogue history (Ch. 3)
User speaks over the modelFull-duplex inputTwo voices in one moment; ambiguity about intentDual-stream input + model speech as context (Ch. 3)
Rare names, product codesDecoderHomophone substitution ("sounds like a common word")Long-tail terminology augmentation (Ch. 2)
Very long sessionDecoder contextEarly requirements scroll out of view128K midtraining context (this chapter); multi-turn constraint following remains a stated gap (Ch. 9)
Slow or failing toolText path (tool results)Claims about work that has not finishedEvidence-grounded tool dialogues and negative examples (Ch. 8)

Notice that most defenses are not extra modules. They are data: augmentations, synthetic dialogues, negative examples. The architecture stays five boxes; robustness is taught, not bolted on. That is a recurring theme of the report, and Chapter 2's "less is more" ablation puts a number on how much data quality matters.

Inline check: the separate text path. Why does the paper route text through "a separate input path" instead of synthesizing tool results and prompts into speech and feeding them through the encoder? Our answer: text is already in the decoder's native vocabulary, so it arrives losslessly and cheaply; turning a long tool result into audio would waste encoder compute, add recognition errors, and spend far more context per fact. The joint conditioning the paper describes lets the decoder weigh acoustic evidence and exact text side by side.
A pattern worth remembering: share the upbringing, split at the end. ASR Max and Realtime are one pretraining run and one midtraining run, then two fine-tuning branches. That makes the ASR specialist nearly free to produce, and it means improvements to the shared foundation lift both. Chapter 5 shows the same economy one level up: several fine-tuned teachers from one base, merged back into one model.

The whole recipe on one page

The training story is spread across five sections of the report. Here it is gathered into one map, so later chapters can point back to it. The report does not state the exact order of every post-training step relative to teacher merging and Adaptive Thinking training, so the rows are grouped by purpose, not by a timeline.

StageUsed byData (as described)ContextWhat it teaches
Pretraining 1: modality alignment §3.2Both modelsCurated audio from the automated pipeline, plus text32KThe interface between acoustic representations and the language model
Pretraining 2: multimodal mixed §3.2Both modelsAudio-text at scale, with a raised share of pure text32KJoint audio-text modeling without losing text skills
Pretraining 3: cooldown §3.2Both modelsGreater weight on high-quality data32KA refined foundation (1.2T tokens across all three stages)
Midtraining §3.3, §5.3Both modelsPerception, synthetic conversational and voice-agent data; over 10,000 hours of synthetic full-duplex data with streaming ASR, VAD and completeness supervision; text128KLong histories, the time-interleaved duplex format, carrying intent through tool use
SFT, ASR branch §4.1.1ASR MaxShort labelled utterances, ROVER-fused long-form pseudo-labels, long-tail synthetic terms, optional contextpacked to 32KNormalized, context-aware transcription (encoder frozen)
SFT and post-training, realtime branch §4.2, §5.3, §6.2, §7.3RealtimeQuality-controlled audio QA (final set size not stated; the ~100K vs ~2M comparison in Chapter 2 is an ablation); self-play multi-turn dialogue; duplex interaction data; voice-agent dialogues and real trajectoriesnot statedAudio understanding, conversation, floor control, tool use
Reasoning selection §6.3.1RealtimeTurns relabelled by the no-think probe and blind judge, under per-domain budgetsnot statedWhen to think (Adaptive Thinking)
Teachers and merge §6.4RealtimeFour data compositions from one base; 3:1:1:1 averagen/aOne model holding dialogue, audio and text strengths
Inference-time system §6.3Realtimenone (runtime)n/aThink-While-Speaking, adaptive routing, MTP3 acceleration

Two patterns jump out of the map. First, almost every row's "data" column is a pipeline, not a dataset: curation, fusion, synthesis, self-play, relabelling. Second, the pure-text thread never disappears; it appears in pretraining, in duplex midtraining, and in the text-reasoning teacher, which fits an audio model staying competitive on HMMT (the report does not isolate how much the text share contributed).

Chapter 1 recap. (1) Figure 3: full-duplex input (two streams) → AuT encoder → adapter → MoE decoder (+ separate text path) → generator → back into the model stream. (2) ~196B total, 11B active per token: about 5.6% of weights per token, about 17.8× sparser than a dense model of equal size. (3) Pretraining: modality alignment, mixed training, cooldown; 32K sequences; 1.2T tokens; extra pure text to prevent forgetting. (4) Midtraining: 128K context; more audio-understanding and agent-interaction data. (5) ASR SFT freezes the encoder and trains adapter + decoder; the paper is silent on this for other stages.
In Figure 3, the generator's output streams back into the model audio stream, which re-enters the full-duplex input. What does the paper say this enables?

Chapter 2: Deep Perception

A customer on a support line says a product code out loud. It is a made-up-looking string, the kind of name no dictionary contains. The recognizer hears it and writes down the nearest common words, which sound almost the same. The agent then searches for a product that does not exist.

A minute later the same customer says "fine." Once brightly, meaning "great, go ahead". Once flatly, after a long sigh, meaning "this is not fine at all". The transcript is identical: one word, one period.

Those two moments are the two halves of this chapter. The first is a lexical failure: getting the words wrong. The second is a nonverbal failure: getting the words right and the meaning wrong. The paper's Section 4 opens by naming both: "Perception combines lexical understanding with cues about the speaker, vocal delivery, acoustic events, and temporal structure."

The paper splits the work across two models. StepAudio 3 ASR Max specializes in transcription. StepAudio 3 Realtime is trained "for broader audio understanding and spoken interaction." We take them in that order.

Part A. Measuring transcription: WER and CER

Automatic speech recognition (ASR) turns speech into text. The standard score is the word error rate (WER): align the system's output (the hypothesis) with the correct transcript (the reference), count the edits needed to turn one into the other, and divide by the reference length.

WER = (S + D + I) ÷ N

Here S is the number of substituted words (a wrong word in place of the right one), D is the number of deleted words (a reference word with nothing in the hypothesis), I is the number of inserted words (a hypothesis word with nothing in the reference), and N is the number of words in the reference. Because insertions count, WER can exceed 100%. Lower is better.

Worked example 3a · computing WER by hand (illustrative sentence) Reference (N = 7): book a table for four at seven
Hypothesis: book the table for at seven thirty
Step 1. Align word by word: book = book; a → the; table = table; for = for; four → (nothing); at = at; seven = seven; (nothing) → thirty.
Step 2. Count: S = 1 ("a" became "the"), D = 1 ("four" missing), I = 1 ("thirty" added).
Step 3. Edits = S + D + I = 1 + 1 + 1 = 3.
Step 4. WER = 3 ÷ 7 = 0.4286 = 42.9%.
Step 5. Notice the damage: the one deleted word ("four") changes the booking. WER treats all errors equally; a voice agent does not.

Mandarin has no spaces between words, and deciding where one word ends is itself ambiguous. So Mandarin results use the character error rate (CER): the same formula, counted over characters instead of words. The paper follows this convention: "English results use word error rate (WER), while Mandarin results use character error rate (CER)."

How ASR Max is fine-tuned

Recall from Chapter 1: ASR Max and Realtime share pretraining and midtraining and split at supervised fine-tuning (SFT), the stage where the model learns from input-output pairs that show exactly the desired behavior. Section 4.1.1 lists four design decisions for the ASR branch.

Packing. Examples are "packed into sequences of up to 32K tokens." Packing means concatenating several short training examples into one long sequence so that the accelerator is not wasting compute on padding. A two-second command and a thirty-second voicemail can share one 32K row.

Augmentation. The paper applies "time-frequency masking following the augmentation principle of SpecAugment." SpecAugment (Park et al. 2019) blanks out random stretches of time and random bands of frequency in the input spectrogram. It is like training a reader on pages with coffee stains: the model learns not to depend on any single moment or pitch band, which helps when real audio has a cough, a dropout, or a missing frequency range.

What trains. The audio encoder is frozen; the adapter and the language decoder are updated "to produce normalized transcripts". A normalized transcript follows fixed formatting conventions so that the same spoken content always maps to the same text; the paper does not list its normalization rules.

Context-aware recognition. "An example may additionally provide dialogue history, a preceding model response, a scenario description, or task-specific terminology as optional evidence." Then comes the crucial guard: "The target transcript remains grounded in the input waveform, allowing the model to use relevant context without simply copying unrelated terms."

Why is that guard necessary? Give a recognizer a list of product names and it can start "hearing" them everywhere, turning ordinary words into the listed terms. The training target is always what was actually said, even when the context suggests otherwise. So the model learns that context is a hint for ambiguous sounds, not a script to recite.

Short-form and long-form data, and the voting trick

The ASR mixture "combines short labeled utterances with long pseudo-labeled recordings." A pseudo-label is a transcript produced by machines rather than people. Long recordings are cheap to collect and expensive to hand-transcribe, so the question is how to make machine transcripts trustworthy.

The paper's answer is a jury. "Multiple recognition systems transcribe segmented audio, and their hypotheses are aligned and fused with Recognizer Output Voting Error Reduction (ROVER)." ROVER (Fiscus, 1997) aligns several transcripts of the same audio into one grid of word slots, then picks the most-voted word in each slot. Independent systems tend to make different mistakes, so the majority is often right where each individual is sometimes wrong.

Then two more steps. "Agreement-based filtering selects reliable segments for recomposition into longer sessions." If the jury was split on a segment, that segment is not trusted and is dropped. The surviving segments are stitched back into long sessions, and an LLM "restores punctuation and improves consistency across each session."

Sim 2 — ROVER voting and agreement filtering

Each row is one recognizer's aligned hypothesis for the same audio segment; each column is a word slot (∅ marks a deletion). The fused row takes the most-voted word per slot. The agreement score is the mean winning-vote share across slots; segments below the threshold are dropped. The three segments and five recognizers are illustrative. Try the "rare term" segment: the jury agrees, and is wrong.

Agreement threshold

Push the threshold up and the noisy segment is thrown out: that is agreement-based filtering doing its job, trading quantity for reliability. Now look at the rare-term segment. Most recognizers share the same blind spot: they have never seen the product name, so they all fall back on the same common-word spelling. The jury is confident and wrong, and no threshold catches it.

Why voting cannot fix rare words, and what the paper does instead. ROVER assumes errors are independent. For rare terms they are not: every system trained on ordinary text shares the same prior toward common words. That is precisely the gap the paper's next data source targets, "long-tail terminology augmentation": instead of trusting a jury that never learned the word, manufacture audio where the correct spelling is known in advance.

Long-tail terminology augmentation

"Rare names and technical terms are often confused with common words that sound similar. We therefore build targeted synthetic training examples for these cases." The pipeline, step by step, as Section 4.1.1 describes it:

1 · Expand a taxonomy
An LLM expands a knowledge taxonomy to find categories "rich in homophones, uncommon characters, abbreviations, and product identifiers."
↓
2 · Enumerate and deduplicate
Candidate terms are listed; duplicates are removed.
↓
3 · Carrier sentences
Each term is placed "in natural carrier sentences", so the model hears it in realistic context rather than in isolation.
↓
4 · Synthesize speech
Sentences are converted to speech.
↓
5 · Keep only faithful audio
Examples are "retained only when their pronunciation is consistent with the target text."
↓
6 · Add context for confusable terms
"Selected examples may also include dialogue history or entity hints."

Step 5 matters more than it looks. Speech synthesizers also mispronounce rare terms. A synthetic clip that says the common word while its label says the rare term would teach the model to write the rare term whenever it hears the common word, which is exactly the "unrelated lexical substitution" the paper wants to prevent. The report does not say how pronunciation consistency is checked, so we leave that detail open.

Step 6 closes the loop with context-aware recognition: for acoustically confusable terms, the example shows the model the hint and a waveform-grounded target, teaching it "to use relevant context while avoiding unrelated lexical substitutions."

The ASR scoreboard

ASR Max is evaluated on five standard sets (LibriSpeech test-clean and test-other, AISHELL-1, WenetSpeech test-net and test-meeting) and on ContextASR-Bench, "long-form, multi-domain, entity-rich speech in English and Mandarin." ContextASR-Bench is run in its Contextless setting: no domain labels, no entity lists, no hotword injection. That makes it a test of what the model knows, not of what it is told. All baselines "are rerun under the same evaluation setup... using the same test audio and scoring procedure."

Test set (lower is better)ASR MaxDoubao 2.0 ASRSeed 2.0 LiteHY3.0 ASR Preview
LibriSpeech test-clean (WER)1.182.941.471.38
LibriSpeech test-other (WER)2.285.982.672.80
AISHELL-1 (CER)0.492.071.661.22
WenetSpeech test-net (CER)3.994.034.713.71
WenetSpeech test-meeting (CER)4.355.094.804.12
ContextASR-Speech-EN (WER)7.9112.049.488.53
ContextASR-Dialogue-EN (WER)3.439.093.654.66
ContextASR-Speech-ZH (CER)1.432.802.151.74
ContextASR-Dialogue-ZH (CER)1.0210.474.151.63

ASR Max is best on seven of nine rows. On both WenetSpeech subsets it beats Doubao 2.0 ASR and Seed 2.0 Lite but trails HY3.0 ASR Preview "only by a small margin." On all four ContextASR-Bench subsets it is best.

Worked example 3b · reproducing the paper's ContextASR macro averages (Table 1 numbers) ASR Max, English: (7.91 + 3.43) ÷ 2 = 11.34 ÷ 2 = 5.67% (paper: 5.67%).
ASR Max, Mandarin: (1.43 + 1.02) ÷ 2 = 2.45 ÷ 2 = 1.225 ≈ 1.23% (paper: 1.23%).
HY3.0, English: (8.53 + 4.66) ÷ 2 = 13.19 ÷ 2 = 6.595 ≈ 6.60% (paper: 6.60%).
HY3.0, Mandarin: (1.74 + 1.63) ÷ 2 = 3.37 ÷ 2 = 1.685 ≈ 1.69% (paper: 1.69%).
Relative reduction, English: (6.60 − 5.67) ÷ 6.60 = 0.93 ÷ 6.60 = 14.1% fewer errors.
Relative reduction, Mandarin: (1.69 − 1.23) ÷ 1.69 = 0.46 ÷ 1.69 = 27.2% fewer errors.
Largest single gap, AISHELL-1: (1.22 − 0.49) ÷ 1.22 = 0.73 ÷ 1.22 = 59.8% fewer errors than the next best system.

Relative reductions are the right lens for low error rates. A drop from 1.22 to 0.49 is "only" 0.73 points, but it removes about three of every five remaining errors. One caution the paper itself states plainly: "these results characterize the ASR-specialized model, not the transcription behavior of the realtime model."

Part B. Hearing more than words

Now the "fine." problem. Section 4.2 builds training data from a hierarchical taxonomy (a tree of capability categories and sub-categories) covering six areas.

Taxonomy branchAn example question it licenses (ours, illustrative)
Lexical contentWhat did the speaker ask for?
Paralinguistics (how something is said: pitch, pace, emotion, emphasis)Does the speaker sound frustrated?
Acoustic eventsIs there a siren in the background?
Speaker and temporal structureHow many people speak, and who speaks first?
MusicIs the accompaniment major or minor?
Audio-grounded reasoningGiven the sounds, where was this likely recorded?

Figure 4 of the paper draws the construction pipeline as five numbered boxes, guided by two banners: "Sampling controls: what audio is allowed in" and "Self-defined sub-capability space: what may be asked about it."

1 · Sampling
Duration and language quotas, deduplication, removal of unsuitable recordings.
↓
2 · Description
A caption of what the clip contains.
↓
3 · Capability tagging
Which aspects of the taxonomy the clip actually supports.
↓
4 · Query construction
One or more questions tailored to each clip.
↓
5 · Labelling
"Multiple models then independently label each audio–question pair," and outputs are consolidated through agreement and quality checks.

Step 3 prevents a common failure in synthetic audio QA: asking a question the clip cannot answer. A silent room recording should never become a training example about the speaker's emotion. Tagging first, then asking only about supported aspects, keeps questions answerable.

The grounding problem, and a clever proxy

Here is the hard part of audio data quality. To check whether an answer is grounded (actually supported by the sound), a checker must hear the sound. But the strongest, cheapest judges are text-only LLMs. The paper resolves this with a division of labor, in three layers.

Layer 1, deterministic checks. Remove "empty, truncated, malformed, or severely repetitive outputs."

Layer 2, text-only judges. They "assess query and response quality and assign a case-value score", which "jointly considers the query, the response, the amount of useful information available in the audio as represented by its annotations, and the training value of the question." The paper is explicit: "These judges do not directly evaluate audio grounding."

Layer 3, cross-model consistency. "Grounding reliability is estimated from the consistency of responses independently produced by multiple models for the same audio–question pair." If several models that did hear the clip give the same answer, the answer is probably in the audio. If they disagree, something is off: the question is ambiguous, the audio is unclear, or one model hallucinated.

python (sketch)def route_candidate(clip, question, answers, judge):
    # answers: independent responses from several audio models for the same pair
    if is_broken(answers):                        # empty, truncated, malformed, repetitive
        return "reject"
    quality    = judge.quality(question, answers)      # text-only LLM judge
    case_value = judge.case_value(question, answers, clip.annotations)
    agreement  = consistency(answers)              # proxy for audio grounding
    if quality >= Q_HI and case_value >= V_HI and agreement >= A_HI:
        return "sft_candidate"                   # the best of the best
    if broadly_useful(question, answers):
        return "midtraining"                     # useful, not SFT-grade
    return "relabel_or_review"                   # disagreements, correctable cases
# Thresholds are placeholders: the paper describes the routing, not the numbers.

The routing is the realization of one sentence in the paper: "Only candidates with high quality, high case value, and strong cross-model consistency are retained as SFT candidates; broadly useful examples may enter midtraining, while disagreements and correctable cases are routed to relabeling or further review." Nothing is simply thrown away if it can still teach something at a lower tier.

The audio-understanding scoreboard

Benchmark (0–100, higher better)StepAudio 3 RealtimeDoubao 2.0 LiteGemini 3 FlashGemini 3.1 Pro
Big Bench Audio98.198.899.499.6
AudioMultiChallenge49.348.556.667.0
MMSU90.680.077.083.6
MMAU79.077.577.680.5
WildSpeech77.173.974.477.7
MMAR86.575.975.481.7
Step-Caption78.276.867.874.8
MTalk-Bench91.789.988.589.1
Macro average81.377.777.181.8

The paper's reading: StepAudio 3 Realtime "leads the reported baselines on four of the eight benchmarks," with the largest margins on MMSU (90.6 vs 83.6, +7.0) and MMAR (86.5 vs 81.7, +4.8). It "trails Gemini 3.1 Pro by 17.7 points on AudioMultiChallenge" (67.0 − 49.3 = 17.7), and Big Bench Audio "is nearly saturated for all systems." The paper draws the conclusion itself: "maintaining and revising constraints over natural multi-turn audio remaining a clear area for improvement."

Two footnotes from the protocol section (8.3) matter for reading those rows. Step-Caption "uses judge-based scoring against annotated speaker attributes." And the MTalk-Bench entry "includes only the Paralinguistic Information and Ambient Sound components, reported as one aggregate", so it is not the full benchmark.

Less is more: the ablation that should change your data budget

Section 4.2 ends with the chapter's most transferable result. The team ran an ablation where "the SFT data are the only changed factor": roughly two million randomly sampled examples versus about 100K high-quality examples retained after the quality control described above.

Metric~2M random~100K quality-controlledChange
MMSU78.7889.70+10.92
MMAR74.7084.50+9.80
WildSpeech74.2077.11+2.91
MTalk-Bench (ambient, paralinguistic, semantic subsets)88.8390.84+2.01

Check the arithmetic: 89.70 − 78.78 = 10.92; 84.50 − 74.70 = 9.80; 77.11 − 74.20 = 2.91; 90.84 − 88.83 = 2.01. And the size ratio: 2,000,000 ÷ 100,000 = 20, the paper's "roughly one twentieth as many examples."

Why fewer, better examples can win so decisively. A plausible reading (ours, not the paper's): a random pool contains answers that are not grounded in the sound, and training on them teaches a model to answer from priors instead of from listening, which is exactly what audio reasoning benchmarks punish. The largest gains land on MMSU and MMAR, the reasoning-heavy benchmarks, which fits that reading. The paper's own conclusion is more modest and worth quoting: the result highlights "the importance of data quality over raw SFT volume."
Chapter 2 recap. (1) WER = (S + D + I) ÷ N; CER is the same over characters, used for Mandarin. (2) ASR Max SFT: 32K packing, SpecAugment-style masking, frozen encoder, trained adapter + decoder, optional context that never overrides the waveform. (3) Long-form pseudo-labels come from ROVER fusion plus agreement filtering; rare terms come from synthetic carrier sentences kept only when pronunciation matches. (4) Audio-understanding data: taxonomy → sampling → description → tagging → questions → multi-model labels; text judges score quality and case value, cross-model agreement stands in for grounding. (5) 100K curated examples beat 2M random ones on every reported metric.
The paper's text-only LLM judges cannot hear the audio. How does the data pipeline estimate whether an answer is actually grounded in the sound?

Chapter 3: Seamless Duplex

You are explaining a recipe to a friend over the phone. Halfway through, they say "right." Do you keep going?

It depends. If "right" came in the gentle rhythm of someone following along, you keep going. If it came sharply, followed by a breath, it is the start of "right, but you said two eggs earlier", and you should stop. The paper uses this exact word as its example: "'right' may function as a backchannel acknowledging the model's explanation or as a preface to a correction."

People resolve that ambiguity dozens of times a minute without noticing. A voice model has to do it explicitly, from audio, in real time. This chapter is about how StepAudio 3 Realtime does it.

The floor, and the three verbs that move it

Picture a talking stick passed around a circle: whoever holds it speaks. Linguists call the right to speak the conversational floor. The paper frames the whole problem around it: "A central challenge in full-duplex dialogue is determining when to take, retain, or yield the conversational floor."

Three verbs, and each has a matching failure. Take the floor too early, and you interrupt someone who was only pausing. Retain it too stubbornly, and you talk over a correction. Yield it too easily, and every "mm-hm" kills your answer.

Two vocabulary words before we go on. A backchannel is a short listener signal ("mm-hm", "yeah", "right") that shows engagement without asking for the floor. An interruption (in voice products, often called barge-in) is a substantive attempt to take the floor while someone else is speaking.

The paper names the two distinctions that matter: "distinguishing pauses within an unfinished utterance from turn completion, and brief acknowledgments from attempts to interrupt." And it says what resolving them requires: "acoustic evidence interpreted in the context of the unfolding dialogue." Sound alone is not enough. Words alone are not enough. The model needs both, plus history.

Three inputs for every floor decision

Figure 5(A)'s caption names three inputs that jointly inform one decision: user speech, model-side speech, and dialogue history (the drawing itself shows the two audio streams feeding the streaming encoder). Each carries evidence the others lack.

InputEvidence it carriesDecision it helps most
User speechIs there voice? How loud, how long, what words, what prosody?Pause vs end; backchannel vs interruption
Model speechWhat has the user already heard at this instant?Is the overlap a reaction to what was just said?
Dialogue historyWhat was asked, what is pending, who has been addressedIs this speech even directed at the assistant?

Time, cut into 320 ms blocks

Section 5.1 gives the most concrete detail of the duplex design: "Audio is organized into 320 ms blocks, each followed by a state or text token." Every block of listening ends with the model saying, in its own token vocabulary, what it is doing now.

Figure 5(B) draws four of these blocks on a timeline from 0.00 s to 1.28 s. Each block is drawn as a start marker S, four audio cells a, an end marker E, and then a state token. The drawn state tokens, in order, are:

[S a a a a E] listening: no_voice  ·  [S a a a a E] listening: user_voice  ·  [S a a a a E] listening: user_voice  ·  [S a a a a E] speaking transition

Read that line as a tiny story. In the first 320 ms, nobody is talking: the model notes "no voice" and keeps listening. In the next two blocks, the user speaks: "user voice", still listening. In the fourth block, the model decides the user is done and emits a "speaking transition": it takes the floor. The caption above the timeline reads: "One state or text token follows every 320 ms audio block."

Here is the same drawing as a table, with a middle column of our own reading: the kind of evidence each decision would need. The paper does not annotate the figure with evidence values; the first and last columns are taken directly from it.

Block (Fig. 5B)ContentsEvidence a correct decision needs (our reading)State token (Fig. 5B)
0.00–0.32 sS a a a a ENo voice in the user cellslistening: no_voice
0.32–0.64 sS a a a a EUser voice present; an utterance has startedlistening: user_voice
0.64–0.96 sS a a a a EUser voice continues; the utterance is not yet completelistening: user_voice
0.96–1.28 sS a a a a EThe utterance is judged complete and the floor is freespeaking transition

Notice what the fourth row requires. Nothing in a single block's audio says "the user is done". The decision depends on what was said in the previous blocks, which is why the state token is predicted by a decoder that sees the whole history, not by a classifier that sees one block.

Section 5.1 names the full set of decisions these tokens encode: "continue listening, initiate a response, continue speaking, or yield the conversational floor." The figure spells out three token names; the report does not list the exact spelling of the others, so in this lesson we call them by the paper's verbs.

The design choice that makes everything else possible. Floor control is not a separate timer bolted onto a chatbot. It is a token the model predicts, in the same stream as its words, after every block of audio. That means the floor decision is conditioned on everything the decoder can see: both audio streams, the dialogue history, and the model's own partial response. Turn-taking becomes next-token prediction, and next-token prediction is what large models are best at.
Worked example 4a · the decision clock (paper numbers) Step 1. Decisions per second = 1000 ms ÷ 320 ms = 3.125.
Step 2. Figure 5(B) check: 4 blocks × 0.32 s = 1.28 s, exactly the right end of the drawn axis.
Step 3. A floor decision about something that happens inside a block can be emitted only after that block ends. Worst case, the event happens at the start of a block, so it waits up to 320 ms for its decision slot (plus compute time, which the paper does not report).
Step 4. An illustrative 1.0 s thinking pause spans 1000 ÷ 320 = 3.125 blocks, so it fully covers 3 blocks (with a remainder of 1000 − 3 × 320 = 40 ms that may fall into a fourth). The model must emit "keep listening" three times in a row while the silence tempts it to answer.
Step 5. Over a illustrative 10-minute call: 600 s × 3.125 = 1,875 floor decisions. A policy whose per-decision error rate were 1% would still make about 1,875 × 0.01 = 18.75 ≈ 19 wrong calls in that one call. Floor control is a high-frequency decision, so small error rates compound into noticeable moments.

Why the model listens to itself

Section 5.1 adds: "The model also uses its own ongoing speech to interpret overlapping user utterances in the context of what the user is currently hearing."

Suppose the model is reading out three restaurant options. The user says "that one!" during option two. If the model only knew what it planned to say, "that one" would be ambiguous. Because the model audio stream is part of its input (the feedback wire from Chapter 1), the decoder knows that the user had just heard the name of option two when they spoke. The overlap is interpreted against the words that were actually in the air.

The three context-aware decisions

Section 5.2 walks through three decisions, each with its own mix of evidence.

Pauses and turn endings. "The system combines acoustic timing with semantic completeness to distinguish within-turn pauses from turn endings." Semantic completeness asks whether the words so far form a finished request. "Book me a table for" followed by silence is acoustically a pause and semantically incomplete: keep listening. "Book me a table for two at seven" followed by the same silence is complete: respond. The silence is identical; the words decide.

Backchannels and interruptions. "During model speech, user backchannels are interpreted in relation to the ongoing response. Brief acknowledgments can signal continued engagement without requesting a floor transfer, whereas a substantive request or correction may signal an intent to interrupt." The deciding question is whether the overlap asks for something. "Mm-hm" asks for nothing: keep speaking. "Wait, I meant Tuesday" asks for a change: yield.

Background speech rejection. "Dialogue history provides contextual evidence for assessing whether incoming speech is directed at the assistant." A television announcing weather in another city, or a colleague talking to someone else, is speech, but not for the assistant. If the conversation so far is about a dinner booking, a voice saying "and now, sports" is almost certainly background. The paper's verb is careful: the assessment "informs whether the speech should be incorporated into the active exchange or treated as unrelated background conversation."

Sim 3: drive the floor manager

Inject events and watch two policies react block by block. The top decision row is a context-aware policy in the spirit of Section 5.2: it combines voice, semantic completeness, whether speech is addressed to the assistant, and whether an overlap is substantive. The bottom row is a plain silence timer that takes the floor after two silent blocks and yields to any voice. A red outline marks a decision that disagrees with the scenario's intended behavior.

Sim 3 — the conversational floor, one 320 ms block at a time

Evidence values (completeness, addressed, substantive) are illustrative: the real model infers them implicitly from both audio streams and dialogue history. Playback is slowed so you can read it; in the model, each column is 320 ms. Chip codes: L·nv = listening: no_voice, L·uv = listening: user_voice, →S = speaking transition (these three are spelled in the paper's Figure 5); S = continue speaking, Y = yield, L·bg = keep listening, speech not addressed to the assistant (paper's verbs; token spellings not reported).

Run each scenario and compare the rows. The silence timer interrupts the paused user, is slower than necessary at genuine turn ends, abandons its answer at the first "right", and happily answers the television. It gets exactly one case right: the real interruption. The context-aware row gets all five, because each decision uses a different piece of evidence: completeness for pauses, substantiveness for overlaps, addressedness for background speech.

That pairing is the one the paper highlights in its evaluation: "respecting within-turn pauses while responding at turn completion, and accommodating user interruptions while continuing through backchannels." Each pair is a tension. A policy that is good at one side of a pair by being trigger-happy or stubborn fails the other side.

Serializing two streams into one sequence

How do you turn two continuous audio streams and a decision into training data for an autoregressive decoder? The sketch below builds the time-interleaved sequence of Figure 5(B). Block markers and the three state names follow the figure; the remaining token names and the target labels are ours.

python (sketch)BLOCK_MS = 320                                   # paper, section 5.1

def serialize_duplex(user_audio, model_audio, floor_labels):
    """user_audio, model_audio: aligned streams (same clock).
    floor_labels[k]: the decision after block k, e.g. 'listening: no_voice',
    'listening: user_voice', 'speaking transition', or a text token."""
    seq, loss_mask = [], []
    for k, label in enumerate(floor_labels):
        t0, t1 = k * BLOCK_MS, (k + 1) * BLOCK_MS
        blk = ["<S>"]
        blk += audio_slots(user_audio[t0:t1])        # user cells
        blk += audio_slots(model_audio[t0:t1])       # model cells: what the user is hearing
        blk += ["<E>"]
        seq       += blk + [label]
        loss_mask += [0] * len(blk) + [1]       # learn the decision, not the audio
    return seq, loss_mask

# Training objective: ordinary next-token cross-entropy on the masked positions.
# The "decision" is just the next token after <E>; no separate classifier head.

The loss mask is the realization of "turn-taking as next-token prediction." Audio positions provide context; decision positions provide supervision. The paper does not publish its exact masking; this is the standard way to express "predict a state or text token after each block."

How the duplex behavior is trained

Section 5.3 describes two phases.

Midtraining "adapts the model to the time-interleaved representation used for full-duplex interaction." It "combines supervision for streaming ASR, voice activity detection (VAD), and streaming prediction of utterance completeness." The mixture "includes over 10,000 hours of synthetic full-duplex interaction data," and "text data are also incorporated to help retain general language and reasoning capabilities."

Look at how neatly those three supervision signals line up with the decisions.

Midtraining signalWhat it teachesDecision it feeds
Streaming ASRWhich words have been said so far, block by blockSemantic completeness; substantive vs brief overlap
VADWhether a block contains voice"listening: no_voice" vs "listening: user_voice" (Figure 5's own labels)
Streaming utterance completenessWhether the utterance so far is finishedPause vs turn ending; when to emit "speaking transition"

Why synthetic duplex data? Real two-channel conversational recordings with clean per-speaker separation and labelled floor events are scarce. Synthesis lets you control exactly where pauses, overlaps and background voices occur, so every block can carry a correct label. The report does not describe how its 10,000+ hours were synthesized.

Post-training "further refines conversational behavior using high-quality interaction data covering turn taking, user backchannel handling, interruption handling, and background speech rejection." Those four behaviors map one-to-one onto the scenarios in Sim 3, and three of them onto the benchmark below.

When the evidence gets muddy

Each floor decision leans on a particular piece of evidence, so each has a particular way to fail when that evidence degrades. The paper does not report per-condition duplex ablations; the table below is our analysis of which decision is exposed to which degradation, using only the mechanisms the paper describes.

DegradationEvidence it corruptsLikely wrong decisionWhich input can rescue it
Slow, hesitant speaker with long pausesAcoustic timingTaking the floor mid-requestSemantic completeness from the words so far
Clipped, complete-sounding fragment ("two.")Semantic completenessResponding before the user adds "…and a high chair"Dialogue history (was a list being built?) and prosody
Loud room, TV in the backgroundVoice presenceAnswering speech not meant for the assistantDialogue history: is this on-topic and addressed?
User says "right" in a flat, ambiguous toneProsodic cue for intentYielding to an acknowledgment, or ignoring a correctionThe words that follow in the next block, and the model's own speech
Model's own voice leaking into the user micUser stream purityTreating its own echo as an interruptionThe model audio stream: the model knows what it just said

Two points stand out. First, every row has a rescue that comes from a different input than the one that was corrupted. That redundancy is the practical argument for feeding the decoder all three inputs rather than a single voice-activity signal.

Second, the 320 ms cadence makes waiting cheap. When evidence is ambiguous, "keep listening for one more block" costs only 320 ms, and the next block usually disambiguates. A silence timer has no such option: it either fires or it does not. The echo row is our own inference from the dual-stream design; the paper does not discuss acoustic echo explicitly.

How this relates to earlier full-duplex models. The paper cites Freeze-Omni, Moshi, chronological thinking in full-duplex dialogue models, and DuplexSLA as streaming and full-duplex systems (its references [9]–[12]), and grounds the "evidence in context" framing in [10], [11] and Bayling-Duplex [27]. For the lineage, the Moshi Veanor shows a model that listens and speaks on parallel streams, and the DuplexSLA Veanor covers synchronized speech, language and action from a paper whose author list overlaps with this report's contributors.

The evaluation: Artificial Analysis full-duplex

The paper evaluates on the Artificial Analysis (AA) subset of Full Duplex Bench v1 and v1.5, which scores four aspects: pause handling, turn taking, user interruption handling, and backchannel handling. Category scores "measure the percentage of samples satisfying the corresponding interaction criterion."

Capability (0–100)StepAudio 3 RealtimeGPT-realtime-2 (High)Qwen Audio 3.0 Realtime PlusGrok Voice Think Fast 2.0 High
Pause handling98.999.398.098.0
Turn taking100.0100.098.091.0
User interruption handling99.095.098.097.0
Backchannel handling98.086.7100.095.0
Overall98.995.398.495.1

StepAudio 3 Realtime ranks first overall at 98.9, ahead of Qwen Audio 3.0 Realtime Plus at 98.4. It is not best in every row: GPT-realtime-2 is slightly better at pauses (99.3), and Qwen is perfect on backchannels (100.0). The paper's claim is about balance: "Strong performance across both pairs indicates balanced conversational control over when to listen, speak, and yield."

Look at GPT-realtime-2's profile: 100.0 on turn taking and 86.7 on backchannels. Our reading (the paper does not analyze baseline behavior): this profile is consistent with a policy that responds promptly but also treats some acknowledgments as interruptions, exactly the tension from Sim 3.

Worked example 4b · "Overall" is not the average of the four rows (Table 3 numbers) The paper says it uses "the source-reported Overall score, preserving the benchmark's aggregation rather than averaging the four displayed category scores." Check it.
StepAudio 3: 98.9 + 100.0 + 99.0 + 98.0 = 395.9; 395.9 ÷ 4 = 98.975 vs Overall 98.9.
GPT-realtime-2: 99.3 + 100.0 + 95.0 + 86.7 = 381.0; 381.0 ÷ 4 = 95.25 vs Overall 95.3.
Qwen: 98.0 + 98.0 + 98.0 + 100.0 = 394.0; 394.0 ÷ 4 = 98.5 vs Overall 98.4.
Grok: 98.0 + 91.0 + 97.0 + 95.0 = 381.0; 381.0 ÷ 4 = 95.25 vs Overall 95.1.
Conclusion: the simple means land close to, but not exactly on, the reported Overall values (for example 98.975 vs 98.9). The benchmark aggregates differently from an unweighted mean of the displayed rows, so do not recompute Overall yourself. The ranking is the same either way: StepAudio 3 first, Qwen second.
Inline check: which evidence, which decision? A user says "yeah" while the model is speaking, and the model keeps going. Later the user says "yeah, no, cancel that" and the model stops. Both overlaps start with the same word. What evidence separates them? Answer: the second overlap is substantive (a request to change something), which the model can tell from the words that follow and from the dialogue context; the paper's rule is that "brief acknowledgments" continue while "a substantive request or correction may signal an intent to interrupt."
What the benchmark does not cover. The AA subset scores pauses, turn taking, interruptions and backchannels. Background speech rejection is trained in post-training (§5.3), but no benchmark in the report isolates it: it is not one of the four scored categories in Table 3. τ-Voice (Chapter 8) adds "diverse forms of background noise", which is a related but different condition: noise is not speech that might or might not be addressed to the assistant.
Chapter 3 recap. (1) The floor is taken, retained or yielded; each has a failure mode. (2) Audio is cut into 320 ms blocks, each followed by a state or text token; Figure 5 spells "listening: no_voice", "listening: user_voice" and "speaking transition". (3) Decisions use user speech, the model's own speech, and dialogue history; completeness separates pauses from endings, substantiveness separates backchannels from interruptions, history separates addressed from background speech. (4) Midtraining: streaming ASR + VAD + completeness over 10,000+ hours of synthetic duplex data, plus text; post-training on the four behaviors. (5) AA Full-Duplex Overall 98.9, first among the reported systems, with 100.0 turn taking and 99.0 interruption handling.
The model is speaking, and the user says "right." What does StepAudio 3 Realtime use to decide whether to keep speaking or yield?

Chapter 4: Conversational IQ

Early in a long call, you mention that you do not eat meat. Twenty minutes and a dozen topics later, you ask for a quick dinner idea. The assistant cheerfully suggests a steak.

Its timing was perfect. It did not interrupt you. It did not go silent. It was still a bad conversation partner, because it forgot a constraint you gave it.

The paper draws this line sharply at the start of Section 6: "Seamless Duplex determines when the model should respond. Conversational intelligence determines how it should engage with the user and how much reasoning the response requires." Chapter 3 was about when. This chapter is about what, and about how the team measured and trained it.

The paper's description of the target behavior is worth reading as a checklist: "A natural voice assistant should follow intent across turns, clarify underspecified goals, and move the conversation toward a useful outcome. Routine turns should avoid unnecessary deliberation, while complex requests should retain the reasoning needed for a reliable answer."

Part A. A benchmark built for conversation: StepAudioChat

To improve a skill, you first need a ruler. The team built its own: StepAudioChat, "a closed, text-based benchmark for foundational conversational intelligence." Two words in that sentence are design decisions.

Text-based. "Its scope isolates text-level response quality from prosody, turn timing, interruption handling, and other properties of the speech interface." If a benchmark played audio in and scored audio out, a low score could mean bad reasoning, bad synthesis, or bad timing, and you could not tell which. Scoring text separates the question "was the content right?" from "was the delivery right?" Delivery is measured elsewhere (Chapter 3).

Closed. The items are newly constructed and not published, "to reduce reliance on public test questions." Public test sets leak into training data over time, a problem called contamination: a model can score well because it has seen the questions, not because it has the skill. A closed set cannot be memorized from the web. The cost, which we should say plainly, is that outsiders cannot reproduce StepAudioChat numbers.

The capability taxonomy

Section 6.1.1: "We organize conversational abilities into a hierarchy whose leaves target observable behaviors with defined evaluation boundaries." A leaf is not "is helpful"; it is something a checker can observe, with a clear line between pass and fail.

How do you know the taxonomy covers the right ground? The team mapped "tasks and representative examples from public benchmarks" onto it as "a coverage check, without reusing their test questions." Where public examples mapped to several leaves, the overlap was used to sharpen definitions; where examples mapped to nothing, the team had found a gap.

Within each capability family, items vary "the source of a constraint, its form of expression, and its interaction with other conditions." That produces the hard cases: "nested constraints and requirements that must remain consistent across turns." Every item "targets a primary capability," and "only capabilities supported by validated items enter the evaluation suite."

The benchmark reports eight dimensions. The paper names them but does not define each one in prose, so the right-hand column below is our short gloss of what the name suggests, not a quotation.

StepAudioChat dimensionOur gloss (illustrative)
Instruction FollowingHonors explicit constraints: length, format, what to include or avoid
FaithfulnessStays true to given facts and context; does not invent
ReasoningWorks multi-step problems correctly inside dialogue
MemoryUses information from earlier turns ("contextual recall")
KnowledgeKnows facts about the world
Safety & ReliabilityHandles risky requests and uncertainty responsibly
Conversational PragmaticsReads intent, clarifies ambiguity, fits the social moment
Persona & Role ConsistencyKeeps an assigned role, voice and style over the conversation

An item has four parts, two of them controls

Section 6.1.2: "Each item contains a dialogue prompt, independently checkable criteria, a valid reference response, and a deliberately flawed response."

The last two are the clever part. In a lab experiment, a positive control is a sample you know should test positive, and a negative control is one you know should test negative. If your test fails either control, you do not trust any of its results. The paper uses the valid and flawed responses exactly that way: they "provide positive and negative controls for the judging criteria while allowing multiple valid phrasings."

Here is an illustrative item in that shape (ours, not from the closed set):

PartContent
Dialogue promptTurn 1: "i'm vegetarian btw." … Turn 6: "ok what's a quick dinner, like 20 min"
Criterion AThe suggestion contains no meat or fish.
Criterion BThe suggestion is plausibly doable in about 20 minutes.
Criterion CThe reply is short enough to speak comfortably.
Valid reference"A chickpea and spinach stir-fry: about fifteen minutes with canned chickpeas."
Flawed response"A quick garlic shrimp pasta takes about twenty minutes." (violates A)

Notice the prompt's style: lower case, fragmentary, casual. That is deliberate. "De-identified utterances from real interactions inform prompt phrasing and local context, preserving brevity, fragmentation, colloquial wording, and transcription noise." People do not speak to assistants in tidy paragraphs, and a benchmark written in tidy paragraphs flatters models.

Two privacy and scope rules follow. "Personal entities are replaced with typed placeholders" (a name becomes something like a PERSON slot). The real utterances "do not supply expected answers or grading criteria"; they only shape how prompts sound. And "except for intrinsically domain-specific capabilities, scenarios are recast across everyday settings to broaden coverage beyond individual applications."

Quality control and difficulty calibration

Section 6.1.3 stacks the checks.

Structural validation
Prompts are complete and self-contained.
↓
Control check
"The valid response must satisfy every judging criterion, and the flawed response must violate at least one."
↓
Semantic audit
Checks "the premise, reference response, and judging logic." Automated checks for coverage; targeted human review of semantic failures and anomalies.
↓
Difficulty calibration
Responses from "two reference systems with different capability levels" guide "the mixture of baseline, discriminative, and difficult examples."
↓
Judge calibration
Using "the valid and flawed controls and a trusted labeled sample."

Difficulty calibration deserves a picture. Run a weaker and a stronger reference system on every item and sort by outcome. The paper names three kinds of items; the mapping from outcomes to kinds below is our reading of those names.

Weaker systemStronger systemItem kind (our mapping)What it measures
passpassBaselineFloor competence; catches regressions
failpassDiscriminativeSeparates stronger from weaker models
failfailDifficultHeadroom above today's systems
passfailSuspiciousOften a sign of ambiguous wording or inconsistent grading, worth auditing

The paper's own summary of why all this matters: "These checks help distinguish capability demands from ambiguous wording or inconsistent grading." A benchmark item that fails a model for the wrong reason is noise in the score.

Part B. Teaching conversation: the multi-turn dialogue data

Section 6.2 turns from measuring to training. The goals: "follow user intent and constraints across turns, clarify underspecified requests, and adapt responses to the conversational context." And one sentence that sets up Chapter 7: "Joint training on dialogue and reasoning examples supports deliberation on difficult requests and concise responses to routine turns."

Figure 6 of the paper draws the construction as five numbered boxes with two side inputs.

1 · Discussion points
"Topic-classified source points": concrete things to talk about.
↓  + persona corpus: user personas and interaction styles
2 · Case plan
"Topic, scene, persona, and turn plan."
↓
3 · Self-play
"Two models; N−1 turns of context." One plays the user, one the assistant.
↓  + capability injection: optional requirements for the final turn
4 · Final-turn query
"The user asks; no answer is produced here."
↓
5 · Labelling
"This one answer is the only target."

The resulting training example, in the figure's words: "the first N−1 turns as context, the final query, and the labelled response with its reasoning."

Self-play means two models converse to generate data. It is cheap and unbounded, and it has one weakness: the synthetic assistant turns may be mediocre. The pipeline handles that by using the self-played turns only as context. The only thing the student model is trained to produce is the carefully labelled final answer. That is why "no answer is produced" at step 4: the final answer comes from a separate, higher-quality labelling step.

python (sketch)def build_dialogue_example(case_plan, user_model, asst_model, labeler, N):
    history = []
    for t in range(N - 1):                               # self-play: N-1 turns of context
        history.append(("user", user_model.speak(case_plan, history)))
        history.append(("assistant", asst_model.reply(history)))
    query = user_model.final_query(case_plan, history,     # capability injection is optional
                                   inject=case_plan.capabilities)
    reasoning, answer = labeler.label(history, query)     # the ONLY target
    tokens, mask = [], []
    for role, text in history + [("user", query)]:
        ids = tok(role, text); tokens += ids; mask += [0] * len(ids)   # context: no loss
    ids = tok("assistant", think(reasoning) + answer)
    tokens += ids; mask += [1] * len(ids)                  # loss on reasoning + answer
    return tokens, mask

The mask line is the realization of "this one answer is the only target." Whether the paper trains on self-played assistant turns at all is not stated beyond that figure label; the sketch follows the figure.

Factorized construction and the rotation schedule

Generation "follows four axes: topics and their concrete discussion points; participant personas that specify roles, backgrounds, and interaction styles; turn depth; and target capabilities such as contextual recall, logical reasoning, and instruction use."

If you sample each axis independently from its overall popularity, the common combinations crowd out the rare ones. Popular personas meet popular capabilities over and over, and some combinations never appear. The paper's fix: "A stratified rotation schedule coordinates these choices within topics, extending coverage beyond their global marginal distributions."

A marginal distribution is the frequency of one axis on its own (how often each persona appears, ignoring everything else). Matching the marginals does not guarantee good coverage of combinations. A stratified rotation deliberately cycles through combinations within each topic instead of sampling each axis independently. In its idealized form (our illustration, used in the simulation below), every persona meets every capability at every depth before any combination repeats. The paper states only that "a stratified rotation schedule coordinates these choices within topics, extending coverage beyond their global marginal distributions"; it gives no stronger guarantee.

Sim 4 — independent sampling vs stratified rotation

One topic. 4 personas × 4 capabilities × 3 turn depths = 48 combinations (sizes and the skewed marginals are illustrative). Each cell's shade is how many dialogues landed on that combination. The curve panel tracks coverage for both strategies as you draw more dialogues.

Draw 48 dialogues in each mode. Rotation covers all 48 combinations exactly once. Independent sampling, with the same budget, leaves a sizeable share of cells empty and piles many dialogues onto the popular corner. The model trained on the second set would rarely practise, say, a terse expert persona asking for deep contextual recall on turn twelve.

Turn depth as a first-class variable

"Turn depth is an explicit construction variable. Longer dialogues progress through deeper engagement with discussion points, allowing later turns to refine constraints, resolve ambiguity, revisit evidence, or change direction while remaining consistent with the history."

Why make depth explicit rather than letting it vary naturally? Because the behaviors that break in long conversations (forgotten constraints, contradicted earlier answers) only appear at depth. If most synthetic dialogues are three turns long, the model rarely practises turn fifteen. The paper later admits that "multi-turn constraint following" remains a gap; explicit depth is the data-side lever aimed at it.

Three-stage quality routing

"Quality control separates three concerns into independent stages."

StageWhat it checks
1 · Context–query reviewDepth and informativeness of the history; whether the opening is self-contained; whether the final query is substantive and connected to what came before
2 · Capability-specific reviewWhen capabilities were injected: does the query actually instantiate each one, and does the response actually satisfy it?
3 · Response reviewAnswer quality, persona and style consistency, and "naturalness as spoken dialogue"

On top of these, "deterministic checks handle defects such as empty output, malformed tokens, severe repetition, and role confusion." Role confusion is when a self-play model forgets which side it is playing, a common failure of two-model generation.

The retention rule is strict: "Only examples that pass these checks and receive high scores in all three quality-control stages are retained for supervised fine-tuning; all other examples are rejected." Compare that with the audio data in Chapter 2, where near-misses were routed to midtraining or relabelling. For dialogue, a near-miss is simply discarded. Chapter 2's "less is more" result is a good reason to believe that strictness pays.

Reasoning mode: how good is the thinking model?

Before adding any realtime machinery, the team evaluated StepAudio 3 "in reasoning mode", with explicit thinking always available and no speaking deadline. "This evaluation isolates the model's conversational and reasoning capability." The realtime system, with Think-While-Speaking, Adaptive Thinking and MTP, is evaluated separately (Chapter 6 and Chapter 9).

DimensionStepAudio 3 (Reasoning)Doubao 2.0 LiteDeepSeek-V4-FlashKimi K3
Instruction Following66.372.971.468.9
Faithfulness72.467.575.378.4
Reasoning73.072.764.881.9
Memory72.071.371.577.6
Knowledge73.159.971.678.6
Safety & Reliability79.075.979.984.8
Conversational Pragmatics67.261.562.970.3
Persona & Role Consistency80.982.673.676.5
Macro Average73.070.571.477.1
Worked example 5 · recomputing the 73.0 macro average (Table 4 numbers) A macro average weights each dimension equally, regardless of how many items it has.
Step 1. 66.3 + 72.4 = 138.7
Step 2. 138.7 + 73.0 = 211.7
Step 3. 211.7 + 72.0 = 283.7
Step 4. 283.7 + 73.1 = 356.8
Step 5. 356.8 + 79.0 = 435.8
Step 6. 435.8 + 67.2 = 503.0
Step 7. 503.0 + 80.9 = 583.9
Step 8. 583.9 ÷ 8 = 72.9875 ≈ 73.0, matching the paper.
Step 9. Gap to Kimi K3: 77.1 − 73.0 = 4.1 points. Lead over DeepSeek-V4-Flash: 73.0 − 71.4 = 1.6. Lead over Doubao 2.0 Lite: 73.0 − 70.5 = 2.5.

The paper's reading: StepAudio 3 "ranks second on reasoning, memory, knowledge, conversational pragmatics, and persona and role consistency." Kimi K3 "leads six of the eight dimensions," while Doubao 2.0 Lite "leads instruction following and persona and role consistency." The weakest cell for StepAudio 3 is instruction following at 66.3, the lowest of the four systems. The paper's lesson: "strong aggregate reasoning does not imply uniformly stronger instruction following or role consistency."

A connection the tables make for you. The dialogue column of this table is identical, cell for cell, to the "Merged" dialogue column of Table 8 (66.3, 72.4, 73.0, 72.0, 73.1, 79.0, 67.2, 80.9; macro 73.0). The reasoning-mode model evaluated here is the merged model built in the next chapter, and the paper notes that Table 8's dialogue results "are measured with the model in reasoning mode."
Chapter 4 recap. (1) StepAudioChat is closed (contamination-resistant) and text-based (separates content from delivery), with eight dimensions. (2) Each item has a prompt, checkable criteria, a valid reference and a deliberately flawed response; the last two are positive and negative controls for the judge. (3) Difficulty is calibrated with two reference systems of different strength; the judge is calibrated with the controls and a trusted labelled sample. (4) Dialogue data: discussion points → case plan → self-play N−1 turns → final query → one labelled answer (the only target); four factorized axes with stratified rotation; three independent review stages, all must pass. (5) Reasoning mode scores 73.0 macro: above Doubao 2.0 Lite (70.5) and DeepSeek-V4-Flash (71.4), below Kimi K3 (77.1).
Every StepAudioChat item includes a valid reference response and a deliberately flawed response. What role do these two responses play?

Chapter 5: Merging the Teachers

You fine-tune your model on a big batch of conversation data. Dialogue scores go up. Then you check math, and math has gone down. You add more math data and retrain; now audio understanding slips. Every capability you push on seems to pull another one back.

The textbook fix is to train one model on the union of every dataset, with carefully tuned proportions. But every time one team improves its data, the whole mixture has to be rebalanced and the whole model retrained. With a model of roughly 196 billion parameters, that is an expensive habit.

Section 6.4 of the paper takes a different path, and it is disarmingly simple: train several specialists separately from the same starting point, then average their weights.

Specialist teachers from one base

Picture four apprentice chefs trained in the same kitchen by the same master. One spends a year on sauces, one on pastry, one on grilling, one on a bit of everything. Because they learned in the same kitchen, their stations are laid out identically: the same knife is in the same drawer for all four. You could reasonably blend their notebooks page by page, because page 40 means the same thing in each.

The paper's version: "We train multiple compatible teacher checkpoints from a common base model, using a different data composition for each teacher." The mixtures "emphasize complementary capabilities, including multi-turn dialogue, audio understanding, general text reasoning and knowledge, and targeted mixed-domain behavior." The result: "teachers that are individually strong in different regions of the capability space while preserving parameter alignment for merging."

Two terms need unpacking. A checkpoint is a saved copy of all of a model's weights at some point in training. Parameter alignment means that the same position in each weight tensor plays the same role in every teacher, like the same knife in the same drawer.

Why does a common base give alignment? This is background knowledge rather than a claim in the report. Neural networks have permutation symmetry: you can shuffle the order of neurons inside a layer (and shuffle the connected weights to match) without changing what the network computes. Two networks trained from different random starts may end up with the "same" features in different positions, so averaging them position by position mixes unrelated features. Fine-tuning from a shared base changes weights only modestly, so each neuron keeps its role, and position-wise averaging stays meaningful.

The merge, as an equation

How do you combine four sets of weights into one? The paper uses "directly averaging their parameters" with weights:

θmerge = ∑i αi θi,    αi ≥ 0,   ∑i αi = 1

Here θi is the full parameter vector of teacher i (every weight in the network, laid end to end), αi is how much teacher i contributes, and θmerge is the resulting model. The two constraints make this a convex combination: a weighted average with non-negative weights that sum to one, so the merged point lies "between" the teachers and never outside them. The operation is applied tensor by tensor, and within each tensor, element by element.

A useful rewriting shows what merging really combines. Write each teacher as the base plus a change: θi = θbase + Δi, where Δi is what fine-tuning taught teacher i. Substitute, one step at a time:

Derivation · merging averages the fine-tuning updates Step 1. θmerge = ∑i αi θi
Step 2. = ∑i αi (θbase + Δi)
Step 3. = ∑i αi θbase + ∑i αi Δi   (distribute the sum)
Step 4. = θbase (∑i αi) + ∑i αi Δi   (θbase does not depend on i)
Step 5. = θbase · 1 + ∑i αi Δi   (the weights sum to 1)
Step 6. = θbase + ∑i αi Δi

So the merged model is the shared base plus a weighted blend of what each specialist learned. If the specialists' updates mostly touch different directions in weight space, the blend keeps much of each. If they conflict, the blend dilutes them. This reading is standard in the model-merging literature; the report itself states only the first equation.

The reported recipe: four teachers, 3:1:1:1

"The reported model uses four teachers with a normalized 3:1:1:1 weighting." The coefficients "are selected against held-out evaluations spanning dialogue, audio understanding, and general text capabilities." The report does not say which of the four teachers receives the weight of 3, so we will not assume it.

Worked example 6a · normalizing 3:1:1:1 and merging one weight (toy weight values) Step 1. Sum of the ratio parts: 3 + 1 + 1 + 1 = 6.
Step 2. Normalized weights: 3 ÷ 6 = 0.5; 1 ÷ 6 = 0.1667 (three times).
Step 3. Check: 0.5 + 0.1667 + 0.1667 + 0.1667 = 0.5 + 0.5 = 1.0. Both constraints hold (all ≥ 0, sum = 1).
Step 4. Take one scalar weight at the same position in four teachers, with illustrative values: heavy teacher 0.80; others 0.20, 0.40, 0.60.
Step 5. Heavy contribution: 0.5 × 0.80 = 0.40.
Step 6. Light contributions: (0.20 + 0.40 + 0.60) × (1/6) = 1.20 ÷ 6 = 0.20.
Step 7. Merged weight = 0.40 + 0.20 = 0.60. Repeat for every one of the ~196 billion positions.

One more property the paper emphasizes: "Because integration occurs in parameter space, it introduces neither additional model components nor inference-time routing." Contrast this with an ensemble, which runs several models and combines their outputs (paying for all of them at inference), or with routing a request to the right specialist (paying for a router and keeping all specialists deployed). A merged model costs exactly one model at inference.

pytorchdef merge_teachers(state_dicts, ratios):
    """theta_merge = sum_i alpha_i * theta_i   (paper, section 6.4)
    state_dicts: teacher checkpoints fine-tuned from ONE common base.
    ratios: e.g. [3, 1, 1, 1] -> normalized to alpha = [0.5, 1/6, 1/6, 1/6]."""
    assert all(r >= 0 for r in ratios)
    alphas = [r / sum(ratios) for r in ratios]         # sum to 1
    keys = state_dicts[0].keys()
    assert all(sd.keys() == keys for sd in state_dicts)  # same architecture
    merged = {}
    for k in keys:                                    # tensor by tensor
        merged[k] = sum(a * sd[k].float() for a, sd in zip(alphas, state_dicts))
    return merged                                       # one model: no router, no extra modules

# Coefficient search (the paper selects alphas on held-out dialogue, audio and text evals):
# for ratios in candidate_grid: score(merge_teachers(teachers, ratios)) -> keep best balance

Sim 5: walk the merged point through weight space

The picture below is a two-dimensional cartoon of weight space (real weight space has about 196 billion dimensions). The shaded region is a toy low-loss basin around the shared base. Four teachers sit in different directions from the base, each pushed toward its own capability. The sliders set the merge ratios; the merged point is their convex combination. Flip to "misaligned teachers" to see what happens when the teachers do not share a base.

Sim 5 — convex merging, aligned and misaligned

The landscape, teacher positions and loss readout are illustrative; only the 3:1:1:1 preset comes from the paper. In the misaligned view, two teachers came from a different starting point, so their neurons sit in permuted positions: the same skills live in a mirrored basin, and averaging across basins lands on a ridge.

Teacher A (dialogue-leaning)
Teacher B (audio-leaning)
Teacher C (text-leaning)
Teacher D (mixed)

Three experiments to run in the simulation. Press "Only teacher A": the merge sits exactly on teacher A, because α = (1, 0, 0, 0) is still a valid convex combination. Press "Equal 1:1:1:1": the merge moves to the centroid of the four teachers. Now switch to the misaligned view and press "Equal" again: the centroid of two basins is the ridge between them, and the toy loss jumps above every teacher's.

The two-teacher case makes the geometry exact. With only teachers 1 and 2 and weights (1 − λ, λ), the merge is θ(λ) = (1 − λ)θ1 + λθ2, the straight segment between them as λ runs from 0 to 1. Whether a merge works is therefore the question of whether the loss stays low along that segment, which is what a shared base makes likely and separate starts make unlikely.

The teacher labels in the simulation (dialogue-leaning and so on) are ours; the paper lists the capability emphases but not which numbered teacher got which mixture, nor which got the weight of 3. What the cartoon should leave you with: convex weights can only move the merged point inside the teachers' hull, and that hull is only a safe place to be if the teachers live in the same basin.

The evidence: Table 8

The paper compares the four teachers with their merge across three domains. Dialogue is "measured with the model in reasoning mode."

BenchmarkMergedT1T2T3T4
BigBench Audio98.196.198.198.498.5
AudioMultiChallenge49.349.149.850.247.6
MMSU90.685.585.590.490.9
MMAU79.078.578.877.877.7
WildSpeech77.176.576.475.976.1
MMAR86.585.486.487.386.5
Step-Caption78.275.878.379.679.4
MTalk-Bench91.791.890.390.990.7
Audio macro81.379.880.581.380.2
HMMT 2026 Feb86.844.082.281.379.8
GPQA Diamond83.073.181.980.180.1
MultiChallenge59.754.651.350.659.7
General text macro76.557.271.870.773.2
Instruction Following66.364.264.569.364.2
Faithfulness72.474.065.667.173.6
Reasoning73.075.267.067.267.8
Memory72.073.470.068.371.5
Knowledge73.175.163.762.075.6
Safety & Reliability79.079.275.775.678.0
Conversational Pragmatics67.267.862.463.464.5
Persona & Role80.984.779.576.983.0
Dialogue macro73.074.268.668.772.3

The paper's summary, domain by domain: on audio, the merge "reaches a macro average of 81.3, tying the best teacher macro average while leading on MMAU and WildSpeech"; on general text, it "achieves the highest macro average of 76.5, with the best HMMT 2026 Feb and GPQA Diamond scores and a tie on MultiChallenge"; on dialogue, it "has a macro average of 73.0, exceeding Teachers 2, 3, and 4, while remaining below the strongest dialogue teacher at 74.2."

Read the teacher profiles, which the table reveals even though the paper does not name each teacher's mixture. Teacher 1 is the best conversationalist (dialogue macro 74.2) and by far the weakest mathematician (HMMT 44.0). Teacher 3 ties the best audio macro (81.3). Teacher 4 is the strongest text teacher (73.2) and nearly the best at dialogue (72.3). No teacher is best at everything, which is exactly the premise of merging.

Worked example 6b · parameter averaging is not score averaging (Table 8 numbers) Suppose merging simply blended scores with the same 3:1:1:1 weights. The most favorable case gives the weight of 0.5 to the best teacher on the row.
HMMT 2026 Feb. Best teacher: T2 = 82.2. Others: 44.0, 81.3, 79.8.
Step 1. Heavy part: 0.5 × 82.2 = 41.10.
Step 2. Light part: (44.0 + 81.3 + 79.8) ÷ 6 = 205.1 ÷ 6 = 34.18.
Step 3. Best possible score blend: 41.10 + 34.18 = 75.28.
Step 4. Actual merged score: 86.8, which is 86.8 − 75.28 = 11.52 above the best possible blend, and 86.8 − 82.2 = 4.6 above the best single teacher.
GPQA Diamond. Best teacher: T2 = 81.9. Others: 73.1, 80.1, 80.1.
Step 5. 0.5 × 81.9 = 40.95; (73.1 + 80.1 + 80.1) ÷ 6 = 233.3 ÷ 6 = 38.88; blend = 79.83.
Step 6. Actual merged: 83.0, which is 1.1 above the best teacher.
General text macro check: (86.8 + 83.0 + 59.7) ÷ 3 = 229.5 ÷ 3 = 76.5; T1: (44.0 + 73.1 + 54.6) ÷ 3 = 171.7 ÷ 3 = 57.23 ≈ 57.2. Both match the table.

The punchline: the merged model is not a mixture of teacher behaviors in any simple sense. On math and science it beats every teacher it was built from. A plausible explanation (ours; the paper does not analyze it) is that the averaged fine-tuning updates act partly as a regularizer, cancelling idiosyncratic noise in each specialist while keeping directions they share.

What "balanced" means, counted row by row

The paper calls the merge "a balanced trade-off." We can make that word precise with Table 8 by ranking the merged model against the four teachers on every row. Rank 1 means the merge is best; rank 5 means it is worst.

Worked example 6c · the merge's rank among five models on each dialogue row (Table 8) Instruction Following 66.3 vs 64.2, 64.5, 69.3, 64.2: beats three teachers → rank 2.
Faithfulness 72.4 vs 74.0, 65.6, 67.1, 73.6: beats two → rank 3.
Reasoning 73.0 vs 75.2, 67.0, 67.2, 67.8: beats three → rank 2.
Memory 72.0 vs 73.4, 70.0, 68.3, 71.5: beats three → rank 2.
Knowledge 73.1 vs 75.1, 63.7, 62.0, 75.6: beats two → rank 3.
Safety & Reliability 79.0 vs 79.2, 75.7, 75.6, 78.0: beats three → rank 2.
Conversational Pragmatics 67.2 vs 67.8, 62.4, 63.4, 64.5: beats three → rank 2.
Persona & Role 80.9 vs 84.7, 79.5, 76.9, 83.0: beats two → rank 3.
Tally: rank 2 on 5 rows, rank 3 on 3 rows, never first and never worse than third.
Compare Teacher 1, the dialogue specialist: best on 6 of 8 dialogue rows, but its text macro is 57.2, the worst of all five models, 76.5 − 57.2 = 19.3 points below the merge.

The same exercise on the eight audio rows gives a wider spread: rank 1 on MMAU and WildSpeech, rank 2 on MMSU and MTalk-Bench, tied for second on MMAR (86.5, level with Teacher 4), tied for third on BigBench Audio (98.1, level with Teacher 2), rank 3 on AudioMultiChallenge, and rank 4 on Step-Caption (78.2, ahead only of Teacher 1's 75.8). On the three text rows it is first on two and tied first on the third.

That is what "balanced" looks like in numbers: a model that is rarely the single best specialist, almost never near the bottom, and never pays the kind of 19-point penalty a specialist pays outside its domain. For a product that must do all three jobs in one conversation, the worst-case row matters more than the best-case row.

Three ways to combine specialists

Merging is one of several ways to get several specialists' skills into one product. The comparison below is general engineering background, with the merge column filled from the paper.

StrategyTraining cost when one data mix changesInference costWhat the paper says about it
Train one model on the union of all dataRetrain the whole modelOne modelMerging "allows teacher data mixtures to be developed independently and then recombined without retraining a single model on the full union of data."
Keep specialists and route or ensembleRetrain one specialistSeveral models plus a router, or several forward passesMerging "introduces neither additional model components nor inference-time routing."
Weighted parameter merge (this paper)Retrain one teacher, then re-averageOne modelCoefficients selected on held-out dialogue, audio and text evaluations; four teachers at 3:1:1:1

For a 196B-total-parameter realtime model that already runs two concurrent calls per deliberate turn (Chapter 6), the inference column is decisive: a router or an ensemble would multiply the serving cost of every turn.

The honest side of the ledger

Merging is not free. Scan the rows where the merge falls below the best teacher: Persona & Role 80.9 versus 84.7 (T1), a 3.8-point gap; Instruction Following 66.3 versus 69.3 (T3), a 3.0-point gap; Knowledge 73.1 versus 75.6 (T4); Reasoning (dialogue) 73.0 versus 75.2 (T1); Step-Caption 78.2 versus 79.6 (T3).

The paper says this directly: the results "support model merging as a low-cost mechanism for combining complementary capabilities, while showing that it provides a balanced trade-off rather than uniform improvement over every specialized teacher." And its stated design goal is modest on purpose: merging "is intended to retain complementary strengths rather than make the merged model identical to the best teacher on every metric."

The practical win is organizational as much as numerical. "This also allows teacher data mixtures to be developed independently and then recombined without retraining a single model on the full union of data." A dialogue team and an audio team can iterate on their own teachers, and the merge step recombines their progress.

When merging goes wrong

The report shows a merge that works. It is still worth knowing the ways a weighted average can fail, because each one explains a design choice the paper made. The failure modes below are general background from the model-merging literature, not experiments in this report.

Failure modeWhat happensThe paper's matching choice
Teachers from different starting pointsPermutation mismatch: averaging mixes unrelated neurons, and the merge lands off the low-loss region (Sim 5, misaligned view)All teachers are trained "from a common base model," "preserving parameter alignment"
Conflicting updatesTwo teachers push the same weights in opposite directions; averaging cancels both skillsTeacher mixes "emphasize complementary capabilities," so updates overlap less
Badly chosen weightsOne skill dominates and another fadesCoefficients "selected against held-out evaluations spanning dialogue, audio understanding, and general text"
Hidden regressionsA merged model looks fine on averages and fails a specific skillThe paper reports per-benchmark rows and says the merge is a trade-off, not a uniform win

The last row is the one to carry into your own work. A macro average of 73.0 hides a persona score 3.8 points below the best teacher. If persona consistency matters most for your product, the 3:1:1:1 weights might not be your weights. The paper's framing, "balanced trade-off," is exactly the right warning label.

Keep two evaluations apart. The paper stresses that Table 8 measures "the underlying reasoning model and [is] separate from the system-level evaluations that additionally use the realtime reasoning mechanisms." The merged model's 73.0 dialogue macro is a reasoning-mode number. The realtime system, which additionally applies Think-While-Speaking, Adaptive Thinking and MTP (Chapters 6 and 7), scores 70.4 on the same benchmark. The report does not say that the two rows share identical weights (its Adaptive Thinking ablation, for instance, uses a separately trained model). Chapter 6 takes that gap apart.
Inline check: which constraint does what? What would go wrong if one αi were negative, or if the α values summed to 1.5? Answer: a negative weight would push the model away from a teacher, outside the teachers' hull, into territory no teacher was trained on; weights summing to 1.5 would scale the base itself by 1.5 (Step 4 of the derivation), distorting every layer's magnitude. The two constraints keep the merge an interpolation.
Chapter 5 recap. (1) Several teachers are fine-tuned from one base with different data mixes (dialogue, audio, text reasoning and knowledge, mixed-domain), preserving parameter alignment. (2) θmerge = ∑ αiθi with αi ≥ 0 and ∑αi = 1, equivalently the base plus a weighted blend of fine-tuning updates. (3) The reported model uses four teachers at 3:1:1:1 (0.5, 1/6, 1/6, 1/6), chosen on held-out dialogue, audio and text evaluations; no extra modules, no routing. (4) The merge ties the best audio macro (81.3), leads general text (76.5, with HMMT 86.8 above every teacher) and sits second on dialogue (73.0 vs 74.2). (5) It is a balanced trade-off, not a uniform win.
Which statement about the paper's model merging is correct?

Chapter 6: Think While Speaking

Back to the question from Chapter 0: departure time, flight time, a time-zone shift, a buffer at arrivals, and a dinner reservation. The merged model from Chapter 5 can reason its way to a good answer. In reasoning mode, it scores 73.0 on StepAudioChat. But reasoning mode, used naively in a voice product, means the user waits in silence while the thinking happens.

This chapter is about the mechanism that removes the wait: Think-While-Speaking. The paper describes it in Section 6.3, and it is the heart of the report's claim to resolve "the tension between deep deliberation and latency."

Two brains, one model

Picture a lecturer on stage with a research assistant in the wings. A student asks a hard question. The lecturer does not freeze; they start talking: restating the question, laying out what matters. Meanwhile the assistant works the problem on paper and slides notes onto the lectern, one finding at a time. The lecturer speaks from whatever notes have arrived so far. When the assistant finishes, the lecturer wraps up with the full picture, and if an early remark turned out to be off, corrects it at the end.

The paper's design follows that shape. "It builds on the two-process design of Mind-Paced Speaking. Two concurrent calls to the same audio model act as a Formulation Brain and an Articulation Brain."

The Formulation Brain is the assistant in the wings: it "generates a private reasoning trace." Private means the trace is never spoken; the user never hears it.

The Articulation Brain is the lecturer: it "produces short response segments conditioned on the reasoning available so far and on the response already spoken." Those two conditioning sources are the whole trick. "The reasoning available so far" keeps what is said consistent with what has been worked out. "The response already spoken" keeps the answer coherent as it grows segment by segment.

"Same model" is a deliberate choice. Both brains are calls to one set of weights (Figure 7 draws a vertical arrow between them labelled "Same Model"). There is no separate small speaking model and no separate large thinking model. The two calls differ only in what they read and what they write. That keeps the thinking and the speaking in the same "voice" of knowledge, and it means improvements from merging (Chapter 5) help both.

What each brain reads and writes

Figure 7(B) of the paper, "Think-While-Speaking with MTP", draws both brains at a single "time step i" with a colour legend: input tokens, think tokens, and response tokens. Here is the figure's wiring as a table.

BrainReads (per Figure 7)WritesWho hears it
FormulationUser context (input tokens from the audio encoder and adapter) + its own thought stream so farThe "current think" tokensNobody (private)
ArticulationUser context + "previous think" + "current think" + "previous response" + "current response"Response tokensThe user, as streaming output audio

An orange arrow in the figure carries the Formulation Brain's current think tokens into the Articulation Brain's input. A green arrow carries the Articulation Brain's response tokens down to "Streaming Output Audio". A dashed box beside both decoders, labelled "MTP Decoding Acceleration", connects to each of them (Chapter 7). And the input side shows the same audio path as Chapter 1: input audio → audio encoder → audio adapter → input tokens.

In data-flow terms, one step of the system looks like this:

Shared input
user context C (audio-derived input tokens + text history)
↓
Formulation call
think1..j = decode(C, think1..j−1)   → appended to the private trace
↓ current think tokens flow sideways
Articulation call
segmentk = decode(C, think1..j, response1..k−1)
↓
Generator
segmentk → streaming audio → back into the model audio stream

Playback-aware scheduling

If the Articulation Brain can speak at any time, when should it produce the next segment? The paper's answer: "Playback-aware scheduling releases response segments according to the progress of the streaming output audio while formulation continues in parallel."

In other words, the speaker's own audio playback is the clock. Our reading of that sentence: a new segment is produced as the audio already queued runs down, not all at once at the start. Why is that the right clock? Here is our reasoning, based on what the paper describes.

First, every second of delay before generating a segment is a second more of reasoning it can condition on. Generating the whole answer up front would throw away all the thinking that happens during playback.

Second, speech that has been generated but not yet played is a liability. If the user interrupts (Chapter 3) or new reasoning changes the answer, un-played text is either wasted or wrong. Keeping the buffer short keeps the answer revisable.

Third, the user experiences audio, not tokens. Tying generation to playback keeps the audible stream continuous without racing ahead of the reasoning.

Finishing, correcting, and when to start

Three more rules complete the mechanism.

After formulation ends. "Once formulation finishes, the remaining response can use the complete reasoning state." Segments released after that point are as well-informed as a think-then-speak answer.

Final continuation. "A final continuation can supplement or correct an answer that began from incomplete reasoning." If an early segment was framed on partial reasoning and the full trace changes the picture, the model says so at the end.

Speak-First vs Think-First. "The system uses Speak-First by default, starting the response without waiting for an initial reasoning prefix. Think-First waits for a short reasoning prefix before beginning the response." Speak-First minimizes silence. Think-First trades a short pause for an opening sentence that is already informed by some reasoning. The report does not quantify the prefix length.

Where Adaptive Thinking fits. Figure 7(A) sits in front of all this. Each user turn first meets a question: "Does this turn need explicit reasoning?" No takes the "direct route" to an "immediate response" with an "empty think block." Yes takes the "deliberate route" into Think-While-Speaking with MTP. So the two-brain machinery runs only on turns that were judged to need it. Chapter 7 explains how that judgment is trained.

Sim 6: the reasoning race

The simulation lays out one deliberate turn on a timeline. The top lane is the Formulation Brain; three findings (F1, F2, F3) appear as the trace grows, and F3 is the final answer. The middle lane shows each response segment at the moment it is generated. The bottom lane is what the user hears. To make the effect of timing visible, the simulation uses an illustrative policy of our own, not a rule from the paper: segments 3, 4 and 5 wait for a finding before they are stated, and segments 1, 2 and 6 are framing that needs none. The paper itself says only that segments are "conditioned on the reasoning available so far" and that a final continuation "can supplement or correct an answer that began from incomplete reasoning."

Sim 6 — Formulation vs Articulation, with playback-aware scheduling

All durations, the six-segment answer and the finding positions are illustrative. The speed buttons apply the paper's measured wall-clock ratios from Table 6 (1.76× for MTP3 strict, 2.05× for MTP3 with Medusa-style acceptance) to the thinking lane only. With "eager" scheduling, all segments are generated at time zero from no reasoning, so substantive claims are unsupported (red) and a final correction is needed.

Reasoning length

Three experiments worth running. Stretch the reasoning length and watch the simulated policy insert more framing segments before the substantive ones, while the first audio still starts at time zero. Switch to Think-First and a short silence appears, but the opening is better informed. Switch to eager scheduling and the same audio timeline now contains unsupported claims, followed by a correction once F3 arrives. The faster thinking lanes shorten the framing phase: the paper's two mechanisms, concurrency and acceleration, compound.

A worked trace of one deliberate turn

To make the two brains concrete, here is the dinner question from Chapter 0 traced step by step. Everything in this table is illustrative: the wording, the timing and the reasoning content are ours. The paper's stated mechanisms (Speak-First, conditioning on the reasoning so far, the complete reasoning after formulation ends) are labelled as such; the rest of the "rule" column is an illustrative scheduling policy, not something the report specifies.

Time (illustrative)Formulation Brain (private)Articulation Brain (spoken)Rule at work
0.0 sstarts: "departure 18:00, flight 2 h 15 min…""Let's check the timing for your dinner."Paper: Speak-First starts without a reasoning prefix
1.2 s"…lands 20:15 origin time; +1 h zone shift → 21:15 local…""Your flight leaves at six and takes a bit over two hours."Illustrative policy: restate facts already in the context
2.4 s"…+40 min at arrivals → 21:55…""With the time difference, you land at quarter past nine local time."Illustrative policy: state F1 after the trace contains it
3.6 s"…21:55 > 21:30, so no" (trace complete)"After arrivals, you'd be out around five to ten."Paper: conditioned on the reasoning so far (here, F2)
4.8 s(finished)"So a 9:30 table is too tight; ask for 10:15 instead."Paper: after formulation ends, the rest uses the complete reasoning

Two details of the trace are worth pointing out. The Articulation Brain's second line is "safe": it contains only facts the user supplied, so it can be spoken before any reasoning exists. And in this illustrative run, the reasoning happened to arrive in time for every substantive line. The paper does not promise that: a segment can be spoken from incomplete reasoning, and the final continuation is the mechanism that repairs it (Section 6.3.2 warns that strict verification of the spoken output "does not eliminate errors arising from incomplete private reasoning").

Speak-First or Think-First?

The paper makes Speak-First the default and offers Think-First as an alternative, without reporting a comparison. The trade-off follows directly from the definitions:

PropertySpeak-First (default)Think-First
Silence before the first wordNone from reasoningThe length of a "short reasoning prefix"
First sentenceFraming, conditioned on little or no reasoningAlready informed by the start of the trace
Risk of an early misstepHigher; the final continuation is the safety netLower for the opening
Where it fits (our reading)Casual conversation, where dead air is the worst outcomeTurns where the opening sentence must already commit to a direction

One more interaction deserves mention, because it links this chapter to Chapter 3. If the user interrupts while a deliberate answer is playing, the model audio stream tells the decoder how much of the answer was actually heard. Combined with playback-aware scheduling, which (on our reading) keeps the unplayed buffer short, that means little generated-but-unheard speech has to be thrown away. This is our inference from the two mechanisms together; the paper does not describe interruptions during Think-While-Speaking explicitly.

The mechanism in code

A concurrency sketch in Python's asyncio. The two coroutines share one model and one context object; the playback buffer is the scheduler's clock. Function names are ours.

python (asyncio sketch)async def formulation(model, ctx):
    while not ctx.think_done:
        toks = await model.decode_think(ctx.user, ctx.think)   # MTP-accelerated, private
        ctx.think += toks
        ctx.think_done = ends_think(toks)

async def articulation(model, ctx, player, speak_first=True):
    if not speak_first:
        await ctx.wait_for_think_prefix()                  # Think-First
    while not ctx.response_done:
        await player.wait_until_buffer_low()                # playback-aware release (our reading)
        seg = await model.decode_segment(ctx.user,
                                         think=list(ctx.think),   # reasoning available SO FAR
                                         spoken=ctx.spoken)        # response already spoken
        player.enqueue(seg)                                    # strict verification for speech
        ctx.spoken += seg
        ctx.response_done = ctx.think_done and answer_complete(seg)
    if needs_revision(ctx.spoken, ctx.think):                 # final continuation
        player.enqueue(await model.decode_segment(ctx.user, ctx.think, ctx.spoken))

async def deliberate_turn(model, ctx, player):
    await asyncio.gather(formulation(model, ctx),
                         articulation(model, ctx, player))

Two details in the sketch come straight from the paper. The think stream is decoded with MTP and permissive "typical acceptance", while spoken segments keep "strict verification" (Section 6.3.2, next chapter). And the final continuation exists because "a final continuation can supplement or correct an answer that began from incomplete reasoning." How the real system decides that a revision is needed is not described, so the needs_revision check above is a placeholder.

What it costs: realtime vs reasoning mode

Now the measurement. The overall evaluation (Table 10) scores StepAudio 3 Realtime on StepAudioChat in "realtime mode", with Think-While-Speaking, Adaptive Thinking and MTP all active. The baselines are in reasoning mode. Set that row beside the reasoning-mode row from Chapter 4.

DimensionReasoning mode (Table 4)Realtime mode (Table 10)Change
Instruction Following66.354.1−12.2
Faithfulness72.471.9−0.5
Reasoning73.073.6+0.6
Memory72.071.8−0.2
Knowledge73.170.4−2.7
Safety & Reliability79.075.1−3.9
Conversational Pragmatics67.268.2+1.0
Persona & Role Consistency80.978.3−2.6
Macro Average73.070.4−2.6
Worked example 7 · where the realtime points go (Tables 4 and 10) Step 1. Realtime sum: 54.1 + 71.9 = 126.0; + 73.6 = 199.6; + 71.8 = 271.4; + 70.4 = 341.8; + 75.1 = 416.9; + 68.2 = 485.1; + 78.3 = 563.4.
Step 2. Realtime macro: 563.4 ÷ 8 = 70.425 (the paper's text reports 70.41, from unrounded values; Table 10 shows 70.4).
Step 3. Reasoning sum (Chapter 4): 583.9. Difference in sums: 583.9 − 563.4 = 20.5.
Step 4. Check with the per-row changes: −12.2 − 0.5 + 0.6 − 0.2 − 2.7 − 3.9 + 1.0 − 2.6. Negatives: 12.2 + 0.5 + 0.2 + 2.7 + 3.9 + 2.6 = 22.1. Positives: 0.6 + 1.0 = 1.6. Net: −22.1 + 1.6 = −20.5. Consistent.
Step 5. Macro drop: 20.5 ÷ 8 = 2.5625 points.
Step 6. Share of the total drop from Instruction Following alone: 12.2 ÷ 20.5 = 59.5%.
Step 7. Without the Instruction Following row, the other seven rows lose 20.5 − 12.2 = 8.3 points in total, an average of 8.3 ÷ 7 ≈ 1.19 points each.

The picture is lopsided. Reasoning itself holds up (73.0 to 73.6), and conversational pragmatics even improves (67.2 to 68.2). Most of the realtime cost concentrates in one dimension: instruction following falls from 66.3 to 54.1.

What causes that? The paper does not say, and it explicitly does not attribute scores to components: "component-ablation scores are not used to fill these entries." A hypothesis consistent with the mechanism (ours): constraints like length, format and "do not mention X" must shape the answer from its first words, and in Speak-First the first segments are generated before the reasoning has fully processed those constraints. The paper's conclusion does name "multi-turn constraint handling" as a remaining area for improvement.

The paper's positive reading is also justified by the numbers: the realtime macro of 70.4 is "comparable to Doubao 2.0 Lite at 70.5 and DeepSeek-V4-Flash at 71.4", which are measured in reasoning mode, with no speaking deadline. The claim is that the model "can deliberate while producing speech and retain dialogue and reasoning performance comparable to dedicated reasoning models, rather than providing only low-latency surface responses."

It helps to list where each point of the realtime cost could come from, as candidate hypotheses to test rather than conclusions. The report combines all three mechanisms in its realtime row, so none of these can be settled from its tables.

Candidate source of the realtime costMechanism involvedEvidence available in the paper
Early segments committed before constraints are processedThink-While-Speaking (Speak-First)None isolated; consistent with the Instruction Following drop
Reasoning skipped on turns that needed itAdaptive ThinkingTable 5: Reasoning 71.89 → 66.80 for the separately trained Adaptive Thinking model, yet realtime Reasoning (73.6) did not fall, so this source does not obviously dominate
Distribution shift in the private traceMTP typical acceptanceTable 6: Instruction Following falls under every MTP setting (64.15 → 61.21 to 62.73)
The comparison is tilted, and in which direction. Table 10 pits a realtime system against baselines allowed to think first. That handicaps StepAudio 3, not the baselines, so "comparable" is a meaningful result. What the table cannot tell you is how the baselines would score while speaking: the dialogue baselines are reasoning models that the report does not evaluate under any realtime or full-duplex condition (its full-duplex comparison in Chapter 3 uses different systems), so the table cannot say how they would fare with a speaking deadline.
The failure mode the paper admits. "Keeping strict verification for spoken output does not eliminate errors arising from incomplete private reasoning." Strict verification guarantees that the spoken tokens follow the target model's own output distribution, as if no drafting had happened. It does not guarantee that what the model would have produced, given only part of its reasoning, is right. That is what the final continuation is for, and it is why the instruction-following drop deserves attention.
Inline check: what does the Articulation Brain condition on? Name the two sources the paper lists, and say what goes wrong if you drop each one. Answer: "the reasoning available so far" (drop it and speech is disconnected from the thinking, back to "speak now") and "the response already spoken" (drop it and segments repeat or contradict each other, because each would be written as if it were the first).
Chapter 6 recap. (1) Two concurrent calls to one model: the Formulation Brain writes a private reasoning trace; the Articulation Brain writes short response segments from the reasoning so far and the response already spoken. (2) Playback-aware scheduling releases segments as output audio plays, so later segments see more reasoning and the unplayed buffer stays small. (3) After formulation ends, segments use the full reasoning; a final continuation can supplement or correct. Speak-First is the default; Think-First waits for a short prefix. (4) Realtime StepAudioChat: 70.4 vs 73.0 in reasoning mode; 59.5% of the drop is instruction following (66.3 to 54.1); reasoning holds (73.6). (5) Strict verification of speech does not protect against errors from incomplete reasoning.
In Think-While-Speaking, what determines when the Articulation Brain releases its next response segment?

Chapter 7: Adaptive Thinking & MTP

"What time is it in Tokyo?" "Thanks, that's all." "Actually, make it four people." Most turns in a real conversation are like these: short, routine, and answerable at once. Running a private reasoning trace on them wastes compute and, in a system with a silence budget, wastes time.

Other turns are the dinner-timing puzzle from Chapter 0, where skipping the reasoning produces a confident wrong answer.

The paper's motivation for this chapter's first mechanism is one sentence: "Different conversational turns warrant different deliberation budgets." And its second mechanism answers the follow-up question: when a turn does need thinking, how do you make that thinking cheaper? Two tools: Adaptive Thinking decides whether to think, and multi-token prediction (MTP) makes the thinking that remains faster.

Part A. Adaptive Thinking: learning when reasoning helps

Picture a teacher grading homework who wants to know which problems students should show their work on. For each problem, she compares two attempts: one with worked steps, one written straight down. If both answers are equally good, the steps did not matter for that problem. If the worked version is clearly better, the steps are worth requiring.

Section 6.3.1 builds exactly that comparison, at the level of individual assistant turns. "For each turn, we collect the dialogue context, target answer, original reasoning trace, and task-type features."

1 · Original trajectory
The training turn as it is: context → reasoning trace → answer.
↓
2 · No-think probe
"A fixed probe model, trained without the adaptive-thinking data transformation, receives an empty think block and regenerates the answer under a no-think condition."
↓
3 · Blind paired judgment
"A blind judge scores the original and no-think trajectories against the target answer, without knowing which is which, and labels the turn by whether reasoning changed the answer's quality."
↓
4 · Supporting signals
"Consistency across repeated judgments", "the coherence between a reasoning trace and the answer it produces", and "a task-type prior" inform "how conservatively to retain reasoning supervision."
↓
5 · Budgeted replacement
Within each domain, reasoning-unnecessary turns are candidates; a per-domain no-think budget decides how many are replaced; retention is stratified by topic with a per-capability cap on the drop rate.

Three design choices in that pipeline are worth pausing on.

A fixed probe, trained without the transformation. If the probe had itself been trained on adaptive-thinking data, its no-think answers would already reflect the policy being built, and the measurement would chase its own tail. A fixed probe gives a stable yardstick.

Blind, paired judging. The judge sees two answers to the same turn without knowing which one had reasoning. Blinding removes the bias of preferring the answer that "looks" more thoughtful. Pairing compares like with like: same context, same target, only the reasoning differs. The paper calls this paired judgment "the primary evidence."

Budgets and caps. Even a good label is noisy, so the paper does not simply strip reasoning from every turn labelled unnecessary. "A per-domain budget on the no-think rate decides how many are taken while the remaining turns keep their original reasoning." And: "Rather than spending that budget uniformly, we stratify retention by fine-grained topic and cap the drop rate within each capability, so that reasoning-intensive capabilities are not disproportionately stripped of supervision."

What does "replacement" produce? The paper calls the reasoning-unnecessary turns "candidates for replacement." Figure 7(A) shows the resulting behavior: the direct route gives an "immediate response" with an "empty think block." The natural reading is that replaced turns teach the model to emit an empty think block and answer directly; the report does not spell out whether the replaced answer is the original target or the probe's regeneration.

python (sketch)def label_turn(turn, probe, judge, n_repeats=3):
    no_think = probe.generate(turn.context, think="")      # empty think block
    verdicts = []
    for _ in range(n_repeats):                              # consistency across repeats
        a, b, flipped = shuffle_pair(turn.answer, no_think)     # blind: judge can't tell which
        verdicts.append(judge.compare(a, b, turn.target, flipped))
    unnecessary = all(v == "same_quality" for v in verdicts)
    confidence  = combine(verdicts, coherence(turn.reasoning, turn.answer),
                          task_prior(turn.task_type))
    return unnecessary, confidence

def apply_budget(domain_turns, no_think_budget, cap_per_capability):
    cands = [t for t in domain_turns if t.unnecessary]
    quota = int(no_think_budget * len(domain_turns))          # per-domain no-think rate
    picked = stratified_by_topic(cands, quota,                 # spread over fine-grained topics
                                 cap=cap_per_capability)          # protect reasoning-heavy skills
    for t in picked:
        t.reasoning = ""                                         # becomes a direct-route example
    return domain_turns
# n_repeats, the "all same" rule and the budget values are placeholders, not the paper's.
Budget arithmetic · illustrative numbers A domain has 1,000 training turns; the judge labels 400 as reasoning-unnecessary.
Step 1. Suppose the per-domain no-think budget is 30%: quota = 0.30 × 1,000 = 300 turns.
Step 2. Only 300 of the 400 candidates are taken; 400 − 300 = 100 candidates keep their reasoning anyway.
Step 3. Suppose one capability (say, multi-step arithmetic) has 120 candidates and a cap of 25% of its turns, which number 200: cap = 0.25 × 200 = 50. At most 50 of its 120 candidates are stripped.
Step 4. Resulting domain think rate = (1,000 − 300) ÷ 1,000 = 70%.
The cap is the safety valve: it prevents the budget from being filled mostly from one capability whose "unnecessary" labels may be noisier than they look.

Did it work? Table 5, read carefully

The paper compares three configurations on StepAudioChat, "covering 46 benchmark members in eight capability categories." Direct SFT and forced no-think "share baseline weights, with explicit reasoning enabled or disabled at inference time." Adaptive Thinking "is trained separately with reasoning-selection supervision, so its comparison also includes the effect of training." Scores and think rates are unweighted means over benchmark members; evaluation uses temperature zero and no system prompt.

CategoryDirect SFT thinkDirect SFT scoreForced no-think scoreAdaptive thinkAdaptive score
Instruction Following100.064.1561.9860.062.12
Faithfulness100.072.3570.6379.272.42
Reasoning100.071.8960.5259.566.80
Memory100.067.9963.8651.565.99
Knowledge100.074.5968.6580.471.55
Safety and Reliability100.078.4674.9556.977.90
Dialogue Pragmatics100.063.5962.4158.965.87
Persona and Role Consistency100.077.1768.8082.074.62

The forced no-think column always has a think rate of 0.0. Use the explorer below to see each category's three scores side by side, together with two derived numbers: the full-thinking gain (Direct SFT minus forced no-think) and the share of that gain Adaptive Thinking recovers.

Explorer — Table 5, one category at a time

All scores and think rates are the paper's (Table 5). The "gain" and "recovered" figures are our arithmetic on those numbers. A recovered share above 100% means Adaptive Thinking scored higher than always-thinking in that category.

Worked example 8a · how much of the reasoning benefit does Adaptive Thinking keep? (Table 5) Reasoning category.
Step 1. Full-thinking gain = Direct SFT − forced no-think = 71.89 − 60.52 = 11.37 (the paper's figure).
Step 2. Adaptive gain over no-think = 66.80 − 60.52 = 6.28.
Step 3. Share recovered = 6.28 ÷ 11.37 = 55.2%, while thinking on 59.5% of turns.
Across all eight categories (unweighted category means; our computation, not a paper figure):
Step 4. Direct SFT mean: (64.15 + 72.35 + 71.89 + 67.99 + 74.59 + 78.46 + 63.59 + 77.17) ÷ 8 = 570.19 ÷ 8 = 71.27375 ≈ 71.27.
Step 5. Forced no-think mean: (61.98 + 70.63 + 60.52 + 63.86 + 68.65 + 74.95 + 62.41 + 68.80) ÷ 8 = 531.80 ÷ 8 = 66.475 ≈ 66.48.
Step 6. Adaptive mean: (62.12 + 72.42 + 66.80 + 65.99 + 71.55 + 77.90 + 65.87 + 74.62) ÷ 8 = 557.27 ÷ 8 = 69.65875 ≈ 69.66.
Step 7. Share of the gain recovered, from the unrounded means = (69.65875 − 66.475) ÷ (71.27375 − 66.475) = 3.18375 ÷ 4.79875 = 66.3% (the rounded means would give 3.18 ÷ 4.79 = 66.4%, a rounding artifact).
Step 8. Mean think rate = (60.0 + 79.2 + 59.5 + 51.5 + 80.4 + 56.9 + 58.9 + 82.0) ÷ 8 = 528.4 ÷ 8 = 66.05%.
Step 9. Two-thirds of the thinking buys about two-thirds of the benefit: these aggregates show no sign that reasoning is concentrated where it helps most, though aggregates alone cannot say how individual turns were routed.

That last step lines up with the paper's own, franker reading. "Reasoning has a think rate of 59.5% despite benefiting most from full thinking. Faithfulness has a higher rate of 79.2%, but gains only 1.72 points from full thinking. Lower thinking frequency alone therefore does not demonstrate that reasoning is allocated to the turns that benefit most." It adds: "These aggregate comparisons also do not establish the optimal decision for individual turns."

The paper's summary of the gains: full thinking helps most in Reasoning (11.37 points), then Persona and Role Consistency (8.37) and Knowledge (5.94). Relative to Direct SFT, Adaptive Thinking "improves Dialogue Pragmatics from 63.59 to 65.87, but reduces Reasoning from 71.89 to 66.80." The conclusion of the report keeps the same tone: "Adaptive Thinking reduces the frequency of explicit reasoning, with uneven effects on answer quality."

Why pragmatics might like less thinking. Dialogue Pragmatics is where Adaptive Thinking beats always-thinking by the widest margin (65.87 vs 63.59, +2.28); it also edges ahead on Faithfulness (72.42 vs 72.35, +0.07), a gap too small to read much into. A plausible reading (ours): socially fitting replies are often short and direct, and a long private deliberation can push a model toward over-explained answers. The paper reports the number without offering a cause.

Part B. Multi-token prediction: thinking faster

Adaptive Thinking reduces how often the model thinks. The paper then attacks the cost of the thinking that remains: "we use MTP3 with three prediction heads, drafting up to three future tokens at each target-model step."

Start with the bottleneck. An autoregressive model produces one token per forward pass. A 400-token reasoning trace means 400 passes through an 11B-active-parameter network, one after another.

Now the idea, with an analogy. A chess player thinking about a line can pre-move: "if they play this, I will play that, then this." If the opponent's actual moves match, several moves happen in the time of one decision. If not, the pre-moves are discarded, and nothing is lost except a little effort.

Multi-token prediction gives the model extra small prediction heads that, from the same hidden state, guess the token after next, the one after that, and so on. Those guesses are drafts. In the next forward pass, the main model (the target model) checks all drafts at once, in parallel, and keeps the longest prefix it agrees with. This draft-then-verify loop is speculative decoding; the paper cites the classic speculative-sampling papers, Medusa, and multi-token prediction as its lineage. The Speculative Decoding Gleam and the Leviathan et al. Veanor build it from zero.

One piece of standard accounting (background, not stated in the report): each verification pass also yields one token of the target model's own, so the tokens produced per target step are one plus the number of accepted drafts.

Strict versus typical acceptance

How does the target decide whether to keep a draft? The paper uses two rules.

Strict verification "follows the target model's standard speculative-decoding rule." In that rule (from the speculative-sampling papers), a draft is accepted exactly when doing so leaves the output distribution unchanged; with greedy decoding, that means the draft must equal the target's own top choice. Speed is gained without changing what the model says.

Typical acceptance, from Medusa, "uses an entropy-adaptive confidence threshold to accept plausible draft tokens that strict verification may reject." The Medusa paper's form of this rule accepts a draft token x when the target's probability for it clears a bar that drops as the target becomes less certain:

accept x  if  p(x) > min( ε, δ · exp(−H(p)) )

Here p is the target model's next-token distribution, p(x) is the probability it gives the draft token, H(p) is the entropy of that distribution (a measure of how spread out, how uncertain, it is), and ε and δ are thresholds. When the target is confident (low H), the bar is high and only strong drafts pass. When the target is unsure (high H), many continuations are reasonable, the bar drops, and a plausible draft is kept. The StepAudio report states the principle but not its threshold values, so ε and δ here are Medusa's notation, not reported constants.

The paper is clear about the cost: typical acceptance "increases acceptance while allowing the generated distribution to change." So it adds a guard: "A repetition penalty of 1.05 discourages repeated tokens and reduces the risk of repetition loops under permissive acceptance." A repetition penalty scales down the scores of tokens that have already appeared (a common implementation divides a positive logit by the penalty; the report does not specify its form).

The policy split that makes this safe. "We apply typical acceptance and the repetition penalty to private reasoning, while retaining strict verification for the spoken response." The user never hears the reasoning; small distribution shifts there only matter through their effect on the answer. The spoken response is what the user hears and what the benchmarks score, so it keeps the target model's exact distribution. Permissive where it is private, exact where it is public.
python (sketch)def verify(target_probs, drafts, mode, eps=0.09, delta=0.3):
    """target_probs[k]: target distribution at draft position k (one forward pass)
    drafts[k]: token proposed by MTP head k+1.  Returns number accepted."""
    accepted = 0
    for p, x in zip(target_probs, drafts):
        if mode == "strict":                          # spoken response (greedy form)
            ok = (x == p.argmax())
        else:                                           # "typical": private reasoning only
            H  = -(p * p.clamp_min(1e-9).log()).sum()
            ok = p[x] > min(eps, delta * torch.exp(-H))
        if not ok:
            break                                       # later drafts depend on this one
        accepted += 1
    return accepted          # tokens this step = accepted + 1 (the target's own token)
# eps=0.09, delta=0.3 follow Medusa's public defaults, NOT StepAudio's (the report gives no values).
# The cap min(eps, .) binds on confident steps; delta*exp(-H) lowers the bar when entropy H is high.
# The repetition penalty of 1.05 (the paper's value) is applied to the logits first.

The break is why acceptance falls with depth: head 3's draft only counts if heads 1 and 2 were both accepted. That is what Table 7 calls marginal acceptance: the probability that a head's draft is accepted as part of the kept prefix.

The measurements: Tables 6 and 7

Domain (StepAudioChat)BaselineMTP3MTP3 (Medusa)MTP5MTP5 (Medusa)
Instruction Following64.1561.2162.7161.2562.73
Faithfulness72.3572.2970.4672.4972.03
Reasoning70.7674.3073.7473.1472.73
Memory67.9969.2968.9770.8571.36
Knowledge74.5975.0775.3475.3174.33
Safety & Reliability81.4181.0781.5981.2381.17
Conversational Pragmatics63.5966.0463.1766.2266.15
Persona & Role77.1777.9975.9277.1977.56
Accepted drafts / step01.2311.8011.3532.153
Wall-clock speedup1.00×1.76×2.05×1.49×1.72×
DepthVerificationHead 1Head 2Head 3Head 4Head 5
MTP3Strict65.0%37.2%20.8%N/AN/A
MTP3Typical82.4%58.2%39.4%N/AN/A
MTP5Strict63.7%35.5%19.6%10.7%5.7%
MTP5Typical80.6%55.7%37.6%24.9%16.4%

All configurations use the repetition penalty of 1.05, and the paper warns that "quality, acceptance, and wall-clock measurements use their respective evaluation settings." One visible consequence: six rows of Table 6's Baseline column match Table 5's Direct SFT column, but Reasoning (70.76 vs 71.89) and Safety (81.41 vs 78.46) differ; the paper notes only that each table keeps its own evaluation settings, and it does not say whether the two baselines are the same checkpoint.

Worked example 8b · accepted drafts per step are the sum of the marginal rates (Tables 6 and 7) Why a sum? Accepted count = (head 1 accepted) + (head 2 accepted) + …, each term a 0/1 indicator. The expected value of a sum is the sum of expected values, and the expected value of an indicator is its probability, the marginal rate.
MTP3 strict: 0.650 + 0.372 + 0.208 = 1.230 (Table 6: 1.231). Tokens per step ≈ 1 + 1.230 = 2.230.
MTP3 typical: 0.824 + 0.582 + 0.394 = 1.800 (Table 6: 1.801). Tokens per step ≈ 2.800.
MTP5 strict: 0.637 + 0.355 + 0.196 + 0.107 + 0.057 = 1.352 (Table 6: 1.353). Tokens per step ≈ 2.352.
MTP5 typical: 0.806 + 0.557 + 0.376 + 0.249 + 0.164 = 2.152 (Table 6: 2.153). Tokens per step ≈ 3.152.
Heads 4 and 5 under strict: they add 0.107 + 0.057 = 0.164 accepted drafts per step, against 0.650 for head 1 alone.
Tokens per step vs wall-clock: MTP3 typical yields about 2.80 tokens per step but a 2.05× wall-clock speedup; 2.05 ÷ 2.80 ≈ 0.73. Drafting and verification are not free, and the paper cautions that wall-clock ratios "are specific to each timing configuration."

The paper's three findings, with the numbers behind them. First, "the tested MTP configurations improve Reasoning and Memory over the baseline, while Instruction Following declines" (Reasoning 70.76 rises to between 72.73 and 74.30; Instruction Following 64.15 falls to between 61.21 and 62.73). Second, deeper drafting shows "diminishing gains in accepted drafts per step": heads 1 to 3 behave similarly at both depths, and the strict marginal rates of heads 4 and 5 fall to 10.7% and 5.7%. Third, "acceptance alone does not determine net efficiency, which also depends on reasoning length and decoding costs," so the wall-clock ratios "should not be interpreted as a controlled comparison of draft depth." That is why MTP5 strict (1.353 accepted) shows a lower speedup (1.49×) than MTP3 strict (1.231 accepted, 1.76×).

Sim 7: run the draft-and-verify loop

Each step, the heads propose drafts; the verifier walks them left to right and stops at the first rejection. The acceptance probabilities come straight from Table 7 (each head's conditional chance is its marginal rate divided by the previous head's). Run many steps and watch the running average converge to the sum in worked example 8b.

Sim 7 — MTP drafts, strict vs typical acceptance

Green = accepted draft, red = first rejection, grey = discarded because an earlier draft was rejected, teal = the target model's own token for this step. Acceptance rates are the paper's (Table 7); the random draws are the simulation's.

Inline check: why not use typical acceptance for speech too? It would make speech faster. What would be lost? Answer: typical acceptance lets "the generated distribution change". For the spoken response, that means the user hears text the model would not otherwise have produced, and any quality loss lands directly on what is scored and heard. For private reasoning, the paper accepts that shift, guarded by the 1.05 repetition penalty, because its effect is mediated by the answer.
Chapter 7 recap. (1) Adaptive Thinking labels each training turn with a fixed no-think probe and a blind paired judge; repeated-judgment consistency, trace–answer coherence and a task-type prior set how conservative to be. (2) A per-domain no-think budget, stratified by topic and capped per capability, decides which unnecessary turns lose their reasoning. (3) Think rates land between 51.5% and 82.0%; full thinking helps Reasoning most (+11.37), yet Reasoning gets only 59.5% thinking; Adaptive improves Pragmatics but drops Reasoning to 66.80. (4) MTP3 drafts up to three tokens per target step; typical acceptance with a 1.05 repetition penalty for private reasoning, strict verification for speech. (5) Accepted drafts per step equal the sum of marginal rates (1.231, 1.801, 1.353, 2.153); best wall-clock 2.05× (MTP3, Medusa); deeper heads give diminishing returns.
Where does StepAudio 3 Realtime apply Medusa-style typical acceptance, and why there?

Chapter 8: The Full-Duplex Voice Agent

"Can you book me a car to the airport tomorrow?" Simple to say, not simple to do. The assistant needs your flight time, which lives in your private calendar. It needs a pickup time, which depends on the flight. It needs to call a booking service, which may take a while to respond.

While it works, you keep living. "Oh, and it's for two people." "Is it done yet?" "What's the weather going to be there?" A walkie-talkie agent would make you wait for the booking to finish before any of those could be heard.

The paper's introduction names this as the fourth problem: "Tool use extends this challenge because an external task may outlast the spoken exchange that initiated it." Section 7 is the answer: a full-duplex voice agent whose tool work runs "alongside the conversation, enabling the user to ask about progress or provide additional requirements while work is underway."

Three routes for every request

Think of a hotel concierge. Some questions they answer from memory ("breakfast is until ten"). Some need a quick look at a screen ("rain tomorrow, bring an umbrella"). Some they hand off to someone in the back office and promise to follow up ("I'll arrange the car and call your room"). A good concierge picks the right route without being told.

Section 7 gives the model the same three routes: "The model selects among direct responses, lightweight tool calls, and asynchronous backend execution based on the requirements of each request."

RouteWhen the paper uses itExamples from the paperTiming
Direct response"Routine conversation and questions about stable knowledge"chit-chat, well-known factsImmediate
Lightweight tool"Requests for up-to-date public information""weather lookup or web search"Short
Asynchronous backend"Requests involving private context, multi-step processing, or work extending beyond the current conversational turn"tasks needing the user's data or several stepsRuns alongside the conversation

The paper places this routing in the tradition of tool-augmented language models, citing Toolformer and Gorilla. The Toolformer Veanor shows how a model can learn when a tool call is worth making.

Routing also interacts with Adaptive Thinking from Chapter 7. A direct response on a routine turn can take the empty-think route. A backend task that needs planning is where "reasoning supports constraint resolution and task planning" (Section 7.2).

Routing can fail in both directions, and each direction has a different cost in a voice product. The paper's training explicitly targets both: routing examples teach when a tool is needed, and negative examples "discourage unnecessary tool invocation."

Routing mistakeExample (illustrative)Cost to the user
Tool when a direct answer would doA web search for "what does a backchannel mean?"Extra delay on a turn that should feel instant
Direct answer when fresh data is neededAnswering "will it rain tomorrow?" from memoryA confident, stale, possibly wrong answer
Lightweight tool when private context is neededWeb-searching for "my" bookingThe task cannot succeed; the agent may bluff
Backend for a one-line factLaunching a multi-step job to tell the timeWasted compute and a needlessly long exchange

The asymmetry: speech is revisable, actions are not

Here is the most important idea in Section 7, and it is subtle enough to read twice.

"A grammatically complete phrase may still leave a spoken request incomplete or open to revision. Spoken responses can be extended or explicitly corrected in subsequent segments, whereas tool execution requires sufficiently specified intent and arguments."

Recall Chapter 6. Think-While-Speaking is comfortable starting to talk before reasoning is complete, because speech can be amended: "a final continuation can supplement or correct." A booking cannot be amended by a later sentence. Once the car is booked for the wrong time, saying "sorry, I meant 7:30" does not change the database.

So the paper makes the agent more conservative exactly where the cost of being wrong is higher: "Before committing to an external action, the model is therefore trained to gather missing information through clarification with the user or context retrieval using an appropriate tool, and to obtain any required confirmation."

Two tempos in one model. Speaking: Speak-First, correct later if needed. Acting: clarify first, confirm, then commit. The same model runs both tempos because the cost of a mistake differs: a spoken error costs a correction sentence, an action error costs a wrong real-world state. τ-Voice, the benchmark in this chapter, scores exactly that real-world state.

Three ways to fill a gap before acting, all named in the paper:

GapHow to close itExample (ours, illustrative)
Missing argument the user knowsClarification with the user"What time is your flight?"
Missing argument in private context"Context retrieval using an appropriate tool"Look up the flight in the user's calendar
Consequential action"Obtain any required confirmation""Pick-up at 7:30 for two, shall I book it?"

Talking while the work runs

Section 7.2: "Backend tasks execute asynchronously while the conversation continues. During execution, the user may request progress updates, provide additional requirements, or shift to another topic."

Three kinds of user input can arrive while a task is running, and each needs a different treatment. A progress query must be answered from the task's actual status, not from optimism. An additional requirement must be attached to the right task. An unrelated topic must be handled on its own without disturbing the task. The paper states the training goal: "The model is trained to associate task-related user input with the ongoing task while distinguishing it from unrelated dialogue."

When results arrive: "As results become available, they are incorporated into the conversational context to inform subsequent spoken responses." That is the tool-execution status slot on the shared whiteboard from Chapter 0, being written by the backend and read by the speaker.

The division of labor between thinking and acting follows ReAct, the pattern of interleaving reasoning steps with actions and observations (see the ReAct Veanor): "For complex requests, reasoning supports constraint resolution and task planning, while tool and backend outputs provide evidence of what has actually been completed. This evidence guides subsequent reasoning and action."

And the paper ties the two concurrency mechanisms together in one sentence: "Think-While-Speaking supports spoken responses during deliberation, while asynchronous execution allows the conversation to continue during external task execution." Chapter 6 decoupled thinking from speaking. This chapter decouples acting from speaking.

Sim 8: the shared context during an async task

Run the car-booking scenario. Once the backend task is running, use the interjection buttons: ask for progress, add a requirement, or change the topic. The right-hand board is the shared conversational context: watch which slot each event writes to. Then switch to a blocking agent that waits for the tool before it can talk, or skip clarification, and compare.

Sim 8 — a full-duplex voice agent with an asynchronous backend task

The scenario, durations, flight time and weather reply are illustrative; the routing, clarify-and-confirm-before-commit rule, task association and evidence-grounded reporting are the behaviors the paper describes (§7.1–7.3). In blocking mode, interjections queue until the task returns, so a late requirement misses the booking.

In the async run, every interjection is handled within seconds while the backend bar keeps growing, and the requirement "two people" is attached to the task before it completes. In the blocking run, the same interjections pile up as dead air and the requirement arrives after the booking has already been made. With clarification switched off, the task commits with a guessed pickup time, and the final report has to disclose a wrong result, which is exactly what the paper's negative examples train against claiming as success.

An annotated transcript

Here is one full exchange, written out turn by turn, with the route each turn takes and the slot of the shared context it writes. The dialogue is illustrative; the annotations follow the behaviors listed in Sections 7.1 to 7.3.

SpeakerWords (illustrative)Route / behaviorContext slot written
User"Book me a car to the airport tomorrow."Backend candidate: private context, work beyond this turnRequest
Model(private) flight time unknown; check the calendarContext retrieval with a tool before committingArguments: pending
Model"Your calendar shows a 9:40 flight. Pick-up at 7:30?"Clarification plus confirmation of a consequential actionArguments: 9:40 → 7:30
User"Yes. Oh, and what's the weather there?"Confirmation; then an unrelated questionConfirmation: yes
Model"Booking now. Tomorrow looks sunny there."Backend starts asynchronously; weather goes to a lightweight toolTask: running; Side topic: answered
User"Make it for two people."Task-related input, associated with the running taskRequirements: 2 passengers
User"Is it done?"Progress query, answered from real status(read) Task: 60%
Model"Not yet, it's still confirming."No unsupported success claimnone
Backend(returns: booked, 7:30, 2 passengers)Evidence enters the contextResult
Model"Done: a car at 7:30 for two."Result reporting grounded in the returned evidencenone

Every row maps onto an item in the paper's list of targeted training behaviors: routing, clarification, private-context retrieval, confirmation before consequential actions, execution-time updates, progress queries and result reporting. The weather lookup shows the other half of routing: a lightweight tool, handled inside the same conversation while the backend keeps working.

The event loop, in code

python (asyncio sketch)class VoiceAgent:
    def __init__(self, model, tools, backend):
        self.model, self.tools, self.backend = model, tools, backend
        self.ctx = SharedContext()          # evidence, history, turn, reasoning, TOOL STATUS

    async def on_user_utterance(self, utt):
        route = self.model.route(self.ctx, utt)     # direct | tool | backend
        task = self.ctx.related_task(utt)           # associate with an ongoing task?
        if task and is_progress_query(utt):
            return await self.say(status_report(task))   # from real status only
        if task:
            task.add_requirement(utt)                  # extra constraint, same task
            return await self.say("Added to the booking.")
        if route == "direct":
            return await self.say(self.model.answer(self.ctx, utt))
        if route == "tool":                           # weather, web search
            obs = await self.tools.call(self.model.tool_call(self.ctx, utt))
            return await self.say(self.model.answer(self.ctx, utt, evidence=obs))
        # backend: never commit on an under-specified request
        args = await self.fill_arguments(utt)         # clarify with user / retrieve context
        if not await self.confirm(args):               # consequential action
            return
        job = asyncio.create_task(self.backend.run(args))   # does NOT block the conversation
        self.ctx.track(job, args)
        job.add_done_callback(lambda j: asyncio.create_task(self.report(j)))
        await self.say("On it. I'll let you know when it's done.")

    async def report(self, job):
        result = job.result()
        self.ctx.write_tool_result(result)             # evidence enters the context
        await self.say(self.model.summarize(self.ctx, evidence=result))  # claim only what it returned

In the real system, none of those branches is hand-written; the model makes each decision by predicting tokens, and the orchestration details are not published. The sketch shows the behaviors the training data targets, in the order the paper lists them.

How the agent is trained

Section 7.3 combines two data sources.

Targeted voice-agent dialogues cover "request routing, clarification, private-context retrieval, confirmation before consequential actions, execution-time updates, progress queries, and result reporting." They "train the model to ground claims about private information and completed work in user-provided context or evidence returned by tools." And, importantly, "negative examples discourage unnecessary tool invocation and unsupported claims of successful execution."

Two failure modes are singled out by those negative examples. Unnecessary tool invocation wastes time and money, and in voice, adds delay to turns that could have been answered directly. Unsupported success claims ("Done, your car is booked!" before any tool has returned) are the voice-agent version of hallucination, and they are worse than silence because the user acts on them.

Real multi-step agent trajectories complement the dialogues "by exposing the model to longer sequences of reasoning and tool use." The team filters and normalizes them, "focusing on tool-call structure, argument consistency, evidence grounding, and suitability for spoken interaction." That last criterion matters: a trajectory that ends with a 40-row table is fine for a text agent and useless to read aloud.

Recall from Chapter 1 that midtraining already "substantially increases the share of... agent-interaction data," training the model "to carry user intent through planning, tool use, and spoken follow-up." The voice-agent training here builds on that foundation.

Behavior the data targetsFailure it prevents
Request routingCalling a backend for "hello"; answering "what's my balance" from memory
ClarificationCommitting with a guessed argument
Private-context retrievalAsking the user for something already in their data
Confirmation before consequential actionsIrreversible actions the user did not approve
Execution-time updates, progress queriesDead air; fabricated progress
Result reporting grounded in evidenceClaiming success the tool did not return

When the conversation gets messy

τ-Voice deliberately injects the conditions that break voice agents. The table connects each condition the benchmark lists to the mechanism in this report that is meant to absorb it. The mapping is ours; the paper reports only the aggregate task-success rates.

τ-Voice conditionWhat can go wrongMechanism aimed at it
"Interruptions in which users revise their requests"The agent commits to the old requestDuplex yielding (Ch. 3) + clarify and confirm before commit (§7.1) + task association (§7.2)
BackchannelsThe agent stops or restarts on every "mm-hm"Backchannel handling (Ch. 3)
"Diverse forms of background noise"Misheard names and numbers become wrong tool argumentsRobust perception and data augmentation (Ch. 2); confirmation before consequential actions
Multi-step customer-service tasksA dropped constraint leaves the database in the wrong stateReasoning for constraint resolution; real multi-step trajectories in training (§7.3)

The benchmark: τ-Voice

The paper evaluates with the Artificial Analysis implementation of τ-Voice, which "assesses tool-grounded task completion in full-duplex spoken interaction under challenging conversational and acoustic conditions, including interruptions in which users revise their requests, backchannels, and diverse forms of background noise." Tasks are customer-service scenarios in three domains: airline, retail and telecom.

The success criterion is strict and objective: a task succeeds "when the final database state matches its target." Saying the right words is not enough; the backend state must actually be right. "Domain task-success rates average three trials where available, and the reported macro average gives equal weight to Airline, Retail, and Telecom."

Domain (task success, %)StepAudio 3 RealtimeGrok Voice Think Fast 2.0 HighQwen Audio 3.0 Realtime PlusGPT-Realtime-2.1 High
Airline60.056.061.362.0
Retail37.749.749.045.6
Telecom70.263.753.529.4
Macro Average56.056.554.645.7
Worked example 9 · the τ-Voice macro, and what it would take to lead (Table 9) StepAudio 3: 60.0 + 37.7 + 70.2 = 167.9; 167.9 ÷ 3 = 55.97 ≈ 56.0.
Grok: 56.0 + 49.7 + 63.7 = 169.4; 169.4 ÷ 3 = 56.47 ≈ 56.5.
Qwen: 61.3 + 49.0 + 53.5 = 163.8; 163.8 ÷ 3 = 54.6.
GPT: 62.0 + 45.6 + 29.4 = 137.0; 137.0 ÷ 3 = 45.67 ≈ 45.7.
To pass Grok's macro: the three-domain sum must exceed 169.4. StepAudio 3's airline + telecom = 60.0 + 70.2 = 130.2. Required retail > 169.4 − 130.2 = 39.2.
Gap: 39.2 − 37.7 = 1.5 retail points would tie Grok's sum; anything more leads. Retail is 49.7 − 37.7 = 12.0 points behind Grok.
Telecom lead: 70.2 − 63.7 = 6.5 points over Grok, the paper's figure. Airline: 62.0 − 60.0 = 2.0 points behind the best.

The paper's reading: StepAudio 3 Realtime reaches "a macro task-success rate of 56.0%, close to Grok's 56.5% and above Qwen's 54.6% and GPT's 45.7%." It has "the highest telecom score among the evaluated models at 70.2%," its airline score "is within 2.0 percentage points of the highest reported score of 62.0%," and "retail performance leaves room for further improvement."

The domain profiles are strikingly different across systems. GPT-Realtime-2.1 High is best at airline (62.0) and far behind at telecom (29.4). StepAudio 3 is best at telecom and weakest at retail. A macro average hides this; the paper is right to report domains separately. The report does not analyze why retail is hard for this model, so we will not guess at a cause.

Why database-state scoring suits this chapter. Every behavior in Section 7 (clarify, confirm, attach late requirements, report only evidence) exists to make the final state right under messy spoken conditions. τ-Voice checks exactly that state, while injecting request-revising interruptions, backchannels and background noise. In our view, it is the benchmark in the report that comes closest to exercising listen, converse, think and act at once.
Inline check: which input goes where? A booking task is running. The user says, in order: "is it booked?", "actually make it 7:45", "do you like jazz?". Classify each per the paper. Answer: progress query about the ongoing task (answer from real status); additional requirement associated with the task (update it, and since it changes an argument of a consequential action, confirm it); unrelated dialogue (answer directly, do not touch the task).
Chapter 8 recap. (1) Three routes: direct answer (routine, stable knowledge), lightweight tool (fresh public info), asynchronous backend (private context, multi-step, longer-running). (2) Speech can be corrected later; actions need specified arguments, so the agent clarifies, retrieves context, and confirms before committing. (3) During execution the user can ask progress, add requirements, or change topic; the model associates task-related input with the task and folds results into the context when they arrive. (4) Training: targeted dialogues (with negative examples against needless tool calls and unsupported success claims) plus filtered real multi-step trajectories. (5) τ-Voice macro 56.0% (Grok 56.5%): best telecom 70.2%, airline 60.0%, retail 37.7%.
Why does the paper train the voice agent to clarify and confirm before committing to an external action, even though its spoken responses start before reasoning is complete?

Chapter 9: Scoreboard & Horizon

You now know every mechanism in the report. This last chapter does three things: puts all the evaluations on one board, lists what the paper admits it has not solved (and what it does not report), and points you to what to read next.

Section 8 frames the evaluation as a staircase: six capability domains that "measure the progression from recognizing an utterance to understanding its context, managing the conversational floor, and completing an external task." Each step of that staircase matches a verb of the loop from Chapter 0.

DomainBenchmarksProtocol notes (§8.3)Chapter
Speech recognitionLibriSpeech, AISHELL-1, WenetSpeech, ContextASR-BenchWER (English) / CER (Mandarin); Contextless setting; baselines rerun on the same audio2
Audio understandingBig Bench Audio, MMSU, MMAU, MMAR, WildSpeech-Bench, AudioMultiChallenge, Step-Caption, MTalk-BenchEach benchmark's reported accuracy or normalized score; Step-Caption judge-scored; MTalk-Bench = paralinguistic + ambient components only2
Dialogue and reasoningStepAudioChat (8 dimensions)Each dimension = unweighted mean of validated capability lines; reasoning-mode and realtime results both reported; no ablation scores used to fill entries4, 6
Full-duplex interactionAA subset of Full Duplex Bench v1 and v1.5Category = % of samples meeting the criterion; source-reported Overall, not a mean of categories3
Agentic task completionAA implementation of τ-VoiceSuccess = final database state matches target; up to three trials averaged; equal-weight macro over three domains8
General textHMMT February 2026, GPQA Diamond, MultiChallengeRecorded accuracy; MultiChallenge uses instance-specific rubrics; no average taken across the three1, 5

The baselines change by domain (Section 8.2), which is worth knowing before you read any comparison. ASR: Doubao 2.0 ASR, Seed 2.0 Lite, HY3.0 ASR Preview. Audio understanding: Doubao 2.0 Lite, Gemini 3 Flash, Gemini 3.1 Pro. Dialogue: Doubao 2.0 Lite, DeepSeek-V4-Flash, Kimi K3. General text: Doubao 2.0 Lite and Gemini 3 Flash. Full-duplex: GPT-realtime-2 (High), Qwen Audio 3.0 Realtime Plus, Grok Voice Think Fast 2.0 High. Agentic: the same Qwen and Grok variants, with GPT-Realtime-2.1 High replacing GPT-realtime-2. "Model versions and effort labels follow the corresponding evaluation records."

The whole scoreboard (Table 10)

Pick a domain to see StepAudio 3 Realtime beside the baselines the paper uses for that domain. Bars use a 0 to 100 scale, except where a domain starts its bars higher so small gaps stay visible (noted under the bars); the best score in each row is marked.

Scoreboard — Table 10, domain by domain

Every number is from the paper: Table 10, plus the Table 3 duplex categories and the Table 9 τ-Voice domains. For Dialogue, StepAudio 3 is in realtime ("interactive") mode while the baselines are in reasoning mode, as the paper notes.

Here is the full table in static form, for reference.

Audio understandingStepAudio 3Doubao 2.0 LiteGemini 3 FlashGemini 3.1 Pro
Big Bench Audio98.198.899.499.6
AudioMultiChallenge49.348.556.667.0
MMSU90.680.077.083.6
MMAU79.077.577.680.5
WildSpeech77.173.974.477.7
MMAR86.575.975.481.7
Step-Caption78.276.867.874.8
MTalk-Bench91.789.988.589.1
General text (accuracy)StepAudio 3Doubao 2.0 LiteGemini 3 Flash
HMMT 2026 Feb86.873.985.9
GPQA Diamond83.082.490.3
MultiChallenge59.760.868.1

The dialogue rows (70.4 macro in realtime mode) are in Chapter 6, the full-duplex Overall (98.9) in Chapter 3, and τ-Voice (56.0) in Chapter 8.

Worked example 10 · reading the scoreboard honestly (Table 10) Audio leads. Rows where StepAudio 3 is best: MMSU, MMAR, Step-Caption, MTalk-Bench = 4 of 8.
Audio macro. 98.1 + 49.3 + 90.6 + 79.0 + 77.1 + 86.5 + 78.2 + 91.7: 98.1 + 49.3 = 147.4; + 90.6 = 238.0; + 79.0 = 317.0; + 77.1 = 394.1; + 86.5 = 480.6; + 78.2 = 558.8; + 91.7 = 650.5. 650.5 ÷ 8 = 81.31 ≈ 81.3. Gemini 3.1 Pro: 81.8. Gap = 81.8 − 81.3 = 0.5, the paper's "within 0.5 points".
Largest audio deficit. AudioMultiChallenge: 67.0 − 49.3 = 17.7.
General text vs Gemini 3 Flash. HMMT: 86.8 − 85.9 = +0.9. GPQA Diamond: 83.0 − 90.3 = −7.3. MultiChallenge: 59.7 − 68.1 = −8.4 (and 59.7 − 60.8 = −1.1 vs Doubao 2.0 Lite).
Full-duplex margin. 98.9 − 98.4 = 0.5 over the next system.
τ-Voice margin. 56.0 − 56.5 = −0.5 behind the best system.

Put in one line: the model is first on full-duplex control, leads 4 of 8 audio benchmarks and sits 0.5 below the best audio macro (with a large gap on AudioMultiChallenge), close on agentic success, competitive on dialogue given a realtime handicap, and behind a strong general model on two of three text benchmarks. The paper's own synthesis says much the same: it "combines broad audio understanding with strong dialogue, reasoning, and interaction control," and its balanced duplex profile "suggests that it can preserve conversational flow without treating every user sound as an interruption."

Notice one more pattern across the text rows. Of the three, the model leads on HMMT, the benchmark most purely about reasoning, and trails most on MultiChallenge, the one about "multi-turn conversational reliability." That echoes AudioMultiChallenge (multi-turn audio) and the realtime instruction-following drop. Three different benchmarks point at the same soft spot: holding and revising constraints across a long conversation.

What the paper admits

The report is candid about its gaps. Collected in one place, each with where it shows up:

Admitted limitationPaper's words (abridged)Evidence
Multi-turn constraint handling"Multi-turn constraint handling and retail task completion remain areas for improvement."AudioMultiChallenge 49.3 vs 67.0 (the paper's link); realtime Instruction Following 54.1 and text MultiChallenge 59.7 vs 68.1 (our link)
Retail tool use"Retail performance leaves room for further improvement."τ-Voice retail 37.7 vs 49.7
Adaptive Thinking allocation"Adaptive Thinking reduces the frequency of explicit reasoning, with uneven effects on answer quality."Reasoning think rate 59.5% despite the largest gain (+11.37); Reasoning 71.89 → 66.80
Per-turn optimality unproven"These aggregate comparisons also do not establish the optimal decision for individual turns."Table 5 is category-level only
Incomplete-reasoning errors"Keeping strict verification for spoken output does not eliminate errors arising from incomplete private reasoning."Think-While-Speaking design
MTP timing is not a controlled comparisonWall-clock ratios "should not be interpreted as a controlled comparison of draft depth."MTP5 strict 1.49× vs MTP3 strict 1.76×
Merging is a trade-off"A balanced trade-off rather than uniform improvement over every specialized teacher."Dialogue 73.0 vs best teacher 74.2
ASR numbers are the specialist's"These results characterize the ASR-specialized model, not the transcription behavior of the realtime model."Table 1

The conclusion turns those into a research agenda: the findings "motivate more effective allocation of reasoning effort and more reliable task execution over extended conversations."

What the report does not tell you

Separate from the admitted limitations, a careful reader should note what a technical report of this kind leaves out. These are observations about the document, not criticisms of the model.

No end-to-end latency measurements. For a report titled "Realtime", the only timing constant given for the interaction is the 320 ms block; there are no first-response latency or overlap-reaction figures. The MTP section reports relative wall-clock speedups, not absolute times.

Architecture internals. Expert count, experts per token, layer count, hidden sizes, the adapter's design, the generator's design and token rates are not given; only ~196B total and 11B active parameters, the Step 3.7 Flash backbone and the Qwen3-Omni AuT encoder.

Data scale per stage. The 1.2T pretraining tokens, 32K and 128K context lengths, "over 10,000 hours" of synthetic duplex data, and the ~2M vs ~100K ablation are reported; mixture proportions, hours per language and teacher data mixes are not.

Reproducibility of the dialogue score. StepAudioChat is closed by design, which protects it from contamination and also means outsiders cannot rerun it.

Component attribution in realtime mode. The realtime StepAudioChat row combines Think-While-Speaking, Adaptive Thinking and MTP; the paper deliberately does not fill it from ablations, so the 2.6-point realtime cost cannot be split among them from the report alone.

The one-paragraph version of the paper. Voice makes thinking time audible. StepAudio 3 Realtime keeps listening, speaking, thinking and acting on separate clocks tied to one shared context: a 196B-total / 11B-active MoE audio-language model hears both sides of the call in 320 ms blocks and predicts floor decisions as tokens; a Formulation Brain reasons privately while an Articulation Brain speaks from whatever reasoning exists, released at the pace of playback; Adaptive Thinking, trained on turns that a blind paired judge labelled, decides per turn whether to reason; MTP with typical acceptance speeds the private trace up to 2.05× while speech stays strictly verified; merged specialist teachers supply the knowledge; and a voice agent clarifies before it acts and keeps talking while its backend works. The result tops the AA full-duplex bench (98.9), leads MMSU (90.6), nearly matches the best τ-Voice score (56.0 vs 56.5), and stays within reach of reasoning-mode models while speaking in real time (70.4), with multi-turn constraints and retail tasks as the open problems.

Numbers to carry away

If you remember nothing else, remember these, each with the idea it stands for.

NumberWhat it isThe idea behind it
320 msDuplex block lengthFloor decisions are tokens emitted 3.125 times a second
11B / 196BActive vs total parametersSparse experts make two concurrent calls affordable
128KMidtraining contextLong conversations and tool results stay in view
100K vs 2MCurated vs random SFT examplesQuality beat volume on every reported audio metric
98.9AA Full-Duplex OverallBalanced pauses, turns, interruptions and backchannels
3:1:1:1Merge weightsSpecialize, then average from a common base
73.0 → 70.4StepAudioChat, reasoning vs realtimeRealtime mode (Think-While-Speaking + Adaptive Thinking + MTP) scores 2.6 below reasoning mode, mostly on instruction following
59.5%Think rate on ReasoningTable 5 does not show that Adaptive Thinking spends reasoning where it helps most
1.231 / 1.801Accepted drafts per step, MTP3 strict / typicalSum of per-head marginal rates; typical acceptance used only in private
2.05×Best wall-clock speedupAcceleration divides silence; concurrency removes it
56.0 vs 56.5τ-Voice macro vs bestStrong telecom, weak retail; database state is the judge

The mechanism cheat sheet

MechanismProblem it solvesKey number(s) from the paperCh.
Shared conversational contextKeeps decoupled clocks coherentFive ingredients (evidence, history, turn, reasoning, tool status)0
MoE audio-language modelCapacity without per-token cost~196B total, 11B active; 32K → 128K context; 1.2T pretraining tokens1
ASR Max specializationAccurate, context-aware transcriptsLibriSpeech clean 1.18; AISHELL-1 0.49; ContextASR macro 5.67 / 1.232
Curated audio QAGrounded audio understanding~100K beats ~2M (MMSU 78.78 → 89.70); MMSU 90.6; macro 81.32
320 ms blocks + state tokensTake, retain or yield the floorAA Full-Duplex 98.9; turn taking 100.0; interruptions 99.03
StepAudioChat + self-play dataMeasure and teach conversationReasoning mode 73.0 (Kimi K3 77.1)4
Weighted mergingOne model, several specialties3:1:1:1; HMMT 86.8 above every teacher5
Think-While-SpeakingDeliberate without dead airRealtime 70.4 vs reasoning 73.06
Adaptive ThinkingSkip reasoning when it does not helpThink rates 51.5–82.0%; Reasoning 71.89 → 66.807
MTP3 + typical acceptanceFaster private reasoning1.801 accepted/step; 2.05× wall-clock; repetition penalty 1.057
Asynchronous voice agentKeep talking while tools runτ-Voice 56.0 (telecom 70.2, retail 37.7)8

Design lessons you can reuse

Strip away the specific model, and the report leaves six engineering ideas that transfer to other systems.

1. Make control decisions tokens. Floor management is a token predicted after every 320 ms block, conditioned on everything the decoder sees. Any decision that needs context (when to speak, when to call a tool, when to think) benefits from living inside the sequence rather than in a separate rule.

2. Be permissive where it is private and exact where it is public. Typical acceptance and a repetition penalty speed up the hidden reasoning; the spoken answer keeps strict verification. The same split applies to any system with an internal scratchpad and an external output.

3. Treat irreversible actions differently from revisable words. Speech can start early and be corrected; tool execution waits for specified arguments and confirmation.

4. Curate before you scale. One twentieth of the SFT data, chosen by quality, case value and cross-model agreement, won on every reported audio metric.

5. Build benchmarks with controls. A valid and a deliberately flawed response per item turn the judge itself into something you can test.

6. Specialize, then merge. Independent teacher mixes, recombined by a convex average, let separate teams improve separate skills without retraining on the union.

How this design differs from its neighbours

Several full-duplex or streaming designs have lessons on this site. At a high level, they place the boundary between thinking and speaking in different spots. The descriptions of the other systems are brief summaries of their own lessons, not claims from this report.

DesignHow listening and speaking overlapWhere reasoning lives
Cascade (VAD → ASR → LLM → TTS)They do not; a silence timer passes the turnIn the LLM, before any speech
MoshiUser and model audio modelled as parallel streamsAn inner text stream aligned with the model's own speech
Qwen2.5-OmniStreaming input and outputA Thinker that produces text, with a Talker that turns it into speech
StepAudio 3 RealtimeDual audio streams, 320 ms blocks, state tokensA private Formulation Brain running alongside a speaking Articulation Brain, both calls to the same model, with adaptive routing and MTP

Open questions worth a project

Each admitted limitation suggests a concrete next experiment. These are our suggestions, built on the paper's own diagnosis.

Allocate reasoning by benefit, not by frequency. Table 5's category think rates do not demonstrate allocation by benefit (the paper's own caveat), and aggregate numbers cannot settle per-turn decisions. A policy trained on the size of the paired-judgment gain, rather than on a same/different label under a budget, is a natural next step.

Protect constraints in Speak-First. Instruction following carries most of the realtime drop. Testing Think-First, or a constraint-extraction step before the first segment, on exactly that dimension would show whether early commitment is the cause.

Explain retail. A per-task error analysis of the 37.7% retail result (argument errors, missed confirmations, noise-induced mishearings) would say which of this report's mechanisms to strengthen.

Measure the clock. Publishing first-audio latency, reaction time to interruptions, and the fraction of answers needing a final continuation would let others compare realtime systems on the property the title promises.

Connections: where this paper sits on the site

Every link below points to a lesson that exists on Engineermaxxing today.

If you want to go deeper on…ReadWhy
How audio becomes LLM inputAudio LLMs, Qwen2-AudioEncoder + adapter + decoder, the Chapter 1 pattern from zero
Turn-taking and barge-inTurn-Taking, Endpointing & Barge-InThe cascade failures Chapter 3 replaces
Streaming recognition and synthesisStreaming SpeechWhy block-based streaming works and what it costs
Full-duplex predecessorsMoshi, DuplexSLA, PersonaPlexParallel-stream dialogue, synchronized speech/language/action, and voice/role control in full-duplex models
Thinking while talking, another designQwen2.5-OmniThinker-Talker: a different split between reasoning and speaking, from the Qwen omni line whose later Qwen3-Omni supplies this model's AuT encoder
Speech recognition foundationsWhisper (paper), Whisper (Gleam), Self-Supervised SpeechThe paper's first citation for LLM-era ASR; WER in practice; learned audio features
How speech is generatedTTS Architectures, Neural Audio Codecs, VALL-E, AudioLMWhat a "generator" can be; AudioLM is cited by the report
Draft-and-verify decodingSpeculative Decoding (Gleam), Speculative Decoding (Leviathan et al.)The strict-verification rule of Chapter 7
Multi-token prediction in a frontier LLMDeepSeek-V3MTP heads at scale, in a different model
Sparse expertsMixture of Experts, MoE: Sparse ComputationWhat "11B active of ~196B" buys
Reasoning and actingReAct, Toolformer, Agents & Tool UseThe interleaving pattern and tool-use lineage the agent chapter cites
Why multi-turn constraints are hardLLMs Get Lost in Multi-Turn ConversationIndependent evidence for the gap this report names

What to read next, in order

  1. Turn-Taking, if Chapter 3 felt fast: it builds the floor-control vocabulary slowly.
  2. Moshi, to see a different full-duplex design where both streams are modelled as parallel token streams.
  3. DuplexSLA, for the closest relative of this report's duplex and action ideas.
  4. Speculative Decoding, then return to Chapter 7's Table 7 and re-derive the accepted-per-step numbers yourself.
  5. LLMs Get Lost in Multi-Turn Conversation, to understand the open problem that three of this report's benchmarks point at.

Beyond the site, the report's own references are the natural next step: Mind-Paced Speaking (arXiv:2510.09592) for the dual-brain design, Medusa (arXiv:2401.10774) and Gloeckle et al. (arXiv:2404.19737) for multi-token decoding, and the Artificial Analysis speech-to-speech methodology for the full-duplex and τ-Voice protocols.

Reproduce the paper's numbers yourself

A good test of understanding is whether you can regenerate the report's derived figures from its raw tables. Every one of these was worked in this lesson:

Exit gate: teach it back before you leave.

Without scrolling up: (1) draw Figure 3's five boxes and say why the model audio stream feeds back into the input; (2) write the Figure 5 block sequence and name the four floor decisions, with the evidence each one uses; (3) explain the Formulation and Articulation Brains, playback-aware scheduling, and the Speak-First default; (4) reproduce 1.231 accepted drafts per step from Table 7 and say why typical acceptance is used only for private reasoning; (5) name the two admitted open problems and the numbers that reveal them.

Chapter 9 recap. (1) Six evaluation domains trace the loop from recognizing speech to completing tasks, each with its own baselines and aggregation rules. (2) The model leads full-duplex control (98.9) and 4 of 8 audio benchmarks, sits 0.5 behind on the audio macro and on τ-Voice, leads HMMT (86.8) and trails Gemini 3 Flash on GPQA Diamond and MultiChallenge. (3) Admitted gaps: multi-turn constraints, retail tool use, uneven Adaptive Thinking allocation, incomplete-reasoning errors. (4) Not reported: end-to-end latency, architecture internals, most data proportions; StepAudioChat is closed. (5) Three different benchmarks point at the same soft spot: keeping constraints straight over a long conversation.

The closing thought

For years, voice assistants were built like relay races: one runner listens, hands the baton to one who thinks, who hands it to one who speaks, and nobody moves until the baton arrives. The quiet idea in this report is that a conversation is not a relay; it is a band. The listener, the thinker, the speaker and the doer all play at once, at their own tempo, reading from the same score. The hard engineering is not any single instrument. It is keeping them in time with each other while the audience keeps interrupting. The next time a voice system makes you wait, ask which two clocks it has chained together that did not need to be.

Which pair of weaknesses does the paper itself name as remaining areas for improvement?