The most counterintuitive lesson in voice assistants: the hard part is not understanding speech — it is conversational timing. When you cough mid-sentence, should the assistant keep talking or stop? When music plays in the background, should it interrupt itself? Humans handle these calls instinctively; models need explicit design. In mid-September, Ant Group's Venus team, together with Tsinghua University, open-sourced their answer: Realtime-Venus, a system built around proactive full-duplex interaction and asynchronous delegation. The paper is on arXiv (2609.13814), and weights for its two 9B models are live on Hugging Face under the inclusionAI organization.
Two 9B Frontends, One Causal Timeline
Realtime-Venus does not bet on one do-everything model. It ships two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken dialogue. Both adapt the open-source MiniCPM-o 4.5 (Omni-Flow architecture), with a Qwen3-8B language backbone, SigLIP2 vision encoder, Whisper-Medium audio encoder, and speech generation via discrete S3 speech tokens with a streaming flow-matching decoder. Context window: 40,960 tokens; weights in BF16.
The core design compresses every event onto a shared causal timeline: user inputs, model outputs, and delegation events are all aligned in one-second chunks. Within each chunk the model first decides "listen or speak" (<|listen|> / <|speak|>), perception never pauses during speech, and the unspoken continuation stays revisable — that is the structural basis of talking while listening. Interruption handling is not a crude voice-activity-detection (VAD) binary, but semantic event classification: backchannels keep the response going, floor-taking interruptions stop queued audio, corrections trigger regeneration, redirects suspend the current plan.
The Harness: Chat in the Foreground, Work in the Background
Full-duplex solves "how to talk"; asynchronous delegation solves "how to get work done". The system pairs the frontends with Realtime-Venus-Harness, an execution framework: the frontend emits private delegation requests in-stream (inaudible to the user), the Harness routes tasks to registered backend capabilities asynchronously, and results pass freshness checks before re-entering the conversation at a chosen moment. The foreground never blocks — you can keep chatting while it looks things up.
On the data side, both models share a common post-training corpus of over 2.8 million samples across nine categories, spanning offline understanding, proactive full-duplex trajectories, and delegation workflows. The Omni variant trains on audio-visual plus audio-only data; the Audio variant uses the audio-only subset. The Omni model also carries a training-free long-video memory module: visual memory gating via motion-compensated prediction cost (inspired by AdaCodec's predictive visual coding), retrieval via MaxSim-style fine-grained token matching with an MMR-style relevance-novelty balance, supporting hour-scale video understanding.
Benchmarks: Strong at Not Getting Derailed, Weaker at Reacting Fast
On Full-Duplex-Bench v1.5, Realtime-Venus-Audio's continuation rates — 97% under user backchannels, 88% under speech directed to others, 86% under background speech — exceed GPT-4o and Gemini 3.1 Live on all three metrics; its interruption response rate is 75%, above MiniCPM-o 4.5's 60% but below Joy-Duplex's 88%, GPT-4o's 78%, and Gemini 3.1 Live's 77%. The paper's own conclusion is candid: strong continuation does not imply strong interruption handling — they are different skills.
On understanding, the Omni model tops the evaluated online models on six of eight video benchmarks, including StreamingBench 70.2%, OVO-Bench 64.7%, and Daily-Omni 81.3%; the Audio model leads compared models on MMAU 78.0%, MMAU-Pro 63.2%, Llama Questions 83.8%, and Speech CMMLU 67.8%, while matching the best VoiceBench AlpacaEval score of 4.81. On tool use (FDB-v3): Omni scores 86.0% tool-selection F1, 53.1% argument accuracy, and 43.0% Pass@1 — GPT-Realtime leads all three at 87.6% / 68.0% / 60.0%.
The Cold Water: Delegation Decisions Are Not There Yet
The most valuable section may be the in-house Delegate Benchmark. The Audio model achieves 92.22% delegation recall on external capabilities but only 39.44% non-delegation specificity on routine interaction — six out of ten routine questions get shoved to the backend; the Omni model is the mirror image: 84.44% specificity but 68.33% recall. Each model limps on one leg, with overall routing accuracy of 75.93% vs 68.89%. The benchmark also only evaluates "should this be delegated", not whether delegated tasks execute and integrate correctly — the paper acknowledges that needs further evaluation. Offline baselines like Gemini-3.5-Flash (77.3/71.6/56.5) score higher, but they are reference points without real-time constraints.
So the significance of this paper is not "SOTA" but a complete engineering answer to the "conversation is conversation, computation is computation" route: a 9B frontend plus heterogeneous backend compute plus explicit scheduling semantics. OpenAI's GPT-Live split the dialogue layer from the reasoning layer in July and shipped it in the API this month; Ant's work shows the same split can land at open 9B scale — except that "whether to hand the work over" remains, for now, a 75-point skill.
References
- Paper: arXiv:2609.13814 (https://arxiv.org/abs/2609.13814)
- Model weights: https://huggingface.co/inclusionAI/Realtime-Venus
- Project page: https://realtime-venus.github.io/