NetEase Youdao recently officially released Ziyue LLM version 4.0, announcing the full open-sourcing of its core multimodal model and TTS engine, marking this education-AI veteran as formally stepping into the full-modal era.
The Ziyue 4 multimodal model elevates the math-and-science ability of vision inputs to the industry's top level at the 27B-parameter scale, performing impressively on tough visual math/physics problems that include charts and figures. Pure-text math problem accuracy reaches 81.4%, also industry-leading. More noteworthy, the new model adopts a refined chain-of-thought reconstruction scheme, compressing the reasoning CoT output length by 43.2% — meaning it can give answers with fewer Tokens and a shorter reasoning path, directly reducing inference cost in real business scenarios.
Also open-sourced in this release is the speech synthesis engine, based on a frontier speech-encoder + LLM architecture, supporting 14 languages, capable of completing original-voice cloning in 3 seconds with cloning accuracy above 97% and similarity above 85%. Cross-language cloning doesn't leak accent — a rare feature among domestic TTS open-source solutions.
The Ziyue team also rebuilt the translation model, introducing a multi-expert OPD mode paired with reinforcement-learning-based format reward and language detection mechanisms, improving quality while delivering 80% inference-speed boost. This is highly significant for scenarios requiring high-frequency, high-concurrency translation services.
From the original virtual-person speaking coach Hi Echo to today's Ziyue 4 full-modal open source, Youdao's accumulated expertise in education AI is translating into real open-source competitiveness. For developers and enterprises, this open-source package offers a directly deployable, high-cost-performance choice — the flexibility of open source combined with scene-validated performance guarantees. As the bar for multimodal and speech synthesis drops further, the real productivity revolution may be just around the corner.