Machine translation, a task many declared "solved," is getting crowded again. After Tencent Hunyuan and Cohere open-sourced dedicated translation models, Bilibili's Index LLM team delivered its own answer on September 30: the Index-Translate family, with 2B, 9B, and 35B-A3B (preview) text models released under Apache-2.0 on both Hugging Face and ModelScope (model card).
A translation-specific MoE: 35B total, 3B active
The flagship Index-Translate-35B-A3B-preview is built on Qwen3.5: 35B total parameters with roughly 3B activated per token — a standard sparse MoE recipe. It covers text translation across 150 languages, ships with a 262,144-token context window (official examples serve at 32K), and self-hosts on vLLM.
The technical report describes a three-stage pipeline: first, multilingual mid-training on 167.77B tokens (general, monolingual, and parallel data at 1:1:1 in the constant stage, shifting to a 1:4:2 pivot organization in decay); then specialist SFT and RL for three experts — general translation, instruction following, and meme translation — combining XCOMET-XXL, language-validity, and Rubric-as-Reward signals; finally parameter interpolation merges complementary experts, with multi-teacher on-policy distillation (MOPD) patching remaining weak spots (report). The "divide experts, then merge" playbook now mirrors mainstream post-training for general LLMs.
Constraints as a hard spec
The most interesting part is instruction following. Translation constraints split into two tiers: hard constraints enforce terminology alignment, preserve JSON/CSV/markdown structure, and keep code blocks and variable placeholders intact; soft constraints handle style adaptation (formal, casual, meme-flavored) and domain disambiguation. For localization pipelines this is far more practical than "sounds more human": glossaries stay locked, formats don't collapse, variables don't vanish.
Reading the numbers
On the vendor's self-reported table, the 35B-A3B preview posts the highest FLORES COMET-22 (0.8794) and instTrans IFscore (0.8336) among all compared systems, with WMT26 Judge at 76.76. Against similarly sized peers: Tencent's Hy-MT2-30B-A3B scores 66.81 on WMT26 Judge, TranslateGemma-12B 71.19, and the Qwen3.5-35B-A3B base only 71.33 — the gains from translation-specific training are visible. On low-resource pairs the off-target rate is just 2.4%, versus 14.5% for Hy-MT2-30B-A3B.
The cold water: closed API flagships still lead — GPT-5.6-Sol scores 89.10 on WMT26 Judge, DeepSeek-V4.1-Flash 83.55. Open translation models win on self-hosting, zero call cost, and constraint control, not absolute quality.
Why Bilibili
The family also includes Index-Echo (speech-to-subtitles and speech-to-speech), Index-Homura (dubbing toward a target syllable count), and Index-NativeLong (full-document translation). The playbook points clearly at video subtitling, cross-language dubbing, and community content going global — all Bilibili's own business scenarios. On the MEME translation benchmark it scores 0.7405, clearly ahead of Hy-MT2-30B-A3B's 0.5812. Slang, abbreviations, and meme-laden UGC text are exactly where general models stumble, and Bilibili holds the densest corpus of it anywhere.
The takeaway: while general models race toward trillion-parameter scale, the moat for vertical tasks is shifting to "who has scenario data, and who grinds the constraint engineering fine." Bilibili's open-sourcing is less charity than converting its business moat into industry-standard parts.