2026-10-02 · 13 min · explainer · llm · speech · reinforcement-learning · benchmarks · architecture · open-weights · multimodal
Most "translation model" releases are one checkpoint and a language list. Index-Translate, from Bilibili's Index LLM team and released on 30 September 2026, is four things that share a spine: a text translator, an end-to-end speech translator, a dubbing model that translates to a target syllable count, and a model that translates a whole document in one pass. All of it starts from the same Qwen3.5 base, and the 2B and 9B weights for every head are on Hugging Face under Apache-2.0 — not gated, which I checked against the IndexTeam model list rather than take on faith.
A family is only worth the name if the pieces share something real and the numbers under each one hold up. So this is two questions at once. What does each head actually do, and how are they the same model? And when you read the headline figures — 150 languages, 0.8789 on FLORES, 81.92% of dubs inside their time budget — what exactly is being counted?
- architecture
- Qwen3_5ForConditionalGeneration
- task
- translation
- license
- apache-2.0
- safetensors
- 4 shards
- largest file
- 5.36 GB
- files
- 16
- downloads
- 191
- likes
- 29
repo last modified 2026-10-01
One base, four heads
The shared part is a text model, built once at three sizes — 2B, 9B, and a 35B-A3B mixture-of-experts preview. Every head is a post-training of it.
It starts from the Qwen3.5 base and runs a multilingual mid-training stage: 150 languages, 167.77B tokens (reported — the paper's recipe is 10,000 steps at 4,096 tokens per sequence and 4,096 sequences per batch, which multiplies out to that figure). The data is mixed on purpose — monolingual, ordinary parallel, and pivot-organised data, where a sentence in two languages shares one English anchor, so the model sees alignments it was never given directly.
Then three specialists are fine-tuned off that foundation, each with SFT followed by reinforcement learning tuned to its job:
- a General expert (GRPO against reference-based XCOMET-XXL plus target-language and adequacy judgments) for raw translation quality;
- an Instruction expert (GRPO with Rubric-as-Reward, a judge scoring compliance) for following constraints like "keep this hashtag, translate the rest";
- a Meme expert (judge-based RL with adversarial supervision) for community slang and cultural context.
The clever move is how they are recombined. Instead of training one model on everything, the three experts are merged by parameter interpolation — a model soup, weights 0.8 / 0.1 / 0.1 for General / Instruction / Meme on the evaluated 2B and 9B. A soup is literally the weighted average of the three weight tensors; it works because the three were fine-tuned from the same initialisation and sit in the same loss basin. Whatever the merge still does badly, a final pass of targeted MOPD (multi-teacher on-policy distillation) repairs: domain-matched teachers supervise the student's own rollouts with a reverse-KL objective, so the merged model is pulled back toward each specialist only on the task types where the average lost something.

That text model is the spine. Speech bolts an audio encoder onto its front. Dubbing re-trains it with a length reward. Long-document translation extends its context from 4K to 128K. Same decoder weights underneath, four output interfaces.
The FLORES yardstick, and what "150 languages" doesn't cover
The headline translation number is FLORES COMET-22 = 0.8789 for the 9B (reported). To read it you need to know what each half is.
FLORES-200 is a parallel evaluation set: the same sentences, professionally translated across many languages, so you can score any direction. COMET-22 (wmt22-comet-da) is a learned metric — a model trained on human quality ratings that scores a translation against source and reference on roughly a 0–1 scale. It is the field's default because it tracks human judgment better than BLEU, which counts n-gram overlap and punishes correct paraphrases. Higher is better; differences in the third decimal are small but real at this scale.
Here is where the 150 does not do what the sentence implies. The flagship FLORES table is not 150 languages. It is 126,000 inputs over 420 directions among 21 languages (300 sentences per direction, devtest) — reported, and stated plainly in the paper's Table 3. 420 is 21 × 20: every ordered pair of 21 high-resource languages. The 150 is the mid-training coverage and the instruction-following claim, not the set the headline COMET number is averaged over. Low-resource ability is measured separately, on FLORES_minor_pair: 104,000 inputs across 1,040 directions among 62 languages. Two different sets, two different stories, and only the first produces 0.8789.
With that straight, the comparison on the 21-language set:
Read it honestly. On raw FLORES the field is tight — a 9B open model at 0.8789, a dedicated 30B-A3B competitor (Hy-MT2) at 0.8787, and Index's own 35B-A3B preview at 0.8794 are within five ten-thousandths of each other. The gap that matters is against the base it was built from: Qwen3.5-9B scores 0.8316 cold, so the whole recipe — mid-training, specialists, soup, distillation — is worth about 0.047 COMET on this set (reasoned, from the two reported rows). COMET this close between finished systems is near the metric's noise floor; the family earns its keep not here but on instruction following and the low-resource set, where the spread is wide (the 9B's low-resource off-target rate is 4.0%, against 35.7% for Hy-MT2-7B — reported). FLORES is the yardstick everyone quotes; it is also the one where everyone good looks the same.
Length-controlled dubbing: translating to fit the mouth
The head I find most interesting is Index-Homura, because it solves a problem ordinary translation metrics don't even see.
Dub a line of video and the translation has to end when the speaker's mouth stops. A perfect translation that takes twice as long to say is useless for dubbing — it overruns the shot, the next line is late, and lip-sync falls apart. Natural languages differ in how many syllables carry the same meaning, so "translate faithfully" and "fit the original's timing" pull against each other constantly. What you want is a dial: translate this line into roughly N syllables, where N is set by how long the original took.
Index-Homura is that dial. It takes the text model and re-trains it with GRPO against a length reward — a term that pays the model for landing near the requested syllable count, on top of the usual quality signal. The syllable count goes in the prompt; the model learns to reword, not just truncate, to hit it. The README's demo makes the rewording visible on one source line at three budgets:
| Target / observed syllables | Index-Homura-9B output |
|---|---|
| 10 / 10 | to live for two days. What would that be like? |
| 14 / 14 | What would it be like to live there for two days, I wonder? |
| 18 / 18 | What would it be like to live there for two days, trying to get by somehow? |
Same meaning, three lengths, each hitting its target exactly. The widget below lets you pick the budget and see both bands SandGlass scores — within one syllable, and within 10% of target — redraw around it, and swap between the length-RL model and its SFT parent:
“What would it be like to live there for two days, I wonder?” observed 14, hits the budget
With the length reward on, the realized count lands within 10% of target 81.92% of the time and within one syllable 74.42% of the time — but translation quality falls to 0.7863. Flip to SFT: quality climbs to 0.8581 and the budget is met less than half the time. That gap (81.92% vs 45.75%) is the whole point of the RL head — and its cost, in COMET points, is right there beside it.
The aggregate is the honest part. On SandGlass (3,600 cases per model — 300 subtitle sentences × 4 target languages × 3 length budgets), Index-Homura-9B with the length reward lands within 10% of the target 81.92% of the time and within one syllable 74.42% of the time, with mean relative deviation 0.0693 (all reported). The same 9B without the length RL — the SFT checkpoint — hits within 10% only 45.75% of the time and within one syllable 41.39%, and when it misses it misses wider: mean relative deviation 0.1752, against the RL model's 0.0693 (reported). That is the win.
And here is the cost, which the paper prints right beside it: the length-RL model's translation quality is 0.7863, while the SFT checkpoint scores 0.8581 (reported). Teaching the model to obey a syllable budget made it a measurably worse translator. The trade is real and it is in one direction — you give up about 0.072 of quality to roughly double your in-budget rate (reasoned). For subtitles you read, keep the SFT model. For a dub that has to fit the shot, the RL model's control is the thing you were buying, and its quality cost is the price.

The baselines in that figure are the second half of the point. Strong general models — GPT-5.6-Sol, DeepSeek, Hy-MT2 — keep their translation quality but land in budget under a quarter of the time. Knowing how to translate is not knowing how to translate to a length; that is a trained skill, and it does not come free.
End-to-end speech: Index-Echo vs the cascade
The usual way to translate speech is a cascade: ASR transcribes the audio, a translator translates the text, a TTS speaks the result. It works, but every stage's errors compound, and the pipeline throws away everything in the audio that isn't words — who is speaking, their voice, the timing.
Index-Echo is the end-to-end alternative, and it reuses the text model directly. For speech-to-text (S2TT) it connects a Qwen3-Omni AuT audio encoder and a connector to an Index-Translate 2B or 9B decoder, and trains the whole thing end to end on speech translation — acoustic input in, target-language text out, no intermediate transcript as a bottleneck. On an in-house video set (140 windows × 7 directions — in-house, not a public benchmark, so read it as the publisher's number), Index-Echo-9B scores 0.857 on an MT judge, between Qwen3.8-Omni-Flash (0.887) and Gemini-3.1-Pro with thinking (0.849), and the 2B and 9B post the two lowest ASR and timestamp errors in the comparison (all reported).
Speech-to-speech (S2ST) is the part that earns "end-to-end". Instead of printing text and handing it to a separate TTS, Index-Echo replaces CosyVoice3's text-tokenizer interface with a ~30M-parameter Hidden2CV mapper that reads the translator's final hidden states and feeds the speech generator's semantic layer directly. The mapper is distilled first with both ends frozen, then tuned with DiffRO (differentiable reward optimization), which backpropagates a content-consistency reward through Gumbel-Softmax speech-token samples while the translation backbone stays frozen. The source audio also supplies a reference prompt and a CampPlus speaker embedding, so the output keeps the original speaker's voice — the thing a cascade discards at the ASR step.

The released S2ST package covers six directions — zh to en/es/ja and en to zh/es/ja. The paper's matched comparison against the cascade (Index-Echo-S2TT + CosyVoice3) is the honest read: end-to-end lowers mean content error in 8 of 12 size–direction pairs, with speaker similarity within 0.01–0.02 of the pipeline (reported). Not a rout — a modest, consistent win that also preserves the voice, which is the trade going end to end is supposed to buy. It does not win everywhere, and the paper says so.
Long documents in one pass: Index-NativeLong
The fourth head, Index-NativeLong, is published under the model IDs IndexTeam/Index-Nailong-2B and -9B — the "Nailong" in the weights is the ID, "NativeLong" the descriptive name. It extends training to entire documents, pushing sequence length from 4K to 128K during mid-training decay and post-training on whole books and full-length movie transcripts.
The point of a native long-document model is consistency a chunked pipeline can't hold. Translate a novel in 4K windows and the same character name drifts between renderings, because no window sees the others. The README's example is exactly this: across a ~32K-token fantasy document, the character "王妃" comes out as "Mentor Wang Fei", "Wang Fei", and "the Dean" from a chunked 9B workflow at three positions, while the native full-document 9B keeps the name stable throughout. On the document benchmarks (274 windows, document-SEGALE/COMET), Index-NativeLong-9B scores 0.7891 on GuoFeng, 0.7683 on BWB Track A3, and 0.8848 on Books (reported), ahead of the general-model baselines it is compared against.
What holds, and what I can't check
Every benchmark number here is reported — I read them from the paper, the repo's docs/evaluation.md, and the model cards, but I did not re-run an evaluation, so I am trusting the publisher's harness, judges (GPT-5.6-Sol and Gemini variants do a lot of the scoring), and test sets. Several of the strongest results — the S2TT and S2ST comparisons — are on in-house video sets, not public benchmarks, so there is no independent run to check them against yet. The instTrans, SandGlass-V2, nailong-bench and meme-bench sets are listed as not-yet-released. The 35B-A3B is explicitly a preview; the official version is on the TODO list.
What I can confirm: the weights are real and open. The 2B and 9B checkpoints for all four heads, plus the 35B-A3B preview, are on Hugging Face under IndexTeam, tagged Apache-2.0, not gated — the license is the one thing you don't have to take the paper's word for. And the structural claim — one base, four post-trainings — is consistent across the paper's training figures, the model cards' base-model tags (qwen3_5, qwen3_5_moe, qwen3_5_text), and the shared inference defaults in the repo.
The honest summary: Index-Translate is a well-built family with one genuinely novel member. The text model is competitive but sits in a crowded field where everyone good scores the same on FLORES. The dubbing head is the one that does something others don't — and it is also the one most upfront about what it costs, which is the part I trust most.
If you want the neighbours: Qwen3.8-LiveTranslate is the simultaneous-interpretation take on the same latency-vs-quality trade; Qwen-Audio-3.1 is the API-only audio family next door; speech-to-speech is the cascade Index-Echo is trying to collapse; and Qwen3.8-Omni-Flash is the omni model whose AuT encoder Index-Echo borrows for its front end.