{
  "claim": "The release links Omnilingua-MSpeaker as the evidence for its speaker-attribution and translation-quality wins. The public release of Omnilingua-MSpeaker cannot produce those numbers: it contains no translation references, no Chinese or English audio, and no quality metric. The benchmark that was cited and the benchmark that was released are not the same artifact.",
  "method": "Cloned github.com/QwenLM/Omnilingua-Bench at v202609 and computed every figure in the right-hand column directly from Omnilingua-MSpeaker/Omnilingua-MSpeaker_v202609_public.jsonl — 297 rows parsed, task and track counted, source/target language pairs tallied, media durations summed, source_url values de-duplicated, reference keys enumerated. The left-hand column is what the release post and its figure state.",
  "source": "https://github.com/QwenLM/Omnilingua-Bench",
  "captured": "2026-09-19",
  "note": "The 14 directions the figure claims fit one construction exactly: the six released source languages (es, fr, ja, ko, ru, th) each into English and Chinese, plus en→zh and zh→en. 6×2+2 = 14. So the evaluation looks like the public audio set re-annotated with translation references and extended with a zh/en pair — a reasonable thing to build, and none of it is in the repository. Separately, \"% of pairwise comparisons\" is a judge-scored preference eval, and the release does not say who or what the judge was.",
  "columns": [
    { "key": "field", "label": "" },
    { "key": "cited", "label": "as cited in the release post" },
    { "key": "released", "label": "as released in the repo (measured)" }
  ],
  "rows": [
    {
      "field": "task",
      "cited": "translation — faithfulness, fluency, conciseness + DER",
      "released": "speaker_attributed_asr, 297/297 rows"
    },
    {
      "field": "language directions",
      "cited": "14 directions over 8 languages (zh, en, es, fr, ja, ko, ru, th)",
      "released": "0 translation directions; every row has source == target"
    },
    {
      "field": "languages present",
      "cited": "8, including Chinese and English",
      "released": "6 — es (50), fr (47), ja (50), ko (50), ru (50), th (50). No zh, no en."
    },
    {
      "field": "metrics defined",
      "cited": "faithfulness, fluency, conciseness, DER",
      "released": "tcpWER, cpWER, DER — no quality metric of any kind"
    },
    {
      "field": "reference fields",
      "cited": "—",
      "released": "segments[] of {media_id, start_sec, end_sec, speaker_id, text, language}; 21,046 segments, no translation field"
    },
    {
      "field": "scale",
      "cited": "\"multi-speaker long-audio evaluation set\"",
      "released": "297 items, 105 distinct videos, 26.27 h; 2 speakers median, 20 maximum"
    },
    {
      "field": "media",
      "cited": "—",
      "released": "none shipped — source_url plus start/end offsets into public social-media video, CC BY-NC 4.0, research only"
    }
  ]
}
