Skip to content
Back to release history

Benchmark release mmpisa-ling-3-flash-12lang-2026-09-08

Model accuracy on mmpisa, using multiple-choice-v1. This release records the tested languages, scores, explicit language comparisons and downloadable evidence.

Scores and uncertainty intervals

These scores cover academic multiple-choice questions. They can’t tell you how well a model handles every task in a language. If the gap’s interval includes zero, we can’t call the direction of the difference. The same effort label can also mean different computing budgets across APIs.

How we got these numbers
Scores and uncertainty intervals · mmpisa-ling-3-flash-12lang-2026-09-08
ModelReasoning effortAlbanianCzechChineseEnglishGermanFrenchKoreanKazakhRussianSerbianSlovakSwedish
Ling 3.0 Flash
inclusionAI
High92.0%88.0%92.0%96.0%92.0%92.0%92.0%32.0%92.0%44.0%88.0%96.0%
Ling 3.0 Flash
inclusionAI
Low96.0%92.0%92.0%96.0%92.0%96.0%88.0%36.0%92.0%32.0%92.0%84.0%

Accuracy by repeat

Ling 3.0 Flash · High

Albanian

25 unique questions per language · 1 repeat

  • Repeat 1: 92.0%

Czech

25 unique questions per language · 1 repeat

  • Repeat 1: 88.0%

Chinese

25 unique questions per language · 1 repeat

  • Repeat 1: 92.0%

English

25 unique questions per language · 1 repeat

  • Repeat 1: 96.0%

German

25 unique questions per language · 1 repeat

  • Repeat 1: 92.0%

French

25 unique questions per language · 1 repeat

  • Repeat 1: 92.0%

Korean

25 unique questions per language · 1 repeat

  • Repeat 1: 92.0%

Kazakh

25 unique questions per language · 1 repeat

  • Repeat 1: 32.0%

Russian

25 unique questions per language · 1 repeat

  • Repeat 1: 92.0%

Serbian

25 unique questions per language · 1 repeat

  • Repeat 1: 44.0%

Slovak

25 unique questions per language · 1 repeat

  • Repeat 1: 88.0%

Swedish

25 unique questions per language · 1 repeat

  • Repeat 1: 96.0%

0 refusals · 35 unparseable answers

Selected completed responses: cost unknown. Total run spending is recorded in execution.json.

Ling 3.0 Flash · Low

Albanian

25 unique questions per language · 1 repeat

  • Repeat 1: 96.0%

Czech

25 unique questions per language · 1 repeat

  • Repeat 1: 92.0%

Chinese

25 unique questions per language · 1 repeat

  • Repeat 1: 92.0%

English

25 unique questions per language · 1 repeat

  • Repeat 1: 96.0%

German

25 unique questions per language · 1 repeat

  • Repeat 1: 92.0%

French

25 unique questions per language · 1 repeat

  • Repeat 1: 96.0%

Korean

25 unique questions per language · 1 repeat

  • Repeat 1: 88.0%

Kazakh

25 unique questions per language · 1 repeat

  • Repeat 1: 36.0%

Russian

25 unique questions per language · 1 repeat

  • Repeat 1: 92.0%

Serbian

25 unique questions per language · 1 repeat

  • Repeat 1: 32.0%

Slovak

25 unique questions per language · 1 repeat

  • Repeat 1: 92.0%

Swedish

25 unique questions per language · 1 repeat

  • Repeat 1: 84.0%

0 refusals · 34 unparseable answers

Selected completed responses: cost unknown. Total run spending is recorded in execution.json.

Experiment settings

Protocol: multiple-choice-v1

Dataset version: 22b1caa65980ca1fa295acc31ee5e35200e0c8c0

Download

You can check the math without paying to run the models again. Download all the release files, then use the benchmark CLI to check their checksums and recompute the scores. No model calls or API keys needed.

File checksums · SHA-256

resolved.json
b6dce6dbad551d661b340463b7ae15d86affb776fda70f344648c203cf0531ab
identity.json
7476cfe95c36396dd8d2fef3584d16af6c2ad8af2a886b49351006f0c7ddc0a8
dataset.jsonl
29f03b5906cc233e4b35508a18b39acedeefc3f6da3a04a411a3ce2b5f3e81b3
dataset-manifest.json
81c75ee0967e1ce708e83958d9f53d880de4735d0036144ac3b9463372a5e0ce
items.jsonl
832b3050ab2a19e9b54fe9eaf23fee06f4b35b27639cbf36dc87866f2cc05a90
analysis.json
752f38e6242843537041ad21fcd2f1a96d7fe919112212adae2bb8edd5a10a13
aggregate.json
88af62800d8f2e1f1f9f474fc9b238eaef0d0b58010c7c41d68bf23500b88f00
aggregate.csv
2129162183931b783caf526f3da4a0d8135ab467ab43d701c877bdc3671a5a10
attempts.jsonl
e8fb319841b52a6332df76762995c5d9e2f136bde6d4db9548ceb9f86d5abb0f
execution.json
d031b2dbbb3d300751211325e45305bb8993e098e11d06d61c5b2208b79b9c5e
ATTRIBUTION.md
6993d7ef4b5780ffd7418287354c3796fc5ab02f15fb4a3e0a039fc43f9222a9

Cite this release

Copy the citation for your paper or report, or export it as BibTeX.

Limit 115. Lang Gap: mmpisa. Release mmpisa-ling-3-flash-12lang-2026-09-08, created September 8, 2026. Protocol: multiple-choice-v1. https://llang-gap-web.vercel.app/en/releases/mmpisa-ling-3-flash-12lang-2026-09-08/
BibTeX
@misc{llang-gap-mmpisa-ling-3-flash-12lang-2026-09-08,
  author = {{Limit 115}},
  title = {{Lang Gap: mmpisa}},
  year = {2026},
  howpublished = {Benchmark release},
  note = {Release mmpisa-ling-3-flash-12lang-2026-09-08. Created 2026-09-08. Protocol: multiple-choice-v1.},
  url = {https://llang-gap-web.vercel.app/en/releases/mmpisa-ling-3-flash-12lang-2026-09-08/}
}
Download .bib (Downloads a file)