Skip to content
Back to release history

Benchmark release mmpisa-ling-3-flash-ja-es-2026-09-08

Model accuracy on mmpisa, using multiple-choice-v1. This release records the tested languages, scores, explicit language comparisons and downloadable evidence.

Scores and uncertainty intervals

These scores cover academic multiple-choice questions. They can’t tell you how well a model handles every task in a language. If the gap’s interval includes zero, we can’t call the direction of the difference. The same effort label can also mean different computing budgets across APIs.

How we got these numbers
Scores and uncertainty intervals · mmpisa-ling-3-flash-ja-es-2026-09-08
ModelReasoning effortJapaneseSpanishJA − ES · Gap · ppJA − ES · 95% interval
Ling 3.0 Flash
inclusionAI
High88.0%96.0%-8.0[-24.0, +8.0]
Ling 3.0 Flash
inclusionAI
Low96.0%88.0%+8.0[0.0, +20.0]
Ling 3.0 Flash
inclusionAI
Medium96.0%92.0%+4.0[-8.0, +16.0]

Accuracy by repeat

Ling 3.0 Flash · High

Japanese

25 unique questions per language · 1 repeat

  • Repeat 1: 88.0%

Spanish

25 unique questions per language · 1 repeat

  • Repeat 1: 96.0%

0 refusals · 0 unparseable answers

Selected completed responses: cost unknown. Total run spending is recorded in execution.json.

Ling 3.0 Flash · Low

Japanese

25 unique questions per language · 1 repeat

  • Repeat 1: 96.0%

Spanish

25 unique questions per language · 1 repeat

  • Repeat 1: 88.0%

0 refusals · 1 unparseable answers

Selected completed responses: cost unknown. Total run spending is recorded in execution.json.

Ling 3.0 Flash · Medium

Japanese

25 unique questions per language · 1 repeat

  • Repeat 1: 96.0%

Spanish

25 unique questions per language · 1 repeat

  • Repeat 1: 92.0%

0 refusals · 0 unparseable answers

Selected completed responses: cost unknown. Total run spending is recorded in execution.json.

Experiment settings

Protocol: multiple-choice-v1

Dataset version: 22b1caa65980ca1fa295acc31ee5e35200e0c8c0

Download

You can check the math without paying to run the models again. Download all the release files, then use the benchmark CLI to check their checksums and recompute the scores. No model calls or API keys needed.

File checksums · SHA-256

resolved.json
95744556328f52f5ef573ed7e2658f7609a22e0657fe885ce88cdb552f62fb33
identity.json
4bfcc794d0f3cd66e960887d7d84c7f35fc2f30ff8b81a80ab78529c8ce5e893
dataset.jsonl
1f73a4f3c0f8ea04521d1590de49de9c04a1deab34e48defeafbd1a0cfa91df5
dataset-manifest.json
81c75ee0967e1ce708e83958d9f53d880de4735d0036144ac3b9463372a5e0ce
items.jsonl
96f2dc2edf537378fdb82a078a10d15a77ca66e5fdddfc1ae44716e61943abb3
analysis.json
59e4b3a49da7b81d5e25876ffaf0092407fccda38df6745298cdf97993461722
aggregate.json
878e9eae154b7586b3932d6f14f029079df142cd40eef5ed1f66091b912f58d7
aggregate.csv
a9de96f5b0c765978e14100d63eab2eb2f8464973c5a12d4f5024ca59999d20d
attempts.jsonl
07161e536d004ddb778f23608dc6562ff183dd2d80edc3707d4b8464648b7dab
execution.json
be1edb0bd8b6515d09f5ec3ed50bdd35882f29ded8f560cce9c0a2f90dd11472
ATTRIBUTION.md
6993d7ef4b5780ffd7418287354c3796fc5ab02f15fb4a3e0a039fc43f9222a9

Cite this release

Copy the citation for your paper or report, or export it as BibTeX.

Limit 115. Lang Gap: mmpisa. Release mmpisa-ling-3-flash-ja-es-2026-09-08, created September 8, 2026. Protocol: multiple-choice-v1. https://llang-gap-web.vercel.app/en/releases/mmpisa-ling-3-flash-ja-es-2026-09-08/
BibTeX
@misc{llang-gap-mmpisa-ling-3-flash-ja-es-2026-09-08,
  author = {{Limit 115}},
  title = {{Lang Gap: mmpisa}},
  year = {2026},
  howpublished = {Benchmark release},
  note = {Release mmpisa-ling-3-flash-ja-es-2026-09-08. Created 2026-09-08. Protocol: multiple-choice-v1.},
  url = {https://llang-gap-web.vercel.app/en/releases/mmpisa-ling-3-flash-ja-es-2026-09-08/}
}
Download .bib (Downloads a file)