Benchmark release mmpisa-ling-3-flash-ja-es-2026-09-08
Model accuracy on mmpisa, using multiple-choice-v1. This release records the tested languages, scores, explicit language comparisons and downloadable evidence.
Scores and uncertainty intervals
These scores cover academic multiple-choice questions. They can’t tell you how well a model handles every task in a language. If the gap’s interval includes zero, we can’t call the direction of the difference. The same effort label can also mean different computing budgets across APIs.
How we got these numbers| Model | Reasoning effort | Japanese | Spanish | JA − ES · Gap · pp | JA − ES · 95% interval |
|---|---|---|---|---|---|
| Ling 3.0 Flash inclusionAI | High | 88.0% | 96.0% | -8.0 | [-24.0, +8.0] |
| Ling 3.0 Flash inclusionAI | Low | 96.0% | 88.0% | +8.0 | [0.0, +20.0] |
| Ling 3.0 Flash inclusionAI | Medium | 96.0% | 92.0% | +4.0 | [-8.0, +16.0] |
Accuracy by repeat
Ling 3.0 Flash · High
Japanese
25 unique questions per language · 1 repeat
- Repeat 1: 88.0%
Spanish
25 unique questions per language · 1 repeat
- Repeat 1: 96.0%
0 refusals · 0 unparseable answers
Selected completed responses: cost unknown. Total run spending is recorded in execution.json.
Ling 3.0 Flash · Low
Japanese
25 unique questions per language · 1 repeat
- Repeat 1: 96.0%
Spanish
25 unique questions per language · 1 repeat
- Repeat 1: 88.0%
0 refusals · 1 unparseable answers
Selected completed responses: cost unknown. Total run spending is recorded in execution.json.
Ling 3.0 Flash · Medium
Japanese
25 unique questions per language · 1 repeat
- Repeat 1: 96.0%
Spanish
25 unique questions per language · 1 repeat
- Repeat 1: 92.0%
0 refusals · 0 unparseable answers
Selected completed responses: cost unknown. Total run spending is recorded in execution.json.
Experiment settings
Protocol: multiple-choice-v1
Dataset version: 22b1caa65980ca1fa295acc31ee5e35200e0c8c0
Download
You can check the math without paying to run the models again. Download all the release files, then use the benchmark CLI to check their checksums and recompute the scores. No model calls or API keys needed.
- manifest.json (External website)
- resolved.json (External website)
- identity.json (External website)
- dataset.jsonl (External website)
- dataset-manifest.json (External website)
- items.jsonl (External website)
- analysis.json (External website)
- aggregate.json (External website)
- aggregate.csv (External website)
- attempts.jsonl (External website)
- execution.json (External website)
- ATTRIBUTION.md (External website)
File checksums · SHA-256
- resolved.json
- 95744556328f52f5ef573ed7e2658f7609a22e0657fe885ce88cdb552f62fb33
- identity.json
- 4bfcc794d0f3cd66e960887d7d84c7f35fc2f30ff8b81a80ab78529c8ce5e893
- dataset.jsonl
- 1f73a4f3c0f8ea04521d1590de49de9c04a1deab34e48defeafbd1a0cfa91df5
- dataset-manifest.json
- 81c75ee0967e1ce708e83958d9f53d880de4735d0036144ac3b9463372a5e0ce
- items.jsonl
- 96f2dc2edf537378fdb82a078a10d15a77ca66e5fdddfc1ae44716e61943abb3
- analysis.json
- 59e4b3a49da7b81d5e25876ffaf0092407fccda38df6745298cdf97993461722
- aggregate.json
- 878e9eae154b7586b3932d6f14f029079df142cd40eef5ed1f66091b912f58d7
- aggregate.csv
- a9de96f5b0c765978e14100d63eab2eb2f8464973c5a12d4f5024ca59999d20d
- attempts.jsonl
- 07161e536d004ddb778f23608dc6562ff183dd2d80edc3707d4b8464648b7dab
- execution.json
- be1edb0bd8b6515d09f5ec3ed50bdd35882f29ded8f560cce9c0a2f90dd11472
- ATTRIBUTION.md
- 6993d7ef4b5780ffd7418287354c3796fc5ab02f15fb4a3e0a039fc43f9222a9
Cite this release
Copy the citation for your paper or report, or export it as BibTeX.
BibTeX
@misc{llang-gap-mmpisa-ling-3-flash-ja-es-2026-09-08,
author = {{Limit 115}},
title = {{Lang Gap: mmpisa}},
year = {2026},
howpublished = {Benchmark release},
note = {Release mmpisa-ling-3-flash-ja-es-2026-09-08. Created 2026-09-08. Protocol: multiple-choice-v1.},
url = {https://llang-gap-web.vercel.app/en/releases/mmpisa-ling-3-flash-ja-es-2026-09-08/}
}