Spanish
EspañolYour shortlist. Side by side.
Accuracy by benchmark
Each benchmark measures a different set of questions.
Ling 3.0 Flash · High reasoningMercury 2.5 Preview · High reasoninggpt-oss-120b · High reasoning
Accuracy · 0–100%
The same models, in other languages
Follow a column to see how a model’s accuracy changes with language.
| Language | Ling 3.0 FlashHigh reasoning | Mercury 2.5 PreviewHigh reasoning | gpt-oss-120bHigh reasoning |
|---|---|---|---|
| SpanishThis language | 96.0% | 96.0% | 92.0% |
| Albanian | 92.0% | — | — |
| Chinese | 92.0% | 88.0% | 88.0% |
| Czech | 88.0% | — | — |
| English | 96.0% | 96.0% | 100.0% |
| French | 92.0% | — | — |
| German | 92.0% | — | — |
| Japanese | 88.0% | — | — |
| Kazakh | 32.0% | 92.0% | 88.0% |
| Korean | 92.0% | — | — |
| Russian | 92.0% | 88.0% | 96.0% |
| Serbian | 44.0% | — | — |
| Slovak | 88.0% | — | — |
| Swedish | 96.0% | — | — |
Spanish · Accuracy is the percentage of correct answers. Each available benchmark has equal weight in the overall score.