Russian
РусскийYour shortlist. Side by side.
Accuracy by benchmark
Each benchmark measures a different set of questions.
gpt-oss-120b · High reasoningQwen3.7 Flash · High reasoningQwen3.7 Flash · Low reasoning
Accuracy · 0–100%
The same models, in other languages
Follow a column to see how a model’s accuracy changes with language.
| Language | gpt-oss-120bHigh reasoning | Qwen3.7 FlashHigh reasoning | Qwen3.7 FlashLow reasoning |
|---|---|---|---|
| RussianThis language | 96.0% | 96.0% | 96.0% |
| Albanian | — | — | — |
| Chinese | 88.0% | 100.0% | 100.0% |
| Czech | — | — | — |
| English | 100.0% | 96.0% | 96.0% |
| French | — | — | — |
| German | — | — | — |
| Japanese | — | — | — |
| Kazakh | 88.0% | 92.0% | 84.0% |
| Korean | — | — | — |
| Serbian | — | — | — |
| Slovak | — | — | — |
| Spanish | 92.0% | 92.0% | 92.0% |
| Swedish | — | — | — |
Russian · Accuracy is the percentage of correct answers. Each available benchmark has equal weight in the overall score.