Skip to content

Spanish

Español
Joint highest accuracy
Ling 3.0 FlashinclusionAI · High reasoning
96.0%

2 results share the top score

A close alternative
Mercury 2.5 PreviewInception · High reasoning
96.0%

Matches the highest score

How much scores vary
Lowest score80.0%Highest score96.0%
16.0pp

Between the highest and lowest of 11 published results.

How the models compare

Published results, ranked by accuracy. Higher is better.

Compare side by side
Spanish · How the models compare
#ModelAccuracyGap to top
1Ling 3.0 FlashHigh reasoning
Ling 3.0 Flash: 96.0%96.0%
Top score
1Mercury 2.5 PreviewHigh reasoning
Mercury 2.5 Preview: 96.0%96.0%
Top score
2gpt-oss-120bHigh reasoning
gpt-oss-120b: 92.0%92.0%
-4.0 pp
2gpt-oss-120bLow reasoning
gpt-oss-120b: 92.0%92.0%
-4.0 pp
2Ling 3.0 FlashMedium reasoning
Ling 3.0 Flash: 92.0%92.0%
-4.0 pp
2Qwen3.7 FlashHigh reasoning
Qwen3.7 Flash: 92.0%92.0%
-4.0 pp
2Qwen3.7 FlashLow reasoning
Qwen3.7 Flash: 92.0%92.0%
-4.0 pp
3Ling 3.0 FlashLow reasoning
Ling 3.0 Flash: 88.0%88.0%
-8.0 pp
3Solar Pro 4High reasoning
Solar Pro 4: 88.0%88.0%
-8.0 pp
4Mercury 2.5 PreviewLow reasoning
Mercury 2.5 Preview: 80.0%80.0%
-16.0 pp
4Solar Pro 4Low reasoning
Solar Pro 4: 80.0%80.0%
-16.0 pp

The numbers, with the working attached.

Tests, prompts, settings and model responses live on the release pages.

View published releases