Your model speaks your language. Does it get it right?
We test LLMs across datasets and languages, then publish the scores, prompts and answers. Including the wrong ones.
How the models did, language by language
Missing a model?Or a language?Run your own test.
We haven’t tested everything. Pick your models, then a dataset and the languages it supports. We’ll write the command. You run the experiment.
Build a runSet it up here. Run it in your terminal.