We build the multilingual layer for English-first AI. Custom evals, benchmarks, and RL environments across 200+ languages.
Most agent and coding benchmarks ship in English. We build non-English counterparts grounded in language and culture, along with the multilingual environments models train on, so labs and enterprises can measure and improve how their models perform in the languages their users speak.
AURORA ranks frontier models on non-English agentic tasks grounded in language and culture. Every task is built or verified by native-language domain experts, so results reflect how models perform for real users, not how well they handle a translation. Benchmarks at launch, with more coming:
Harness, trial counts, reasoning settings and confidence intervals are published for every run. See the AURORA collection below.
Open releases make it easier for the community to stress-test our work, reproduce our scores, and extend our benchmarks to new languages. We publish each benchmark with its methodology and explicit limitations.
If you use one of our datasets or benchmarks, please cite the paper linked on its dataset card or article.