LILT

We build the multilingual layer for English-first AI. Custom evals, benchmarks, and RL environments across 200+ languages.

Most agent and coding benchmarks ship in English. We build non-English counterparts grounded in language and culture, along with the multilingual environments models train on, so labs and enterprises can measure and improve how their models perform in the languages their users speak.

New: AURORA, the multilingual AI leaderboard

AURORA ranks frontier models on non-English agentic tasks grounded in language and culture. Every task is built or verified by native-language domain experts, so results reflect how models perform for real users, not how well they handle a translation. Benchmarks at launch, with more coming:

Harness, trial counts, reasoning settings and confidence intervals are published for every run. See the AURORA collection below.

Why we publish here

Open releases make it easier for the community to stress-test our work, reproduce our scores, and extend our benchmarks to new languages. We publish each benchmark with its methodology and explicit limitations.

What you'll find here

Links

Citation

If you use one of our datasets or benchmarks, please cite the paper linked on its dataset card or article.