# Evaluation Coverage Expands Beyond English

As AI systems deploy into markets their training data barely covers, purpose-built evaluations in low-resource languages keep exposing capability gaps that aggregate multilingual scores conceal, pushing per-language evidence into procurement requirements.

- Conviction: 36 / 100 (weakening)
- Horizon: Emerging (watchlist)
- Tracking since: 2026-08-25T00:00:00.000Z
- Last updated: 2026-08-27T14:00:31.454Z
- Canonical: https://polylog.news/ai/trends/eval-coverage-beyond-english
- Publisher: Polylog
- Affected regions: Global

## Recent score history

- 2026-08-27: 36
- 2026-08-28: 34

## Recent evidence

- [confirms] Two New Benchmarks Test Languages the Frontier Evaluation Suites Skip Entirely (2026-08-25): Two new benchmarks target languages frontier suites skip: one on emotion, sarcasm and cultural reasoning in Nigerian Pidgin, the other on retrieval in Khmer, where ambiguous word boundaries defeat multilingual embeddings. Both probe failure modes — pragmatics and tokenization — that aggregate multilingual scores average away entirely.
