# Two New Benchmarks Test Languages the Frontier Evaluation Suites Skip Entirely

One targets emotion, sarcasm and cultural reasoning in Nigerian Pidgin, the other retrieval in Khmer, where word boundaries are ambiguous and multilingual embeddings perform poorly.

- Published: 2026-08-25T06:26:21.272Z
- Canonical: https://polylog.news/ai/2026-08-25/two-new-benchmarks-test-languages-the-frontier-evaluation-su
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CL (Wazobia Eval)](https://arxiv.org/abs/2608.21369), [arXiv cs.CL (KSE-Web)](https://arxiv.org/abs/2608.21365)

Model capability claims are reported almost entirely on English benchmarks, and two preprints posted to arXiv today measure what that leaves out. Wazobia Eval targets Nigerian Pidgin, which the authors describe as one of Africa's most widel…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-25/two-new-benchmarks-test-languages-the-frontier-evaluation-su (subscription information: https://polylog.news/pricing).