←── back to feed
/topics/cross-lingual-and-multilingual-llm-benchmarks

Cross-lingual and multilingual LLM benchmarks

3 items1 sourcesupdated 27d agotrend 0

Three new benchmarks address gaps in multilingual and cross-lingual language model evaluation: Wazobia Eval for Nigerian Pidgin emotion and cultural understanding, LützCross for cross-lingual retrieval over Luxembourgish documents, and Bulbul for multi-dialect Arabic speech recognition. Together they expand LLM testing beyond high-resource languages to underrepresented African and low-resource European languages.

  • Wazobia Eval contains 550+ manually annotated Nigerian Pidgin examples covering emotion, sarcasm, and cultural reasoning across 16 categories
  • LützCross enables cross-lingual page-level retrieval in English, French, German, and Luxembourgish over visually rich PDF documents
  • Bulbul dataset spans 275 speakers across 11 Arab countries with structured coverage of dialects, sub-dialects, classical, and modern standard Arabic
  • LützCross combines text-focused and visually grounded QA pairs for PDF-based retrieval-augmented generation tasks
  • Bulbul addresses diglossia and regional dialect variation challenges in Arabic automatic speech recognition