←── back to feed
/topics/cross-lingual-and-multilingual-llm-benchmarks
Cross-lingual and multilingual LLM benchmarks
3 items●1 sources●updated 27d ago●trend 0
Three new benchmarks address gaps in multilingual and cross-lingual language model evaluation: Wazobia Eval for Nigerian Pidgin emotion and cultural understanding, LützCross for cross-lingual retrieval over Luxembourgish documents, and Bulbul for multi-dialect Arabic speech recognition. Together they expand LLM testing beyond high-resource languages to underrepresented African and low-resource European languages.
- Wazobia Eval contains 550+ manually annotated Nigerian Pidgin examples covering emotion, sarcasm, and cultural reasoning across 16 categories
- LützCross enables cross-lingual page-level retrieval in English, French, German, and Luxembourgish over visually rich PDF documents
- Bulbul dataset spans 275 speakers across 11 Arab countries with structured coverage of dialects, sub-dialects, classical, and modern standard Arabic
- LützCross combines text-focused and visually grounded QA pairs for PDF-based retrieval-augmented generation tasks
- Bulbul addresses diglossia and regional dialect variation challenges in Arabic automatic speech recognition