MMTEB: Massive Multilingual Text Embedding Benchmark

Exploring foci of: arXiv (Cornell University) MMTEB: Massive Multilingual Text Embedding Benchmark February 2025 • Kenneth Enevoldsen, Isaac Chung, Imene Kerboua, Márton Kardos, Ashwin Mathur, David Stap, Jay Gala, Wissam Siblini, Dominik Krzemiński, Genta Indra W… Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more comprehensive evaluation, we introduce the Massive Multilingual Text Embedding Benchmark (MMTEB) - a large-scale, community-driven expansion of MTEB, covering over 500 quality-controlled evaluation tasks across 250+ languages. MMTEB includes a diverse set of challenging, novel tasks such as instruction following, long-document retrieval, a… Open Article Page

Benchmark (Surveying) Computer Science Artificial Intelligence Geography Cartography Open Article