Technology

mmBERT

mmBERT is an open-source, massively multilingual encoder-only language model trained on 3 trillion tokens across 1,833 languages.

Developed by Johns Hopkins University, mmBERT updates the aging XLM-RoBERTa architecture by bringing modern transformer optimizations to encoder-only models (1.1.1, 1.2.6). Built on the high-performance ModernBERT architecture, it delivers 2 to 4 times faster inference speeds and natively supports an expanded 8,192-token context window (1.1.1, 1.2.4). The core innovation is its annealed language learning training strategy: a three-phase schedule that prevents overfitting on high-resource languages and ensures robust representation for low-resource languages (1.1.1, 1.1.7). This approach makes mmBERT a highly efficient, production-ready standard for multilingual classification, retrieval, and semantic search (1.1.1, 1.2.5).

https://github.com/JHU-CLSP/mmBERT

What builders pair with mmBERT

Projects using both technologies. Select a pairing to see a project.

Pairing: Cohere c4ai-command-a-03-2025

Photo from the event
Event photo

A Guardrail for the Hardest Conversations: (Bilingual) Youth Crisis Detection

Montreal · June 17, 2026

Recent Talks & Demos

Showing 1-1 of 1

Members-Only

Sign in to see who built these projects