Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

ICLROral2026

Authors
Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas Hinostroza, Mattia Nee, Eliot Krzysztof Jones, Irène Girard, David Mach, Anastasia Stasenko, Ivan P. Yamshchikov
Affiliation
Pleias
Venue
ICLR 2026
Track
Oral

TL;DR

We assemble and release the largest truly open multilingual dataset for LLM pre-training consisting of 2 trillion tokens…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

pre-training dataset llm

← All ICLR 2026 Oral papers · Browse the whole archive