Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
ICLROral2026
TL;DR
We assemble and release the largest truly open multilingual dataset for LLM pre-training consisting of 2 trillion tokens…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
pre-training dataset llm