Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3ABKCY2XIR" target="_blank" >RIV/00216208:11320/26:BKCY2XIR - isvavai.cz</a>
Result on the web
<a href="https://aclanthology.org/2025.acl-long.123/" target="_blank" >https://aclanthology.org/2025.acl-long.123/</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.18653/v1/2025.acl-long.123" target="_blank" >10.18653/v1/2025.acl-long.123</a>
Alternative languages
Result language
angličtina
Original language name
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
Original language description
Recent English Common Crawl datasets like FineWeb-Edu and DCLM achieved significant benchmark gains via aggressive model-based filtering, but at the cost of removing 90% of data. This limits their suitability for long token horizon training, such as 15T tokens for Llama 3.1. In this paper, we show how to achieve better trade-offs between accuracy and data quantity by a combination of classifier ensembling, synthetic data rephrasing, and reduced reliance on heuristic filters. When training 8B parameter models for 1T tokens, using a high-quality subset of our data improves MMLU by 5.6 over DCLM, demonstrating the efficacy of our methods for boosting accuracies over a relatively short token horizon. Furthermore, our full 6.3T token dataset matches DCLM on MMLU, but contains four times more unique real tokens than DCLM. This unlocks state-of-the-art training over a long token horizon: an 8B parameter model trained for 15T tokens, of which 7.2T came from our dataset, is better than the Llama 3.1 8B model: +5 on MMLU, +3.1 on ARC-Challenge, and +0.5 on average across ten diverse tasks. The dataset is available at https://data.commoncrawl.org/contrib/Nemotron/Nemotron-CC/index.html.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
—
Continuities
—
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
ISBN
979-8-89176-251-0
ISSN
—
e-ISSN
—
Number of pages
17
Pages from-to
2459-2475
Publisher name
Association for Computational Linguistics
Place of publication
—
Event location
Vienna, Austria
Event date
Jan 1, 2026
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—