Common Crawl

commoncrawl.org

Visit Website

Free, open repository of petabytes of web crawl data, updated monthly — the raw material behind many large language models.

Why it is useful

Building a large-scale web dataset from scratch is prohibitively expensive for almost anyone. Common Crawl's free, regularly updated corpus is the foundation many NLP and LLM projects are trained or evaluated on.