Common Crawl
commoncrawl.org
Free, open repository of petabytes of web crawl data, updated monthly — the raw material behind many large language models.
Why it is useful
Building a large-scale web dataset from scratch is prohibitively expensive for almost anyone. Common Crawl's free, regularly updated corpus is the foundation many NLP and LLM projects are trained or evaluated on.