#commonCrawl
Going back to the early days of... 2023... we see chats between employees in using CommonCrawl for "swept up" copyright material and maybe a better way to get LibGen rather than loading it through torrents (especially not on Meta's IP addresses!). /6
January 14, 2025 at 1:43 PM
Don't mind me, just downloading 1 million random web pages from the CommonCrawl dataset for my next blog post.
August 28, 2026 at 5:08 PM
I'll finetune my own model on the commoncrawl data but strip out any instance of Collatz
May 27, 2026 at 4:12 PM
RECOX

Free, simple and quick online tool to find subdomains and endpoints.

Data sources: HackerTarget, URLScanIO, RapidDNS, CertSpotter, JLDC, DNSRepo, crtSH, WayBack, CommonCrawl, AlienWault OTX

recox.hackerz.space
March 19, 2026 at 9:29 AM
Interesting: "Crowd-sourced lists of urls to help Common Crawl crawl under-resourced languages" over at github.com/commoncrawl/...
GitHub - commoncrawl/web-languages: Crowd-sourced lists of urls to help Common Crawl crawl under-resourced languages. See https://github.com/commoncrawl/web-languages-code/ for the code
Crowd-sourced lists of urls to help Common Crawl crawl under-resourced languages. See https://github.com/commoncrawl/web-languages-code/ for the code - commoncrawl/web-languages
github.com
November 30, 2024 at 10:20 AM
CommonCrawlは小規模な非営利団体ですが、無料で利用できる最大のWebクロールデータ ソースとして、生成AIの開発に重要な役割を果たしています。反面多くのモデルが偏った有害なデータや著作権で保護された素材でトレーニングされることにもつながりました。
LAION-5BもCommonCrawlのデータで製作されており、そもそもCSAMはそこに入っていたわけです。
foundation.mozilla.org/en/blog/Mozi...
Mozilla Report: How Common Crawl’s Data Infrastructure Shaped the Battle Royale over Generative AI
Mozilla investigates Common Crawl’s influence as a backbone for Large Language Models: its shortcomings, benefits, and implications for trustworthy AI
foundation.mozilla.org
April 22, 2025 at 3:44 AM
🤔 Ever wondered how prevalent some type of web content is during LM pre-training?

In our new paper, we propose WebOrganizer which *constructs domains* based on the topic and format of CommonCrawl web pages 🌐

Key takeaway: domains help us curate better pre-training data! 🧵/N
February 18, 2025 at 12:31 PM
They are not trained on “stolen info.” The training corpuses are a combination of public data (eg CommonCrawl that’s been used in academia forever) and libraries of books that they paid for and scanned to use internally.

Some early groups torrented small collections, but that was years ago and
May 18, 2026 at 1:06 PM
URLFINDER

#go URL discovery tool:
- different sources (alienvault,commoncrawl etc)
- filter by extensions/regex
- very fast (122000+ URLs in 30 sec):
github.com/projectdisco...

Creator x.com/pdnuclei

#osint #cybersecurity #bugbounty
April 10, 2025 at 4:32 AM
We're very happy to release cc-downloader, a new CLI tool to download Common Crawl data 📚🚀🧑‍💻

‍cc-downloader is still under active development, so if you find any issues or would like to submit a feature request, please visit its GitHub repository at github.com/commoncrawl/....
January 21, 2025 at 11:57 PM
1/3 The genesis web annotation in commoncrawl looks useful only for 5 topics from terabytes or i am mistaken?
While i haven't dug harder yet. Has any one managed to use the genesis wen annotation in commoncrawl meaningfully?
#commoncrawl #crawler #web
April 18, 2026 at 3:01 PM
I hear more people / publishers they don't want their data being used for LLM training; luckily you can check if you are potentially alreay being used by querying the commoncrawl dataset; index.commoncrawl.org And you can block their crawler.
Common Crawl Index Server
index.commoncrawl.org
November 29, 2024 at 9:45 PM
[1/6] 🚨 We’re releasing LAION-BVD: a 10-million-hour open video dataset for multimodal pre-training.

- 1.3B video URLs from CommonCrawl
- 80M downloaded videos
- 10M video hours
- 55M captioned clips
- 300M frame-caption pairs

🌐: projects.laion.ai/bvd/

🧵👇
August 26, 2026 at 4:21 PM
does anyone sell copies of commoncrawl on ultrum tapes? downloading it takes forever and you probably have to buy more storage to store it anyway (it's on order 100TB)
September 30, 2026 at 12:44 AM
Meta used my book chapter to train its language model. Google, OpenAI, and others used CommonCrawl to scrape my dissertation research (link to check in thread).

Authors, find out if your work has been pirated here:

www.theatlantic.com/technology/a...
Search LibGen, the Pirated-Books Database That Meta Used to Train AI
Millions of books and scientific papers are captured in the collection’s current iteration.
www.theatlantic.com
March 20, 2025 at 3:18 PM
Nothing is stolen, and most deployed ML models are trained on similar datasets (like commoncrawl)
December 1, 2025 at 11:08 PM
C5: Web-Scale Creative Commons Web Data

- Effort to collect CC-licensed web data in one place
- 147M documents explicitly CC-licensed
- License detection from HTML links & metadata
- Preserves exact license type & version
- Maps to original URLS & crawl dates

huggingface.co/datasets/Bra...
BramVanroy/CommonCrawl-CreativeCommons · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
May 6, 2025 at 9:03 AM
The May 2026 crawl archive (CC-MAIN-2026-21) is now also available on our HF bucket. 🤗

huggingface.co/buckets/comm...
commoncrawl/commoncrawl - Storage Bucket
Storage bucket commoncrawl/commoncrawl on Hugging Face
huggingface.co
June 29, 2026 at 11:56 PM
This contamination is not intentional: we identified websites that reframed splits of MMLU as user-friendly quizzes
These websites can then be found in CommonCrawl dumps that are generally used for pretraining data curation...
November 7, 2025 at 9:11 PM
the SIO report (at least the abstract which is all i have read so far) says they just aimed photodna at the commoncrawl corpus, which like, why didn't google do that

(or indeed why didn't commoncrawl do that) purl.stanford.edu/kh752sm9123
Identifying and Eliminating CSAM in Generative ML Training Data and Models
Generative Machine Learning models have been well documented as being able to produce explicit adult content, including child sexual abuse material (CSAM) as well as to alter benign imagery of a cl...
purl.stanford.edu
December 20, 2023 at 4:09 PM
i mean, there is already commoncrawl, but like yeah, commoncrawl with moar
June 12, 2025 at 6:00 PM
There's really only one ur-dataset powering all this: CommonCrawl.

A little Dutch nonprofit run by a family trust that's been crawling and snapshotting the public web since 2008 and making the results freely available publicly for research.
December 30, 2025 at 6:42 AM
#BradParscale runs 11 websites for #Israel under a $46.5M contract.
#DropSiteNews found that between Jan & Jun, 10 of these websites were crawled 912 times by #CommonCrawl.
Common Crawl is the backbone of the #Al industry-a massive repository used to train large language models.

#AUSPOL #Palestine
August 2, 2026 at 1:16 AM
Anyone know if we can search CommonCrawl the way we search the Way-back machine? I just read an article mentioning CommonCrawl, and supposedly it has been scraping the internet forever? I lost my backups to my circa 2010 web pages and would kill to get them back (esp. hundreds of guitar photos).
Common Crawl - Get Started
Dive into Common Crawl: your guide to accessing vast web data. Start here to harness the web's potential effortlessly.
commoncrawl.org
April 20, 2025 at 10:55 PM