#commonCrawl
does anyone sell copies of commoncrawl on ultrum tapes? downloading it takes forever and you probably have to buy more storage to store it anyway (it's on order 100TB)
September 30, 2026 at 12:44 AM
Perso dès que j'ai su j'ai modifié mon robots.txt + case à cocher Cloudflare. Mais bon, Claude en web search viole tjs à ce jour mon robots.txt et mes CGU. CommonCrawl a scrapé mon site pendant des années sans rien me demander (illégal là aussi en droit français). J'ai écrit à CC, ils ont stoppé.
1/
September 15, 2026 at 9:14 PM
Good of Amazon to make themselves liable for AI-LLM copyright abuse through CommonCrawl. ;) aws.amazon.com/blogs/public...
How Common Crawl and AWS Open Data built the foundation for the AI revolution | Amazon Web Services
The Amazon Web Services (AWS) Open Data Sponsorship Program has hosted the Common Crawl open repository of web data at no cost since January 2012. It has become one of the most important sources of tr...
aws.amazon.com
September 14, 2026 at 9:15 AM
TIL the Common Crawl did a sampling of llms.txt in August! Nearly 600K sites!

https://huggingface.co/spaces/commoncrawl/llms.txt-report has info and https://huggingface.co/spaces/commoncrawl/llms.txt-report has some very reasonably sized parquet versions of the WARC files that I may hack on […]
Original post on mastodon.social
mastodon.social
September 5, 2026 at 2:39 AM
80 million videos pulled out of CommonCrawl, captioned for picture and sound. Research use only, and the copyright question sits right underneath.

courionai.com/news/2026-08-30-laion-big-video-dataset

#AI #TrainingData #OpenSource
September 1, 2026 at 9:02 AM
LAION just dropped a 10‑million‑hour video dataset—think massive video‑to‑text training data scraped from CommonCrawl. Perfect fuel for the next wave of video‑language models. Dive in to see what InternVid can unlock! #LAION #VideoDataset #VideoLanguageModels

🔗 aidailypost.com/news/laion-r...
August 29, 2026 at 10:14 AM
Don't mind me, just downloading 1 million random web pages from the CommonCrawl dataset for my next blog post.
August 28, 2026 at 5:08 PM
The second twist is that you don’t even have to provide data: a meta-task of autollm is to detect relevant documents from large corpora (e.g. CommonCrawl). (this twist is yet to be implemented :))
August 28, 2026 at 2:10 PM
[6/6] We hope this helps make large-scale multimodal pre-training more open and accessible.

Huge thanks to grass.io for providing us access to its data infrastructure in support of this release and @hf.co for storage!

🌐: projects.laion.ai/bvd/
📄: arxiv.org/abs/2608.24845
LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with...
arxiv.org
August 26, 2026 at 4:21 PM
[3/6] We built the dataset from CommonCrawl WAT files using a large-scale distributed pipeline:

- Extracted 1.3B safe/platform-specific video URLs
- Downloaded ~80M videos using thousands of distributed workers
- Processed videos into clips, audio segments, and frames
August 26, 2026 at 4:21 PM
[1/6] 🚨 We’re releasing LAION-BVD: a 10-million-hour open video dataset for multimodal pre-training.

- 1.3B video URLs from CommonCrawl
- 80M downloaded videos
- 10M video hours
- 55M captioned clips
- 300M frame-caption pairs

🌐: projects.laion.ai/bvd/

🧵👇
August 26, 2026 at 4:21 PM
(and I recognize that "pirated" is similar to the JSTOR case, but the point is that I don't think I've seen cases ruling that training on eg CommonCrawl violates copyright?)
August 21, 2026 at 5:58 PM
Bu oran, Ocak 2021 ile Temmuz 2026 arasındaki döneme ait web sayfaları üzerinden belirlendi.

#PewResearchCenter #CommonCrawl #ChatGPT #YapayZekaYazarlığı #Teknoloji #Webrazzi

Daha fazlası: https://t.me/ensonhbr
August 21, 2026 at 7:12 AM
Pew Study Finds Over a Third of New Web Content Is AI-Generated
#.com #.gov #.org #10% #35% #chatgpt #commonCrawl #edu #emDashes #openPangram #oxfordCommas #pewResearchCenter
#nile1
Pew Study Finds Over a Third of New Web Content Is AI-Generated
A new Pew Research Center report analyzing half a million webpages reveals that over 35% of content published after ChatGPT shows signs of AI authorship.
nile1.com
August 20, 2026 at 5:26 PM
I tried to clean Commoncrawl
Yes, I understand. At first, I went through many iterations of the code, frequently stopping, modifying, and rewriting parts of the pipeline. One of the main problems was that I didn’t have a reliable way to resume processing from an exact position within a WARC file. Because of that, some WARC records were processed again when a job was restarted, which is what caused the duplicate rows you found. The ~5.5k duplicate records make sense in that context. I agree that this should be handled independently at the final assembly stage as well. I’ll add a global `record_id` uniqueness check so that even if something goes wrong during processing or resuming, duplicate source records cannot make it into the final published dataset. The English/non-English filtering was another weakness in my earlier pipeline. I didn’t have a reliable language-identification stage; instead, I was relying on filters that were intended primarily for English content. As a result, some non-English websites could still pass through the pipeline. Your LID results make it clear that this needs to be handled much more explicitly, by using the language metadata from the WARC files for language identification Your other points about separating main-content extraction from structure preservation are also very useful. I’ll go through those next and document the pipeline stages and settings more clearly. Thanks for taking the time to actually run these checks — this is exactly the kind of audit that helps me identify the weak points in the pipeline and make the dataset more reliable.
discuss.huggingface.co
August 15, 2026 at 8:23 AM
I tried to clean Commoncrawl
For now, I did a quick check in Colab: * * * I think the basic direction here is worth exploring. Preserving some of the original Web document structure instead of immediately flattening everything to plain text is not just a cosmetic choice: recent work on web-corpus construction suggests that the HTML extraction step can change which pages survive downstream filtering, and can matter especially for structured content such as tables and code. At the same time, I would be careful not to jump from that to “Markdown is inherently better than plain text” — that stronger claim still needs a controlled comparison. I ran a couple of lightweight checks against the current `CC-FilteredCorpus` snapshot, mainly to see what kinds of failure modes are actually visible rather than guessing from the pipeline description. The short version is: Check | What I observed | How I would use it ---|---|--- Exact source-record duplication | 153,470 rows, 147,927 unique WARC `record_id`s → **5,543 redundant rows (~3.61%)** | Add a final global uniqueness check after shard assembly English filtering | In a deterministic 10k sample, two independent LID models both flagged **~8.87%** as non-English | Treat this as a review-candidate rate, then re-check the LID stage/model/threshold Raw HTML ↔ output spot-check | 10/10 selected pages were recovered from Common Crawl with matching payload digests; several showed structure loss and/or boilerplate retention | Evaluate main-content extraction separately from Markdown serialization There was also a positive result: in the first 5k sample, Markdown-like structure was present in a substantial fraction of the data — headings in ~61%, lists in ~34%, table-like syntax in ~11%, and fenced code in ~1.3%. So I would **not** describe the Markdown conversion as simply failing across the dataset. It looks more like there are particular boundaries worth tightening. If I were iterating on this, my default order would probably be: 1. **Do one final whole-dataset`record_id` uniqueness check after all shards are assembled.** 2. **Re-check/document the language-ID stage if the intended corpus is English-only.** 3. **Separate “did I keep the main content?” from “did I preserve useful structure inside the main content?”** 4. Only after those cheap checks, consider a more formal extractor/downstream comparison. ## 1. The clearest mechanical issue I found: exact WARC records repeated across shards On the snapshot I checked: * total rows: **153,470** * unique `record_id`: **147,927** * redundant rows beyond one copy per `record_id`: **5,543** * redundant-row share: **~3.61%** * duplicate `record_id` clusters: **5,543** * every duplicate cluster contained exactly two copies * every duplicate cluster crossed a Parquet-shard boundary * within each duplicate cluster, URL, payload digest, and extracted text were the same That is a cleaner signal than ordinary “same URL” or “same text” deduplication. `WARC-Record-ID` is defined by the WARC 1.1 specification as a mandatory identifier that is globally unique for its period of intended use. So seeing the same WARC record ID twice in the derived dataset is good evidence that the **same source record has entered the final assembled dataset twice** , rather than merely two similar pages being classified as duplicates. I would not infer the exact cause from the outside. It could be something around shard overlap, resume/checkpoint boundaries, regenerated shards, final concatenation/publishing, or simply the scope at which exact deduplication is applied. But it seems cheap to guard against at the end of the pipeline, independently of whatever near-dedup method you use. Conceptually, something as simple as: SELECT COUNT(*) AS rows, COUNT(DISTINCT record_id) AS unique_record_ids, COUNT(*) - COUNT(DISTINCT record_id) AS extra_rows FROM final_dataset; and then requiring `extra_rows == 0` before publishing would catch this class of issue. I would treat this as a **final assembly invariant** , not as a reason to redesign the MinHash/near-dedup logic. More detail on the duplicate pattern (click for more details) ## 2. I would take another look at the English-language filter I also ran two independent lightweight language-ID checks on a deterministic 10,000-row sample: * fastText `lid.176.ftz`: **8.97% non-English candidates** * `langid.py`: **9.58% non-English candidates** * both models said non-English: **8.87%** * exact predicted language-code agreement between the two models: **98.91%** * fastText non-English with confidence ≥ 0.8: **7.70%** The largest consensus groups included Portuguese, German, Spanish, Polish, Italian, Czech, and French, and manual inspection of the high-confidence examples showed plenty of pages that are plainly written in those languages. So if the target is still an **English-only / English-focused corpus** , I think the LID stage is worth revisiting. But I would **not** call 8.87% the dataset’s true language-error rate. Real Web language identification is substantially messier than clean benchmark LID: short pages, boilerplate, multilingual pages, code, named entities, and mixed-script content all complicate classification. The recent CommonLID benchmark was created specifically around noisy real Common Crawl text and shows that systems evaluated on cleaner datasets can look much better there than they do on actual Web data. So I would interpret the 8.87% as: > a fairly large, high-information candidate set that is worth checking, not: > 8.87% of the dataset is definitely mislabeled. A few practical things that would make this easier to reason about are documenting: * the LID model * the model revision if relevant * the threshold * whether LID runs before or after main-content extraction * whether classification is document-, chunk-, or line-level * how short pages are handled * how mixed-language pages are handled If LID currently runs on raw/pre-extraction content, one useful branch would be to compare it with classification on the **actual text that will enter training**. That can reduce cases where menus, language selectors, footer text, etc. influence a decision about the main document. For reference, the fastText language-identification models are useful lightweight baselines, but I would still use a small human sample for the final decision rather than treating any LID model as ground truth. Language-ID sample details (click for more details) ## 3. I think “Markdown quality” is easier to reason about if it is split into two stages This was probably the most interesting part to me. A useful conceptual split seems to be: raw HTML | v main-content decision | +-- content that belongs to the document | | | +-- was the content retained? | | | +-- were useful structures retained? | (headings, lists, tables, code, links, formulas...) | +-- navigation / footer / sidebar / cookie UI / boilerplate | +-- was it removed? In other words, there are at least three different metrics hiding inside “clean Markdown”: 1. **Main-content recall** Did the extractor accidentally remove useful document content? 2. **Main-content structure fidelity** For content that should stay, did headings/lists/tables/code/links survive in a useful representation? 3. **Boilerplate leakage** Did menus, cookie notices, related-post widgets, service-area lists, login UI, etc. survive into the training text? Those should not be collapsed into one score. For example, raw HTML may contain 200 links, but if 190 are navigation links, dropping them is a **success** , not a structure-preservation failure. Conversely, flattening the actual article’s section headings or code blocks is a different failure. A very close precedent is SWEb. Their pipeline deliberately treats **HTML → Markdown conversion** and **main-content extraction** as separate stages. They note that converting HTML to Markdown does not by itself remove menus, advertisements, and other unwanted page content, so they subsequently use a dedicated extractor. Their Markdown extractor is also published on Hugging Face. That separation seems useful here even if your implementation is completely different from SWEb. ### What I saw in archived-page spot-checks I selected 10 high-information examples and recovered the corresponding original records from the Common Crawl index. All 10: * were successfully recovered, * had a Common Crawl payload digest matching the dataset’s `payload_digest`. So those comparisons were against the archived source payload rather than whatever happens to be live at the URL today. Among those **selected case studies** : * 7/10 had HTML headings while the final output had zero Markdown heading markers * 8/10 had HTML list items while the final output had zero Markdown list markers * 10/10 had HTML links while the final output had zero Markdown links * one showed a table-preservation discrepancy * one showed a code-preservation discrepancy * several retained obvious navigation/boilerplate-like material These are **not corpus-wide failure rates**. The cases were selected because they looked informative, so they are intentionally not a representative random sample. But they do demonstrate that both failure modes exist: * useful structure can be flattened, * unwanted page chrome can survive. The Common Crawl CDXJ index makes this kind of check fairly reproducible because it exposes the WARC filename, byte offset, length, and digest for a capture. One extra thing I would check specifically: links (click for more details) ## 4. At the same time, a lot of the Markdown structure is clearly making it through I do not want the spot-check above to give the impression that the whole conversion is flattening everything. In a separate deterministic 5k sample, simple syntax checks found: * any Markdown-like structure: **67.72%** * heading markers: **61.14%** * list markers: **33.58%** * table-like Markdown syntax: **10.86%** * fenced-code syntax: **1.34%** These are only syntax-presence measurements — they do not tell us whether the semantics are correct — but they are a useful positive control. So my current interpretation would be: > the structured representation is doing useful work, but some page/extraction paths appear to flatten structure or retain boilerplate. That is a much more encouraging problem than “the Markdown conversion does not work”. It also lines up with the broader extractor literature. For example, Beyond a Single Extractor finds that changing the HTML extractor can substantially change which Web pages survive a fixed downstream filtering pipeline, and that the extractor choice matters more on structured content such as tables and code than on ordinary language-understanding benchmarks. The important caveat from that paper is also useful here: **structure preservation does not imply that one particular Markdown serialization is universally optimal**. The thing to optimize is useful retained information, not Markdown punctuation for its own sake. ## 5. A little more pipeline metadata would make this much easier for other people to evaluate The dataset is already useful to inspect because it preserves URLs, WARC metadata, scores, etc. The next thing I would find most useful is a compact “recipe” section in the Dataset Card. Something like: Extraction - HTML -> Markdown tool: - exact version / commit: - important options: - main-content extractor: - main-content options: Language - LID model: - threshold: - applied before/after extraction: - document/chunk/line level: - short/mixed-language policy: Deduplication - exact-dedup key: - near-dedup algorithm: - n-gram / similarity parameters: - dedup scope: per shard / per crawl / final corpus - which member of a duplicate cluster is retained: Quality signals - perplexity model/tokenizer: - document/chunk aggregation: - filtering threshold, or metadata only: - FineWeb-Edu model/revision: - threshold, or metadata only: Stage counts raw records -> parse success -> extraction success -> LID pass -> quality pass -> exact-dedup survivors -> near-dedup survivors -> final published rows This is less about documentation for its own sake and more about making each observed failure assignable to a stage. For comparison, FineWeb’s Dataset Card documents its major processing stages and links to a working implementation, and the public DataTrove FineWeb pipeline makes the extractor, language filtering, quality filters, and MinHash configuration inspectable. One particularly useful pattern there is saving removed documents via `exclusion_writer`. I like that pattern a lot for iteration: even keeping a **small sample** of rejected/borderline pages from each stage makes it much easier to detect the opposite failure mode — good documents being removed. For example: audit_samples/ language_rejected/ quality_rejected/ dedup_removed/ borderline/ This does not have to be part of the public release or contain everything. Even a small deterministic sample can make threshold changes much easier to audit. ## 6. If you do a small manual extraction audit, I would stratify it by page type I would not start with a huge benchmark. A small manually reviewed set is probably more informative at this stage, but I would avoid using only random article/blog pages. Something like: articles/blogs forums documentation/code product/e-commerce pages listings/directories and then, for each page, score only: main content retained? main content hallucinated/added? # should always be no heading hierarchy useful? lists retained? tables retained? code retained? links retained as intended? boilerplate remaining? The reason for page-type stratification is that ordinary article extraction is comparatively easy. Recent work such as WCXB reports much wider extractor differences on structured/non-article page types than on ordinary articles. So even 10–20 pages per type can reveal more than a larger sample dominated by blog posts. If you eventually want to make a stronger research claim (click for more details) A few checks that did *not* produce a strong concern (click for more details) Related work / why I think the direction is interesting (click for more details) ## Where I would stop for now For the current stage of the project, I do not think you need a large GPU experiment just to make the dataset more defensible. The highest-information / lowest-cost route looks more like: 1. final global record_id uniqueness 2. LID recipe + small labeled sanity sample 3. separate: main-content recall structure fidelity boilerplate leakage 4. keep small rejected/borderline samples 5. document the recipe Then, if the goal later becomes demonstrating that this representation improves model training, move to matched extractor/serialization ablations. So overall: **I think the structure-preserving direction is interesting, and the quick audit did show that it is already retaining a lot of Markdown structure. The main things I would tighten first are final cross-shard exact-record deduplication, the English LID stage, and the boundary between main-content extraction and structure preservation.** Those all look fixable/testable without changing the core idea.
discuss.huggingface.co
August 15, 2026 at 2:23 AM
"Popular datasets like CommonCrawl and LAION 5B contain copyrighted works and personal data used without explicit permission or compensation for data holders."
wp.oecd.ai/app/uploads/...
August 11, 2026 at 1:56 PM
#BradParscale runs 11 websites for #Israel under a $46.5M contract.
#DropSiteNews found that between Jan & Jun, 10 of these websites were crawled 912 times by #CommonCrawl.
Common Crawl is the backbone of the #Al industry-a massive repository used to train large language models.

#AUSPOL #Palestine
August 2, 2026 at 1:16 AM
auf CommonCrawl gelistet
--> also sind sie in Trainingsdaten
--> also beeinflussen sie Antworten
--> also beeinflussen sie Millionen Nutzer

Keine dieser drei Folgerungen in der Kette trägt inhaltlich wirklich.
Es wird auch ständig "wird als Suchergebnis gezeigt"

⬇️
July 31, 2026 at 5:13 PM
The May 2026 crawl archive (CC-MAIN-2026-21) is now also available on our HF bucket. 🤗

huggingface.co/buckets/comm...
commoncrawl/commoncrawl - Storage Bucket
Storage bucket commoncrawl/commoncrawl on Hugging Face
huggingface.co
June 29, 2026 at 11:56 PM
New Commoncrawl index published: commoncrawl.org/blog/june-20...

If you made any changes regarding CCBot crawl recently, time to check the results :-)
Common Crawl - Blog - June 2026 Crawl Archive Now Available
We are happy to announce the release of the June 2026 crawl archive, consisting of 2.10 billion web pages, or 354.59 TiB of uncompressed content.
commoncrawl.org
June 26, 2026 at 8:39 AM
Full command:
~$ gau --threads 5 --providers wayback,commoncrawl,otx,urlscan $target | \
gf redirect | \
qsreplace 'https://attacker[.]com' | \
sort -u | \
httpx -silent -follow-redirects -location -status-code -match-string 'attacker[.]com' -o open_redirects.txt
June 23, 2026 at 4:41 PM