Free, simple and quick online tool to find subdomains and endpoints.
Data sources: HackerTarget, URLScanIO, RapidDNS, CertSpotter, JLDC, DNSRepo, crtSH, WayBack, CommonCrawl, AlienWault OTX
recox.hackerz.space
Free, simple and quick online tool to find subdomains and endpoints.
Data sources: HackerTarget, URLScanIO, RapidDNS, CertSpotter, JLDC, DNSRepo, crtSH, WayBack, CommonCrawl, AlienWault OTX
recox.hackerz.space
LAION-5BもCommonCrawlのデータで製作されており、そもそもCSAMはそこに入っていたわけです。
foundation.mozilla.org/en/blog/Mozi...
LAION-5BもCommonCrawlのデータで製作されており、そもそもCSAMはそこに入っていたわけです。
foundation.mozilla.org/en/blog/Mozi...
In our new paper, we propose WebOrganizer which *constructs domains* based on the topic and format of CommonCrawl web pages 🌐
Key takeaway: domains help us curate better pre-training data! 🧵/N
In our new paper, we propose WebOrganizer which *constructs domains* based on the topic and format of CommonCrawl web pages 🌐
Key takeaway: domains help us curate better pre-training data! 🧵/N
Some early groups torrented small collections, but that was years ago and
Some early groups torrented small collections, but that was years ago and
#go URL discovery tool:
- different sources (alienvault,commoncrawl etc)
- filter by extensions/regex
- very fast (122000+ URLs in 30 sec):
github.com/projectdisco...
Creator x.com/pdnuclei
#osint #cybersecurity #bugbounty
#go URL discovery tool:
- different sources (alienvault,commoncrawl etc)
- filter by extensions/regex
- very fast (122000+ URLs in 30 sec):
github.com/projectdisco...
Creator x.com/pdnuclei
#osint #cybersecurity #bugbounty
cc-downloader is still under active development, so if you find any issues or would like to submit a feature request, please visit its GitHub repository at github.com/commoncrawl/....
cc-downloader is still under active development, so if you find any issues or would like to submit a feature request, please visit its GitHub repository at github.com/commoncrawl/....
While i haven't dug harder yet. Has any one managed to use the genesis wen annotation in commoncrawl meaningfully?
#commoncrawl #crawler #web
While i haven't dug harder yet. Has any one managed to use the genesis wen annotation in commoncrawl meaningfully?
#commoncrawl #crawler #web
- 1.3B video URLs from CommonCrawl
- 80M downloaded videos
- 10M video hours
- 55M captioned clips
- 300M frame-caption pairs
🌐: projects.laion.ai/bvd/
🧵👇
- 1.3B video URLs from CommonCrawl
- 80M downloaded videos
- 10M video hours
- 55M captioned clips
- 300M frame-caption pairs
🌐: projects.laion.ai/bvd/
🧵👇
Authors, find out if your work has been pirated here:
www.theatlantic.com/technology/a...
Authors, find out if your work has been pirated here:
www.theatlantic.com/technology/a...
- Effort to collect CC-licensed web data in one place
- 147M documents explicitly CC-licensed
- License detection from HTML links & metadata
- Preserves exact license type & version
- Maps to original URLS & crawl dates
huggingface.co/datasets/Bra...
- Effort to collect CC-licensed web data in one place
- 147M documents explicitly CC-licensed
- License detection from HTML links & metadata
- Preserves exact license type & version
- Maps to original URLS & crawl dates
huggingface.co/datasets/Bra...
huggingface.co/buckets/comm...
huggingface.co/buckets/comm...
These websites can then be found in CommonCrawl dumps that are generally used for pretraining data curation...
These websites can then be found in CommonCrawl dumps that are generally used for pretraining data curation...
(or indeed why didn't commoncrawl do that) purl.stanford.edu/kh752sm9123
(or indeed why didn't commoncrawl do that) purl.stanford.edu/kh752sm9123
A little Dutch nonprofit run by a family trust that's been crawling and snapshotting the public web since 2008 and making the results freely available publicly for research.
A little Dutch nonprofit run by a family trust that's been crawling and snapshotting the public web since 2008 and making the results freely available publicly for research.
#DropSiteNews found that between Jan & Jun, 10 of these websites were crawled 912 times by #CommonCrawl.
Common Crawl is the backbone of the #Al industry-a massive repository used to train large language models.
#AUSPOL #Palestine
#DropSiteNews found that between Jan & Jun, 10 of these websites were crawled 912 times by #CommonCrawl.
Common Crawl is the backbone of the #Al industry-a massive repository used to train large language models.
#AUSPOL #Palestine