#data-deduplication
Kopia is a Cross-platform open source backup tool for Windows, macOS & Linux with fast, incremental backups, client-side end-to-end encryption, compression and data deduplication. CLI and GUI included for backups and restore.

github.com/kopia/kopia
GitHub - kopia/kopia: Cross-platform backup tool for Windows, macOS & Linux with fast, incremental backups, client-side end-to-end encryption, compression and data deduplication. CLI and GUI included.
Cross-platform backup tool for Windows, macOS & Linux with fast, incremental backups, client-side end-to-end encryption, compression and data deduplication. CLI and GUI included. - kopia/kopia
github.com
July 30, 2025 at 4:14 AM
Added the label a while ago because I'd forget about the deduplication and compression and think I was missing data.
December 13, 2025 at 3:13 AM
data deduplication is definitely a thing, but this ain't it
February 9, 2025 at 6:42 PM
also that is *not* what deduplication means! deduplication refers to multiple copies of *data* - when for example you back up something multiples times, a deduplicated storage scheme will only add references to the original data rather than copies. So this is a matter of saving disk space not money!
February 9, 2025 at 6:23 PM
Kopia est un projet open-source sous licence Apache 2 pour de gérer vos sauvegardes. Il vise Linux, macOS et Windows, gère la sauvegarde incrémentale, la compression, la déduplication, ... et propose une interface graphique et en ligne de commande ⬇️

github.com/kopia/kopia/
GitHub - kopia/kopia: Cross-platform backup tool for Windows, macOS & Linux with fast, incremental backups, client-side end-to-end encryption, compression and data deduplication. CLI and GUI included.
Cross-platform backup tool for Windows, macOS & Linux with fast, incremental backups, client-side end-to-end encryption, compression and data deduplication. CLI and GUI included. - kopia/kopia
github.com
August 1, 2025 at 5:59 AM
@sergiodxa.com wrote about the underlying data transport used by @reactrouter.com. It's a vendor'd version of turbo-stream, a library inspired by @rich-harris.dev's devalue. sergiodxa.com/tutorials/le...
How to Leverage React Router's Built-in Data Deduplication by sergiodxa
Learn how React Router's built-in deduplication system uses references to eliminate duplicate data transmission when combining promises in y
sergiodxa.com
October 7, 2025 at 12:48 AM
Are you still using npm transpile services like esm.sh and unpkg.com?
❌ dependency deduplication
❌ install hooks and native add-ons
❌ loading data files

Here's why we recommend importing npm packages natively via npm specifiers 👇

deno.com/blog/not-usi...
February 13, 2025 at 5:05 PM
Are you struggling with messy datasets? Look no further than Matasoft's cutting-edge fuzzy data matching, deduplication, and entity resolution services, expertly provided by Zlatko Matić.

#entity-resolution #EntityResolution #FuzzyMatch #fuzzy-match #fuzzy-matching

matasoft.hr/qtrendcontro...
December 30, 2024 at 7:05 PM
i love too proudly display that i do not understand the difference between layout (deduplication) and data (unique constraint)
February 9, 2025 at 6:42 PM
Data deduplication
August 17, 2026 at 8:24 PM
Deduplication exists, but it isn't what Musk seems to think it is. It's used in storage systems to save room by only keeping a single copy of identical data chunks that appear multiple times in the system.
February 9, 2025 at 10:25 PM
It's the deduplication part. Deduplication has nothing to do with it. It's a method for storing data, often to save space. It has nothing to do with the data itself.
February 9, 2025 at 6:19 PM
If you are in Python then this piece is mandatory to read-
towardsdatascience.com/surprisingly...

#python #coding #datascience #data #ai
Surprisingly Effective Way To Name Matching In Python
Data Matching, Fuzzy Matching, Data Deduplication
towardsdatascience.com
September 13, 2023 at 5:23 PM
the "old" email would still be *noise*. I would probably use deduplication to cut down on disk space, and that'd have a *lot* bigger impact. Migration of older data to cold-line storage would be another step I might take (need to finish speccing out that f'bsd project).
July 15, 2025 at 12:53 PM
also that is *not* what deduplication means. deduplication refers to multiple copies of *data* - when for example you back up something multiples times, a deduplicated storage scheme will only add references to the original data rather than copies. So this is a matter of saving disk space not money!
February 9, 2025 at 6:19 PM
🔒Enhancing #FederatedLearning with #Privacy Preserving Deduplication

Dr Aydin Abadi of @computingnewcastle.bsky.social introduces a protocol that boosts machine learning model accuracy & efficiency without compromising data privacy
Read more ➡️ tinyurl.com/ywr8w94e
#AI #MachineLearning #DataScience
Privacy-preserving deduplication to enhance federated learning
Data privacy in machine learning models has never been more critical - our experts talk about the challenges of solving data deduplication in LLMs and AI.
tinyurl.com
March 18, 2025 at 4:28 PM
Two datasets won’t join because there’s no unique ID? Matasoft performs fuzzy data matching to link and merge records. matasoft.hr/qtrendcontro...
Data Matching Services
Data matching, linking, merging, cleansing, deduplication, consolidation and other data processing tasks on your business data, such as customer contact, real estate or product lists. Using powerful Q...
matasoft.hr
May 21, 2026 at 9:14 PM
hell, it was just a few days ago when he showed the world he has no idea what data deduplication is
February 11, 2025 at 7:35 PM
Funny you mentioned me on the Haxl, I'm actually working on a competitor on that front github.com/iand675/sofe...
GitHub - iand675/sofetch: Automatic batching and deduplication of concurrent data fetches for Haskell
Automatic batching and deduplication of concurrent data fetches for Haskell - iand675/sofetch
github.com
February 14, 2026 at 5:23 PM
the architecture was simple: 6 emphemeral hetzner auction dedis as crawler machines and a vps that served both the database and the website. each dedi had a ledger in sqlite making sure we don't lose any data; the ClickHouse side was handled with the magic of ReplaceMergeTree for deduplication
June 26, 2026 at 10:49 PM
OpenZFS 2.3 is here, with RAID expansion and faster dedup
OpenZFS 2.3 is here, with RAID expansion and faster dedup
Coming soon to April's TrueNAS SCALE release, dubbed 'Fangtooth' The latest version of OpenZFS offers RAID expansion, plus faster data deduplication donated by iXsystems. The code will be available very soon in the beta of TrueNAS SCALE 25.04.…
dlvr.it
January 23, 2025 at 11:07 AM
18.71 billion documents. 25.25 trillion tokens. 152 data sources, deduplicated, decontaminated, and sorted into 200 buckets before a single training run.

A new blog from Will Held @williamheld.com on Datakit, the Marin team's pretraining data pipeline:

openathena.ai/blog/marin-d...
September 14, 2026 at 8:22 PM