#LLM-Memorization
accidentally releasing the coca-cola recipe through llm memorization leaking after i opened up the pdf with adobe
May 1, 2024 at 6:57 AM
I keep saying McKinsey is going to destroy a bunch of companies by pitching it to management and its absolutely going to happen/a business is going to accidentally publicize a bunch of trade secrets/pii through llm memorization
Folks don’t seem to understand. It doesn’t MATTER that AI is shitty at your job. Your CEOs will happily deliver a 100x shittier AI version of your job if it means they don’t pay your salary.

This is true no matter your pay level. But it’s coming sooner if you’re underpaid.

Neutrality is complicity
Said it before and I’ll say it again: we in SAG-AFTRA and/or the WGA know they are not going to stop with OUR jobs.

If your job CAN be replaced with AI, it WILL be.
August 9, 2023 at 6:32 PM
ハリーポッターの文章を最大95.8%(Claude)LLMから抽出できたという論文
非常にカス
arxiv.org/abs/2601.02671
Extracting books from production language models
Many unresolved legal questions over LLMs and copyright center on memorization: whether specific training data have been encoded in the model's weights during training, and whether those memorized dat...
arxiv.org
January 15, 2026 at 8:01 AM
llm-memorization

Give your local LLM a real memory with a lightweight, fully local memory system — just like a human recalling past discussions. 100% offline. 100% under your control.

https://github.com/victorcarre6/llm-memorization
June 27, 2025 at 8:15 AM
Hubble is finally out! We used 200k GPU hours from NAIRR and NVIDIA to build a comprehensive resource for the scientific study of LLM memorization. Fully open-source models & data up to 8B params + 500B tokens with controlled data insertion to study memorization risks 🔭✨
Announcing 🔭Hubble, a suite of open-source LLMs to advance the study of memorization!

Pretrained 1B/8B param models, with controlled insertion of texts designed to emulate key memorization risks: copyright (e.g., book passages), privacy (e.g., synthetic biographies), and test set contamination
October 24, 2025 at 6:36 PM
Want to know what training data has been memorized by models like GPT-4?

We propose information-guided probes, a method to uncover memorization evidence in *completely black-box* models,

without requiring access to
🙅‍♀️ Model weights
🙅‍♀️ Training data
🙅‍♀️ Token probabilities 🧵 (1/5)
Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
High-quality training data has proven crucial for developing performant large language models (LLMs). However, commercial LLM providers disclose few, if any, details about the data used for training. ...
arxiv.org
March 21, 2025 at 7:08 PM

Further strong evidence against the "stochastic parrot" / "mere memorization" hypothesis for explaining LLM reasoning.

arxiv.org/abs/2411.06198
December 9, 2024 at 1:37 PM
Our workshop on LLM Memorization is coming to ACL 2025! The call for papers is out, please submit both archival and non-archival (work in progress or already published) papers
📢 The First Workshop on Large Language Model Memorization (L2M2) will be co-located with
@aclmeeting.bsky.social in Vienna 🎉

💡 L2M2 brings together researchers to explore memorization from multiple angles. Whether it's text-only LLMs or Vision-language models, we want to hear from you! 🌍
January 27, 2025 at 11:23 PM
Where does training data memorization live inside models (e.g., LLM)? How is it stored? How much is it involved in different tasks?

They examine all of these questions using loss curvature, which is s like PCA, but for loss curvature instead of variance:
November 6, 2025 at 7:22 PM
The "AI Training creates a compressed copy" argument maps really well to the existence of "Memorization of training data" in LLM discussions: Models trained without strong safeguards against memorization absolutely DO contain compressed copies, and should be thrashed on those grounds.
July 28, 2026 at 2:41 PM
broke: burning a bunch of sources because my comsec relied on honoring a robo.txt

bespoke: burning sources due to llm memorization issues
The CIA is planning to roll out a ChatGPT-style tool across the 18-agency US intelligence community to give analysts better access to intelligence (Bloomberg)

Main Link | Techmeme Permalink
September 26, 2023 at 4:54 PM
“Judge Alsup found that training an LLM is transformative use—even when there is significant memorization,” says one IP lawyer. The ruling may reverberate across the dozens of other AI copyright lawsuits winding through US courts.
Anthropic Scores a Landmark AI Copyright Win—but Will Face Trial Over Piracy Claims
While the startup has won its "fair use" argument, it potentially faces billions of dollars in damages for allegedly pirating over 7 million books to build a digital library.
www.wired.com
June 24, 2025 at 4:47 PM
This is for anyone who thinks they don't need to at least practice due diligence that output from an LLM is not copyright infringement or (obvious) plagiarism. (I contend it should be understood as both, even when the emitted text is not near-verbatim, because the mechanism is identical.)
January 8, 2026 at 4:09 PM
March 1, 2024 at 4:32 AM
Even before our current level of LLM technology emerged autocomplete had already destroyed at least half of my spelling capability. That's not so terrible as it's mostly memorization of exceptions to general rules, but the way ppl use LLMs to reason or decide for them is scary
No. The chief problem is, AI is literally A TOOL FOR UNLEARNING. The more you use it, the worse you get at whatever it is you’re using it for. It’s the opposite of educating, it’s uneducating. 2/2
November 19, 2025 at 10:56 PM
3-minute explanation of my relationship to LLM memorization research

m.youtube.com/watch?v=unfz...
ABBA - Mamma Mia (Official Music Video)
YouTube video by AbbaVEVO
m.youtube.com
January 6, 2026 at 9:12 PM
Safeguards aren't perfect but filters and techniques to prevent memorization of complete passages should be enough to get LLM companies out of trouble.

Retrieval+LLM might be a different story. Not sure yet.

Image and music draw from radically case law, it'll be a tough place for AI
November 18, 2024 at 1:02 AM
Came to #NeurIPS2024 for the research news, but staying for these incredible views. I am presenting some recent works that (I think) significantly advance the discourse on LLM memorization, training data detection; & a study on hallucinations x model collapse in diffusion models.
December 10, 2024 at 10:59 PM
Alignment isnt only thing LLMs are faking. Reasoning is another one that they are good at faking. Reading paper on LLM performance on reasoning tasks of doctors. Just started reading but either going to be:
1. Memorization or
2. Priming or
2. Confirmation prompting

www.anthropic.com/research/ali...
Alignment faking in large language models
A paper from Anthropic's Alignment Science team on Alignment Faking in AI large language models
www.anthropic.com
December 18, 2024 at 8:19 PM
amusingly existing LLMs memorize a lot more than we usually think they do, and mixture of experts models specifically are more memorization-prone, and this usually improves performance. so i am not actually sure an ideal LLM doesn't have long stretches of text memorized under current designs
March 27, 2025 at 5:33 PM
This is completely dependent on training data.

And an LLM doesn't know about truth, lying, or memorization. It is not thinking.
July 3, 2025 at 6:51 AM
How do these phases relate to LLM behavior?

- Entropy-seeking: Correlates with short-sequence memorization (♾️-gram alignment).

- Compression-seeking: Correlates with dramatic gains in long-context factual reasoning, e.g. TriviaQA.

Curious about ♾️-grams?
See: bsky.app/profile/liuj...
🧵4/9
October 31, 2025 at 4:19 PM
New auditing technique catches unauthorized LLM training on protected data even when models hide it through behavioral changes rather than memorization—critical for enforcing data privacy in agentic workflows.

https://arxiv.org/abs/2604.22191

#AI #MachineLearning
April 28, 2026 at 10:00 PM
Interesting, "GPT-style models have a fixed memorization capacity of approximately 3.6 bits per parameter."
venturebeat.com/ai/how-much-...
#ai #memorization #llm
How much information do LLMs really memorize? Now we know, thanks to Meta, Google, Nvidia and Cornell
Using a clever solution, researchers find GPT-style models have a fixed memorization capacity of approximately 3.6 bits per parameter.
venturebeat.com
June 7, 2025 at 8:07 AM