#TimeCapsuleLLM
TimeCapsuleLLM: A language model trained from scratch exclusively on data from certain places and time periods to reduce modern bias and emulate the voice, vocabulary, and worldview of the era.
GitHub - haykgrigo3/TimeCapsuleLLM: A LLM trained only on data from certain time periods to reduce modern bias
A LLM trained only on data from certain time periods to reduce modern bias - haykgrigo3/TimeCapsuleLLM
github.com
July 26, 2026 at 12:32 AM
Okay, now try a protein LLM trained on only ancestral proteins to reduce modern bias (a simpler time without so many protein fitness influencers) github.com/haykgrigo3/T...
GitHub - haykgrigo3/TimeCapsuleLLM: A LLM trained only on data from certain time periods to reduce modern bias
A LLM trained only on data from certain time periods to reduce modern bias - haykgrigo3/TimeCapsuleLLM
github.com
January 12, 2026 at 4:38 PM
Das Modell TimeCapsuleLLM wurde ausschließlich mit Texten aus dem London des 19. Jahrhunderts trainiert.

Das Modell verhält sich so, als würde es sich im 19. Jahrhundert befinden. Der Ansatz könnte relevant für Gessellschafts- und Geschichtsforschung sein.

arstechnica.com/information-...
College student’s “time travel” AI experiment accidentally outputs real 1834 history
Hobbyist training AI on Victorian texts gets an unexpected history lesson from his own creation.
arstechnica.com
January 26, 2026 at 8:49 AM
There are a couple toy models that try to achieve something similar (though obviously no RLHF from Victorian era people)
huggingface.co/haykgrigoria...
huggingface.co/bahree/londo...
haykgrigorian/TimeCapsuleLLM-v2-llama-1.2B · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
February 14, 2026 at 2:28 PM
This is interesting. One dev is training an AI from scratch on books from 1800s London.

It's called TimeCapsuleLLM, not a fine-tuned modern model, but one trained entirely on historical data. No modern language or context.

Built on nanoGPT by Karpathy. github.com/haykgrigo3/...
GitHub - haykgrigo3/TimeCapsuleLLM: An LLM trained only on data from certain time periods to reduce modern bias
An LLM trained only on data from certain time periods to reduce modern bias - haykgrigo3/TimeCapsuleLLM
github.com
July 14, 2025 at 7:30 AM
Do you actually want this or do you just think you have an unassailable talking point to hurl?

github.com/haykgrigo3/T...

huggingface.co/PleIAs/Pleia...

huggingface.co/blog/Pclangl...
GitHub - haykgrigo3/TimeCapsuleLLM: A LLM trained only on data from certain time periods to reduce modern bias
A LLM trained only on data from certain time periods to reduce modern bias - haykgrigo3/TimeCapsuleLLM
github.com
February 12, 2026 at 4:23 AM
Your critique is of the behavior of a select few VC-backed speculative companies, and not of the underlying technology, which can certainly be trained on material that nobody living should have any vested interest in github.com/haykgrigo3/T...
GitHub - haykgrigo3/TimeCapsuleLLM: A LLM trained only on data from certain time periods to reduce modern bias
A LLM trained only on data from certain time periods to reduce modern bias - haykgrigo3/TimeCapsuleLLM
github.com
February 12, 2026 at 3:36 AM
TimeCapsuleLLM: LLM trained only on data from 1800-1875 [Discussion]
TimeCapsuleLLM: LLM trained only on data from 1800-1875
TimeCapsuleLLM: LLM trained only on data from 1800-1875
github.com
January 12, 2026 at 5:52 PM
⚡ Hackernews Top story: TimeCapsuleLLM: LLM trained only on data from 1800-1875
GitHub - haykgrigo3/TimeCapsuleLLM: A LLM trained only on data from certain time periods to reduce modern bias
A LLM trained only on data from certain time periods to reduce modern bias - haykgrigo3/TimeCapsuleLLM
github.com
January 12, 2026 at 5:18 PM
TimeCapsuleLLM: LLM trained only on data from 1800-1875
View Article | Join the HN Conversation

Summary of HN discussion 🧵👇
GitHub - haykgrigo3/TimeCapsuleLLM: A LLM trained only on data from certain time periods to reduce modern bias
A LLM trained only on data from certain time periods to reduce modern bias - haykgrigo3/TimeCapsuleLLM
github.com
January 13, 2026 at 8:00 AM
https://github.com/haykgrigo3/TimeCapsuleLLM
TimeCapsuleLLMは、特定の場所と時代に限定したデータで学習された言語モデルです。
現代のバイアスを減らし、当時の声、語彙、世界観をエミュレートします。
まるでAIモデルが歴史的な存在であるかのように振る舞います。
GitHub - haykgrigo3/TimeCapsuleLLM: A LLM trained only on data from certain time periods to reduce modern bias
A LLM trained only on data from certain time periods to reduce modern bias - haykgrigo3/TimeCapsuleLLM
github.com
January 13, 2026 at 10:24 AM
There's good stuff happening in the community.

This person has been training an LLM on data up to the year 1875 only. A school kid recently reported training a capable small model for $1200. Allen Institute produces OLMO which is only trained on ethically-sourced data.

github.com/haykgrigo3/T...
GitHub - haykgrigo3/TimeCapsuleLLM: A LLM trained only on data from certain time periods to reduce modern bias
A LLM trained only on data from certain time periods to reduce modern bias - haykgrigo3/TimeCapsuleLLM
github.com
February 19, 2026 at 4:00 PM
연구차원에서 1800~1875년 데이터로만 학습된 언어모델(AI), TimeCapsuleLLM

news.hada.io/topic?id=25780

흥미롭네요. 조선왕조실록만 학습시킨 AI는 완전 조선식 사고방식을 갖게 되는걸까요.
만약 이 LLM이 그 당시에 없는 과학적 발견을 혼자 할 수 있다면 학습(자료 먹임) 없이도 인공 지성이 발전할 수 있다는게 아니냐는 의견.
TimeCapsuleLLM: 1800~1875년 데이터만으로 학습된 대형 언어 모델 | GeekNews
TimeCapsuleLLM은 특정 시기(1800~1875년)의 자료만으로 학습된 대형 언어 모델(LLM) 로, 현대적 편향을 최소화하고 당시의 언어와 세계관을 재현하는 목적모델은 런던 지역의 역사적 문서, 서적, 신문, 법률 문서 등으로 구성된 데이터셋을 사용해 시대별 언어 스타일과 어휘를 반영초기 버전은 nanoGPT, 이후 버전은 Microsoft Ph
news.hada.io
January 15, 2026 at 7:20 AM
TimecapsuleLLM - die #KI, die 'denkt', dass sie im 19. Jhdt. lebt: Die #KünstlicheIntelligenz wurde ausschließlich mit Daten des viktorianischen London trainiert. Historische LLMs könnten künftig nützlich für die Forschung werden #KIAssistent #LLM #ArtificialIntelligence #Geschichte #history #AI
Die KI, die "denkt", dass sie im 19. Jahrhundert lebt
TimecapsuleLLM hat ausschließlich von Daten des viktorianischen London gelernt. Historische LLMs könnten künftig nützlich für die Forschung werden
www.derstandard.at
January 26, 2026 at 8:18 AM
TimeCapsuleLLM: LLM trained only on data from 1800-1875
January 12, 2026 at 6:20 PM
#history #AI #LLM

'For the past month, [Hayk] Grigorian has been developing what he calls TimeCapsuleLLM, a small AI language model (like a pint-sized distant cousin to ChatGPT) which has been trained entirely on texts from 1800–1875 London.'

arstechnica.com/information-...
College student’s “time travel” AI experiment accidentally outputs real 1834 history
Hobbyist training AI on Victorian texts gets an unexpected history lesson from his own creation.
arstechnica.com
August 23, 2025 at 12:29 PM
💡 Summary:

haykgrigo3による「TimeCapsuleLLM」リポジトリは、19世紀のテキスト(1800年から1875年)だけを用いて訓練された、歴史的なロンドンに焦点を当てた言語モデルのためのデータセットとスクリプトを提供しています。選択的時系列学習(STT)を用いることで、現代の偏りを最小限に抑えつつ、その時代に忠実な出力を生成することを目指しています。モデルは16Mから700Mパラメータまでのさまざまなアーキテクチャ(nanoGPT、Phi 1.5、Llamaなど)で構築されており、 (1/2)
January 13, 2026 at 1:43 AM
-Is it ethical to use open weight LLMs (not paying rents to firms training on pirated data)?
-Is it ethical to use TimeCapsuleLLM (all training data in public domain)?
September 10, 2026 at 4:54 AM
Apache-2.0 on a 300M Llama-compatible model is rare; CPU-class inference makes it a genuine edge tinkerer's toy.

No benchmarks, no data card, and the Ascend NPU claim is unverified—"eval1" suggests the authors know.

Hard to see this outrunning a tuned Qwen2.5-0.5B until provenance is answered....
Tiny but Mighty: TimeCapsuleLLM-v2mini Arrives on Modelers.cn
aichina.news
August 8, 2026 at 1:15 PM
🤖 TimeCapsuleLLM: 160GB dataset of 1800s texts trains Victorian-era AI

A 500M parameter model learns from 40B tokens of 1800-1875 English data.

https://theneuralfeed.com/share/post/1QbThzRV

#AINews #TechNews

Read the full story →
theneuralfeed.com
July 12, 2026 at 7:29 AM
🤖 TimeCapsuleLLM: 160GB dataset of 1800s texts trains Victorian-era AI

A 500M parameter model learns from 40B tokens of 1800-1875 English data.

https://theneuralfeed.com/share/post/1QbThzRV

#AINews #TechNews

Read the full story →
theneuralfeed.com
July 11, 2026 at 7:49 AM
August 23, 2025 at 12:01 AM