#Pleias
“They said it could not be done”. We’re releasing Pleias 1.0, the first suite of models trained on open data (either permissibly licensed or uncopyrighted): Pleias-3b, Pleias-1b and Pleias-350m, all based on the two trillion tokens set from Common Corpus.
December 5, 2024 at 4:39 PM
Breaking: Pleias releases a new generation of small reasoning models for RAG and source synthesis. Pleias-RAG-350M and Pleias-RAG-1B come with built-in support for source citation, SOTA performance and an accuracy comparable to models ten times their size. huggingface.co/PleIAs/Pleia...
April 24, 2025 at 3:29 PM
I'm very happy to announce a strategic partnership between @wikimediafoundation.org enterprise and Pleias for open, ethical and trustworthy AI innovation. enterprise.wikimedia.com/blog/pleias-...
February 18, 2025 at 4:16 PM
And new Pleias release in partnership with GSMA: CommonLingua, a 2.35M parameters model for language detection, currently performing best on the CommonLID benchmark by a wide margin. huggingface.co/PleIAs/Commo...
April 28, 2026 at 10:27 AM
this reasoning format looked really interesting to me, and i tried to find out what it is, and i have bad news

it's used by a model called the baguettotron

i wish i was joking huggingface.co/PleIAs/Bague...
August 25, 2026 at 2:17 AM
This release is not just about open and ethical data. We’re publishing two models for knowledge retrieval with unprecedented performance: Pleias-Nano (based on Pleias-1b) Pleias-Pico (based on Pleias-350m), competitive with best existing models (Llama, Qwen) ten times their size.
December 5, 2024 at 4:46 PM
Pleias has released a model to do OCR correction on cultural heritage material. They pretrained it from scratch on 18 billion tokens of (mostly) late 19c and early 20c material. #machinelearning h/t @dorialexander.bsky.social huggingface.co/PleIAs/OCRon...
PleIAs/OCRonos-Vintage · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
August 6, 2024 at 1:26 PM
And new paper out: Pleias 1.0: the First Family of Language Models Trained on Fully Open Data

How we train an open everything model on a new pretraining environment with releasable data (Common Corpus) with an open source framework (Nanotron from HuggingFace).

www.sciencedirect.com/science/arti...
September 27, 2025 at 11:44 AM
New LLM model release: Pleias 1.0, "trained exclusively on open data, meaning data that are either non-copyrighted or are published under a permissible license"

This is a really big deal: we finally get to try a model that wasn't trained on unlicensed scraped data!
simonwillison.net/2024/Dec/5/p...
New Pleias 1.0 LLMs trained exclusively on openly licensed data
I wrote about the [Common Corpus](https://simonwillison.net/2024/Mar/20/releasing-common-corpus/) public domain dataset back in March. Now Pleias, the team behind Common Corpus, have released the firs...
simonwillison.net
December 5, 2024 at 5:22 PM
Common Corpus, a multilingual collection of 500B words, released yesterday from PleIAS. #MLSky #blueskAI explanatory blog in next post +
Common Corpus - a PleIAs Collection
The largest public domain dataset for training LLMs.
huggingface.co
March 21, 2024 at 12:11 PM
the Baguettotron posts described the process pretty well

basically, aggressive quality filtering & synthetic rephrasing

if you rephrase text many times over, it discourages rote memorization and encourages learning the concepts

huggingface.co/PleIAs/Bague...
PleIAs/Baguettotron · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
huggingface.co
September 17, 2026 at 9:22 PM
Breaking: Pleias and Nvidia release the first open synthetic dataset for personas in Europe: Nemotron-Personas-France. 1M synthetic French persons, with rich imaginary lives grounded on (complex) demographic distribution. huggingface.co/datasets/nvi...
March 17, 2026 at 10:21 AM
Meta and the other big corps' AI is built on this. Pleias demonstrated just recently that this is entirely unnecessary and you can actually build quite effective AI products on open-source info (and they're partnering with the Wikimedia Foundation to do more of that).
February 22, 2025 at 2:55 AM
Here is in exclusivity the first demo of Pleias knowledge retrieval model: SPQR·LLM. The first ever LLM stuck in the antiquity, using only Latin and Greek sources written from 450 bc to 400 ad.
December 11, 2024 at 11:43 AM
Very happy to see that Pleias multilingual data processing pipelines have contributed to the largest open pretraining project in Europe.

From their tech report: huggingface.co/swiss-ai/Ape...
September 2, 2025 at 4:46 PM
Very fittingly, one of the smallest model Anthropic ever trained is on Common Corpus and Pleias 1.2B tokenizer: 2.9M model artificially expanded to 331M to study weights interference for the new Transformer Circuits. transformer-circuits.pub/2026/interfe...
August 22, 2026 at 4:43 PM
Following on our partnership with Wikimedia Enterprise, very happy to see Pleias featured in Wikipedia 25th anniversary post. wikimediafoundation.org/news/2026/01...
January 16, 2026 at 6:59 PM
Pleias is one: pleias.fr

Olmo is another: allenai.org/olmo
December 22, 2025 at 5:33 AM
And new technical blogpost by Pleias application team on deploying small reasoning models for edge devices : featuring cache context management on Rasperry, designing system orchestration under constraints (reranker, chunking) and model specialization. pleias.ai/blog/local-a...
July 8, 2026 at 4:11 PM
Pleias board meeting with Doria and Langlais.
March 10, 2025 at 7:19 PM
Announcing the release of marginalia, a python library to perform corpus analysis and retrieve structured annotations with open LLMs like Mistral Open-Hermes-2.5. github.com/Pleias/margi...
February 12, 2024 at 7:25 PM
this is objectively good
September 27, 2025 at 3:33 PM
All Pleias models including our SLMs, artifacts/classifiers and Common Corpus are on HuggingFace. Much more to come in the next weeks are we're going to release multiple things for the AI Summit. huggingface.co/PleIAs
January 28, 2025 at 4:45 PM
Pleias on an open source AI panel with Jimmy Wales (Wikimedia) and OpenUK tomorow.
February 17, 2026 at 3:11 PM
I believe Pleias (of Baguettotron fame) is open-weights, open-data
January 3, 2026 at 9:23 PM