#tokenizer
Claude’s Tokenizer May Use Just 15,000 Vocabulary Entries

They reverse engineered Claude’s tokenizer closely enough to match Claude’s token counts exactly across 500+ human languages and 22 programming languages.
August 15, 2026 at 1:49 AM
Well...
September 8, 2026 at 10:40 PM
I have a new blog post about the so-called “tokenizer-free” approach to language modeling and why it’s not tokenizer-free at all. I also talk about why people hate tokenizers so much!
September 25, 2025 at 3:14 PM
It suggests Claude’s tokenization scheme is more compact than previously estimated.

Article: www.tokenize.rs/claude
Blog: tokencontributions.substack.com/p/on-the-bio...
Reconstructing Claude's tokenizer
A near-exact reconstruction of Claude's tokenizer, running offline with no API calls.
www.tokenize.rs
August 15, 2026 at 1:49 AM
Opus 4.7 has a new tokenizer.
This means it's also a new base model.
Glory days of pretraining still very much going.
April 16, 2026 at 2:45 PM
Apparently in the next version of Fable, Anthropic changed their tokenizer based on real world usage and made “Let me check rather than answer from memory” into a single token, reportedly saving 43% of all token generation

very interesting

tinyurl.com/58zf9d3c
August 27, 2026 at 7:30 PM
I’m going to start calling my keyboard my “tokenizer” 🤣
December 1, 2024 at 10:44 PM
Claude bemused
December 12, 2025 at 1:08 PM
Sampling with any temp other than 1 is privileging the world view of the tokenizer and the tokenizer is (comparatively) dumb as bricks.
January 6, 2025 at 12:15 AM
crap, people are saying it has GLM’s tokenizer
September 16, 2026 at 6:12 PM
… oh my god. There is no standard tokenizer across models, so what a “token” is varies from provider to provider, with different pricing schemes…

… which means they’re non-fungible…
Stripe acquired OpenRouter because it believes "tokens are the central currency for companies building with AI" and companies will want to optimize across models not go all-in on any particular frontier model.

Notably OpenRouter was founded by one of the co-founders of NFT marketplace, OpenSea.
Why a Payments Giant Is Paying $7 Billion for the ‘Stripe of AI’
The Stripe deal for OpenRouter, founded by an NFT entrepreneur, is a bet on a future where users turn to a mix of AI models.
www.wsj.com
August 20, 2026 at 4:37 PM
i'm gonna be famous when i make the first whale BPE tokenizer
New paper out in Proceedings of the Royal Society B: we apply linguistic tools to sperm whale vowels.

The result: sperm whale vowels do not just look like human vowels. They also behave like them.

We found several parallels. Like in Latin, whales have short and long vowels.
April 15, 2026 at 6:53 AM
Cheapo models that aren't made to perfectly parse individual characters still get this wrong for tokenizer reasons (long story short, they process input into their own unique token alphabet (each token is just a serial number, basically, designating a unique character strong)).
September 23, 2026 at 9:13 PM
RL’d LLMs are getting gold in the IMO and people are still fixating on tokenizer artifacts (they don’t know what tokenizers are)
August 8, 2025 at 8:31 AM
This example largely informs why “it’s the tokenizer, stupid” explanations for counting the letters in strawberry or 5.11-5.9=0.21 fall flat to me. If the tokenizer put such hard limits on what it can “see” then there would be no way to do this translation.
May 26, 2026 at 4:56 PM
Gemini I just wanted you to give me the link, not blow smoke up my ass. You wouldn't know a seminal paper if it bit you in the tokenizer
May 27, 2026 at 5:08 AM
SambaNova's EvaByte

The open-weight tokenizer-free language model. Their 6.5B byte-level LM—-EvaByte matches modern tokenizer-based LMs with 5x less data & 2x faster decoding!
January 22, 2025 at 2:45 AM
Let's Build the GPT Tokenizer with Andrej Karpathy
Let's build the GPT Tokenizer
The Tokenizer is a necessary and pervasive component of Large Language Models (LLMs), where it translates between strings and tokens (text chunks). Tokenizer...
www.youtube.com
February 23, 2024 at 3:11 AM
#coding: Habe den C Code Generator so weit vorbereitet. Nun werde ich Ecmascript Tokenizer und Parser schreiben. Hurra :D
November 12, 2025 at 12:49 AM
My paper "Tokenization as Finite-State Transduction" was accepted to Computational Linguistics.

This was my final PhD degree requirement :)

The goal was to unify the major tokenization algorithms under a finite-state automaton framework. For example, by encoding a BPE tokenizer as a transducer.
August 15, 2025 at 7:25 AM
Es geht weiter. Der Tokenizer ist fertig, nun baue ich den Lexer.
November 12, 2025 at 5:03 PM
Ragebaiting anti-ai compiler devs by calling the lexer the tokenizer, parser decoder and code gen inference.
February 8, 2026 at 3:55 AM
It's Sunday morning so taking a minute for a nerdy thread (on math, tokenizers and LLMs) of the work of our intern Garreth

By adding a few lines of code to the base Llama 3 tokenizer, he got a free boost in arithmetic performance 😮

[thread]
November 24, 2024 at 11:05 AM
i am strongly in favor of blaming the tokenizer and the rlhf pass
December 29, 2025 at 7:22 PM
who needs tool calls it's tokenizer math time, babyyy!!!
May 6, 2026 at 5:07 AM