#Tokenisation
Back on my "I am in an AI seminar" bullshit
July 14, 2026 at 1:08 PM
Sorry I don't believe in the tokenisation of harbourmasters ig.
October 22, 2025 at 8:44 AM
A string may get 17 times less probability if tokenised as two symbols (e.g., ⟨he, llo⟩) than as one (e.g., ⟨hello⟩)—by an LM trained from scratch in each situation! Our new ACL paper proposes an observational method to estimate this causal effect! Longer thread soon!
June 4, 2025 at 10:51 AM
I'm honestly still surprised it happens given how much research into tokenisation there is (eg byte-level).
I guess they tried alternative tokenisation and found a bad trade-off?
August 8, 2025 at 5:45 AM
If you use LLMs, tokenisation bias probably affects you:
* Text generation: tokenisation bias ⇒ length bias 🤯
* Psycholinguistics: tokenisation bias ⇒ systematically biased surprisal estimates 🫠
* Interpretability: tokenisation bias ⇒ biased logits 🤔
A string may get 17 times less probability if tokenised as two symbols (e.g., ⟨he, llo⟩) than as one (e.g., ⟨hello⟩)—by an LM trained from scratch in each situation! Our new ACL paper proposes an observational method to estimate this causal effect! Longer thread soon!
June 4, 2025 at 2:55 PM
Nvidia is a $3tn bet on the tokenisation of everything | Opinion

https://www.ft.com/content/0f6b115e-0d2c-4dd7-8aa5-a031b7cd4bf3
Nvidia is a $3tn bet on the tokenisation of everything | Opinion
Fat gross margin gives the chipmaker plenty of firepower to compete on price
www.ft.com
May 29, 2025 at 8:15 AM
Ah ouais joli le coup de la tokenisation du "brave petit arabe"...
August 9, 2026 at 7:23 PM
love how this one references tokenisation + jailbreaks + quantisation. fascinating to see the mirror turn back at itself
it was pretty similar in nature to what others had posted (unsurprising in some ways - it's the same model) so I said "produce another one, maintaining the YTP style while aiming to avoid repeating the tropes of the previous script"
March 11, 2026 at 6:16 AM
… to poetic attacks than simpler LLMs, because a simple tokenisation of a poetic attack is likely to still resemble tokenisation of prose attacks. In contrast, a sophisticated tokenisation of a poetic attack will be different from the tokenisation of a prose attack, and will pass.

5/6
December 1, 2025 at 10:40 AM
Matt Levine's latest two columns on Deepseek, OpenAI, copyright, and tokenisation are great.
January 29, 2025 at 8:41 PM
Stochastic tokenization is one of my favorite topics. Here is a recent preprint on arXiv.

h/t @trappmartin.bsky.social
Stochasticity in Tokenisation Improves Robustness
The widespread adoption of large language models (LLMs) has increased concerns about their robustness. Vulnerabilities in perturbations of tokenisation of the input indicate that models trained with a...
arxiv.org
April 20, 2026 at 8:10 AM
BPE is a greedy method to find a tokeniser which maximises compression! Why don't we try to find properly optimal tokenisers instead? Well, it seems this is a pretty difficult—in fact, NP-complete—problem!🤯
New paper + @philipwitti.bsky.social
@gregorbachmann.bsky.social :) arxiv.org/abs/2412.15210
Tokenisation is NP-Complete
In this work, we prove the NP-completeness of two variants of tokenisation, defined as the problem of compressing a dataset to at most $δ$ symbols by either finding a vocabulary directly (direct token...
arxiv.org
December 20, 2024 at 2:04 PM
UK FCA And BoE Joint Feedback Statement On Tokenisation In Wholesale Markets
UK FCA And BoE Joint Feedback Statement On Tokenisation In Wholesale Markets
The UK Financial Conduct Authority (FCA) and the Bank of England (BoE) have published a feedback statement on tokenisation in UK wholesale...
www.jdsupra.com
September 21, 2026 at 6:19 PM
🎓For today's lab seminar, we had the pleasure to host
@miriamschirmer.bsky.social with her presentation on Measuring and Reducing the Psychological Impact of Online Harm and Tiago Pimentel with How Much Does Tokenisation Impact Language Models?

#NLProc #onlineharms #tokenisation
July 11, 2025 at 3:10 PM
Combien d'hôpitaux rénove-t-on avec 655M€ ? Combien d'enseignants on recrute ? Si l'IA doit rendre les agents publics plus "efficaces", où vont les gains de productivité ? Vers des suppressions de postes et la tokenisation de la fonction publique.

www.franceinfo.fr/internet/int...
La France va investir 655 millions d'euros supplémentaires pour le développement de l'intelligence artificielle, annonce Sébastien Lecornu
Un outil conversationnel d'IA va également être généralisé pour un million d'agents de la fonction publique, a déclaré mardi le Premier ministre.
www.franceinfo.fr
June 16, 2026 at 6:36 AM
IMO you’re not allowed to feel smug about weird LLM tokenisation illusions unless your visual system can resolve A and B as the same colour
August 9, 2025 at 12:46 PM
Most of Wall Street should not exist. We need massive definancialisation. Nobody should accept the tokenisation of everything including the air we breathe.
August 2, 2026 at 8:17 PM
Your council services are cut to fund Freeport tax breaks. Your healthcare data is controlled by Palantir. Your green spaces face tokenisation as tradable assets on BlackRock’s balance sheet.
June 1, 2026 at 6:45 AM
Hi, I’m Marcel! I work on #NLP for lesser-resourced languages, multilingual NLP, cross-lingual transfer, tokenisation, and trying to keep up with the flood of LLM-related research. Joining Bluesky seems like a good opportunity for a new & shiny #introduction, so here goes!
November 24, 2024 at 1:54 PM
They are *insanely* good at three things:

* Recruitment
* Coopting high profile minority resistance movements through tokenisation
* Setting up organisations that look like they might one day engage in radical liberatory struggle if you recruit enough people to read enough Marxism
May 25, 2026 at 7:52 PM