#Multilinguality
We all know about the curse of multilinguality. We know that empirical performance degrades as you add languages to a model.
But *in theory*, does it have to?
Let’s talk about the theoretical curse of multilinguality for embedding space structure.
September 23, 2026 at 1:17 PM
Teaching young kids importance of multilinguality and why Lisowica is a perfect organism
February 7, 2026 at 5:24 AM
My thoughts on multilinguality in SFF is always thus: if seasoned SFF readers can absorb conlang, complex structures and twisty plots without much exposition, why can't they absorb earth languages that might be foreign to them?
September 19, 2026 at 7:11 AM
Multilinguality aims for two things: shared high-quality monolingual semantics, and cross-lingual alignment. We define these as our conditions for perfect multilinguality.
September 23, 2026 at 1:17 PM
April 30, 2025 at 11:18 PM
The above scaling is mild. We show that it is possible to fit 7000 languages in 30 extra dimensions while maintaining perfect multilinguality. Thus, we argue there is no *theoretical* curse of multilinguality for embedding space structure.
September 23, 2026 at 1:17 PM
First, we need to define the goal: perfect multilinguality for embedding spaces. Then, we'll see what perfect multilinguality costs in theory as we scale up language coverage. If the capacity cost is high, we have a theoretical curse. If not, no theoretical curse!
September 23, 2026 at 1:17 PM
Really fascinating post by Harry Josephine Giles on multilinguality and the translation thereof, minority languages and code-switching (partly also about Krzysztof Bartnicki's translation of her lyric novel Deep Wheel Orcadia into Polish).

hjosephinegiles.substack.com/p/in-transla...
In Translation
Notes on multilinguality
hjosephinegiles.substack.com
December 29, 2025 at 9:48 PM
We would be so culturally impoverished, not to mention multilinguality has only positive effects on brain.

No borders, sure. But language rights are, I'd say, worth dying for.
In a perfect world there would be one language, no borders one humanity
August 24, 2026 at 6:58 AM
Happy to say that our paper "Beyond Literal Token Overlap: Token Alignability for Multilinguality" will be presented at #NAACL2025!

This is work with @tomlim.bsky.social, @jlibovicky.bsky.social, and Alex Fraser.

arxiv.org/abs/2502.06468

#newpaper #NLP #NLProc
Beyond Literal Token Overlap: Token Alignability for Multilinguality
Previous work has considered token overlap, or even similarity of token distributions, as predictors for multilinguality and cross-lingual knowledge transfer in language models. However, these very li...
arxiv.org
March 3, 2025 at 5:04 PM
📢Life update📢

🥳I'm excited to share that I've started as a postdoc at Uppsala University NLP @uppsalanlp.bsky.social, working with Joakim Nivre on topics related to constructions and multilinguality!

🙏Many thanks to the Walter Benjamin Programme of the DFG for making this possible.
September 15, 2025 at 3:10 PM
I'm in Suzhou to present our work on MultiBLiMP, Friday @ 11:45 in the Multilinguality session (A301)!

Come check it out if your interested in multilingual linguistic evaluation of LLMs (there will be parse trees on the slides! There's still use for syntactic structure!)

arxiv.org/abs/2504.02768
November 6, 2025 at 7:08 AM
We then look the empirical curse of multilinguality for embedding space structure on small scale models for 100 languages for 8 training and evaluation configurations and find some non-cursed settings.
September 23, 2026 at 1:17 PM
Our paper pushes for stating what we want from multilinguality, for being precise in defining our famous curse, for thinking about theoretical solutions to this curse, and for questioning the gap between theory and practice.
September 23, 2026 at 1:17 PM
There is No Theoretical Curse of Multilinguality For Embedding Space Structure arxiv.org/abs/2608.17088
September 24, 2026 at 2:27 PM
Modern LLMs "speak" hundreds of languages... but do they really?
Multilinguality claims are often based on downstream tasks like QA & MT, while *formal* linguistic competence remains hard to gauge in lots of languages

Meet MultiBLiMP!
(joint work w/ @jumelet.bsky.social & @weissweiler.bsky.social)
✨New paper ✨

Introducing 🌍MultiBLiMP 1.0: A Massively Multilingual Benchmark of Minimal Pairs for Subject-Verb Agreement, covering 101 languages!

We present over 125,000 minimal pairs and evaluate 17 LLMs, finding that support is still lacking for many languages.

🧵⬇️
April 8, 2025 at 12:27 PM
Our paper 'Beyond Literal Token Overlap: Token Alignability for Multilinguality' will be at #NAACL2025! We show that token alignability is a stronger predictor of cross-lingual transfer than literal token overlap.

Read it here: arxiv.org/abs/2502.06468
March 10, 2025 at 3:48 PM
🥳I'm excited to introduce everyone to my group at Leipzig University, the LION (Linguistically-Oriented NLP) Lab! 🦁

Looking forward to exciting work on computational linguistics, interpretability, and multilinguality!

Find out more on our website lionlabnlp.github.io
June 23, 2026 at 1:51 PM
Let's talk about eval (automatic or human) and multilinguality at #EMNLP in Suzhou! 🇨🇳

- Efficient evaluation (Nov 5, 16:30, poster session 3)
- MT difficulty (Nov 7, 12:30, findings 3)
- COMET-poly (Nov 8, 11:00, WMT)

(DM to meet 🌿 )
October 28, 2025 at 9:45 AM
#NAACL2025 ended more than a week ago & @ufal-cuni.bsky.social folks were there:
Main conf: @kathaem.bsky.social presented joint work w/ @tomlim.bsky.social, @jlibovicky.bsky.social and Alex Fraser: Beyond Literal Token Overlap: Token Alignability for Multilinguality aclanthology.org/2025.naacl-s...
May 16, 2025 at 12:07 PM
Paper 👉Beyond Literal Token Overlap: Token Alignability for Multilinguality👈 by @kathaem.bsky.social, @tomlim.bsky.social, @jlibovicky.bsky.social and Alex Fraser will appear at #NAACL2025! arxiv.org/abs/2502.06468 Congratulations to all authors! 🥳
March 10, 2025 at 3:52 PM
Bi- or multilinguality is a gift. That is why Reform Ltd wants to take it away.
December 4, 2025 at 6:41 AM
Multilinguality ♥️
August 16, 2025 at 7:23 PM
Highlights from multilingual #NLP and machine translation papers I found on arXiv in November are now on my blog: jlibovicky.github.io/2024/12/05/M...
Highlights from Machine Translation and Multilinguality in November 2024
Mitigating Metric Bias in Minimum Bayes Risk Decoding
jlibovicky.github.io
December 6, 2024 at 5:07 PM
In a different universe:

Why do we publish English only resources?
linguistic findings?❌ weak due to focus on a single language
support model multilinguality? ❌so why this specific morphologically poor language?
Spoken by a large population?❌ why don't work on Mandarin instead?
December 4, 2024 at 5:42 PM