colab-links.bsky.social
@colab-links.bsky.social
Warning: a rather long read, but I loved it: The Emergent Symbolic Structure
of Artificial Neural Networks, McCoy et alia stars (2026)

https://arxiv.org/pdf/2608.29530
— The Data Therapist
https://arxiv.org/pdf/2608.29530
arxiv.org
September 9, 2026 at 9:45 PM
Old, but worth a look, reminding of the MoE approaches to separate languages
https://aclanthology.org/2022.naacl-main.255.pdf
— LChoshen
https://aclanthology.org/2022.naacl-main.255.pdf
aclanthology.org
August 21, 2026 at 9:43 PM
language-specific to appear EMNLP 2026
layers https://arxiv.org/abs/2605.26735

cc <@165490503570292736> <@430824383100092426>
— The Data Therapist
Rethinking the Multilingual Reasoning Gap with Layer Swap
Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, even when prompted in non-English languages. Prior work suggests that forcing the CoT to remain in the input language (\emph{native reasoning}) substantially degrades performance relative to allowing the model to reason in English before answering in the input language (\emph{English-pivoted reasoning}). However, most studies of this native reasoning gap rely on inference-time interventions or limited native-language training data. We revisit this comparison at a larger scale and under comparable supervision. We construct long multilingual reasoning datasets across six languages (English, French, German, Spanish, Chinese and Swahili); fine-tune specialists in both native and English-pivoted regimes on top of \texttt{Qwen/Qwen3-8B-Base}, and evaluate across mathematics, science, general knowledge, and code. In this setting, the average native reasoning gap shrinks to 1.9--3.5\% across the f
arxiv.org
August 21, 2026 at 2:02 PM
Hop on the discussion here if you are interested (both authors respond and in general I really like their work so, one can learn from them).
https://x.com/LChoshen/status/2089425036448792628?s=20
— LChoshen
Leshem (Legend) Choshen 🤖🤗 @ACL @ICML (@LChoshen) on X
@yoavartzi @nthngdy I agree, what I say is that you talk about how much you lose of the gradient you compute, and I say, but if you just compute less, the result would lose less and still be equivalent or worse probably? More steps=more information that's true, but a bit of a separate battle right?
x.com
August 17, 2026 at 7:02 PM
I like this paper, gives a new problem (need to think how strong it is, but interesting, and possibly also related to the other bottlenecks we are discussing in the compartmentalization settings.
https://arxiv.org/pdf/2603.10145
— LChoshen
https://arxiv.org/pdf/2603.10145
arxiv.org
August 17, 2026 at 7:02 PM
Interesting discussion,
Tldr Claude has much less tokens (15k?!) than any other and less as versions go, it appears to have a shift rather than capitalization which it learns as a separate character (basically morphology for English) https://x.com/i/status/2082882703024927088
— LChoshen
Tomasz Limisiewicz (@TomLimi) on X
@JumeletJ @magikarp_tokens For instance this paper shows that byte ByT5 (small vocab) outperforms subword mT5 (large vocab) on low resource translation (including Georgian pairs) https://t.co/U4DXWBLaiW
x.com
August 14, 2026 at 12:00 PM