Omer Moussa
omermosa.bsky.social
Omer Moussa
@omermosa.bsky.social
Neurosci x AI PhD Student @the Max Planck Institute for Software Systems, supervised by @mtoneva.bsky.social; CS@MaxPlanck -- ML and CogNeurosci Enthusiast.
13/

We're making everything public.

📄 Paper: arxiv.org/abs/2607.05171
🌐 Interactive website + demos: bridge-ai-neuro.github.io/rabbit/
💻 Code: github.com/bridge-ai-ne...

We would love to hear from you after trying it!

Huge thanks to @mtoneva.bsky.social for making this work possible.
RABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brain
Language understanding in the brain is context-dependent, varying across experimental stimuli and individuals, which makes it difficult to build computational models that generalize across both. This ...
arxiv.org
October 1, 2026 at 2:08 PM
12/

Third, Shared–Idiosyncratic Decomposition (SID) captures what responses share and how individuals differ. It gives us a population starting point, then lets us adapt only the small idiosyncratic heads. During few-shot adaptation, these idiosyncratic heads are the only part tuned in the model.
October 1, 2026 at 2:08 PM
11/

Second, our Temporal Brain Transformer lets each of the cortical regions learn what to attend to in the speech output. These regional representations feed a readout predicting ~41K cortical surface vertices. Speech goes in; a detailed brain prediction comes out
October 1, 2026 at 2:07 PM
10/

We brought together three ideas to make this work. First, brain-tuning: fine-tune a pretrained speech model directly on paired audio and fMRI from CNeuroMod Friends. Brain data helps shape the speech representations from which we predict cortical responses.
October 1, 2026 at 2:07 PM
9/

Second, each brain region in RABBiT learns its own representation. If we start in the primary auditory cortex and move to the most similar region, the model reconstructs the cortical progression: primary auditory → belt → STG/STS → temporal → frontal language, with no anatomical supervision.
October 1, 2026 at 2:06 PM
8/

Beyond accuracy, we also asked: does RABBiT organize speech and language information in a brain-like way? The answer is “sounds like it does”.

First, it reproduces the classic left-lateralized language network, without being trained on any non-naturalistic experiment.
October 1, 2026 at 2:05 PM
7/
The few-shot gains are concentrated in higher-order language regions, regions that we found to be most idiosyncratic: IFG, angular gyrus, supramarginal gyrus, mPFC, MFG/DLPFC. These are regions that no strong zero-shot predictor will predict well because they are not shared across the population.
October 1, 2026 at 2:04 PM
6/

That tiny update beats voxel-wise ridge regression while updating roughly 1000× fewer parameters. Few-shot brain prediction becomes lightweight, fast, personalized, and realistic for settings where collecting hours of fMRI per person simply isn't an option.
October 1, 2026 at 2:03 PM
5/

The few-shot result is where things get exciting for us. With only 5-10 minutes of fMRI, RABBiT personalizes its predictions to totally new subjects and stimuli.

Not by retraining the whole model, or fitting a huge voxel-wise model, but by updating a compact pathway (only ~115K parameters).
October 1, 2026 at 2:02 PM
4/

The zero-shot result is the first big gain. Across 324 unseen participants from 15 held-out studies, RABBiT reaches the inter-subject consistency estimate. RABBiT also outperforms the current state-of-the-art, TRIBEv2, across auditory and language regions (despite RABBiT being much smaller).
October 1, 2026 at 2:00 PM
3/

That gives one model two modes natively.

🟦 Zero-shot: no fMRI. Just speech in, and RABBiT predicts the reliable group-level brain response shared across people.

🟥 Few-shot: give it only ~10 minutes of the new person's fMRI, and RABBiT learns a tiny personalized correction.
October 1, 2026 at 1:59 PM
2/

The key problem is that not all brain regions’ responses are shared across people.

Early auditory cortex is highly consistent, while higher-order language regions are more idiosyncratic. So RABBiT doesn't force one solution everywhere. It separates what's shared from what's idiosyncratic.
October 1, 2026 at 1:57 PM
1/

Foundation models changed AI because they learn strong priors and can adapt quickly to different tasks.

Brain encoding models never had this; existing models either produce shared predictions or need to be fit fully per participant.
We wanted the equivalent of a this for language in the brain.
October 1, 2026 at 1:55 PM