We're making everything public.
📄 Paper: arxiv.org/abs/2607.05171
🌐 Interactive website + demos: bridge-ai-neuro.github.io/rabbit/
💻 Code: github.com/bridge-ai-ne...
We would love to hear from you after trying it!
Huge thanks to @mtoneva.bsky.social for making this work possible.
We're making everything public.
📄 Paper: arxiv.org/abs/2607.05171
🌐 Interactive website + demos: bridge-ai-neuro.github.io/rabbit/
💻 Code: github.com/bridge-ai-ne...
We would love to hear from you after trying it!
Huge thanks to @mtoneva.bsky.social for making this work possible.
Third, Shared–Idiosyncratic Decomposition (SID) captures what responses share and how individuals differ. It gives us a population starting point, then lets us adapt only the small idiosyncratic heads. During few-shot adaptation, these idiosyncratic heads are the only part tuned in the model.
Third, Shared–Idiosyncratic Decomposition (SID) captures what responses share and how individuals differ. It gives us a population starting point, then lets us adapt only the small idiosyncratic heads. During few-shot adaptation, these idiosyncratic heads are the only part tuned in the model.
Second, our Temporal Brain Transformer lets each of the cortical regions learn what to attend to in the speech output. These regional representations feed a readout predicting ~41K cortical surface vertices. Speech goes in; a detailed brain prediction comes out
Second, our Temporal Brain Transformer lets each of the cortical regions learn what to attend to in the speech output. These regional representations feed a readout predicting ~41K cortical surface vertices. Speech goes in; a detailed brain prediction comes out
We brought together three ideas to make this work. First, brain-tuning: fine-tune a pretrained speech model directly on paired audio and fMRI from CNeuroMod Friends. Brain data helps shape the speech representations from which we predict cortical responses.
We brought together three ideas to make this work. First, brain-tuning: fine-tune a pretrained speech model directly on paired audio and fMRI from CNeuroMod Friends. Brain data helps shape the speech representations from which we predict cortical responses.
Second, each brain region in RABBiT learns its own representation. If we start in the primary auditory cortex and move to the most similar region, the model reconstructs the cortical progression: primary auditory → belt → STG/STS → temporal → frontal language, with no anatomical supervision.
Second, each brain region in RABBiT learns its own representation. If we start in the primary auditory cortex and move to the most similar region, the model reconstructs the cortical progression: primary auditory → belt → STG/STS → temporal → frontal language, with no anatomical supervision.
Beyond accuracy, we also asked: does RABBiT organize speech and language information in a brain-like way? The answer is “sounds like it does”.
First, it reproduces the classic left-lateralized language network, without being trained on any non-naturalistic experiment.
Beyond accuracy, we also asked: does RABBiT organize speech and language information in a brain-like way? The answer is “sounds like it does”.
First, it reproduces the classic left-lateralized language network, without being trained on any non-naturalistic experiment.
The few-shot gains are concentrated in higher-order language regions, regions that we found to be most idiosyncratic: IFG, angular gyrus, supramarginal gyrus, mPFC, MFG/DLPFC. These are regions that no strong zero-shot predictor will predict well because they are not shared across the population.
The few-shot gains are concentrated in higher-order language regions, regions that we found to be most idiosyncratic: IFG, angular gyrus, supramarginal gyrus, mPFC, MFG/DLPFC. These are regions that no strong zero-shot predictor will predict well because they are not shared across the population.
That tiny update beats voxel-wise ridge regression while updating roughly 1000× fewer parameters. Few-shot brain prediction becomes lightweight, fast, personalized, and realistic for settings where collecting hours of fMRI per person simply isn't an option.
That tiny update beats voxel-wise ridge regression while updating roughly 1000× fewer parameters. Few-shot brain prediction becomes lightweight, fast, personalized, and realistic for settings where collecting hours of fMRI per person simply isn't an option.
The few-shot result is where things get exciting for us. With only 5-10 minutes of fMRI, RABBiT personalizes its predictions to totally new subjects and stimuli.
Not by retraining the whole model, or fitting a huge voxel-wise model, but by updating a compact pathway (only ~115K parameters).
The few-shot result is where things get exciting for us. With only 5-10 minutes of fMRI, RABBiT personalizes its predictions to totally new subjects and stimuli.
Not by retraining the whole model, or fitting a huge voxel-wise model, but by updating a compact pathway (only ~115K parameters).
The zero-shot result is the first big gain. Across 324 unseen participants from 15 held-out studies, RABBiT reaches the inter-subject consistency estimate. RABBiT also outperforms the current state-of-the-art, TRIBEv2, across auditory and language regions (despite RABBiT being much smaller).
The zero-shot result is the first big gain. Across 324 unseen participants from 15 held-out studies, RABBiT reaches the inter-subject consistency estimate. RABBiT also outperforms the current state-of-the-art, TRIBEv2, across auditory and language regions (despite RABBiT being much smaller).
That gives one model two modes natively.
🟦 Zero-shot: no fMRI. Just speech in, and RABBiT predicts the reliable group-level brain response shared across people.
🟥 Few-shot: give it only ~10 minutes of the new person's fMRI, and RABBiT learns a tiny personalized correction.
That gives one model two modes natively.
🟦 Zero-shot: no fMRI. Just speech in, and RABBiT predicts the reliable group-level brain response shared across people.
🟥 Few-shot: give it only ~10 minutes of the new person's fMRI, and RABBiT learns a tiny personalized correction.
The key problem is that not all brain regions’ responses are shared across people.
Early auditory cortex is highly consistent, while higher-order language regions are more idiosyncratic. So RABBiT doesn't force one solution everywhere. It separates what's shared from what's idiosyncratic.
The key problem is that not all brain regions’ responses are shared across people.
Early auditory cortex is highly consistent, while higher-order language regions are more idiosyncratic. So RABBiT doesn't force one solution everywhere. It separates what's shared from what's idiosyncratic.
Foundation models changed AI because they learn strong priors and can adapt quickly to different tasks.
Brain encoding models never had this; existing models either produce shared predictions or need to be fit fully per participant.
We wanted the equivalent of a this for language in the brain.
Foundation models changed AI because they learn strong priors and can adapt quickly to different tasks.
Brain encoding models never had this; existing models either produce shared predictions or need to be fit fully per participant.
We wanted the equivalent of a this for language in the brain.