Andrew Saxe
banner
saxelab.bsky.social
Andrew Saxe
@saxelab.bsky.social
Professor at the Gatsby Unit and Sainsbury Wellcome Centre, UCL, trying to figure out how we learn
Pinned
Why don’t neural networks learn all at once, but instead progress from simple to complex solutions? And what does “simple” even mean across different neural network architectures?

Sharing our new paper @iclr_conf led by Yedi Zhang with Peter Latham

arxiv.org/abs/2512.20607
Reposted by Andrew Saxe
Understanding how learners conclude “X laughed Y” is incorrect is an age-old question, with several hypotheses, some of which have been ~impossible to disentangle!

@tomyxw.bsky.social, @fredashi.bsky.social, and I use controlled rearing to shed light on this in our new EMNLP paper:

1/n
September 9, 2026 at 4:40 PM
Reposted by Andrew Saxe
Confused about all this talk of compositionality in neuroscience? Read our new perspective www.nature.com/articles/s41... with authors Reidar Riveland and Alex Pouget.
The compositionality continuum as a principle for studying the neural basis of intelligence - Nature Neuroscience
Compositionality exists on a continuum of increasing complexity rather than as a binary trait. Studying its implementation in simpler biological and artificial systems is the most tractable path towar...
www.nature.com
August 8, 2026 at 9:21 PM
Reposted by Andrew Saxe
If an action results in error, each neuron requires an individualized teaching signal that guides change in its output. This is the credit assignment problem of learning. Are there neurons in the brain that can compute such a sophisticated teaching signal? Yes.
www.biorxiv.org/content/10.6...
Climbing fibers encode the gradient of a loss function for the cerebellum
Neurons in the brain are often many synapses away from motoneurons, yet if a movement results in error, each distant neuron needs a teacher that considers its specific contribution to production of th...
www.biorxiv.org
July 28, 2026 at 3:21 PM
Reposted by Andrew Saxe
Pretraining + fine-tuning powers modern ML, but we lack a theoretical understanding of how pretraining actually shapes downstream learning.

In our new @icmlconf.bsky.social paper, we address this gap!

📅 July 9th, Poster #4502 Session 8!
🧵

arxiv.org/pdf/2602.20062
July 8, 2026 at 4:11 PM
Reposted by Andrew Saxe
The lab is at @fens.org, with 5 posters — spatial & auditory working memory, prefrontal decision-making, dopamine in statistical learning, and a 2-photon rat-brain atlas for BrainGlobe.
Unfortunately, I couldn't make it this year, but come find the rest of the lab on Tue, Thurs, & Fri!

#FENS2026
July 6, 2026 at 11:30 AM
Reposted by Andrew Saxe
Is Muon as good as they say? We looked beyond training speed and found a hidden cost: Muon loses the simplicity bias of older optimizers like gradient descent — and this matters for generalization.

Led by Sara Dragutinović and advised by Rajesh Ranganath

arxiv.org/abs/2603.00742
July 2, 2026 at 6:09 PM
Reposted by Andrew Saxe
Ever wondered how the hippocampal cognitive map is read-out?

@changmin-yu.bsky.social @zilong-ji.bsky.social with Jake & John, show that, during navigation, theta sweeps indicate remembered goal-directions (cf. current/next movements or perceptual targets) 1/2

www.nature.com/articles/s41...
July 1, 2026 at 9:35 AM
Reposted by Andrew Saxe
1/7 Excited to share my last PhD article, just accepted to ICML 2026! In it, we (me, Alexandre Payeur, Guillaume Lajoie) used dynamical systems theory to study "local" learning in linear recurrent neural networks. See link for the paper, and thread for a brief summary. arxiv.org/abs/2606.00243
Dynamics and Representation Structure of Local Approximations to Gradient-Based Learning in Linear Recurrent Neural Networks
Biological and neuromorphic recurrent neural networks (RNNs) are subject to spatial and temporal locality constraints on the information that can plausibly be used during learning. A common strategy t...
arxiv.org
June 10, 2026 at 12:27 PM
Reposted by Andrew Saxe
Come chat about this @iclr-conf.bsky.social!
Friday 3:15 PM, Pavilion 4, Poster #4216
Why don’t neural networks learn all at once, but instead progress from simple to complex solutions? And what does “simple” even mean across different neural network architectures?

Sharing our new paper @iclr_conf led by Yedi Zhang with Peter Latham

arxiv.org/abs/2512.20607
April 23, 2026 at 6:04 PM
Reposted by Andrew Saxe
We’ve got an exciting new thing to share! We have causal evidence (using TMR) that memory reactivation during sleep promotes abstract understanding of underlying structure, allowing transfer learning in a new domain with zero superficial feature overlap with the learned one.
Super excited to share this preprint! How do we disentangle underlying structure from the particular features of a learning episode to benefit future learning? We find that memory reactivation during sleep promotes this structure abstraction process.

www.biorxiv.org/content/10.6...
www.biorxiv.org
April 13, 2026 at 3:53 PM
Reposted by Andrew Saxe
New preprint! 🧠
How do RNNs learn abstract rules from sequences, independent of specific stimuli?

By Vezha Boboeva, with Alberto Pezzotta & George Dimitriadis

"From sequences to schemas: low-rank recurrent dynamics underlie abstract relational representations"
www.biorxiv.org/content/10.6...
April 13, 2026 at 3:54 PM
Reposted by Andrew Saxe
Two Analytical Connectionism-related updates:

1. ⏰ 1 week left to apply! Interested in language + AI & cognition? Don’t miss it: www.analytical-connectionism.net/school/2026/

2. 📜 Lecture notes from the first two editions are finally out: proceedings.mlr.press/v320/
April 10, 2026 at 3:09 PM
Reposted by Andrew Saxe
We’re hiring a Group Leader!

Join us to lead a transformative initiative in human systems neuroscience.

Find out more and apply ⤵️

www.sainsburywellcome.org/content/curr...
February 13, 2026 at 1:41 PM
Postdoc opening!

Come work with us on deep learning theory relevant to AI safety

Deadline: 7 Apr 2026
Details and application: www.ucl.ac.uk/work-at-ucl/...
UCL – University College London
UCL is consistently ranked as one of the top ten universities in the world (QS World University Rankings 2010-2022) and is No.2 in the UK for research power (Research Excellence Framework 2021).
www.ucl.ac.uk
April 2, 2026 at 9:23 AM
Very excited by this year's Analytical Connectionism Summer School!

A dream lineup of speakers on the topic of language acquisition in minds and machines

Bursaries available to cover costs

Aug 17 – Aug 28, 2026 Gothenburg

Details: www.analytical-connectionism.net//school/2026/
April 2, 2026 at 9:17 AM
Reposted by Andrew Saxe
A great entry into the proposals available for physiologically plausible gradient descent!

I think the way they use dendrite targeting inhibition in this model is particularly elegant.

Time to start testing these ideas folks!!!

#neuroscience 🧪 #NeuroAI
Our latest publication grapples with how the brain could implement gradient descent by sending learning targets top-down, gating plasticity with dendritic inhibition, and updating synaptic weights with biologically observed learning rules like BTSP.

www.cell.com/cell-reports...
March 27, 2026 at 1:17 PM
Reposted by Andrew Saxe
The First 1,000 Days (1kD) Project - Collecting and Analyzing an Ultra-Dense Naturalistic Dataset of Human Baby Development https://www.biorxiv.org/content/10.64898/2026.03.19.712982v1
March 23, 2026 at 10:15 AM
Reposted by Andrew Saxe
Looking for alternatives to quadratic functions for closed-form analysis in optimization? This post explores matrix Riccati dynamics and their applications to neural networks. francisbach.com/closed-form-...
March 5, 2026 at 3:36 PM
Reposted by Andrew Saxe
Here's a lovely #blueprint on a new study from our lab led by @royeyono.bsky.social.

tl;dr: it implies that there may be interneurons whose role is to normalize credit assignment signals during learning.

#neuroscience 🧪
March 19, 2026 at 4:36 PM
Reposted by Andrew Saxe
A new Department of Cognitive Science is being created at Bocconi University in Milan, Italy.

Here is the call for a cluster hire (for around 10 faculty) in all areas of cognitive science, at both junior and senior levels:

www.unibocconi.it/en/faculty-a...

Deadline: May 4th, 2026
Open Rank Faculty Cluster Hire Search for the New Department of Cognitive Science at Bocconi - Bocconi University
www.unibocconi.it
March 12, 2026 at 12:49 PM
Reposted by Andrew Saxe
Poster tonight at #cosyne26 (1-079)!

@wanqingjiang.bsky.social & @noehamou.bsky.social show that mice learn hidden community structure in a 15-odour graph even when transition statistics are flat.

Fun collaboration with @saxelab.bsky.social that started with East London coffees ☕!
March 12, 2026 at 4:57 PM
Reposted by Andrew Saxe
Really neat work by Fountas and colleagues at UCL:
arxiv.org/abs/2603.04688
They propose that consolidation reflects a form of "predictive forgetting" that aids generalization.
Why the Brain Consolidates: Predictive Forgetting for Optimal Generalisation
Standard accounts of memory consolidation emphasise the stabilisation of stored representations, but struggle to explain representational drift, semanticisation, or the necessity of offline replay. He...
arxiv.org
March 11, 2026 at 5:58 PM
Reposted by Andrew Saxe
Thanks @natmesanash.bsky.social for covering our new work, in @thetransmitter.bsky.social!
March 10, 2026 at 7:59 PM
Reposted by Andrew Saxe
📢📢 Announcing this year's conference on the Mathematics of Neuroscience & AI (Rome, 9-12th June). We’ve got a stellar line-up and venue, and invite everyone to join:

www.neuromonster.org
March 5, 2026 at 8:31 AM
Reposted by Andrew Saxe
📢 Job alert - Deep Learning Theory & AI Safety
Applications open for a postdoc fellow (@saxelab.bsky.social lab) to study artificial deep networks using techniques from applied maths & stat physics.

⏰ Deadline: 26 Mar 2026
🤝 In collaboration with @stefsm.bsky.social
ℹ️ www.ucl.ac.uk/life-science...
February 25, 2026 at 11:18 AM