#backprop
You Can Backprop Anything You Want. It's Just The Chain Rule. Jacobians Don't Care What Kind Of Algebraic Bullshit You're Jamming Into Them.
March 13, 2026 at 6:21 PM
A/B testing is far more dangerous than backprop.
People fear LLMs will soon distort human perception of the world, and meanwhile some shallow optimization method trying to maximize click-rate has been wreaking havoc for years.
July 5, 2025 at 1:14 PM
Introducing PC-ALM: a local-learning alternative to backprop that trains 1000-layer neural nets using only local dynamics.

Blog: pub.sakana.ai/pc-alm/
September 14, 2026 at 8:09 PM
Oh, worst bit is, it's not even a complicated mathematical function. If you've done chain rule, you can do backprop. If you can do backprop, you can build a vanilla perceptron. If you can do that, transformers and llms are impressively boring.
August 8, 2026 at 11:33 AM
please tell me they do backprop via trebuchet
June 19, 2026 at 10:50 PM
Artificial neural networks use backprop, which attributes errors starting at the output layer, then works backwards. Cerebellum is also a multilayer network, but learning does not rely on backprop.
How does the cerebellum guide learning in all of its layers? David Ehrlich explains.
May 10, 2025 at 5:47 PM
“AI expert” to a computer scientist: you’re working on optimizing model architectures and improving backprop.

“AI expert” for real: you’re gluing transformers together with Python and hoping you get better results.
October 2, 2025 at 12:40 AM
because look at it. it's a concrete box. it's nearly atheoretic. "we don't need any specialized structures or steps, just build the whole thing with gradient-descent backprop"
September 19, 2026 at 8:10 PM
backprop
July 8, 2026 at 1:02 AM
building my own mlp implementation from scratch in numpy, including backprop, remains one of the most educational exercises I’ve done
November 30, 2024 at 4:18 AM
you just drop a token labeled with a slot (that's inference) in a slot at the top, see what slot it ends up in, and iterate back up the board moving the pegs (that's backprop) such that it would end up in the correct slot if you dropped it again. repeat several gazillion times.
June 11, 2025 at 2:05 AM
Excited to share our new research on local RL without backprop!
CDS Assistant Professor of Computer Science and Data Science @mengyer.bsky.social and co-author Frank Wu introduce ARQ, a new learning algorithm that skips backpropagation in favor of a more biologically plausible and computationally efficient method.

nyudatascience.medium.com/ditching-bac...
Ditching Backpropagation: A New Method Mimics How Brains Learn to Make Decisions
Mengye Ren and Frank Wu introduce a biologically-inspired learning algo that outperforms traditional methods without using backpropagation.
nyudatascience.medium.com
December 12, 2025 at 12:45 AM
Someone should take the idea of money as backprop gradients seriously.
November 18, 2024 at 3:57 PM
Great write up. I made NNs in excel in the 90s as a teen, and similar. I know backprop, AD, a bunch of metaheuristic algorithms etc.
AI enthusiasts look down on me saying im anti AI and I should trust that they have superior knowledge about all things.
AI is great.
This (gestures around) is not.
September 23, 2026 at 1:02 AM
This isn’t uncommon at all, thankfully. Many technical people who actually touch these systems, and many neuroscientists, don’t think transformer-based backprop models are able to get there because of how inefficient they are.
August 13, 2025 at 10:31 PM
Interviewer: Explain backprop to me

Applicant: PyTorch
November 7, 2025 at 2:43 PM
How does the thing that talks work? Nobody knows because no human wrote it, and the code you get back is continuous (so it can be optimized by the optimizer which implements a giant string of calculus called backprop) and polysemantic, meaning you can't isolate parts to see how they work easily.
August 23, 2025 at 8:48 PM
Introducing 🥚EGGROLL 🥚(Evolution Guided General Optimization via Low-rank Learning)! 🚀 Scaling backprop-free Evolution Strategies (ES) for billion-parameter models at large population sizes

⚡100x Training Throughput
🎯Fast Convergence
🔢Pure Int8 Pretraining of RNN LLMs
November 21, 2025 at 5:56 PM
It would be cool if models could generate strings they want to remember, and then instead of just appending to a memory file, the harness would actually backprop those strings.
December 12, 2025 at 7:05 PM
How's that with backprop nowadays? A meme I created about 10 years ago when that discussion was raging.
February 9, 2026 at 4:42 PM
Is the cerebellum going to turn out to effectively be something like a backprop device
September 16, 2026 at 8:17 PM
Remember how I said backprop will die soon? @sakanaai.bsky.social has begun that process, and this isn't the only promising avenue that's doing it:
Augmented Lagrangian Predictive Coding: training 1000-layer networks without backpropagation
A local alternative to backpropagation. PC-ALM trains residual MLPs up to 1000 layers, nearly matching backprop's performance despite using only layer-local dynamics. PC-ALM equips each layer with a f...
pub.sakana.ai
September 14, 2026 at 3:32 PM
Hot take: this (aka backprop) was the most important scientific ingredient necessary to enable deep learning. The other main drivers are more technical: huge annotated datasets, massively parallel processors.
Oh, and for ChatGPT we can add exploitation of unpaid labor and plundering of art works.
The Cheap Gradient Principle (Baur—Strassen, 1983) states that computing gradients via automatic differentiation is efficient: the gradient of a function f can be obtained with a cost proportional to evaluating f
January 16, 2025 at 7:31 AM
Pretty much none of the truly simple methods in ML scale well. SVM, kNN, random forests are some of the simplest methods out there, and they don't scale at all. Meanwhile "train a transformer via backprop and gradient descent" is a very high-entropy method,
April 17, 2026 at 8:13 AM
u like can't backprop this tho because your path length increases without limit
August 11, 2026 at 5:34 PM