Sakana AI introduces PC-ALM, a local alternative to backpropagation
A new training method uses layer-local dynamics to match backprop performance on deep networks, mimicking how biological brains might solve credit assignment.
What to know
- Sakana AI introduced PC-ALM, a method that trains deep neural networks using layer-local dynamics instead of backpropagation, achieving near-equivalent performance on networks up to 1000 layers.
- The approach addresses how biological brains might solve credit assignment—distributing learning signals across layers—without backprop's requirement for strict timing coordination between forward passes, backward passes, and weight updates.
- Beyond neuroscience, PC-ALM could enable energy-efficient deep learning on neuromorphic hardware where simulating dynamical systems is cheaper than GPU computation.
Sakana AI Research organization
How it unfolded 1 development · click the chart to see its coverage articlesposts
-
1
PC-ALM uses feedback control and layer-local recurrence
The method equips each layer with a feedback control dynamical system and introduces dual neurons (Lagrange multipliers) per layer, making each layer's local recurrence a PI feedback controller that distributes supervision signals across the network.
“Each layer is coupled only to its neighbors. Instead of forward-then-backward, we run each layer forward in time. When run to convergence, the dynamics of the whole system distribute supervision credit signals quickly and accurately across the entire network.”
— Sakana AI -
Oh wow the theoretical implications in neuroscience exite me here - is this a potential model of Fristons Markov Blanket concept“ Probably the most ambitious and all-encompassing version of the ‘Bayesian turn’ in cognitive science is the free energy principle (FEP). The FEP is a mathematical framework, developed by Karl Friston and colleagues…
2 more of the top 3 · 9 posts in this stretch
-
There is a lot of interesting research into predictive coding as an alternative means to solve the credit assignment problem that might be a more plausible model of what happens in the brain.I really liked this paper that showed using a predictive coding learning rule leads to the exact same gradients as backprop in arbitrary networks:Predictive…
-
That was genuinely a strong signal that whatever learning algorithm the brain uses isn't "magic", and probably can be approximated with the ML tools we have.It also pointed at the possibility that the learning algorithms brain uses might be, like the paper has demonstrated, less compute-optimal and data-optimal than backprop - but far easier to…
-
-
background
Method addresses brain's credit assignment problem — The paper frames PC-ALM as a solution to how biological brains distribute learning signals across layers without backpropagation's requirement for strict phase locking—where neurons must wait for error signals before updating weights.
-
background
Sakana AI publishes PC-ALM training method — Sakana AI released a research paper introducing PC-ALM, a local alternative to backpropagation that trains residual MLPs up to 1000 layers nearly matching backprop's performance using only layer-local dynamics.
What people are saying 5 voices from 1 site · best of 9 · verbatim
- Sep 15
-
The brain doesn't need to do back propagation. Back propagation is a simulation of gradient descent which the universe achieved naturally by its own devices. Indeed if you tie the loss function to some sort of excited energy state in such a way then leave it to be, the universe will naturally drive the system to its local minima.
-
This paper uses the "fixed prediction assumption" so I think has caused some confusion (i.e. it's not PC, but PC with a small bandaid). It's a great paper though just the title is misleading somewhat. But all of Beren Millidge's papers are quite good and he was so prolific in the space. Cool to have seen him on the Dwarkesh podcast recently as…
-
Nice introduction to a simple but useful idea! The Lagrangian works like a time-smoothed optimizing direction state, but it can be placed on any wire, even at non-differentiable boundary! Can it be better than existing training methods for discrete components like argmax, MoE or VQ-VAE? Maybe networks can be composed by a lot of learnable discrete…
-
I wonder if you could take a traditional backprop trained LLM and apply this approach to finetuning it (presumably needs less memory and compute?). It could be another entry in the spectrum between LORA and full fine tuning.
- Sep 14
-
Could this relate to continual learning? It lets you update without pausing the entire system.