PC-ALM uses feedback control and layer-local recurrence
1 Sep 14 2:03 PM · 13d ago · 1 article · 2 posts · 3 sources · development 1 of 1
The method equips each layer with a feedback control dynamical system and introduces dual neurons (Lagrange multipliers) per layer, making each layer's local recurrence a PI feedback controller that distributes supervision signals across the network.
“Each layer is coupled only to its neighbors. Instead of forward-then-backward, we run each layer forward in time. When run to convergence, the dynamics of the whole system distribute supervision credit signals quickly and accurately across the entire network.”
Sakana AISakana AI Research organization
The whole story articlesposts the bright band is this development · numbered dots are the others · click one to jump
Reported in the same hours no headline names this development itself — these 1 claim were published in its stretch
-
first by HN Frontpage, 13d ago
What people said 8 voices · best of 9 · verbatim
-
Oh wow the theoretical implications in neuroscience exite me here - is this a potential model of Fristons Markov Blanket concept“ Probably the most ambitious and all-encompassing version of the ‘Bayesian turn’ in cognitive science is the free energy principle (FEP). The FEP is a mathematical framework, developed by Karl Friston and colleagues…
-
There is a lot of interesting research into predictive coding as an alternative means to solve the credit assignment problem that might be a more plausible model of what happens in the brain.I really liked this paper that showed using a predictive coding learning rule leads to the exact same gradients as backprop in arbitrary networks:Predictive…
-
That was genuinely a strong signal that whatever learning algorithm the brain uses isn't "magic", and probably can be approximated with the ML tools we have.It also pointed at the possibility that the learning algorithms brain uses might be, like the paper has demonstrated, less compute-optimal and data-optimal than backprop - but far easier to…
-
Nice introduction to a simple but useful idea! The Lagrangian works like a time-smoothed optimizing direction state, but it can be placed on any wire, even at non-differentiable boundary! Can it be better than existing training methods for discrete components like argmax, MoE or VQ-VAE? Maybe networks can be composed by a lot of learnable discrete…
-
This paper uses the "fixed prediction assumption" so I think has caused some confusion (i.e. it's not PC, but PC with a small bandaid). It's a great paper though just the title is misleading somewhat. But all of Beren Millidge's papers are quite good and he was so prolific in the space. Cool to have seen him on the Dwarkesh podcast recently as…
-
The brain doesn't need to do back propagation. Back propagation is a simulation of gradient descent which the universe achieved naturally by its own devices. Indeed if you tie the loss function to some sort of excited energy state in such a way then leave it to be, the universe will naturally drive the system to its local minima.
-
I wonder if you could take a traditional backprop trained LLM and apply this approach to finetuning it (presumably needs less memory and compute?). It could be another entry in the spectrum between LORA and full fine tuning.
-
Could this relate to continual learning? It lets you update without pausing the entire system.
All 1 developments of Sakana AI introduces PC-ALM, a local alternative to… →
Hacker NewsMastodonNewswires