By treating pyramidal neurons as three-compartment systems — where sensory input, expected value, and error are spatially separated — Lv and colleagues recently demonstrated that multilayer neural networks can be trained without a global error signal, reaching 97.57% on MNIST against backpropagation’s 98.62% while never computing one.

The constraints of biological learning

To train deep neural networks today, we rely overwhelmingly on backpropagation. It is the workhorse of modern deep learning, but its biological plausibility has been heavily debated in the NeuroAI community. The criticisms, as outlined in Lv et al.’s 2025 paper, boil down to three primary, stubborn requirements that real brains simply do not meet.

First, backpropagation requires weight symmetry. The backward pass has to use the same numbers as the forward pass, mirrored: every feedback connection carrying the exact strength of its forward counterpart. Real synapses are biochemical cascades, not shared memory, and have no obvious way of keeping two such copies in step — the “weight transport” problem. Retrograde signalling does exist, but nothing about it looks like transporting a weight matrix.

Second, backpropagation relies on a global error signal. To update a connection in the first layer of a deep network, you must calculate the exact error at the final output and carefully pass that specific error backward through every intermediate layer via the chain rule. A biological synapse, by contrast, is believed to change strength on what it can sense where it sits. A synapse deep in the visual cortex does not have access to a bird’s-eye view of the entire brain’s performance.

Third, backpropagation demands a two-stage training process. There is a distinct forward pass (inference) followed by a distinct backward pass (updating). The brain does not pause its experience of the world to run a backward pass. Whatever its equivalent of those two phases turns out to be, they have to be able to run at the same time.

While alternative algorithms like Feedback Alignment or Predictive Coding address some of these issues, Lv et al. note that most either fail to meet all three at once or, where they do meet them, fall apart on anything complex.

Why the gap matters

It is worth being precise about why anyone should care, because two quite different motivations get bundled together here and they carry different burdens of proof.

The first is scientific. If we want to claim we understand how cortex learns, “it runs backpropagation” is not available to us, and something has to take its place. That is a question about brains, and the standard of evidence is whether the proposed rule matches what neurons are observed to do.

The second is engineering. Local rules avoid shipping an error signal across an entire network, which maps far better onto neuromorphic hardware, where moving data costs more than computing on it. That is a question about chips, and the standard is whether it holds up at scale.

These come apart more often than the literature admits. An algorithm can be excellent engineering and poor neuroscience, or the reverse, and DLL should be judged on both counts separately.

A word on the energy argument, which usually turns up around here and which I have leaned on myself. The brain runs on roughly 20 watts; a large training run consumes many orders of magnitude more. True, but not a like-for-like comparison: it sets one brain doing inference against a cluster training a model that will then serve millions of people. The defensible version is narrower and still worth having. Within a fixed piece of hardware, avoiding a global backward pass saves data movement, and data movement is where most of the energy goes. That is a genuine advantage. It is not grounds for claiming the brain is a millionfold more efficient than a GPU at the same task, because nobody has measured the same task.

DLL also arrives into a crowded field, and it is worth being accurate about what is already solved. Satisfying all three criteria at once is not new: Hebbian learning, R-STDP, Forward-Forward and Equilibrium Propagation all manage it, as does vanilla target propagation. The trouble is that they tend to collapse on anything harder than a toy benchmark — Forward-Forward CNNs sit at chance on CIFAR-10. So the interesting claim here is not that local learning is possible, nor even that all three constraints can be met simultaneously. It is that they can be met while the network still learns something.

Dendritic Localized Learning (DLL)

To address this, the authors propose Dendritic Localized Learning (DLL). Instead of attempting to modify backpropagation incrementally, they drew inspiration directly from the pyramidal neuron.

Pyramidal neurons are the most common neuron in the cerebral cortex — Lv et al. [1] put them at roughly 70-85% of all cortical neurons, citing DeFelipe and Fariñas [2], though you will see the range quoted a little lower elsewhere — and they are central to accounts of cortical learning and memory. Following Sacramento and colleagues [3], the authors divide the artificial pyramidal neuron into three distinct compartments: a soma (the cell body), an apical dendrite, and a basal dendrite. The compartmental scheme is inherited rather than invented here; what DLL contributes is the learning rule that runs on it.

In the DLL framework, the local error is not handed down from a global supervisor. Instead, it is computed locally within the soma. The sensory input (denoted as $u_i$) — what the neuron is currently receiving from the layer below — is stored in the basal dendrite. The expected value (denoted as $x_i$) — the target it is trying to reach, or the “backpropagated activity” — is stored in the apical dendrite.

A neuron in layer $i$ then takes its own error to be the difference between the two values it is holding:

\[e_i = x_i - u_i\]

To visualise this compartmentalised architecture, consider the structure of a pyramidal neuron adapted for the DLL framework:

Three-compartment pyramidal neuron. The apical dendrite at the top holds the expected value x_i, which descends to the soma via trainable backward weights Theta. The basal dendrite below holds the sensory input u_i, which ascends via forward weights W. The triangular soma between them computes the local error e_i = x_i - u_i.

The two quantities the neuron needs in order to compute its own error arrive at physically different places: the target comes down the apical dendrite, the input comes up the basal dendrite, and the soma sits between them holding the difference.

The total loss $\mathcal{L}$ for an $L$-layer network is then the sum of these local squared errors across all layers. Written as a quantity to minimise, that is:

\[\mathcal{L} = \frac{1}{2}\sum_{i=1}^{L} \sum_{j} (x_{i,j} - u_{i,j})^2\]

(The paper states the same objective with the opposite sign and ascends it; I have flipped it here so the update rules below read as descent. The two are equivalent.)

Asynchronous, local updates

Because the information is spatially separated into different compartments, the forward and backward phases can occur simultaneously. To solve the weight symmetry problem (Criterion 1), DLL introduces a special trainable backward weight matrix, $\Theta$, replacing the traditional forward weight transpose ($W^T$) during the backward parameter update.

During the update phase, the expected value $x$ propagates along the apical dendrite via the path defined by $\Theta$, which is independent of the forward weight transpose $W^T$ that backpropagation would have used. (Not independent of $W$ altogether — as the equation below shows, the forward activations still enter through $f^{\prime}$.) The calculation for the change in the expected value, $\Delta x_i$, relies strictly on localised variables:

\[\Delta x_i = (x_i - u_i) - \Theta_i^T [ (x_{i+1} - u_{i+1}) \odot f'(W_i u_i) ]\]

$\odot$ is the elementwise product; $f^{\prime}$ is the derivative of the activation function. Once $\Delta x_i$ is computed, the expected value is updated iteratively ($x_i \leftarrow x_i - \eta_x \Delta x_i$).

Each neuron uses this locally available information to update the synaptic weights connecting it to the adjacent layers. Nothing about the architecture changes; only the weights do. Over many such updates the output of each neuron is drawn towards its own local target, and the network as a whole redistributes synaptic strength without any point in it, at any instant, holding a global view.

The mechanics of convergence, and where I get stuck

The hardest part of abandoning backpropagation is not computing a local error. It is making sure the local errors add up to anything. A global signal explicitly steers every weight towards reducing the final output error; remove it and you risk each layer diligently optimising its own target while the network as a whole goes nowhere.

DLL’s answer is the structured propagation of the expected value $x$ back along the apical dendrites through the trainable backward weights $\Theta$. Both $W$ and $\Theta$ are updated as training proceeds, and empirically the local targets stay aligned with the global objective well enough for the networks to learn. That alignment is a finding rather than a guarantee — demonstrated across a range of architectures, never shown to be necessary. There is no convergence proof here, and the paper does not claim one.

This is where I get stuck, and I would put it to the authors as a genuine question rather than an objection. If the alignment between local targets and the global objective is itself maintained by learning $\Theta$, then what supplies the pressure that shapes $\Theta$? Locality in this framework is a claim about where a computation happens — each soma uses only quantities physically present in its own compartments. It is not obviously a claim about what information is ultimately required, since the target descending the apical dendrite has to have come from somewhere. To be fair to the authors, the update they give for $\Theta$ uses only quantities from a layer and its immediate neighbour, so it is local in form as well as in siting. My unease is about provenance rather than form: those adjacent-layer quantities are themselves shaped, over training, by what happened at the output. The same tension sits underneath Feedback Alignment and I have never seen it fully resolved. It may be that smearing the global signal out into a slow, learned, structural channel is exactly what brains do, and that this is the answer rather than a dodge. But it is the joint I would push on first.

A perfectly plausible algorithm that cannot learn anything hard is a curiosity.

Empirical performance

Does this purely local, compartmentalised approach actually work?

The researchers tested DLL across MLPs, CNNs and RNNs. On MNIST, MLPs trained with DLL reached 97.57% accuracy against 98.62% for backpropagation on the same architecture. On CIFAR-10 — where two of the biologically plausible methods sit at chance outright — DLL managed 70.89% against backpropagation’s 75.10% on the same CNN.

Read those two gaps carefully, because the framing matters. Backpropagation satisfies none of the three biological criteria; DLL satisfies all three. So this is not DLL losing narrowly to a fair competitor, it is DLL paying about one point on MNIST and four on CIFAR-10 for constraints its rival simply ignores. The paper’s own claim is state of the art among algorithms that meet the criteria, not parity with backpropagation, and that is the comparison that means something.

The time-series results need more hedging than they usually get. On the Electricity dataset, RNNs trained with DLL edge out backpropagation on both MSE and MAE — but the margins sit well inside the error bars across three seeds, and Predictive Coding beats DLL there anyway. Electricity is also DLL’s best case: it loses to backpropagation on Metr-la and Pems-bay, and badly on character-level prediction, where it manages 33.7% against backpropagation’s 51.9%. “Competitive on some tasks, well behind on others” is the honest summary, and the character-prediction gap is the one I would want explained.

The limits of local learning

DLL converges on tasks where other biologically plausible methods struggle, and that comparison — against its actual peers rather than against backpropagation — is the result worth keeping.

The deeper limit is not the accuracy gap, though. It is that the three criteria are necessary conditions, not sufficient ones. Satisfying weight asymmetry, locality and single-phase training establishes that an algorithm is not obviously impossible for a brain to run. It does not establish that a brain runs it. Real pyramidal neurons are not three tidy compartments holding three scalars: they have active dendritic conductances, spike-timing dependence, neuromodulatory gating, and a surrounding inhibitory circuit that this model treats as absent. Any one of those could turn out to be load-bearing rather than detail.

So: a plausibility proof, and a good one. Not yet evidence about cortex.

What I think this changes

For machine learning, not a great deal yet, and I do not mean that as a criticism. Backpropagation is not under threat from a gap of a few points on CIFAR-10. What this buys is optionality. If neuromorphic hardware ever becomes economically interesting, the field will need learning rules that do not require a global backward pass, and it is far better to have them worked out in advance than to start looking when the chips arrive.

For neuroscience it sharpens something rather than starting it. Sacramento and colleagues [3] had already built a three-compartment pyramidal model whose plasticity is driven by local dendritic errors, runs continuously in time, and needs no separate phases — most of what this paper means by biological plausibility. Lv et al. take the compartmental scheme from them and say so plainly. What is new here is reach: one rule pushed through MLPs, CNNs and RNNs, across images, text and time series, and benchmarked head to head against the other biologically plausible algorithms — on the image tasks, at least; the sequence tables only compare against backpropagation and predictive coding. Worth noticing that Sacramento is not among those benchmarks — the paper inherits its architecture from a model it never measures itself against, and that is the comparison I would most want to see.

The consequence is the same either way, and it is the part I find worth the read: if the apical dendrite really does carry a target while the basal dendrite carries the input, the two should play asymmetric roles during learning. So experiments that perturb them separately, rather than treating the dendritic tree as one integrating device, become the relevant test. That is a harder experiment than a benchmark, and a much more informative one.

The paper’s own ablation is the most interesting thing in it, and it sits behind the benchmark tables. Freeze $\Theta$ at random values — Feedback Alignment style, so the compartments are all still there but the backward path no longer learns — and run it on five of the paper’s task-dataset pairs. The two image benchmarks barely notice: two tenths of a percentage point on MNIST, about one on CIFAR-10. The two time-series datasets it covers, Electricity and Pems-bay, see their errors rise by roughly a tenth in relative terms across both metrics — modest-sounding, until you notice it is enough to flip Electricity from DLL narrowly beating backpropagation to DLL-FA falling behind it. And character-level prediction does not degrade so much as disintegrate, from 33.7% to 0.71%, which the authors mark as a failure to converge at all.

There is a gradient in those numbers, though I should be clear it is my reading and not theirs. The authors conclude that updating $\Theta$ is critical “in both classification and sequential modeling tasks”, and they count the one-point CIFAR-10 difference as superior performance. Looking at the same table I see something more graded: the learned backward pathway buys very little on static images, something real but not decisive on time series, and the whole difference between learning and not learning on character prediction. If that gradient is real, the question worth asking is what character-level prediction demands that the regressions do not. I am guessing when I reach for long-range structure as the answer — the paper reports no sequence length for that dataset. It is also, awkwardly, the task where DLL trails backpropagation by its widest margin anywhere in the paper.

Either way, the question has moved somewhere more useful: from whether a local rule can meet all three constraints and still learn something hard, to what specifically about this arrangement is doing the work.

References and resources

[1] Lv, C., Xu, J., Lu, Y., Wang, X., Wang, Z., Xu, Z., Yu, D., Du, X., Zheng, X., & Huang, X. (2025). Dendritic Localized Learning: Toward Biologically Plausible Algorithm. Proceedings of the 42nd International Conference on Machine Learning (ICML), PMLR 267, 41682-41700. Preprint: arXiv:2501.09976.

[2] DeFelipe, J., & Fariñas, I. (1992). The pyramidal neuron of the cerebral cortex: morphological and chemical characteristics of the synaptic inputs. Progress in Neurobiology, 39(6), 563-607.

[3] Sacramento, J., Ponte Costa, R., Bengio, Y., & Senn, W. (2018). Dendritic cortical microcircuits approximate the backpropagation algorithm. Advances in Neural Information Processing Systems 31 (NeurIPS 2018).