I recently saw a question on X: can mathematics prove that attention is completely unnecessary in an SNN?
It is an attractive question that is easy to overstate. My feeling is that explicit , Softmax, and attention matrices need not be the only way for a spiking network to perform content selection. Yet attention as a computational function—selecting and aggregating relevant history in response to the current query—does not disappear merely because the named module is removed.
A more defensible research question is:
Under explicit assumptions on tasks, timescales, and resources, can attention computations be compiled into recurrent excitatory–inhibitory spiking dynamics? What are the approximation error, stability cost, and capacity cost?
A module may disappear while its function remains
Standard causal attention can be written as
It performs at least three operations: compute query relevance, make historical candidates compete, and read their weighted information. An SNN may avoid an explicit matrix product and an explicit Softmax. If it still selects history according to the current input, it implements some form of attention in the functional sense.
Existing work supports this cautious distinction. Spikformer builds a spiking self-attention mechanism from spike-form queries, keys, and values without Softmax. Spike-driven Transformer further reformulates relevant operations as masks and sparse additions. These results show that the implementation can change substantially; they do not prove that a fixed-size recurrent SNN can losslessly replace attention for every content-addressing task.
Softmax as a competitive equilibrium
For relevance drives , Softmax is the unique solution of
where
This creates an entry point for an excitatory–inhibitory interpretation. Local excitation maintains candidates that match the current query; shared inhibition limits total activity and induces normalized competition. One continuous-time candidate is
Under ideal conditions, its equilibrium is Softmax; as decreases, competition approaches winner-take-all behavior. A meaningful theorem would need more than equality at equilibrium. It should provide a tracking error of the form
The terms represent transient convergence, query variation, finite-spike approximation, and synaptic delay.
Replace nested spike sums with spike measures
SNN descriptions often expand every layer, time step, and past spike, quickly producing unreadable nested sums. A more compact starting point is to represent neuron ‘s spike train as a counting measure
Synaptic current becomes
The convolution compresses “iterate over every historical spike” into a causal operator. A layer may be abstracted as
and depth becomes operator composition:
Feedback can be written as a fixed-point problem
If , the contraction theorem gives uniqueness. Near a linearization, the resolvent
collects repeated trips through the feedback loop into one expression. This language can provide a compositional interface for existence, stability, approximation error, and depth.
Boerlin, Machens, and Denève’s balanced spiking network offers an important precedent: a spiking network can be derived from a prediction-error objective and made to implement a linear dynamical system. Thus, the idea that spiking dynamics can correspond to optimization is not a new conjecture; new work must address query-dependent dynamic competition rather than merely restating that SNNs can optimize.
Context, capacity, and continual-learning limits
Composing temporal kernels may enlarge a receptive field, but a longer receptive field is not lossless memory. In
the effect of past input is controlled by . Eigenvalues far from the imaginary axis give stability but rapid forgetting; near-critical dynamics retain memory longer while increasing noise amplification and stability risk.
There also seems to be a capacity problem. Standard attention preserves many historical objects that can be queried separately, whereas a fixed-dimensional state compresses them. If a task requires exact indexed retrieval over arbitrarily long sequences, a finite-precision, fixed-size SNN cannot store all history for free. It must add neurons, synaptic state, timing precision, external memory, or approximation error.
From this perspective, context and continual learning are not the same thing either. Fast membrane potentials and synaptic currents may hold context; continual learning changes slower weights while protecting old knowledge. Feedback and inhibition may route inputs into different neural populations, but they do not automatically eliminate catastrophic forgetting.
A credible theory needs positive and negative results
Whether attention is obsolete is not the particular point that concerns me most, but the question did make me think about several problems:
- Which E–I spiking networks approximate entropy-regularized, sparse, or hard attention?
- How does error depend on neuron count, firing rate, delay, noise, and query speed?
- How does depth alter the effective memory kernel and temporal receptive field?
- What trade-off between stability and context length is unavoidable?
- Which sparse, low-rank, or finite-window tasks permit an event-complexity advantage?
- Which indexed or associative-recall tasks force state capacity to grow with sequence length?
- How should fast context state be separated from slow continual learning?
What this direction may really unbind is a certain mode of description: placing spikes, feedback, competition, and memory inside a language of composable causal operators, then explaining when attention can be compiled, when it can be approximated, and when fixed state cannot replace it at all.
References
- Zhaokun Zhou et al., Spikformer: When Spiking Neural Network Meets Transformer.
- Man Yao et al., Spike-driven Transformer.
- Martin Boerlin, Christian K. Machens, and Sophie Denève, Predictive Coding of Dynamical Variables in Balanced Spiking Networks, PLOS Computational Biology 9 (2013).
If you enjoyed this, leave a comment~