Mini-batch gradients are noisy: one batch may push parameters in one direction and the next may push them back, so a local training trace can resemble
This resembles the zigzag posets I studied previously, but resemblance alone may not be enough to establish a mathematical connection. The first necessary correction is that neural-network parameters live in a high-dimensional space, whereas an order relation requires scalar quantities.
Choose a comparable observable first
Let . There is no canonical order between high-dimensional parameter vectors, so expressions such as have no immediate meaning. We must first select an observable, for example
where may be a fixed random direction, a Hessian eigendirection, or a function-space direction defined by validation examples. Only then does
become exactly the order event associated with an alternating oriented path—a zigzag or fence poset.
After a quick search, I found that methods retaining relative order while discarding amplitude are known as ordinal-pattern methods in time-series analysis. Bandt and Pompe systematically used neighboring rank patterns to define permutation entropy. Analyzing SGD through rises and falls is therefore not a new concept by itself. A potentially new contribution would have to connect optimization-specific parameters to the complete distribution of fence patterns.
From gradient signs to local extrema
Ordinary SGD obeys
For a fixed direction , define . Then
Consequently, is equivalent to two consecutive projected gradients switching from negative to positive. Local extrema in the projected trajectory record directional sign changes rather than serving as decorative zigzags.
A length-three turning rate tells us only how frequently direction reverses. Longer fence words can distinguish runs, short oscillations, and repeated record-setting patterns. Two trajectories may have the same number of sign changes while exhibiting very different decay and record structures.
My earlier work on fence order polynomials provides one combinatorial background. For the fence defined by an orientation , a greedy-record statistic satisfies
That public result concerns the combinatorial objects themselves. Applying it to training trajectories requires an additional probabilistic model, statistical estimators, and an optimization interpretation that can be tested. Symbolic similarity alone does not complete the transfer.
A local model that can be calculated exactly
Along a fixed Hessian eigendirection, under an idealized locally quadratic and approximately stationary model, SGD takes the form
With independent Gaussian innovations and , this is a stationary AR(1) process. Consecutive increments have correlation , and the Gaussian arcsine relation yields
Formally, this can be inverted as
This small calculation shows that estimating local dynamics from trajectory shape is not merely hand-waving. It does not, however, amount to estimating a real network’s Hessian. Whether remains fixed, the noise is Gaussian, curvature changes slowly, and a window is approximately stationary all matter.
For a second-order recurrence with momentum, a three-point window often reveals only a combination of parameters and cannot separately identify the learning-rate–curvature product and momentum. Four points provide additional sign correlations and may be the smallest identifiable window in the exact model. This remains a developing theoretical direction. Before robustness outside the model and empirical checks are complete, it should be treated as a clue from a tractable model rather than a universal theorem about modern networks.
What it might provide
Ordinal statistics are attractive because they are invariant under monotone rescaling and do not require absolute scales to match across layers. They may help to:
- separate sustained drift from frequent directional reversal;
- detect changes in training regimes without relying on amplitude;
- estimate local correlation structure with very little online state;
- compare trajectory patterns across learning rates, batch sizes, and momentum values.
But ordinal data deliberately discard amplitude. Multiplying an entire trajectory by a positive constant changes no ordinal pattern, so noise scale cannot be recovered from ordinal observations alone. Finite windows also face ties, quantization, dependence between overlapping samples, and bias introduced by choosing directions after observing the trajectory.
Edge-of-Stability experiments have shown that full-batch gradient descent on neural networks can exhibit nonmonotone loss over short timescales while still decreasing over long ones. This establishes that oscillation is not identical to immediate failure. It does not prove that fence patterns in mini-batch SGD arise from the same mechanism. A credible study must compare curvature-driven, noise-driven, and momentum-driven patterns under controlled conditions.
For me, the most interesting goal is to establish a chain with explicit failure conditions:
If this chain survives only in a quadratic Gaussian model, it is still a clean mathematical exercise. If it admits stable error bounds under slowly changing curvature, non-Gaussian noise, and mixed projections, it may become a useful diagnostic for training dynamics.
References
- Christoph Bandt and Bernd Pompe, Permutation Entropy: A Natural Complexity Measure for Time Series, Physical Review Letters 88 (2002).
- Jeremy M. Cohen et al., Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability, ICLR 2021.
- Pyuyi Chufeng Huang, Greedy Records and Bernstein Transfers for Fence and Circular-Fence Order Polynomials.
If you enjoyed this, leave a comment~