Some Thoughts on Monte Carlo Control Algorithms and Random Walks
A route from Basic MC, Exploring Starts, and ε-greedy control to initial distributions, biased random walks, hitting probabilities, and natural starts.
Pyuyi's log
Lines of words light up one by one; only the night ever answers.
Blog category · 5
Ideas, conjectures, and research questions that are still taking shape.
A route from Basic MC, Exploring Starts, and ε-greedy control to initial distributions, biased random walks, hitting probabilities, and natural starts.
Separating attention modules from attention functions, then exploring a constrained route through excitatory–inhibitory competition, spike measures, and causal operators.
Separating data, token, and reasoning-trajectory sampling in large models, and asking what a Gibbs energy view can genuinely contribute.
Projecting stochastic-gradient trajectories into ordinal patterns, and examining what fence posets may reveal—and fail to reveal—about local optimization dynamics.
As models learn to plan and use tools dynamically, workflows do not disappear; they become executable, recoverable, and verifiable control structures inside a harness.
Blog category · 0
Reading notes, methods, and knowledge gathered along the way.
The Learning shelf is still taking shape. Its first essay will appear here when it is ready.