I previously worked with a senior colleague at HKUST on numerical simulations, which involved learning about several samplers. That made me wonder about the relationship between samplers used in simulation and those used in deep learning.
“Which sampler do large models use?” sounds like a simple question, but the word sampling hides several different random objects. The prior question should always be: what exactly are we sampling?
Distinguish five kinds of sampling
At least five layers appear in a large-model system:
- Parameters and continuous noise: initialization, latent variables, or perturbations may use Gaussian and other continuous distributions.
- Training examples: selecting documents for a batch is usually discrete index sampling.
- Data mixtures: assigning training budget to web, code, mathematics, and different languages.
- The next token: drawing a discrete token from a categorical distribution over the vocabulary.
- Reasoning trajectories: generating several paths or rollouts for one problem, then judging, voting on, or training from them.
When the object is a vocabulary token, there is no need to draw a continuous value from first. Given logits , temperature sampling uses
This is already a discrete probability distribution. As , it approaches greedy selection, while as it becomes flatter. But “more random” does not automatically mean “more meaningfully diverse.” If candidates are structurally similar, raising temperature may add errors rather than new solution strategies.
Training-data sampling decides where the model spends its time
Suppose a corpus contains web text, code, mathematics, and papers. Sampling in proportion to raw size lets the largest source dominate gradients; sampling every source equally can repeatedly overuse a small corpus. A data mixture chooses weights
and samples from
These weights are not merely data-loader settings. Under a fixed token budget, they allocate learning opportunities. Recent work uses small training runs or scaling laws to predict mixtures for larger runs, but the answer still depends on target tasks, model scale, deduplication, and quality filtering.
Curricula, hard-example sampling, and dynamic mixtures all change what is seen when. Their real comparison is not whether the distribution sounds novel, but whether it improves desired capabilities under equal compute without causing overfitting, forgetting, or instability.
Token sampling already has a Gibbs form
Define an energy
Softmax becomes
which has exactly the form of a discrete Gibbs or Boltzmann distribution. Temperature here is not physical temperature; it controls concentration. This equivalence seems useful because it expresses sampling as entropy-regularized optimization. For a probability vector , consider
with . Its optimizer is softmax: the score term favors strong candidates, while entropy prevents immediate collapse.
This gives a partial answer to my initial conjecture about applying the Maxwell–Boltzmann distribution to LLMs. The Boltzmann/Gibbs exponential-energy form is already present in softmax. The Maxwell speed distribution adds a factor arising from the volume of three-dimensional velocity space. An LLM has no natural three-dimensional velocity space, so a physical name alone gives no reason for a better token sampler.
If a model actually contains a -dimensional isotropic Gaussian perturbation, its radius naturally follows a distribution; Maxwell is simply the case. A distribution should follow from the geometry of the sampled object, not be selected first and justified afterwards.
Reasoning-trajectory sampling is the more interesting case
Self-consistency samples several reasoning paths for one problem and chooses a consistent answer. Work such as DeepSeek-R1 also makes rollouts an important object in reasoning-model training. The cost is no longer one token draw but an entire trajectory
Trajectories can be redundant, needlessly long, plainly wrong, or structurally novel. Sampling a fixed number for every prompt is unlikely to be optimal: an easy problem may need one path, while a difficult one may benefit from a larger exploration budget.
As a research model, we can write an entropy-regularized objective
Here is verifiable quality, is token or latency cost, and penalizes duplication with trajectories already seen. The formal optimum is
This is a modeling language, not a universally validated sampler. Before generating a trajectory we do not know accurately; a diversity score may reward superficial variation; and a learned verifier can be exploited systematically.
From fixed temperature to state-dependent control
At present, it seems that a richer question than “Should temperature be 0.7 or 1.0?” is to let temperature and sampling budget depend on the prompt and current uncertainty:
When the model is confident and existing paths agree, generation may stop early. When candidates conflict or verification is unstable, exploration can expand. Such a method should be evaluated at fixed accuracy or utility by measuring whether it reduces total rollout tokens, latency, and training compute.
The most useful gift of statistical mechanics to large-model sampling is therefore not the name of a ready-made distribution. It is a way to put quality and cost into one energy, exploration into entropy, and multiplicity into a density of states.
References
- DeepSeek-AI, DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.
- Xuezhi Wang et al., Self-Consistency Improves Chain of Thought Reasoning in Language Models, ICLR 2023.
- Apple Machine Learning Research, Scaling Laws for Optimal Data Mixtures.
If you enjoyed this, leave a comment~