Skip to main content
Back to timeline
arXivSource publication:

Site-level trade-off between stochastic and nearest rounding in low-precision transformer inference: SR in the MLP with RN at the head brings perplexity to 1.10x the full-precision reference

Related research and updates

Synopsis

The work extends the PRISM vectorized rounding library to arbitrary virtual precision via a variable-precision stochastic rounding (VPSR) algorithm and, holding the numerical format fixed while varying only the rounding rule at individual operation sites, emulates DistilGPT-2; it supports the site-level trade-off with a probabilistic forward-error bound for linear projections and a second-order decomposition of expected cross-entropy at the softmax, finding that at t=6 significand bits SR in the MLP gives 1.15x the full-precision reference perplexity versus 2.21x for RN, that the ordering reverses at the language-model head, and that a mixed-precision configuration (MLP output at t=6, SR in the MLP, RN at the head) brings perplexity within 1.10x, a 28% reduction over matched-bit RN.

Source-provided article image: Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2
Figure 1 ·

Figure 1: Complete DistilGPT-2 architecture and evaluation path. Token and position

arXiv · Page 7

Interpretation

The best rounding rule depends on the specific site in the network rather than on a single global policy. Prior low-precision inference work typically treats the rounding rule as a global setting; here the numerical format is held fixed and only the rounding rule at individual operation sites is varied, isolating the site effect. On DistilGPT-2 at t=6 significand bits, SR in the MLP raises perplexity to 1.15x the full-precision reference versus 2.21x for RN, and the ordering reverses at the language-model head.

The forward-error envelope for linear projections grows with reduction length n as O(sqrt(n) u) for SR versus O(n u) for RN, a gap that widens rapidly at low precision. It provides a probabilistic forward-error bound that writes the SR-versus-RN difference as an explicit function of reduction length and identifies the long MLP down-projection as where the gap is most pronounced. Derived from the paper's probabilistic forward-error bound, with the order and the most pronounced site reported in the abstract.

The expected cross-entropy change at the softmax decomposes into signed drift, drift curvature, and a Fisher-weighted variance penalty, explaining why the two sites behave oppositely. A second-order decomposition splits the output-layer loss change into interpretable terms, showing that MLP noise is predominantly a uniform logit shift to which softmax is invariant so SR's variance is largely discounted, whereas head noise is non-uniform across the vocabulary and is not. From the paper's second-order decomposition of expected cross-entropy loss change at the output softmax.

Assigning rounding rules by site in a mixed-precision configuration substantially improves perplexity. In a mixed-precision configuration with MLP output at t=6, assigning SR to the MLP and RN to the head brings perplexity within 1.10x of the full-precision reference, a 28% reduction over matched-bit RN. Emulation observations on DistilGPT-2, with the 1.10x and 28% figures reported in the abstract.

Perspective

The result is aimed at numerical designers of low-precision transformer inference: under a fixed numerical format with only the rounding rule at individual operation sites varied, site-level assignment of rounding rules is an actionable lever. The VPSR extension makes experiments at freely chosen virtual precisions possible, so follow-up work can test this trade-off across wider precision ranges, more models, and more operation sites; the mixed-precision configuration (MLP output at t=6, SR in the MLP, RN at the head) provides a directly reproducible starting point.

A careful reader would still watch whether the site-level conclusion holds for larger models, other precisions, and other operation sites; whether the non-uniform variance of SR at the head changes with vocabulary size or task; and that the 1.10x and 28% figures were observed on DistilGPT-2 at t=6, so extrapolation to other settings requires new experiments. This reading is at the abstract level, so the figures and the full experimental matrix in the body are not included, leaving the complete conditions and ablation details behind those numbers as open questions.

Sources