Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference
NeurIPSOral2025
TL;DR
This paper demonstrates that low precision causes non-reproducible LLM inference across different setups, proposing a hybrid-precision method, LayerCast, that computes in FP32 to achieve determinism while saving memory.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
memory llm