Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference

NeurIPSOral2025

Authors
Jiayi Yuan, Hao Li, Xinheng Ding, Wenya Xie, Yu-Jhe Li, Wentian Zhao, Kun Wan, Jing Shi, Xia Hu, Zirui Liu
Affiliation
Rice University
Venue
NeurIPS 2025
Track
Oral

TL;DR

This paper demonstrates that low precision causes non-reproducible LLM inference across different setups, proposing a hybrid-precision method, LayerCast, that computes in FP32 to achieve determinism while saving memory.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

memory llm

← All NeurIPS 2025 Oral papers · Browse the whole archive