FP4 All the Way: Fully Quantized Training of Large Language Models

NeurIPSSpotlight2025

Authors
Brian Chmiel, Maxim Fishman, Ron Banner, Daniel Soudry
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

We demonstrate, for the first time, fully quantized training (FQT) of large language models (LLMs) using predominantly 4-bit floating-point (FP4) precision for weights, activations, and gradients on d…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model llm

← All NeurIPS 2025 Spotlight papers · Browse the whole archive