Why FP4 pretraining should use a uniform grid, not E2M1
- Source
- Ant Ling
- Date
We recently released a paper showing that UFP4, our uniform-grid FP4 training recipe, stays closer to BF16 than strong E2M1 baselines across Dense 1.5B, MoE 7.9B, and MoE 124B long-run pretraining. The key insight: FP4 training quality is not only about bit width, but also grid geometry.


The problem we investigate is Shrinkage Bias in E2M1-based FP4 training. Because E2M1 uses a non-uniform 4-bit grid, its asymmetric RTNE rounding bins can systematically shrink magnitudes toward zero instead of behaving like zero-mean quantization noise.

This bias matters during pretraining because small negative rounding errors can accumulate across GEMMs and layers. RHT helps reduce outliers, but with E2M1 it can also move tensors into a local-resolution-limited regime where this shrinkage effect becomes more visible.


UFP4 addresses this by using a uniform E1M2/INT4-style grid. With the grid-level bias removed, RHT can be applied more broadly across training GEMMs, improving quantization quality while keeping the recipe practical for FP4 pretraining.

Context
Ant Ling reports that FP4 training quality depends on the geometry of the 4-bit grid as well as its bit width. E2M1, an FP4 layout, spaces its representable values unevenly, and Ant Ling says its asymmetric rounding bins can systematically shrink magnitudes toward zero instead of behaving like zero-mean noise. Small negative rounding errors can then accumulate across matrix multiplications and layers during . RHT, a technique that reduces outliers, can make the shrinkage more visible with E2M1.
Ant Ling's recipe, UFP4, uses a uniform 4-bit grid instead. It says that with the grid-level bias removed, RHT can be applied more broadly, and that UFP4 stayed closer to BF16 (a 16-bit format) than strong E2M1 baselines across a 1.5B dense model and 7.9B and 124B long-run pretraining. It adds that E2M1 should remain useful for range-limited workloads. These are the company's own reported results, and the paper's derivations were not reviewed here.
- pretraining — The first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.
- mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
Checking sign-in…
Loading comments…




