PowLU: a SwiGLU replacement that tames FP8 loss spikes
- Source
- Ant Ling
- Date
SwiGLU is everywhere in modern LLMs — but for large inputs it behaves like x². That quadratic blow-up inflates activations, amplifies outliers, and makes deep network or low-precision (FP8/FP4) training prone to loss spikes. We propose PowLU, a drop-in activation built for stable large-scale pre-training. 🧵

The idea: replace the quadratic tail with a rational power function. For x>0, PowLU = x · x^(m/(√x + 1)) · σ(x). As x grows, the exponent → 1, so growth smoothly decays from quadratic toward linear — reining in outliers while keeping nonlinearity. For x≤0 it's identical to SwiGLU. One hyperparameter (m). We use m=3.
Does the smoother tail cost you anything? Almost nothing. Scaling-law fits for SwiGLU vs PowLU nearly overlap (exponents -0.0659 vs -0.0661) from 26M to 368M activated params. Across 17 benchmarks at 7.9B (600B tokens) and 124B (800B tokens), PowLU stays competitive — sometimes ahead.

Where it really shows: FP8 training stability. SwiGLU and SwiGLU-Clip both hit loss spikes around step ~77k. PowLU holds a smooth trajectory near 1.32 throughout, with visibly tighter activation & gradient ranges and far fewer outlier channels. Stability without giving up expressiveness.

Context
SwiGLU, an activation function used in many modern language models, grows roughly like the square of its input for large values. Ant Ling says that quadratic growth inflates activations and amplifies outliers, making deep or low-precision (FP8 and FP4) training prone to loss spikes. Its proposed fix, PowLU, is a drop-in replacement that is identical to SwiGLU for inputs at or below zero and uses a rational power function in place of the quadratic tail for positive inputs, so growth eases from quadratic toward linear.
The company reports that scaling-law fits for SwiGLU and PowLU nearly overlap from 26 million to 368 million activated parameters, so the smoother tail costs little. In its FP8 comparison, SwiGLU and a clipped variant both hit loss spikes around step 77,000, while PowLU held a smooth trajectory near 1.32 with tighter activation and gradient ranges and fewer outlier channels. These are Ant Ling's own results from its thread; the linked paper was not reviewed.
postSix untouched Ling-3.0 base checkpoints released for continued pretrainingAnt Ling- postGenerating LoRA weights per input instead of training one adapterTencent Hy
articlePretraining and adapting a language model on a dependency-free stack: GPT-2 124M from random weights, reproduced against llm.c, and a clinical adapter for Qwen3-0.6BThang Tran (CloudKites AI Lab, New South Wales, Australia), Lan Dang (Monash Business School, Monash University, Victoria, Australia)
Checking sign-in…
Loading comments…



