Vibeleaderboard
← All Intel
Intel / post

PowLU: a SwiGLU replacement that tames FP8 loss spikes

Source
Ant Ling
Date
Ant Ling@AntLingAGI
Thread · 5 parts

SwiGLU is everywhere in modern LLMs — but for large inputs it behaves like x². That quadratic blow-up inflates activations, amplifies outliers, and makes deep network or low-precision (FP8/FP4) training prone to loss spikes. We propose PowLU, a drop-in activation built for stable large-scale pre-training. 🧵

The idea: replace the quadratic tail with a rational power function. For x>0, PowLU = x · x^(m/(√x + 1)) · σ(x). As x grows, the exponent → 1, so growth smoothly decays from quadratic toward linear — reining in outliers while keeping nonlinearity. For x≤0 it's identical to SwiGLU. One hyperparameter (m). We use m=3.

Does the smoother tail cost you anything? Almost nothing. Scaling-law fits for SwiGLU vs PowLU nearly overlap (exponents -0.0659 vs -0.0661) from 26M to 368M activated params. Across 17 benchmarks at 7.9B (600B tokens) and 124B (800B tokens), PowLU stays competitive — sometimes ahead.

Where it really shows: FP8 training stability. SwiGLU and SwiGLU-Clip both hit loss spikes around step ~77k. PowLU holds a smooth trajectory near 1.32 throughout, with visibly tighter activation & gradient ranges and far fewer outlier channels. Stability without giving up expressiveness.

Read the full thread on X

Context

SwiGLU, an activation function used in many modern language models, grows roughly like the square of its input for large values. Ant Ling says that quadratic growth inflates activations and amplifies outliers, making deep or low-precision (FP8 and FP4) training prone to loss spikes. Its proposed fix, PowLU, is a drop-in replacement that is identical to SwiGLU for inputs at or below zero and uses a rational power function in place of the quadratic tail for positive inputs, so growth eases from quadratic toward linear.

The company reports that scaling-law fits for SwiGLU and PowLU nearly overlap from 26 million to 368 million activated parameters, so the smoother tail costs little. In its FP8 comparison, SwiGLU and a clipped variant both hit loss spikes around step 77,000, while PowLU held a smooth trajectory near 1.32 with tighter activation and gradient ranges and fewer outlier channels. These are Ant Ling's own results from its thread; the linked paper was not reviewed.

More from Ant Ling
Recommended reads
Comments

Checking sign-in…

Loading comments…