
Parcae shows that adding recurrence (looping a smaller model) can match a twice its size, offering a compute- and memory-efficient path to quality — useful if you're weighing model size against cost. It provides the first scaling laws for looping, giving concrete guidance on when recurrence beats simply adding parameters or data.
“Parcae creates a new medium to scale quality by increasing recurrence rather than purely scaling data, opening up an efficient frontier for training memory-constrained on-device models.”
Together AI
“Our 770M Parcae matches the quality of a 1.3B parameter transformer trained on the same data, achieving the same performance with roughly half the parameters.”
Together AI
“Parcae achieves up to 6.3% lower validation perplexity than previous large-scale looped recipes.”
Together AI
“We establish the first scaling laws for looping , finding that compute-optimal training requires increasing looping and data in tandem .”
Together AI
“We personally found them to suffer from residual state explosion and loss spikes.”
Together AI
Checking sign-in…
Loading comments…