What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis
Source
Chengwen Qi, Deheng Ye, Yatao Bian
Author
Chengwen Qi, Deheng Ye, Yatao Bian
Date
Key takeaways · AI-distilled
The best of seven tested Transformers scored 79.6% on a standard held-out test but dropped to 55.3% on TranSGrid's unified reasoning testbed, and just 15.8% on its hardest subset.
The gap held even for problems within the training length range, showing that 'productivity' (composing longer sequences) alone isn't a sufficient test of systematic generalization.
Reintroducing either of two common simplifications -- near-linear action composition, or making goals explicit about which action to take -- on its own restored solve rates to roughly the ordinary test-set level.
That means either simplification alone quietly turns a systematic-generalization benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → back into an ordinary held-out test -- evaluating all three reasoning types together is what makes TranSGrid harder.
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Shows current benchmarks overstate systematic generalization: reasoning-combined tasks crater model performance even within the training length range, a concrete caution when trusting generalization claims.