← All IntelIntel / article 
Safe Evolution with Circuit Anchors
- Source
- arxiv.org
- Author
- Yan Liu, Jie Fu, Tsung-Yi Ho
- Date

Why it matters
If you run self-improving model loops, capability gains can quietly erode safety behavior; constraining a small mechanistically identified circuit preserved it more cheaply than reward-based penalties.
Read the source arxiv.org
Recommended reads
articleOne Recipe, Many Harnesses: What Self-Evolution Encodes Across Languages and ModelsSiqi Yang, Qianlan Yang, Yu-Xiong Wang, Saurabh Pujar, Martin Hirzel
articleSelf-Evolving Embodied Agents via Skill-Harness EvolutionPeidong Wang, Zhiming Ma, Ying Chang, Xufang Luo, Xiaocui Yang, Shi Feng, Yuqing Yang, Dongsheng Li
articleWhere Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM AgentsMichael Nguyen, Wei Chen Tan, Nurul Aisyah Hassan, Arvind Raman, Li Hua Lim, Ahmad Faiz Razak
Comments
Checking sign-in…
Loading comments…