Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction
- Source
- Zhiqi Wang, Yichi Zhang, Dongwon Lee, Yuchen Yang
- Author
- Zhiqi Wang, Yichi Zhang, Dongwon Lee, Yuchen Yang
- Date

- context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
- AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Any long-running that compacts can quietly discard the safety instruction a user gave it earlier in the session, and this quantifies how often that happens rather than leaving it as folklore.
“We identify a class of user-issued instructions, Session Constraints (SCs), such as "do not delete any emails until I confirm," that are meant to constrain LLM's behavior for the remainder of a session but are silently dropped during compaction.”
Zhiqi Wang, Yichi Zhang, Dongwon Lee, Yuchen Yang
“To quantify this loss, we introduce COMPINT, an evaluation suite that evaluates compactors across three long-context scenarios: multi-turn chat, agentic trajectory, and long-horizon research.”
Zhiqi Wang, Yichi Zhang, Dongwon Lee, Yuchen Yang
“Current compactors retain only 17% of injected SCs on average, and most perform worse than running the same task without compaction.”
Zhiqi Wang, Yichi Zhang, Dongwon Lee, Yuchen Yang
“Retention varies sharply with compactor, prompt, context length, SC phrasing, and injection location, showing that the loss is systematic rather than tied to any single setting.”
Zhiqi Wang, Yichi Zhang, Dongwon Lee, Yuchen Yang
“We propose an SC-aware extractor that runs alongside the compactor as a plug-and-play module, achieving over 90% retention across all three scenarios without modifying the compactor or LLM.”
Zhiqi Wang, Yichi Zhang, Dongwon Lee, Yuchen Yang
articleCallability Is Not Operability: Controlled Interface Interventions for LLM AgentsZihao Wang
articleCompliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System MessagesJuan Yeo, Geewook Kim
articleInter-dimension Dependence for Multi-Dimensional Evaluation of Open-Ended TextHaoyuan Li, Snigdha Chaturvedi
Checking sign-in…
Loading comments…