Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety
Source
Charlie Summers, Prajwal Raghunath, Aaditya Pai, Mayur Kulkarni, Zhuo Zhang, Oliver Kennedy, Eugene Wu
Author
Charlie Summers, Prajwal Raghunath, Aaditya Pai, Mayur Kulkarni, Zhuo Zhang, Oliver Kennedy, Eugene Wu
Date
Key takeaways · AI-distilled
The authors argue current AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → defenses, which constrain agents before execution, modify tool inputs and outputs, or rely on LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → judges, can depend on model behavior or block unsafe actions without helping the agent recover.
Environment Steering models the agent and agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition → execution state as database tables, tracks record-level data flows between them, and checks those flows against declarative policies while the agent runs.
When a policy violation is detected, the environment returns policy- and context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition →-specific feedback that steers the agent toward a safe alternative, rather than only blocking the action.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Why it matters
Shows enforcing safety in the execution environment by tracking record-level data flows against declarative policies and feeding the agent corrective feedback, reporting 0% attack success on AgentDyn without hurting task success.