Transcript
Alex Shaw and Ryan Marten present a rollout-centered view of evaluating and improving AI agents. Drawing on their work on Harbor, Terminal-Bench, and OpenThoughts-Agent, they connect sandboxed environments, agent evaluations, and optimization workflows into a practical framework for generating and learning from rollouts. Speakers: Alex Shaw — Member of Technical Staff, Laude Institute Alex is the creator of Harbor, a framework for evaluating and optimizing agents and language models in sandboxed environments. Ryan Marten — Member of Technical Staff, Laude Institute Ryan builds Harbor and works on research-to-production efforts including Terminal-Bench and OpenThoughts-Agent. Harbor: https://www.harborframework.com/ GitHub: https://github.com/harbor-framework/harbor