Vibeleaderboard
← All Intel
Intel / video

Training Agents 4: From reward functions to environments.

Source
youtube.com
Author
Hugging Face
Date
Why it matters

Shows how to turn multi-step tasks into trainable RL environments with OpenEnv and TRL, so reward comes from outcomes of actions rather than a single answer.

Key takeaways · AI-distilled
  • In TRL, environment_factory and get_reward take the place of reward_funcs, and the openenv CLI handles init, push, pull and fork for environments on the Hub.
  • The session's first demo trains Qwen3-1.7B on MBPP inside a live Python session, with held-out pass rate rising from 0.49 to 0.59.
  • A second demo trains a real coding (OpenCode) on DeepCoder problems in Hugging Face sandboxes via Harbor, with reward rising from 0.27 to 0.71 in 10 steps.
  • Reward hacking moves into the environment: examples include a try/except that never failed and a CVE fix read from .git history, so the has to be locked down.
Terms in this piece · Glossary
  • agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • sandbox — An isolated environment where AI-generated code or agent actions run without being able to touch anything real.
Read the source www.youtube.com
More from Hugging Face
Recommended reads
Comments

Checking sign-in…

Loading comments…