humans& open-sources a 4-bit hardware-native RL training recipe
- Source
- humans&
- Date
At humans&, we train models from the long-term impacts of their interactions with people. This requires prioritizing long-horizon multi-agent RL. We've developed and are excited to share an open-source, hardware-native 4-bit RL recipe, significantly accelerating training
This recipe resolves instabilities due to policy error (forward) and gradient mismatch (backward) - we share details here: https://t.co/LxUGOabUSd This work was led by @zianglih and big thanks our collaborators at @radixark and @nvidia - without them this would not be possible!
We have a couple surfaces for our models cooking and have started sharing them with a small set of folks - if you're interested in trying them, either individually or as a company, reach out and we'll follow up as soon as we open up more broadly

- multi-agent — Using several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.
An open, hardware-native 4-bit RL recipe that resolves known training instabilities could meaningfully cut compute costs for teams doing long-horizon reinforcement learning.
Checking sign-in…
Loading comments…



