Vibeleaderboard
← All Intel
Intel / post

humans& open-sources a 4-bit hardware-native RL training recipe

Source
humans&
Date
humans&@humansand
Thread · 3 parts

At humans&, we train models from the long-term impacts of their interactions with people. This requires prioritizing long-horizon multi-agent RL. We've developed and are excited to share an open-source, hardware-native 4-bit RL recipe, significantly accelerating training

This recipe resolves instabilities due to policy error (forward) and gradient mismatch (backward) - we share details here: https://t.co/LxUGOabUSd This work was led by @zianglih and big thanks our collaborators at @radixark and @nvidia - without them this would not be possible!

We have a couple surfaces for our models cooking and have started sharing them with a small set of folks - if you're interested in trying them, either individually or as a company, reach out and we'll follow up as soon as we open up more broadly

paint strokes starting continuous and becoming discrete going up and to the right
Terms in this piece · Glossary
  • multi-agent — Using several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.
Why it matters

An open, hardware-native 4-bit RL recipe that resolves known training instabilities could meaningfully cut compute costs for teams doing long-horizon reinforcement learning.

More from humans&
Recommended reads
Comments

Checking sign-in…

Loading comments…