← All IntelClip / AI AgentsRL asks for four scarce things at once
From Taking Reinforcement Learning Cross Datacenter — Nan Jiang, Modal · ≈2:27
Frames RL post-training as a capacity problem rather than an algorithms problem: the default loop demands the single hardest GPU shape to obtain.
What’s in it
- Frames RL post-training as a capacity problem rather than an algorithms problem: the default loop demands the single hardest GPU shape to obtain.
Clip transcript
regions different price and different availability. There's still a lot of capacity out there, but it's not one perfect RDM in the island. This is the mismatch. Available computer is distributed, but the default IO loop ask one tightly coupled cluster. And that cluster is exactly the hot part hard to get. IO wants all four of these at the same time. Enough GPU, same region, fast fabric, and available now. Any of these like is manageable, but all four of them that are pretty hard to get at the same time. ADMIC capacity is not elastic in the way the inference capacity is elastic. You cannot assume you can grow the trainer cluster uh halfway through a run just because roll out wants more nodes for a trainer. So if the whole out loop has no live like has to live inside the the one cluster rollout inherents the hottest part capacity constraint in here. So that leads to the key question does the whole out loop actually need
Comments
Sign in to comment.
Loading comments…