← All IntelClip / AI AgentsOpen questions: Muon, fully async RL, and other training stages
From Taking Reinforcement Learning Cross Datacenter — Nan Jiang, Modal · ≈18:44
Honest scoping — the absorption argument depends on Adam's bounded step, so providers moving to Muon and stages beyond RL are unvalidated.
What’s in it
- Honest scoping — the absorption argument depends on Adam's bounded step, so providers moving to Muon and stages beyond RL are unvalidated.
Clip transcript
Last section we have some uh ongoing explorations. So we can see a lot of model providers such as moonshot and also deepc4 they have like they're adopting muan in their post training. Um does the spicy still hold for muan because a lot of thing we discussed previously only for Adam. Second question is async RL at the scale right now we can use the compute across the globe then how like how scalable is the fully async RL this is a very open-ended question there and uh third question is like does it generalize pass because like we have pre-training mid training and SFT like do we have can we apply same paradigm there
Comments
Checking sign-in…
Loading comments…