NVIDIA LPU supports 3 types of disaggregated inferencing: 1. Rubin Prefill + LPU Decode for the fastest interactivity 2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN for the middle of the curve 3. Rubin Prefill + Rubin Decode Verification + LPU Drafter for the middle-left of the curve For low interactivity, raw Rubin still takes the win. Looking forward to seeing Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.

GPTOSS 2T is a proxy model, a made-up model config (layer counts, hidden dims, expert counts, etc.) created by scaling up the GPTOSS 120B architecture so that hardware teams can use it as a stand-in for frontier closed models they can't actually run, like GPT-5 or Gemini. NVIDIA can't benchmark on the real weights. This is standard practice for hardware architecture & bring-up teams.
NVIDIA's LPU supports three prefill and decode splits with different interactivity tradeoffs, and raw Rubin still wins at low interactivity. The headline GPTOSS 2T results come from a made-up scaled architecture, not real frontier weights.
postInside Huawei's 1,024-NPU Atlas 950 SuperPoD and its LPO rack fabric
postAMD's inference gap is a software and tooling problem, not a silicon one
articleMost Neoclouds Suck At Security OpenAI vs HuggingFace, Container Escapes,…Checking sign-in…
Loading comments…