Runway’s Bet Beyond Video: World Models, Robotics, and the Neural OS — Anastasis Germanidis
Source
youtube.com
Author
Latent Space
Date
Why it matters
Explains how video models are being turned into simulators for robotics and into interfaces rendered directly by a network without HTML or React. It signals where generative video and AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → environments are heading.
Key takeaways · AI-distilled
Runway made an early bet on a cluster of 1,000 A100s to kickstart its video research.
Gen-2 grew out of a weekend experiment combining text, depth and video, and after OpenAI released Sora the team scaled Gen-3 by roughly 10x in a three-month push.
Germanidis argues third-person internet video could be the largest training source for robotic intelligence, and describes World Action Models that turn video models into robot policies.
He proposes a "Lucid Dream Test" for judging when world models are good enough, and says real-time video generation is inevitable.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.