Foundation Models Are Eating Robotics From the Bottom Up

Creative Robotics
Foundation Models Are Eating Robotics From the Bottom Up

Something shifted this week in robotics, and it wasn't another humanoid demo.

Google released Gemini Robotics ER 2, a high-level reasoning system for robots. Generalist announced their GEN-1 foundation model now supports everything from five-fingered hands to specialized tools, trained on half a million hours of real interaction data. NEURA Robotics partnered with RWTH Aachen to build dedicated facilities—"NEURA Gyms"—specifically for training physical AI models.

Three separate companies, three different approaches, one underlying message: foundation models aren't the future of robotics anymore. They're the present.

For years, we've heard promises about foundation models transforming robotics the way they transformed natural language processing. The analogy was always compelling but distant—robots operate in physical space with consequences for mistakes, unlike chatbots that can hallucinate without breaking anything. The data requirements seemed impossibly large. The computational demands appeared prohibitive for real-time operation.

But look at what Generalist is claiming: a single base model that works across 9,000 gripper variations. That's not narrow task-specific programming. That's genuine generalization at a scale that would have seemed absurd two years ago. Google's ER 2 handles real-time spatial reasoning and multi-robot coordination—tasks that previously required entirely separate systems.

The NEURA announcement is particularly telling. They're not building one gym. They're building ten globally. That's infrastructure investment at the scale you see when an approach has moved from "promising" to "essential." You don't build a global network of specialized training facilities for a technology you think might work. You build it for technology you know you'll need at volume.

What makes this moment different from previous waves of AI-in-robotics hype is the specificity. These aren't vague promises about robots that will "understand" their environment someday. Google is shipping video understanding for self-correction now. Generalist is demonstrating actual multi-gripper generalization with concrete performance metrics. NEURA is training models that "perceive, decide, and act" in real-world environments—the full perception-to-action loop that's always been the hardest problem in robotics.

The implications ripple outward in unexpected ways. If foundation models become the default intelligence layer for robots, then the competitive dynamics of robotics shift dramatically. Success stops being primarily about mechanical engineering excellence or control system optimization. It becomes about who has the best training data, the most compute, and the most sophisticated model architectures.

That's a very different game, and it's one where traditional robotics companies don't necessarily have advantages over well-funded AI labs. It's also a game where the winner-take-most dynamics of software platforms start applying to physical machines. If one foundation model becomes significantly better than alternatives, why wouldn't everyone use it?

We're watching the software-ification of robotics happen in real time. The question isn't whether foundation models will power the next generation of robots. Based on this week's announcements, that question is already answered. The question is what happens to the robotics industry when intelligence becomes a commodity you can download rather than expertise you have to build.