In June 2026, researchers and engineers released thirteen new foundation models for embodied AI and humanoid robots. Thirteen in a single calendar month. That works out to roughly one new “robot brain” every forty-eight hours.
If you blinked, you missed several of them.
This is not a story about robots. It is a story about software, and about how the real battle in humanoid robotics quietly shifted while most people were still arguing over whether the hardware worked. The hardware works. That argument is mostly settled. The question now is which software brain will win, and what happens to everyone who bets on the wrong one.
The Hardware War Is Over (The Software War Just Started)
Three years ago, the bottleneck in humanoid robotics was mechanical. Building a bipedal machine that could walk without falling was genuinely hard. The actuators were heavy, the sensors were expensive, and the control systems required teams of PhDs writing custom code for every task.
That era ended faster than most people expected. Figure, Unitree, 1X, Boston Dynamics, and a dozen others solved the locomotion problem well enough to deploy real units in real warehouses. Tesla’s Optimus line went from demo-reel novelty to a production-floor pilot program in select facilities. The bodies got good enough.
But a robot that can walk is not a robot that can work. Walk-and-carry is a solved problem. Walk-and-figure-out-what-to-do-next is not. The gap between the two is where thirteen new foundation models showed up in June.
The analogy that helps here: think about the early smartphone era. By 2009, several manufacturers had touchscreen hardware that worked fine. The question was which operating system would run on it. Android and iOS won. Palm and Blackberry, despite solid hardware, did not. The hardware became roughly interchangeable. The software became everything.
Humanoid robotics is in that moment right now.

What a Foundation Model for Robots Actually Does
When most people hear “foundation model,” they think of language models: systems trained on text that can answer questions, write code, or summarize documents. Embodied AI foundation models are the physical-world equivalent. Instead of predicting the next word in a sentence, they predict the next action in a manipulation sequence.
These models learn from video demonstrations, sensor data, and physical feedback. They generalize. A model trained on thousands of hours of warehouse picking tasks should, in theory, handle a novel object it has never seen before, because it has internalized the physics and geometry of grasping rather than memorizing specific grasp patterns for specific objects.
Physical Intelligence, the San Francisco startup known as Pi, has been the most visible player in this space. Their Pi-0 model, released in late 2024, demonstrated cross-task generalization that impressed people who are hard to impress. Since then, the field has moved fast enough that Pi-0 is no longer the most recent reference point. It is barely even a useful baseline anymore.
Unitree, the Chinese robotics company whose G1 and H1 units have appeared in seemingly every viral robot video from the past two years, released their own foundation model work this year. Boston Dynamics, now backed by Hyundai, has been building toward a software layer for Atlas. Figure closed a $1 billion-plus Series C in September 2025 at a $39 billion post-money valuation, bringing total funding to roughly $1.9 billion, and is building its own model stack in-house rather than licensing from third parties. Everyone is doing their own thing, and no one yet has a commanding lead.
That fragmentation is the story.
Why Thirteen Models in One Month Is Actually a Warning Sign
In the large language model space, the competitive picture is reasonably legible. You have GPT-4o and its successors from OpenAI, Claude from Anthropic, Gemini from Google, and a tier of strong open-weight models from Meta and others. There is a clear hierarchy. Buyers know what they are getting.
Embodied AI has no such hierarchy. Thirteen models in thirty days, a figure drawn from trade coverage tracking releases across the sector, is not a sign of a mature, productive field. It is a sign of a field in the middle of a Cambrian explosion, one where the rules of natural selection have not yet sorted out the survivors.
This matters for anyone making capital allocation decisions in or around robotics. Embodied-AI-specific venture funding hit $5.7 billion in the first half of 2026 alone, according to Crunchbase. That figure does not include the broader robotics category, which logged $18.8 billion across all-robotics deals year-to-date. Much of that capital is flowing into companies building on top of foundation models, or into integrators deploying robots in enterprise settings. Those downstream bets are only as good as the underlying model layer. If the model layer is unsettled, which it is, then everything downstream carries model-selection risk.
The parallel that keeps coming to mind is the database wars of the 1990s. Before Postgres and MySQL consolidated the open-source side of the market, there were dozens of competing systems. Companies built products on top of those systems. When the consolidation happened, the products built on the winners scaled. The ones built on the losers got stranded.
Robot integrators who pick the wrong foundation model will find themselves maintaining a private software stack that gets no upstream improvement, no community fixes, no network effects from a growing ecosystem. That is an expensive place to be.
The Shape of the Race

So what does winning look like in this race?
The companies best positioned to win the model layer are the ones with the most physical data. Training an embodied AI model requires robot-hours: real machines, doing real tasks, generating real feedback data. Language models could be trained on text scraped from the internet. There is no internet equivalent for physical manipulation data. You have to generate it yourself, which means you need a fleet.
This creates a compounding advantage for companies that already have robots deployed in volume. Every Figure or Unitree unit operating in a warehouse is generating training data. That data feeds back into the next model version. The next model version makes the robots more capable. More capable robots get deployed in more places. More places generate more data.
Tesla understands this dynamic better than most. Their advantage, if they execute, is not that Optimus is the best robot on the market today. It is that Tesla has the manufacturing scale and the data collection infrastructure to outrun everyone else over a long enough horizon. The question is whether the model layer moves fast enough for raw data volume to be the deciding factor, or whether architectural innovation closes that gap.
Physical Intelligence is betting on the architecture side. Their approach has been to build the most generalizable model possible, one that does not need to be retrained from scratch for each new robot hardware platform. If they succeed, a single Pi model could run on Figure hardware, Unitree hardware, and new platforms that do not exist yet. That would make Pi the Android of humanoid robots: not a hardware company at all, but an OS layer that everything else runs on.
That is the highest-value position in the stack, and it is wide open.
The Question Nobody Can Answer Yet
June 2026 will probably be remembered as an inflection point in this race, assuming anyone pauses long enough to remember it. Thirteen models in one month will look, in retrospect, like the moment the field stopped being about whether robots could work and started being about whose software made them work best.
What nobody can answer yet is what “winning” actually looks like in practice. Does one foundation model become the dominant runtime, the way iOS and Android split the smartphone market between them? Does the space fragment by vertical, with one model dominating logistics, a different one dominating manufacturing, a third dominating elder care? Or does the winning move turn out to be open-weight: a Meta-style release that floods the market, prevents any single company from charging for the model layer, and pushes value downstream into integration and services?
Each of those outcomes implies a radically different investment thesis. Each of them is plausible right now.
Thirteen models in thirty days is not a sign of stability. It is a sign that the race is genuinely on, that the field is wide open, and that the next eighteen months will determine which robot brains get to think at scale and which ones quietly disappear.
The hardware worked. Now comes the hard part.
Chris Meredith writes about AI, technology, and what it actually means for real people.