People frequently ask me how quickly robot capabilities are advancing. The honest answer is I don’t know — and I’m not sure anyone else does either. A big reason is that today’s best robotics benchmarks are far less informative than benchmarks for conventional frontier models.
Last week, as I was figuring out whether GPT-6 Astra was good at robotics, I carefully looked into the model’s performance on the RoboDojo benchmark. RoboDojo is a benchmark made by a consortium of academics from the University of Hong Kong, UC Berkeley, Tsinghua, and others.
RoboDojo seemed to indicate that yes, Astra is good at robotics. The model was (briefly) the top-scoring model on the leaderboard, beating out specialized robotics models from companies like Physical Intelligence. Yet when I looked closer at Astra’s performance, I realized that the story was more complicated than the headline number suggested.
RoboDojo’s simulation benchmark is a collection of 42 tasks spread across five categories: generalization, precision, long-horizon, memory and open tasks. Some tasks require dexterous movements and precision, like inserting tubes into a rack.

Other tasks only test basic manipulation, but they require some “book smarts.” For instance, another task asks the robot to arrange number tiles from left to right to form the largest possible number.
Astra struggled with the precise manipulation tasks but aced the “book smarts” tasks. In comparison, the strongest specialized robot models were better at precision but couldn’t handle semantic variation. Enough of the 42 tasks tested these semantic skills that Astra ended up having the highest average score of all the models evaluated.
If the benchmark had included even a couple more tasks Astra was weak at, some other model might have performed better. Astra’s top ranking depended heavily on how the benchmark was set up.1
Astra’s spiky performance on RoboDojo led me to conclude that it is closer to a “system two” model, but this uneven performance also points to a challenge in understanding robots: robustly measuring (and interpreting) the capabilities of generalist robots is incredibly difficult. Robot benchmarking has challenges that LLM benchmarking does not have, and there is far less information comparing different models with each other.
The real world is complicated
The primary reason robotics benchmarking is hard is the same reason robotics is hard: the real world is complicated, unforgiving and expensive to deal with.
Basically every robot startup that has made a robot has a story about how something completely unexpected happened when its team took that robot into a new setting.
Sometimes, it’s a new type of flooring. When Fauna Robotics (now owned by Amazon) was teaching its robot to walk, it could move around on almost all surfaces fine. But one “extremely high-friction, textured carpet” tripped the robot up, said the company’s CTO on a panel. “That motivated us to go back to our office and buy a bunch of carpets.”
Other times, it’s the difficulty of predicting human behavior. At one point, Fetch Robotics (now owned by Skild AI) was testing its warehouse picking robot in simulation. The simulation assumed that humans would always go to the nearest robot, but that turned out not to be true. This slight mismatch in behavior “had a huge impact on the performance of our algorithms.”
It can even be a subtle hardware bug: Agility Robotics had issues where some of its older robots would occasionally shut down while crouching. It turned out to be because of a chip that flexed out of position over time.
The same real-world complexity applies to benchmarks that try to capture a broader range of robot abilities, rather than a single future deployment task. This is true whether the robot is run on real hardware or simulated in a computer.
For real-world benchmarks, the initial setup can sometimes be straightforward. To test how well a model can stack blocks, all that’s necessary is a couple of blocks, a table and a robot.
But running the evaluation itself is difficult and expensive. A human usually needs to reset the task and the robot after every run. Robots can break down or overheat.
And both the number of robots and the number of humans constrain the number of evaluation runs. While an LLM benchmark with 250 questions can be run quickly in parallel, a robotics benchmark with that many tasks would require either a huge number of robots or a good deal of time to complete the tasks one by one.
As a result, most robotics benchmarks today consist of only a small number of tasks with dozens of trials per task. The independent evaluator Robocurve currently benchmarks models on six robotics tasks, for instance, while RoboDojo’s real-world evaluation suite has 18.
Even so, running a couple hundred trials is a substantial effort. When the startup Pantograph benchmarked Astra and Fable 5.1 using six robots, it took a day and a half to complete 160 trials, according to Camille Fassett, the company’s head of operations.
Simulation benchmarks are much easier to run, but they require a lot of work to set up. One must “get the friction parameters right, masses, make sure contact is working properly, et cetera,” observed roboticist Chris Paxton in January. “This only gets harder as I start to scale simulation up.”2
There’s also always the worry that results in simulation don’t necessarily transfer into the real world, something called the sim-to-real gap. This is especially true for tasks that require a lot of complex manipulation, as that is much more difficult to encode into a computer.
So most robotics benchmarking today in both the real world and in simulation covers fairly simple tasks, and evaluation suites tend to have tens of tasks, rather than the hundreds or thousands that an LLM benchmark might have.
Robo-benchmaxxing
That, in turn, makes it difficult to turn individual evaluations into measurements about general robotics progress.
While LLM benchmarking also has this problem, the major LLM developers all optimize for essentially the same thing: coding and agentic tasks. So model capabilities on benchmarks correlate strongly along just one or two dimensions.
This has let evaluators figure out ways to boil down an AI model’s performance into one holistic judgment, either of software engineering ability (like the METR time-horizon benchmark) or general capability (like the Epoch Capabilities Index).
There’s nothing like that yet for robotics. Self-contained tasks like stacking blocks or folding shirts give some information about how capable a robot model is, but less about the trajectory of the entire field.
Worse, robotics companies often optimize robots to perform well on specific tasks. This is often innocuous: most real-world deployments today involve robots performing a narrow task over and over. But it means that solid performance on a task doesn’t tell us much about how well the robot might perform other tasks, even those closely related.
A crisp example of this dynamic comes from the most recent edition of the World Humanoid Robot Games, hosted in Beijing in August. Robot makers competed in a number of events, ranging from traditional athletic activities like soccer to robot-specific competitions like power tool assembly.
Many of these events featured impressive performances — such as when a robot broke the human world record in the 100-meter race. But those performances didn’t necessarily transfer to other, more useful skills. In the 100-meter race, robots didn’t stop like a human would but instead ran into a barrier. One caught on fire.

That isn’t to say that competitions and benchmarks like the Robot Games aren’t useful. On the contrary, the Robot Games are “probably the most comprehensive testbed for humanoid physical abilities that we’ve seen,” as Center for Technology & Statecraft research fellow Amelia Michael points out. This can be helpful for figuring out how good robots might be if AI models dramatically improve over the next year.
But it’s very difficult to predict the trajectory of general-purpose robots from benchmarks with a limited number of tasks.
We don’t know the best US robot model
Perhaps the biggest challenge to public robotics benchmarking today, though, is that there is very little of it among the top US companies.
One of the virtues of the Robot Games is that the event encouraged participation from some of China’s strongest robotics companies. In the US, there are much less of these broad-based competitions. In fact, there is almost no independent evaluation of the top robotics models at all.
On leaderboards such as RoboDojo, almost all of the American models listed are from organizations like the Allen Institute for AI and Nvidia, which have strong incentives (academic and commercial) to release open models. The most prominent exception is from Physical Intelligence, which openly released its π0.5 model in 2025.
So beyond demo videos and blog posts, we have little public information about how capable the models of other American companies are. Models like Generalist’s GEN-1.5, Dyna’s DYNA-2, and Skild’s S1 have all displayed impressive abilities in videos and blog posts. (Ditto for robots from the big makers like Figure or 1X.) But it’s difficult to turn that information into a systematic evaluation.
I expect this will change as robots start to be deployed in the real world (especially homes): we’ll at least be able to see how robots perform in settings where they can be observed independently. But for now, we’re sort of flying blind.
RoboDojo has since updated its simulation leaderboard to separate the scores of vision-language-action (VLA) models and world models from agent-based systems like GPT-6 Astra.
Though the increasing spatial capabilities of frontier LLMs likely helps with this.



These are all highly valid points and the going won't be easy.
The good news is that general purpose reasoning and path planning should scale quite easily in simulation. It is the precise manipulation and close contact where you have to take into account friction and the physics of the surfaces where the computation can be heavy and inaccurate.
But for the latter likely need to delegate to a physics-aware model anway that only worries about grip and not the big picture.
The other issue is that the hardware is expensive and wears out. So, won't be quick, but some of the work on general AI since 2020 or so (and availability of compute) will help a lot.