On September 3, OpenAI released the first GPT-6 model: Astra. As usual, the company described its new release as the “world’s most intelligent and aligned model.” But OpenAI didn’t just list the customary coding and agentic benchmarks. It also claimed that Astra was significantly better than previous LLMs at controlling computers. “Anything you can do on a computer, Astra can do for you. Fast.”
OpenAI released videos of Astra fluently completing varied computer tasks, from formatting legal documents in a word processor to navigating the Texas DMV’s website. Of these demos, one category stood out: 3D modeling. OpenAI showed models that Astra had made of a house, a circuit board, a video game world, and a car transmission.
This jump in spatial reasoning prompted a question: might Astra be good at robotics? The answer was a resounding yes. In a viral Twitter post the following week, OpenAI robotics employee Thijs Simonian set up Astra with a cheap robot arm and a set of paintbrushes and asked it to paint the Golden Gate Bridge from a reference image. After five iterations, the results were astonishingly good.

Simonian wasn’t the only person to notice this skill. Since Astra’s release, the Internet has reverberated with demos of the model completing various robotics tasks: Astra picking and placing robot blocks, Astra slicing cucumbers, Astra solving a Rubik’s Cube in simulation, even Astra driving a car in a parking lot.
Astra is so good at controlling robots, in fact, that it was briefly the top-rated robot-control model on the RoboDojo Benchmark, ahead of specialized robotics models like π0.5 from Physical Intelligence and MolmoAct2 from the Allen Institute for AI.1 Unlike previous general LLMs, Astra seems good at controlling robots directly, specifying the exact positions that the gripper needs to navigate to complete a task, rather than just giving high-level instructions that must be carried out by a lower-level robotics model.
Astra’s capabilities have big caveats at the moment. For instance, it has to pause for seconds at a time to think about its next movement — far too slow to complete most robotics tasks. But Astra’s recent performance has been jarring to roboticists. “Currently it’s not very useful,” said Yale PhD student Haoxiang You. But the capabilities of the new OpenAI model have “kind of shocked people,” he said.
Some researchers think that OpenAI’s and Anthropic’s models might become good enough to completely supplant specialized robotics models. “I feel it’s totally possible to have one large model that does both very good reasoning and a very good action output,” ETH Zurich PhD student Chong Zhang told me.
Others are more skeptical. Astra will not “just take over everything” in robotics, Princeton PhD student Zeyu Shen said in an interview. Astra is too slow and it’s not reactive enough to sudden changes in the environment. Astra is also far too large to run on board a robot, so robots would probably still need a small local model in case the internet connection goes down.
Whether or not Astra-type models eventually control robots directly, they are useful for robotics development today. I expect roboticists to use Astra to generate new simulation environments and training data. All this might significantly accelerate robotics progress.
How many LLMs does it take to control a robot?
Using LLMs to control robots is not a new idea. In fact, it’s older than ChatGPT.
In April 2022, researchers at Google released a system called SayCan that explored how LLMs could be applied to robotics. The researchers realized that LLMs could break complex robotics tasks into a list of small steps. For instance, an LLM could identify that in order to clean a spill a robot would first need to “obtain a rag” and “navigate to the spill.” The LLM’s broad knowledge of the world would help it make better decisions; for instance, it would know that a placemat would not work well as a rag.
However, there was no way at the time for LLMs to directly control the robot. Some other piece of software would need to turn the subtask “pick up a rag” into a series of robot movements that actually pick up a rag.
It turned out that this lower-level challenge of generating movement trajectories, sometimes called the dexterity problem, was the bigger bottleneck to solving many tasks in the real world. While LLMs seemed useful as a “system two” planner, it wasn’t clear whether they could become a good “system one” that can dexterously control a robot.2
The next year, Google researchers found an approach to do just that. The researchers took an existing LLM with vision capabilities and trained it further on 130,000 robot demonstrations to make what they called a vision-language-action (VLA) model. That is, they trained the system to take in a text prompt and an image from the robot’s cameras and output a sequence of numbers: the position and rotation of the gripper at the end of the robot’s arm. (See Tim Lee’s VLA explainer at Understanding AI for more info.)

The first VLA, RT-2, had several versions trained from LLMs of different sizes, ranging from five to 55 billion parameters. The largest version, which was the most capable, had to be run on four AI chips located outside of the robot. RT-2 displayed all sorts of emergent capabilities and generalized fairly well. But the largest version could only output a single command one to three times a second, far too slow many demanding manipulation tasks.
Later VLA releases have been smaller and faster. Physical Intelligence’s first VLA, π0, takes 73 milliseconds to output a one-second chunk of actions — and only has 3.3 billion parameters. The largest VLAs widely in use today only have about 15 billion parameters, far fewer than the estimated trillions of parameters of a model like GPT-6 Astra.
As a result, the small LLMs that form the backbone of VLAs do not have the general robustness and adaptability of frontier LLMs like GPT-6 Astra.3 Because they are trained on copious real-world data, VLAs can create plausible action trajectories. But they can’t necessarily plan over long time horizons or generalize to new settings.
This is fine for many of the current tasks that robots do — moving totes or folding packages does not require sophisticated planning. But to create general-purpose robots, you need a model that can handle chaotic open-world scenarios.
Robotics companies are trying to make their models generalize
So far I’ve been discussing the possibility that a general-purpose model could become good enough at low-level actions to supplant specialized robot models. Another possibility is that today’s special-purpose robotics models will become more generalized over time. That’s the vision of startups like Generalist and Skild, both of which demonstrated models in August that could complete a task from a single human demonstration, a skill called in-context learning.
Both models also displayed some amount of physical intuition. For instance, when Generalist’s GEN-1.5 model was tasked to place a block in a bowl that was covered with a piece of paper, GEN-1.5 removed the piece of paper without being directly asked to.
This progress was exciting; Generalist’s announcement video showed clips of employees freaking out when the model displayed some new generalization capability.

To Generalist’s CEO Pete Florence, the result vindicated the company’s approach of combining LLM training with large amounts of physical data.4 He told me that for the past couple of years, “there’s been this feeling like, ah man, is all of the generalization just coming from the language model-y stuff and we’re just bolting on this very thin layer?” But the results from GEN-1.5 were “washing away” this concern. “It feels like we are indeed now getting real, very hard to ignore levels of generalization” from the physical data, he said.
Meanwhile, Anthropic has been investigating how good its flagship models were at controlling models.
In July, the company found that Claude was pretty good at writing Python scripts to control the robots. For instance, Claude was able to program a quadruped to walk slowly through a maze.
But Claude and its general-purpose rivals were not yet good at directly controlling robots. On Anthropic’s tests with LIBERO, a well-known benchmark of manipulation tasks using a robot arm, the best performing general LLM was Claude Mythos Preview, which succeeded 5.5% of the time. Compare that to MolmoAct, a specialized VLA that scored 86%. General-purpose LLMs were even worse at controlling unstable humanoids or quadrupeds.
Nevertheless, Anthropic found that general models were getting better over time. “Newer models have made real gains in direct manipulation and high-level policy control across the humanoid and quadruped embodiments we tested,” they wrote.
And once an LLM can directly do a task, it comes with a good deal of generalization capabilities. The day after Generalist announced its new model, independent robotics evaluator Robocurve found that Opus 5 could also put a block in a bowl covered by a piece of paper by directly controlling a robot — though it took seven minutes, rather than seven seconds for a specialized model.
Astra is impressive but not deployable
GPT-6 Astra was released about two months after Anthropic’s report and it was substantially more capable at robotics than its predecessors.
Robocurve co-founder Jay Chooi told me that he saw a “huge jump” when Astra came out. On one task, which involved placing a block into a bowl, the success rate jumped from 5% for Fable 5 to 95% for Astra.
It wasn’t just Robocurve. The Twitter demos we saw in the intro showed that Astra could solve all manner of robotics tasks directly, including tasks that almost certainly weren’t in Astra’s training data.5
This has prompted speculation that general LLMs might soon become capable enough to control deployed robots on their own, without help from a smaller VLA-type model.
On September 7, MIT professor Phillip Isola released an essay that argued that LLMs could become “robot-use agents.” In the same way that GPT and Claude have been trained to use tools to control a computer directly, models may soon puppet robots directly. That would dramatically speed the pace and diffusion of robotics progress, Isola argued, since every upgrade of these models would improve the capabilities of many robots simultaneously.
For now, Astra is still far too slow and isn’t great at dealing with complicated hardware, like hands. But inference is getting much faster: Robocurve found that the maximum LLM token generation speed has been increasing by a factor two to seven times each year, so we could see VLA-type speeds from LLMs by the end of the decade.6 And the capabilities of LLMs in this area have been improving rapidly.
Yet even if those predictions hold up, I don’t think LLMs will completely replace system one models like VLAs, at least in the medium term. It seems more likely that LLMs like Astra serve as effective orchestrators — similar in spirit to SayCan — without controlling robots directly. This is for a couple of reasons.
First, developers will almost certainly want a robotics model directly on the robot in case the Internet goes down. Astra today is far too large to run in a GPU that can fit on a robot and OpenAI doesn’t seem likely to let its weights be downloaded directly.
“I have a very strong view based on industry experience that robot brains need to run on the edge. Most manufacturing facilities have very poor internet, and same thing for most warehouses and most of the places where you want to deploy robots,” Leif Jentoft, the Head of AI Solutions at Standard Bots, told me.
One possibility is a hybrid scenario with a big, remote model and a small model running directly on the robot. If the network connection goes down, the onboard model may be able to complete its immediate task and go into a safe position until connectivity is restored.
Second, the abilities of current LLMs still seem to complement VLAs more than replace them. Astra apparently has enough spatial capabilities to brute-force some manipulation tasks, but it seems to lack the intuitive dexterous capabilities of specialist robotics models.
For instance, while Astra was briefly the top robotics model on the RoboDojo leaderboard, its performance was barely correlated with that of specialized robotics models. Astra excelled at physically basic tasks that required semantic knowledge, but it struggled at tasks that required precise coordination or complex movements. VLAs were comparatively better at precision tasks and worse at semantic ones.
So combining Astra with a VLA can get the best of both worlds. Researchers associated with the Chinese startup Galbot found that a hybrid system combining Astra with the VLA π0.5 performed better on RoboDojo tasks than either Astra or π0.5 alone.
This may change in the future. We know that OpenAI is investing heavily in robotics: it is renting a 202,400-square-foot facility for robotics operations in the Bay Area and has “robotic workcells used in ongoing data acquisition operations,” according to a recent job posting. Future versions of Astra might be trained on enough robotics data to be able to complete tasks that VLAs currently do.
But I won’t be surprised if we continue to see a hierarchical division of labor. After all, the brain has its separations: the prefrontal cortex plays a major role in high-level planning, while the cerebellum helps coordinate precise movements. And it would follow a precedent set by Google. The most recent Gemini Robotics release included three models: two VLAs and a higher-level “embodied reasoning model” based on Gemini 3.5 Flash.
Astra is helpful for other reasons
It’s also possible that general LLMs don’t ever really end up controlling robots in production, at least not directly. There are other robotics paradigms, like world models, which don’t rely as heavily on text-based capabilities. Even if that happens, though, Astra might still end up helping train the next generation of robot models by creating valuable training data.
As I covered in September in Understanding AI, robotics companies are desperate for data to train effective robots.
One strategy is to collect millions of demonstrations of a human piloting a robot to complete a task. This is easy to train into a model, but it’s very expensive to collect.
In contrast, training robots in a computer simulation is much cheaper and easier to scale — provided you already have the environments. Developers can train the robot on programmatically generated trajectories in the simulation engine or let the robot train itself by trial and error.
However, it is still very expensive to make new simulated environments to train a robot in. Historically, simulations required substantial manual work where humans had to make assets for the simulation and specify physical properties such as the friction of each object. There have been some approaches to making assets programmatically, but it’s difficult to make a completely new environment from scratch, especially one which mirrors a real-world setting.
Astra’s uncanny ability to create 3D models might be enough to create the millions of simulated worlds necessary to train robots.
And Astra’s strong ability to solve robotics tasks — either directly as we saw above or using Python to code a robotic policy — could help create data for training VLAs. I talked with a group of researchers who released a paper called EmbodiedSWE. They made a benchmark of difficult everyday tasks and found that coding models could write Python code that solved these tasks (VLAs struggled at this). From trajectories coded by Opus 5, they were able to fine-tune π0.5 to unscrew a lightbulb in a real lamp in two out of ten trials.
Waymo started as a deterministic software system. This enabled Waymo to build up the expertise and training data it needed to eventually train a model to drive the car instead. Perhaps Astra will enable future robotics projects to speed-run through this same process in other robotics domains.
This comparison is probably too generous to GPT-6 Astra, for a couple of reasons (some of which I’ll cover later). First, many of the best robotics policies from top startups aren’t available for evaluators to study — I wouldn’t be surprised if e.g., π0.7 would perform better. Second, Astra’s official benchmark result was restricted to simulated environments because real-world testing stopped after it damaged some of RoboDojo’s hardware. Third, Astra’s performance has since been surpassed by specialized robotics models: it is currently seventh in the standings, though the top model is an LLM agent with access to π0.5.
The software controlling robots is often viewed as having a hierarchy of response times. Planning is a high-level task which doesn’t require quick responses. Dexterity is somewhere in the middle — the ideal command rate is somewhere between one to a hundred times a second. Below that is the software that takes joint positions and turns it into actual motor commands. That software needs to be running hundreds to thousands of times a second, and is sometimes referred to as “system zero.”
This is true also of other model architectures which fulfill the same role as a VLA, like world action models. They still have a relatively small number of parameters in order to be efficient to run in real time.
Unlike RT-2 or Physical Intelligence’s models, Generalist doesn’t call its models VLAs, and mostly pre-trains the models from scratch. But there’s almost certainly text and web data alongside the hundreds of thousands of hours the company has collected of UMI data.
Though Twitter does have some selection bias. Shen, the Princeton PhD student, told me that his advisor had tried to use Astra to fold a shirt overnight, but the robot had been very slow, balled the T-shirt up instead of folding it, and moved the robot’s arms to “pathological states.” That experiment was not reported on Twitter, however.
Or even sooner. Robocurve found that the max token generation speeds of models that do as well as Fable on the Artificial Analysis benchmark was doubling every month. If that pace held up, LLMs could catch up to the current VLA pace within three to four months, Chooi said. I’m skeptical that pace can hold, though.




There is a lot to do till robots can do useful stuff. It looks however that no paradigm shift is needed. We'll need high-level thinking distilled from examples, say as done with LLM, VLM, path planning at high level, the precise reflexive real-time adjustments.
Like an octopus. The main brain does the direction, the tentacle figures the rest.
So, robots will be like current AI on a much grander scale, with vision and tactile sensors, new neural nets, and a whole lot of compute.
"Most manufacturing facilities have very poor internet"?? In the US? Am I missing something?