
October 1, 2026 · Reporting and analysis of developments announced September 17–30.
Imagine a hotel testing a robot that folds towels. A successful demonstration answers one question: can the machine do the movement? The morning shift asks several more. Can it cope with a tangled pile, spot damaged linen, recover from a bad grip and finish before the next delivery arrives?
Three recent developments bring that gap into focus: Anthropic’s new study of physical work, Figure’s tests in unfamiliar homes and fresh global robotics figures. Together, they make a case for taking physical AI seriously while asking much more specific questions about where it is ready.
On September 30, Anthropic published a study estimating that robots can perform 74% of US physical work by estimated working time, equivalent to 34% of all working hours. Much of that capability depends on controlled environments. The study measures task exposure, so these figures should not be read as a forecast that one-third of jobs will disappear. Read the research.
The researchers used Claude to assess occupational tasks against documented robot capabilities and estimate time spent on them. Their categories distinguish purpose-built robot environments, structured human workplaces and unstructured settings. Those model-assisted judgments make the result an estimate to scrutinize, rather than a direct observation of every workplace.
The economic gap is striking: under the study’s assumptions, robots are cost-competitive for tasks representing only 0.3% of all working time. Its costing includes deployment and operating expenses. That is a result of this particular estimation framework, not a universal price comparison or a deadline for automation. See the cost methodology.
Figure’s September 17 Helix 2.5 announcement targets a different question: whether learned behaviour transfers to unfamiliar places. The company tested tidying, towel folding and bed making across 30 Bay Area homes, without collecting training data in those homes. The tasks themselves had been taught using data from elsewhere. Read Figure’s evaluation.
In its comparison, Figure reports that pretraining on its Index human-behaviour dataset increased complete-task success from 9% to 56%, with the other experimental conditions held fixed. The company awarded no partial credit and counted safety interventions as failures.
That is a meaningful result within the reported experiment: it tests whether experience carries across settings. It is also a company-run evaluation of three behaviours, and a 56% completion rate still leaves substantial room for improvement. It does not establish reliable performance across household work generally.
The next evidence to look for is repeated performance over longer periods, in more varied homes, with the time spent on resets and interventions made visible. A compelling video cannot supply those answers on its own.
On September 30, the International Federation of Robotics reported almost 250,000 professional service robot shipments in 2025, up 24%. Transportation and logistics accounted for 117,500 units, or 47% of the total. These are newly published figures about last year’s sales, not a count of robots installed this week. Read the service robot figures.
The same release describes current humanoid applications as mostly specialized and often dependent on human teleoperation. It reports approximately 7,000 full-size humanoids sold in 2025 for commercial and professional uses beyond research and entertainment. That distinction matters when a remotely controlled demonstration is presented alongside an autonomous one.
Separately, the IFR’s September 24 industrial report put the worldwide operational stock at about five million robots in 2025, with more than 600,000 new installations during the year. Those industrial totals cover a broad robotics category; they cannot be treated as a measure of adoption of the newest AI models or humanoids. Read the industrial robot report.
Our interpretation is that the strongest commercial signal is specificity. A buyer can evaluate a robot assigned to a defined material flow against an existing operation. A promise to handle whatever happens in a household is much harder to turn into an acceptance test.
Return to the hypothetical hotel. Suppose a robot folds linen neatly but someone must separate every towel, load it in a precise position and rescue it several times an hour. The folding result may be impressive, while the total labour saved remains small. Change the layout or give the machine a better recovery routine, and the same basic capability could become much more useful.
A useful trial would therefore track finished work per shift, human assistance, downtime, damaged items and total operating cost. It would also record which tasks move to staff when the robot takes over folding. The people who handle exceptions need a say in defining success.
September’s announcements give us better ways to ask those questions. The next milestone worth watching is a machine that can deliver useful work repeatedly, under ordinary conditions, with its support requirements clearly reported. That is the evidence that lets a workplace decide what to hand over.
Cover: AI-generated conceptual illustration.