Humanoid Robots Are Leaving Safety Cages. Proving Actual Safety Is Harder
Most humanoid robots today with real paying jobs – or pilot projects – work behind safety fences. That’s going change over the next 18 months, but it’s not clear that the systems for proving a robot is safe around people are keeping up.
Two weeks ago Agility Robotics launched Digit 5 , a 5’11”, 284-pound humanoid built to work right beside people with no safety cage. The company says it has more than $300 million in multi-year orders. Boston Dynamics, which just released an impressive new version of its Atlas humanoid earlier this year, says its production Atlas supports “fenceless guarding.” And in homes, where there are no fences at all, 1X says it start to ship 20,000 pre-ordered Neo robots this year. Note: Neo is much smaller and lighter than Digit 5 or Atlas, so it – and even smaller, wheeled robots like this $3,555 home robot from a San Francisco startup – is less likely to be dangerous.
But the rulebook hasn’t really been written yet.
ISO 25785-1, the first international safety standard aimed at dynamically stable robots like humanoids, is still a committee draft . Agility is helping write it, but as I noted in my Digit 5 story, robots may well reach customers before the standard governing their behavior is finished. The Association for Advancing Automation’s Jeff Burnstein told me earlier this year that until there is a humanoid safety standard, he doesn’t expect fast growth, and that one early accident could freeze the whole market.
That matters more every month, because the volumes are getting real. Humanoid robot shipments hit roughly 22,000 units in the first half of 2026, according to Counterpoint, with more than 50,000 projected for the full year.
( IDC says 25,000 shipped , and growing fast.)
One industry is a little more mature in its safety regulation history: autonomous vehicles. Erikk Cass, a VP of autonomous mobility at iMerit, says there are multiple challenges with robot safety currently. He works with eight of the world’s leading AV companies on data and safety, and is increasingly doing the same kind of work for robotics.
I asked him in an email interview what the hardest safety problems are for humanoids.
“The hardest cases are the ones where several uncertainties stack up at once: a person moves in a way the robot didn’t predict, an object is partly hidden, the lighting or layout changes, and the robot is looking at something it never saw in training,” Cass told me. “Any one of those is manageable. Together they push the model into territory where its confidence should drop, and the critical question is whether it recognizes that and fails safely.”
Given that context, a dancing robot that kicks an onlooker or slaps a kid is probably either unaware of dangers or overconfident in its abilities.
And the fix, Cass says, isn’t just a smarter robot. It’s a robot that can learn.
“A common misconception is that the answer is a smarter model that can account for everything,” he says. “In practice, safety comes from how quickly a system learns from the situations it got wrong, which means every near miss has to become labeled training and validation data, not just an incident report.”
That’s the same lesson self-driving car companies have learned and are likely still learning. The probably is that humanoids have a much bigger problem to solve.
“Humanoids need the same feedback loop the AV industry spent fifteen years and billions of dollars building: simulation, real sensor data, and a disciplined process for turning real-world failures into training and validation data,” Cass says. “An autonomous vehicle mostly needs to understand where things are and where they’re going. A humanoid also has to understand manipulation, human behavior, and the consequences of getting a task wrong, in thousands of environments rather than a road network.”
His estimate of where the industry is on that curve: “roughly where AVs were a decade ago in terms of how much of that data exists. There is no shortcut through it.”
That’s sobering, given how many companies are promising robots in homes in the next year or two.
Factories and warehouses are relatively easy. They’re generally flat, structured, organized spaces with on-duty personal who are working and reasonably alert, which is why Digit, Atlas, Figure and Apptronik robots are starting there.
Homes are something else entirely.
“A home has none of that,” Cass says. “It has children, pets, clutter, stairs, and people who behave unpredictably, and the layout changes every day. The robot has to recognize far more uncertainty and, just as importantly, know when not to act. That ‘do nothing safely’ behavior is much harder to design and validate than it sounds, and it is where I’d expect the first serious problems to show up.”
Collision avoidance is what most people picture when they think about robot safety: the robot sees you, the robot stops. Digit 5 will have that in 2027, according to Agility Robotics. Cass says that’s necessary but nowhere near enough.
“Collision avoidance is necessary, but it is one layer, and it’s the last one,” he says. “If a robot is relying on collision avoidance, something upstream has already failed.”
What’s needed instead is a robot that can anticipate what a person is likely to do next and choose a safe fallback when it’s unsure, “rather than reacting in the final half second.”
There’s also a problem that didn’t exist in the era of traditional industrial robots, which ran the same deterministic code for years: AI models change. A robot that passed a safety evaluation in June may be running a different brain in September. But, the question is, will it get re-tested or just keep on working?
“You stop thinking of testing as a gate and start thinking of it as a loop,” Cass says. “Every meaningful update gets run against a regression set of known failures, edge cases, and realistic simulations before it ships, and every new failure in deployment goes back into that set.”
Humans stay in that loop, he adds, because metrics can flag that something changed but not why.
“Automated metrics can tell you a model’s performance changed,” Cass says. “They can’t tell you why the robot hesitated in a doorway or misjudged a hand reaching toward it.”
So should there be crash tests for humanoids, the way there are for cars? Yes, sort of. Cass supports common baseline benchmarks, particularly for failure behavior and human interaction, but says the analogy only goes so far.
“Crash tests work because every car faces roughly the same physics,” he says. “Physical AI doesn’t have that. A warehouse robot, a home robot, and a surgical robot operate in different environments with very different consequences when something goes wrong, so any standard will need a common core plus domain-specific tests.”
The big question: are we moving too fast?
“Yes, the technology is moving faster than the evaluation infrastructure,” Cass says. “A polished demo tells you what a robot does under expected conditions. It tells you nothing about the rare, ambiguous situations, and those are the ones that determine whether the robot is actually safe around people.”
Which, frankly, describes most of what we see on social media every week:
Cass’s suggested question for any humanoid company is “not ‘can you show me it working’ but ‘can you show me how it fails, and what you did with that data afterward.’”
The long tail is where public trust will be won or lost, he says, echoing Burnstein’s warning.
“A model can handle the vast majority of cases very well, and public trust will still be decided by what happens in the unusual ones, because those are the incidents that get filmed,” Cass told me. “One well-publicized failure can set a program back years regardless of how good the aggregate numbers are, which is why real-world stress testing has to happen before deployment, not after.”
So what would convince him a humanoid is safe?
Consistent performance across normal and adverse conditions, not a highlight reel. Evidence of what happens when sensors degrade, when people act unpredictably, or when the robot hits something outside its training data. And then a documented and verifiable record that video from those failures went back into training.
“That closed loop is the evidence,” Cass says. “A company that can show it is one I’d trust around people. A company that can’t is asking you to take the demo on faith.”
The cages are coming down, and the robots are entering our space.
Now the industry has to prove that’s actually safe … including around children, infants, and pets.