In a way, it’s like reading about something in a book, and having that mental image of what you imagine it looks like.

We’re due to soon see physical AI, in robot bodies, acting in our world. It’s going to be brand new, and, for some of us, right disturbing. And for many of us, it’s as of yet hard to visualize. We might see a few Unitree models doing backflips in Beijing, but as for real, immersive physical AI, we really, as the song goes, ain’t seen nothin’ yet.

I got a sort of sneak preview by moderating most of a panel including two of my favorite AI experts at the Next Endeavor event Sept. 14-15, that we had in Mountain View, California, which Imagination in Action and partners put on (disclaimer: I help to run these events).

Gordon Wetzstein and Marco Pavone gave us some broad visions of what living with physical AI is going to be like, and how to think about the boundaries and features of the new machines. You can hear my excitement as I gave the intros.

“Physical AI,” Wetzstein said. “Let’s think about what that actually is. People think about robotics. People think about autonomous driving. … But there’s a lot more to physical AI than just robotics and autonomous driving.”

“It’s basically bringing the ChatGPT moment from the digital world into the physical world. And that can mean basically having smart things everywhere. That means your lawn mower is going to mow your lawn. You might be able to do really interesting and useful work at home, by just being (remote and having work) done by some robot or some kind of a physically automated thing. And so I think it’s a step change.”

However, he warned, it’s not all cash and ice cream.

“We need to think about the dangers also, not just the opportunities,” Wetzstein said. “There are amazing opportunities, but also dangers, right? GPT-6 Astra has been shown to be able to use tools really well, tools that weren’t part of its training. It basically modifies the input and observes the output, and it’s using these tools in incredible ways to do things such as 3D computer vision, many other things. So replace a tool with a physical robot or any physical thing. Very soon, if not today, we’ll be seeing the impact of that in the real physical world.”

Here’s how Pavone explained it:

“Physical AI is already here,” he said. “Autonomous driving now is a reality. You have more vehicles actually driving around, and that is a prime example of physical AI. And I think there are a lot of lessons learned from the autonomous driving field that will be leveraged by the broader physical AI ecosystem.”

He disagreed with Wetzstein, though, on the “GPT moment.”

“I don’t think that physically, we have something similar to a GPT moment, defined as a moment whereby a technology, almost from day to night, becomes massively pervasive around the world,” he said, citing obstacles like hardware limitations.

I asked the two whether a robot can understand the physical world well without having a rich internal model, whether from detailed data or real experience.

Wetzstein cited what he called a “bitter lesson” in understanding limits.

“We all want to be smart,” he said. “We all want to understand and interpret what’s happening. We all imagine that it’s better to reconstruct the world around us in an explicitly interpretable state. But time and again, it’s been shown that for cluttered, complex environments such as the real world, learned methods that scale well and that scale with data and model size just outperform these, I would say, engineering or heuristic solutions in many scenarios. And I think that’s true for physical AI systems as well.”

“For some tasks,” Pavone added, “you may not need a particularly sophisticated representation of the world, and so it stands to reason that you don’t need to learn a particularly complicated representation of the world. But for the majority of tasks that actually deliver value, as I’m thinking about, say, manufacturing or oil and gas inspection, and so on and so forth, then yes, you need to have a fairly sophisticated understanding of the surroundings, and also how the world may react to what you do.”

In commenting on the overall context of today’s physical AI, Wetzstein commented that there’s “a lot of experimentation going on.”

He gave an example of a situation where you might not need a robot to really develop its own detailed understanding of a physical space: ex: where there’s already a map.

“Imagine a warehouse, and you actually have a map of the warehouse,” he said. “The warehouse is going to be built to safety standards, with a ramp, have a certain maximum incline, and so on and so forth. It’s a very standardized environment. The operator of the warehouse will have a map that tracks every … little robot that’s driving around. You have the map. If you deliver a robot into that environment, why not use that map, right? There’s no reason to learn that from scratch in real time.”

“I think of simulation as a tool that can be used in different ways as part of the autonomy development program,” Pavone said, “for testing, to understand whether whatever policy you have developed for controlling your robot is good, and (whether) it’s worthwhile to actually test it in the real world; for training, to actually train your decision-making policy in a simulation environment, thereby cutting costs for your development program; and for validation. So basically, to really build confidence that your robot is going to do the right thing, and it’s not going to harm anyone. This is particularly important for safety-critical settings such as autonomous driving or industrial robotics.”

Wetzstein mentioned certain categories of risk.

“There are many variables you can’t observe easily,” he said, mentioning things like object color, slipperiness of a floor, weight of an object, etc., if these are not part of a data model.

“That’s part of the challenge, right?” he said. “Estimating the state, explicitly, of all these unobservable physical parameters is going to be prone to error if you’re trying to do it explicitly. And that’s why I think a lot of these traditional methods, even for perception, 3D computer vision, and so on and so forth, trying to estimate all these parameters, you’re going to make mistakes, and then making decisions based on faulty estimates, that’s challenging.”

Here’s more of what he had to say about complex systems:

“You basically think about the raw sensor data as your representation of the state, and you’re making decisions end-to-end based on those,” he said. “Building better sensors that give you better quality data for whatever task you’re doing is absolutely critical.”

Citing the work of Madeleine L’engle, I asked whether robots need additional sensing capabilities to get to that next level.

“I don’t know if there’s a magic bullet, some north star that we’re shooting for,” Wetzstein said in response. “So, I think a robot doesn’t just have vision sensors. It does have proprioception. So, it does know its own internal state, because typically, you can measure each actuator and what the setting is. You can do computer vision and do egocentric pose estimation. So the robot knows where it is in space, and where all its components are. I would say it can sense its environment mostly through vision, but other senses as well.”

“The key challenge is basically to find the right tradeoffs for the use case,” Pavone said.

Toward the end, I asked each of these experts what they are excited about with AI.

“These days, there are a lot of debates about vision-language-action models, world-action models, and so on and so forth,” said Pavone. “The technology that I’m most enthusiastic about is represented by agentic AI workflows, in part because of the opportunities that are offered in terms of how you architect the autonomy stack, but also, in part, because they represent a key accelerator for autonomy development that is typically extremely expensive.”

So, it’s an autonomy accelerator that’s less costly. And you can architect the autonomy stack.

“I do foresee an explosion of agentic AI applications in the physical world,” Pavone added, “all the way from fine-tuning and adapting workflows to the physical domain to, of course, using, instantiating, and grounding them into autonomy development programs.”

Wetzstein added some of his predictions.

“At the current pace of technology development, I think five years from now, all of this will be old news,” he said. “What I’d like to see ten years from now is an abundance of clean energy and compute hardware that can supply the AI needs of the world, without having to build enormous data centers all over the planet and have a big environmental impact. So that’s a wish list for what I’d love to see 10 years from now: small language models that can go on edge devices that don’t have to go to the cloud to help power these autonomous vehicles, or fundamentally different compute architecture that can do computations in a more compact and more energy-efficient form factor.”

“It will go in waves,” Pavone predicted. “So probably one or two years from now, there will be some consolidation that will happen across the different companies, probably some consolidation that will happen on the technology paths, very similarly to what happened with autonomous driving in the very early stages.”

Wetzstein added this critique of current robotics:

“These humanoid robots with legs have evolved at a pretty impressive pace, but they’re not really doing anything useful right now,” he said. “They look cool, and they’re a novelty item, but what is the utility that they provide today to the consumer, at least? Uh, zero.”

I don’t know for how much longer that will be true. In any case, I had to run upstairs, so I bid by goodbyes and let this panel show themselves out. You can watch the rest in the video. I thought this was a good start on thinking about near-term outcomes for physical AI.