Breaking The Bottlenecks: Rus And Stoica Discuss AI Optimization
One of the best talks at our recent Imagination in Action event, The Next Endeavor, was a talk between Daniela Rus, who needs no introduction here, and Ion Stoica, a UC Berkeley Professor and co-founder of companies Databricks and Anyscale.
First, a little background. The Next Endeavor happened September 14–15 at Google Bay View in Mountain View, California, and was put on by IIA, in collaboration with Stanford HAI. (Disclaimer: I am a main organizer of this event.) The conference was well-attended and key speakers shared a lot of insights about where humanity is at with technology today.
As for Rus, even though she’s familiar to most of my regular readers, I’ll note that she runs the MIT CSAIL (Computer Science and Artificial Intelligence) Lab, and works on liquid AI models, which I am also involved in. So Rus is a friend and colleague of mine, and I enjoyed introducing this segment, and then listening to these two experts hold forth on AI.
Rus asked Stoica questions about what remains in the way of our computers handling the agentic age with confidence.
A lot of the discussion had to do with infrastructure, and the best ways to handle change.
“Sometimes it’s not obvious what is going to change, and why things are going to be hard,” Stoica noted, adding that many key innovations including algorithmic evolutions are driven from the system side. “Training and inference are so expensive, and are growing still exponentially at this stage, that efficiency … is so important, and in order to gain the efficiency, you almost need to innovate at every layer of the stack.”
“I think that one thing you need to pay attention to is memory,” Stoica said. “Why memory? Because it’s still true today, that to store one bit of information, you still need one transistor. There’s not much you can do here, right? You can make the transistor smaller, but you still have this one-to-one mapping.”
Another challenge that Stoica enumerated has to do with the modular design of traditional systems.
“The compute will grow faster than the memory capacity, and presumably, even the bandwidth will grow faster, because it’s easier to scale by increasing the bandwidth,” he said. “Basically, memory is becoming more and more of the bottleneck, and this kind of breaking the layers, the layering and modularity of existing systems in the name of ultimate efficiency.”
Later in the discussion, Stoica also mentioned the threat of recursive self-improvement, and the related challenge of alignment.
“One of the goals should be about using AI to build reliable applications, reliable systems, to align with user intent,” he said. “And this is very, very hard. It’s very hard because fundamentally, when you build something, you have a bunch of requirements, you have some kind of implicit or explicit model of the deployment environment, and then you are going to iterate until it’s going to pass a set of tests, or whatever, and then you deploy it. But there are these fundamental gaps between the user intent and the requirements, and between the real world and this kind of model you have, which are very hard to close, especially in an open, changing world.”
Rus asked Stoica about which types of workloads should be run in the cloud, and what should be run on local devices.
Stoica suggested that local AI on edge devices will help with privacy gaps.
“Maybe you are going to close these gaps,” he said.
He also brought us back to the big data age, explaining how Hadoop and traditional queries gave way to methods that are now much faster. Stoica worked on Spark, a processing technology related to SQL, which, as a database protocol, seems to be becoming obsolete in the AI era.
“We often talk about 100,000 robots running around and doing tasks for us,” Rus said, setting the stage for discussion on physical AI, and these robots have to have ‘capability on device,’ but if we do operate such large fleets of machines that have physical intelligence on them, and that have to operate in the physical world, we need infrastructure.”
Stoica talked about the Hugging Face incident, and the threats of reward hacking and hallucinations, and suggested it’s better, in many cases, to keep operations isolated, and again, address the gaps he identified.
“The gap between requirements and human intent,” he added, “or between the model and the real world, the model of the environment and the real world, if you think about it … there are things which are missing, there are some requirements which are misrepresented, or they are missing - you know, the sandbox shouldn’t be connected to the internet, but is somehow connected to the internet … as long as you build this for humans, the human is going to be part of the loop; because you need to satisfy the human intent … the human can become the bottleneck. We should just make sure that, again, we use a human in the most effective way, to provide feedback, and constantly improve things.”
This, he noted, is not a new problem, not really. Stoica concluded his remarks this way:
“To scale up formal methods, make them more and more powerful, and so forth … you are going to try to close the gap and try to capture, in formal spec, more and more of the human intent, and more and more of the real world.”
I thought this call for clarity was a key thing to remember as we forge ahead with agentic AI. We may not be ready now, but by spending time thinking about what we face, we may one day meet the future with greater confidence. Stay tuned.