Cerebras Co-Founder And Others Talk About AI Hardware
At this year’s The Next Endeavour event at Google Bay View in San Jose, I sat down with Andrew Feldman, co-founder and CEO of Cerebras Systems, along with Atiq Raza, a prominent semiconductor pioneer, and entrepreneur and investor Dave Blundin, also, as I noted, my business partner by day at Link Ventures, to talk about how hardware is advancing in the AI world. I wanted to share some insights that came out of this.
We are at a crossroads with AI hardware. Some consider the proliferation of Nvidia GPUs to be a “bubble.” Some think the AI will kill us all, and the only way to stop it is to constrain it at a hardware level. So this stuff is controversial right now. But if you want extended commentary on the nuts and bolts of innovating in this field, from people who have been doing it a long time, read on.
The Cerebras Dinner Plate
The Cerebras “dinner plate,” or WSE-3 , so named for its 9” surface, has 4 trillion transistors across 900,000 processing cores. It’s a big piece of hardware. And Feldman, for our conversation, actually had one with him.
“We embarked 10 years ago on a big journey,” he explained, suggesting that the people at Cerebras were driven by inspiration, not just money. “We do this because we really don’t know how to not build stuff.”
Now, Feldman said, the company has a $25 billion backlog, with a plan to increase production by 8x this year, and 8x next year.
“It is just an extraordinary time,” he said.
Feldman also spoke to strategy, and longevity in a quickly changing business world.
“The right way to get the timing right in the startup world is to be wrong for a long time, and not be dead,” he said, chronicling some of the Cerebras story. “We had achieved what was one of the holy grails, and absolutely nobody cared, because fast AI didn’t make any sense when the AI was dumb, and now the AI is smart, and getting faster, getting smarter, at an extraordinary rate, and everybody wants answers more quickly.”
Blundin had another take on how the switch happened, from gaming GPUs to a booming market for similar builds that happen to work well for AI.
“We’re definitely running on yesterday’s GPUs that were designed for graphics and video games,” he said. “And you know, it’s funny. When you talk to the people who are now, you know, multi-billionaires … none of them saw it coming. They were busy building the best conceivable GPUs for video games. And it was the AI researchers in the basements that said, ‘Hey, our algorithms will run better on that chip.’ And I think the audience for that was pretty limited at that point in time. They’re like, ‘Yeah, we don’t care about you.’ Boy, did that invert in a hurry.”
Raza thought back to previous decades of innovation, saying he was working on this stuff when they coined the term Moore’s law.
“I’ve competed with and beaten (Moore’s) company when it was a juggernaut,” he said. “Some of you may not be old enough to remember when Intel was so powerful that even Microsoft wouldn’t cross it. With a handful of people, we started out addressing the same problem.”
Skipping ahead to today, he had some suggestions for strategic advancement, with features like in-memory compute.
“Every new innovation addresses a bottleneck in AI, particularly as AI workloads change,” he noted. “The focus was once on training, then shifted to consumer inference. Now it’s moving toward enterprise inference, which is itself undergoing another transformation.”
Together, the crew pondered where the sticking points are on a flywheel that we’re seeing move in real time.
“I think the entire industry is constrained by data center availability right now,” Feldman said. “We’re trying to move at the speed of software, but data centers move at the speed of real estate. That’s holding us all back.”
Fab capacity, he predicted, will be a bottleneck in 3 years.
As for self-inflicted obstacles, here’s what Dave had to say about model engineering:
“What we do in AI inference right now is utterly idiotic,” he said. “The weights don’t need to move, yet we use massive amounts of energy and the world’s most expensive chips to move them from HBM memory (high bandwidth memory) into a GPU, use them for a picosecond, and then discard them. There are solutions that could be a thousand to a million times more efficient. They aren’t here yet, because the supply chain moves slowly, while software changes so quickly that nobody wants to hardwire an approach that might turn out to be wrong. Once we figure out the supply chain—which I expect will happen within three years—we’ll likely run far more efficiently by keeping the weights in place.”
Are We On the Right Track?
Later, I asked this question:
“We’re committing hundreds of billions of dollars to AI infrastructure. Are we building the foundation of the next economy, or overbuilding the biggest technology bubble in history?”
“I don’t think there’s any doubt we’re building the foundation for a fundamentally different economy,” Feldman replied.
Blundin agreed, invoking the example of the internet itself. People freaked out about the internet at its genesis, too, he noted. AI will, similarly, grow over time.
“The internet is the foundation of basically everything we do now,” he said. “It was never a bubble; people panicked, bailed out, and gave up on it. That always happens in America, and it will inevitably happen with AI. Twenty years from now, AI will still be the biggest thing that’s happened in human history, by far. … America always overheats. Capital floods into the wrong nooks and crannies, things collapse, and people say, ‘See, I was right all along.’ But AI is here to stay. It’s going to be massive.”
There’s a lot more in this wide-ranging convo about chips, industrial robots, and AI-native people and systems, with all kinds of insights on design and approach. It’s almost an hour long. Check it out.
Here’s the last part of this exchange, right before we had to go. I asked the panelists what they would say to young career pros.
“Build stuff you like to build,” Feldman said. “We were cool in the mid-to-late ’90s, with the rise of networking. Then, starting around 2001 or 2002, hardware was uncool for 15 or 16 years. Very few hardware companies were started. But I like building hardware. I like the people and the problems. It’s a good fit for my mind. This is what I like doing, and I like surrounding myself with people who enjoy building hardware: chips, printed circuit boards, and systems. Each brings its own set of problems.”
He continued with this animal analogy:
“My advice is: just because the world wants ducks, don’t stop being a leopard. If you’re a leopard, go do leopard stuff. You’ll enjoy your career much more.”
“Innovation has to happen in every dimension,” Raza added. “Throughout my career, I’ve looked at how workloads evolve: what the AI needs to do, how that work is built, and how it’s scheduled. We need good hardware, with memory, interconnects, and compute integrated as tightly as possible. We also need to schedule workloads more effectively across that hardware.”
And then there was Blundin’s input:
“If you can keep AI working productively through a chain of thought, it performs much better,” he pointed out. “That creates immediate opportunities. … AI could have the effective brainpower of billions of people, and think at extraordinary speed. But if every step requires a physical test, progress slows. That’s why simulation and other ways to create feedback loops with data will be especially promising this coming year.”
So interesting! Stay tuned for more of what came out of our September conference.