As artificial intelligence redraws the global computing landscape, former Intel CEO Pat Gelsinger points out bluntly that AI has made chip design simpler, yet physical bottlenecks in manufacturing, memory and energy are triggering an unprecedented hardware renaissance.
On October 9, Silicon Valley venture capital firm a16z aired an in-depth interview. In the conversation, Playground Global general partner and former Intel CEO Pat Gelsinger joined a16z's Raghu Raghuram and Guido Appenzeller for a wide-ranging discussion spanning chip design, memory innovation, optical interconnect, energy constraints and even the infrastructure for AI agents. A chip veteran who personally designed Intel's 386 and 486 processors and witnessed the birth of the modern EDA industry, Gelsinger offered a rare ground-level perspective on the current AI hardware wave.
Gelsinger opened with a central judgment: "For all of us hardware people, this is like a renaissance." In his view, from memory to optical interconnect, from power architecture to network topology, the density of hardware innovation opportunities today is unlike anything he has seen in his career.
Energy capacity equals economic output: data centers may face a wave of defaults
With AI infrastructure construction in full swing, the hardest constraint most easily overlooked by the market is electricity. Gelsinger elevated the energy issue to a macroeconomic level.
"In the AI digital era, energy capacity equals economic output," Gelsinger said. Over the past 15 years, the United States added renewable energy while phasing out coal, leaving overall national energy capacity nearly stagnant, and even the recent surge has annualized growth of only about 4%. That stands in severe misalignment with AI's enormous appetite, which can require million-GPU clusters at a time.
He issued a clear market warning: "Why build new data centers? If I buy a million GPUs and can't power them, you're going to see more and more of these data center projects default, because there simply isn't enough energy." Oracle's earlier remarks about nuclear-powered data centers, in Gelsinger's view, were merely an "initial signal" of the power bottleneck facing the industry.
To solve this, he argued that a fundamental rebuild of the power delivery network is urgent. "800-volt direct current returning to the data center can be called 'Edison's revenge.'" From generation and transmission to one-step voltage step-down using vertical GaN (gallium nitride) technology, as well as upgrades in cooling such as liquid cooling, the entire electrical system is entering a renewal cycle not seen in 25 years.
"HBM is a hideous memory": the biggest memory innovation in 30 years is coming
Regarding HBM (high-bandwidth memory), which the market currently favors heavily, Gelsinger offered a strikingly blunt assessment:
"HBM is a hideous memory. It's just the best we can get right now."
He pointed directly to HBM's physical limitations: extremely poor bit density, edge-limited bandwidth, high power consumption and serious thermal problems (DRAM is very sensitive to heat). Yet it is precisely these pain points, combined with AI's extreme reliance on memory-intensive computing and the $2.5 trillion in market value added by the memory industry over the past four years, that are driving a qualitative change in the industry.
"To be precise, how many major new memory innovations have we really had in the past 30 years? Only DRAM, SRAM and Flash," Gelsinger said. "The first memory innovation in 30 years is already close at hand." He is extremely bullish on new materials such as ferroelectrics for use in non-capacitive, high-density, stackable memory, and revealed that he has already invested in related stealth startups.
On chip stacking technology (3D packaging), he remained rationally restrained. He believes that, constrained by the mathematical rule that yields decline exponentially, "I'm not someone who crazily pursues 16-layer or 32-layer stacking. I think 2 to 4 layers of memory stacking will be the sweet spot."
Three months to design, nine months to manufacture: the "manufacturing time gap" after Moore's Law
AI is dramatically shortening front-end chip design time. Gelsinger believes that today's AI-assisted chip design is as epoch-making as the introduction of modern EDA tools during the 486 era. Using AI tools, a strong team can complete the logic design of an AI accelerator in three months.
But technological progress always shifts the bottleneck elsewhere. "There is no such thing as a pure 'chip' anymore; everything is a rack."
Gelsinger broke down the current manufacturing hell in detail:
"I spent 3 months designing, but I need at least 3 months to get wafers out of the fab, another 3 months for complex 3D advanced packaging, and finally I have to integrate it into a rack-level solution. So I spent 3 months designing, but I actually have to wait 9 months before I can truly start using it."
If you factor in the speed of software deployment and model iteration, this hardware delivery cycle of as long as a year and a half often leads to the awkward situation where "the chip has just landed, but the AI workload has already changed" (as Graphcore once faced).
A hundred AI chip companies in a free-for-all will eventually converge into a handful
Facing nearly a hundred kinds of AI inference accelerator chips emerging in the market, Gelsinger made the assertion that "the world's great trends divide after long unity and unite after long division." He believes the current architectural heterogeneity is only temporary, and the market will ultimately converge to a few dominant platforms.
His logic rests on three points:
Workload evolution: large language models (LLMs) are approaching bottlenecks, and future AI will incorporate more traditional HPC (high-performance computing) characteristics (such as demand for 64-bit high-precision computing for molecular or chemical analysis), making overly specialized chips difficult to adapt to models 2 to 4 years from now.
Capital and scale barriers: building AI computing scale requires tens of billions of dollars in capital expenditure. "Building these things requires enormous capital, which defies the logic that you can sustain 100 different chip architectures at the same time."
Choices of giants: top players such as OpenAI, Nvidia and Anthropic will actively pick winners, because the cost of co-evolving software and hardware ecosystems is extremely high.
Optical interconnect: the death of copper will come eventually
"I've been declaring the death of copper for about 20 to 25 years. Someday I'll be right."
Gelsinger was unequivocal on optical interconnect: all IO functions should migrate to optical interconnect, but the core compute-memory complex itself does not need to go down that path. He believes that the waveguide characteristics of copper cables in current data centers are making copper more expensive at 5 meters than fiber at 100 meters, and the laws of physics are forcing the migration to happen.
On the key timing, he pointed out that 2028-2029 will be the industry turning point—Nvidia has already made clear it will adopt NPO/CPO (near-package optics/co-packaged optics) solutions in that window, and other vendors will follow.
"NVL72 is an engineering miracle and a manufacturing nightmare. If we had had a more mature optical supply chain at the time, we would be more advanced today."
At the network architecture level, Gelsinger supports the move toward optical circuit switching (OCS). He cited network expert Nick McKeown's view:
The traditional internet was designed for "packets that might go anywhere," while AI workloads are "almost exactly the opposite—the traffic is massive and predictable. That looks more like circuit switching than packet switching."
The "new VMware" in the age of AI agents
At the end of the conversation, Gelsinger returned to his management experience from the VMware era to discuss the infrastructure paradigm in the age of AI agents.
"Every innovation leads to the next layer of abstraction." He believes the virtual machine abstraction remains a core paradigm worth preserving in the computing hierarchy, but the object it serves is shifting from "hardware and humans" to "agents."
"Every basic element of virtualization and management needs to be recreated at the next computing layer."
Gelsinger pointed out that a new generation of "agent VMware" needs to solve: how to let agents run efficiently, how to build secure configurations for agents, and how to manage agent performance, while also preserving the ability for humans to set policy, describe a "constitution" (citing Anthropic founder Dario Amodei's concept) and view dashboards.
The full transcript of the interview is as follows (translated with AI assistance).
Chapter 1: Intro
Pat Gelsinger: In the AI digital era, energy capacity equals economic capacity. Why build new data centers? Buy a million GPUs, and without power support, you'll see more and more data center projects default because the electricity simply cannot keep up.
Guido Appenzeller: Whenever you have some technology that makes something easy, the bottleneck shifts elsewhere.
Pat Gelsinger: There is no pure chip anymore; it's already a rack. I spent three months designing it, but I still have to wait another nine months before I can truly start using it. How many entirely new memory technologies have emerged in the past 30 years? The memory sector is about to see its first real innovation in 30 years. I declared the death of copper interconnect about 20 to 25 years ago. Eventually I will be proven right. For all of us hardware people, this is like a renaissance. AI inference accelerator chips—I definitely can't even name them all. Why are you funding so many of these projects?
Raghu Raghuram: You've funded quite a few yourself.
Guido Appenzeller: Never in history has an industry had 100 competing processor vendors at the same time. Is this just temporary? Will it eventually converge to a handful?
Chapter 2: From trade school to joining Intel at 18
Raghu Raghuram: Today I've invited our distinguished guest, dear friend, and my former boss—Pat Gelsinger. Welcome, Pat. Pat is currently a general partner at Playground Global, but he is widely known in the industry for his leadership, having served as Intel's chief technology officer, later led VMware, and held many other roles. Welcome.
Pat Gelsinger: Hey, thank you. Raghu, it's great to be with you and Guido, right? For me, this feels a bit like returning to old ground, you know, like... yeah, it feels like we've overlaid a year of changes, but you two still look the same, and I'm still energetic, so let's get started.
Raghu Raghuram: That's not how you'd describe David. You can't stop either—always so energetic, that's nothing new. But let's start with your career at Intel. Recently, Andreessen Horowitz launched an academy aimed at discovering talented people between 16 and 22, then placing them in a modern educational environment to prepare them to start their own businesses or join today's companies. You had a similar experience—you attended an ordinary vocational trade school, then went straight into Intel at 18 or 19, right?
Pat Gelsinger: Yes, that really was an amazing time. When I was 16, I happened to win a scholarship to a trade school and skipped the last year and a half of high school, and at 18 Intel came to recruit me. You know, I had never flown on a plane at that point, but I had already fallen deeply in love with learning at trade school. The interviewer, his name was Ron Smith, wrote on a piece of paper—I was the 12th person he interviewed that day, and if you interview 12 people in a row, by the end you can't even tell men from women, right? He wrote: smart, motivated, a little arrogant, he'll adapt. And that was it.
Pat Gelsinger: So I was invited to join Intel, and it truly was a brilliant experience. I started as a technician, later joined the design team, took part in wrapping up the 286, became engineer number 4 on the 386, and then architect and design manager for the 486. At the same time, I completed my bachelor's, master's and doctorate. It was simply the perfect career path—learning on the job during the day and putting what I learned into practice at night.
Raghu Raghuram: And it was entirely in the chip field you loved. I'd say the reason what you loved and built there produced such huge breakthroughs is that you were pioneering new territory.
Pat Gelsinger: Yes, in many ways. It really was an incredible time in the industry. I remember in a class at Stanford, a professor proposed a new carry-lookahead design. I was working on carry-lookahead for the 386 at the time, so I thought, let me go try it. Two days later I came back and told the professor: this doesn't work.
Pat Gelsinger: Then I started arguing with my Stanford professor, and the next two classes turned into me arguing with the professor about why his design was unworkable. He had a book about to be published that used that design, which made him very annoyed. We also started looking at his self-test methods. I remember Professor Ed McCluskey at Stanford, who could be called a founder of the field, and we had wonderful arguments about many of his ideas—ideas that weren't practical for real chips. But at the time I was applying built-in self-test to the 386. My thesis advisor was John Hennessy, who later did very well.
Pat Gelsinger: I say, I helped his career, and he later became president, and so on.
Raghu Raghuram: Yeah, we'll talk about the CISC versus RISC debate later, but I want to talk about the 486 first. Sorry, go ahead.
Guido Appenzeller: I had the same experience when I was teaching at Stanford. You walk into class with something you believe is the truth, and then a student working in industry says: well, actually it's not exactly like that. So the next class you come back and say: okay, I checked, you're right, here's the updated version.
Pat Gelsinger: Yes, you know, that's part of the student interaction, and I think that's one of the reasons Silicon Valley is so unique. There are so many companies here, and the entrepreneurial spirit is like it's dissolved in the water—startups, new ideas, challenging professors... To be a professor at a place like Stanford, you have to be prepared to be challenged and enjoy the process.
Guido Appenzeller: That building has a lot of traffic coming and going.
Chapter 3: The 486 and the birth of modern EDA
Raghu Raghuram: Yes, the 486 was your major achievement, and of course many people were involved. You pioneered some new territory in chip design, added capabilities chips should have, and so on.
Pat Gelsinger: Yes, that was also an amazing period, because it was before what you would call the birth of the EDA industry.
Raghu Raghuram: Wow, yes.
Pat Gelsinger: You know, there were a few hints at the end of the 386, but the 486 was truly the first chip to use modern EDA technology. We did high-level design, right? That is, RTL description, but at that time there was no Verilog, so we invented HDL, right? Intel Hardware Description Language, and I wrote a language myself. Of course, you also had to build a compiler for that language, right? So we created a compiler for the language. At that time there was no automatic place-and-route, so we had to invent it. We worked with Alberto Sangiovanni-Vincentelli and some of his students at Berkeley to complete the first place-and-route and the first automatic timing management.
Guido Appenzeller: So they created EDA, in a sense...
Pat Gelsinger: Yes, to a large extent. The 486 was not only a compatible pipelined microprocessor, we also laid much of the foundation for the modern EDA industry. It really was an amazing time for the industry.
Chapter 4: How AI is changing chip design
Guido Appenzeller: If we look at today, when people look back at Blackwell they say this was the last generation of carefully designed chips before AI took over the microprocessor design process. Are we now in a similar transition?
Pat Gelsinger: I think in some ways, yes. So I think when you look at things like Jalapeno today, you'd say this is a way of using AI from first principles in the chip design process, abandoning a lot of tradition. Of course, some parts of design remain very difficult in that sense. A lot of analog work—Circs may be the best example—still can't be done with AI, right? You just need a lot of silicon data to get these things right. So either you're very conservative in analog design so that AI can handle it, or...
Pat Gelsinger: You have to do it the old way, but now a lot of logic functions really can be done with AI tools and techniques in quite incredible ways, right? Your transistor budget is large enough, design tools are getting smarter, and a few experts guide the tools on how to apply them. I really think this will be a bit like the 486 era, right?
Pat Gelsinger: You'll look back and say, yes, that was the beginning of a new era in chip design.
Chapter 5: Why chip manufacturing still takes nine months
Guido Appenzeller: Whenever you have some technology that makes something easy, that means the bottleneck shifts elsewhere. So can we guess where the future bottleneck will be? There are things AI seems to do very well, right? Like handling the software layer, writing kernels, and a lot more... I think a lot of it is being able to specify goals at a higher level and then turn them into very long... So what's the new frontier? What will become difficult in the future?
Pat Gelsinger: You know, look at today. Suppose Guido and Pat, we're going to do a great new journey together. We're going to build a... right? This will be a leading thing, because we understand architecture, understand AI workloads, understand how to combine these, and then the right multiply-accumulate operations and register structures.
Guido Appenzeller: And can perfectly see tomorrow's model workloads.
Pat Gelsinger: All those things. Suppose we unleash all the AI tools—that would be really great. Within three months, we could get an excellent design, right? Of course, we'd be a bit more conservative on the analog components.
Pat Gelsinger: But we still have several major bottlenecks now. One of them is that we still need nine months of silicon processing time. Right? That's just how it is. So I can finish the design in three months, but actually producing the real chip at scale still takes nine months. That's terrible, right? So if we want to drive this kind of innovation, we really have to improve this bottleneck.
Raghu Raghuram: And even after that, we have more waiting cycles.
Pat Gelsinger: You know, first, getting something out of the fab takes at least three months, and that's with a highly optimized design flow. So how do I compress that time? Then it has to go into advanced packaging, and 3D packaging is getting more and more complex, which takes about another three months. Then it has to be integrated into a rack-level solution, because nothing is just a chip anymore—it's a rack. So that adds up to nine months. I spent three months designing, but to really start using it, it takes nine months. So how do we start compressing these parts of design? I think we need new forms of lithography to achieve that, and we need design flows that don't cost $50 million in masks.
Pat Gelsinger: Until I can do prototype testing. You tell me, how do you compress that to one or two months? Because I can finish the design in three months now, but if it takes a year and a half to truly reach volume production and get the software running, then my understanding of the AI workload no longer applies to the chip that finally tapes out, right? So all of this...
Guido Appenzeller: You can see this problem playing out in real time in AI right now.
Pat Gelsinger: Yes. We can look at history—Graphcore's design wasn't bad, but the world moved on, right? So all these aspects offset to some degree the fact that AI makes design easy. All the manufacturing and scaling problems have become huge bottlenecks for me. Of course, another bottleneck is memory, right? Memory bandwidth.
Guido Appenzeller: We know that well. Yes.
Pat Gelsinger: You know, as I described, HBM is a hideous memory, but it just happens to be the best one available.
Raghu Raghuram: We were just about to ask you about that.
Pat Gelsinger: Yes, it's a hideous memory. Poor bit density, bandwidth-limited, plus power and thermal problems. DRAM doesn't like high temperatures.
Guido Appenzeller: Stacking everything in one place.
Pat Gelsinger: For a memory-intensive workload like AI, that's just terrible. So I need to compress the time from development to real scale deployment, and I need to solve the memory problems associated with that. For me, that's another key bottleneck. And then power—power, cooling. Most power simulation tools today are poor; there's no good 3D power modeling tool, no power hotspot analysis, no good voltage input modeling. I'm basically using almost the entire power budget as a guard band, which means I lose about 40% of power to the guard band. So I have to do much better on power modeling, and my cooling technology is also behind.
Raghu Raghuram: The power delivery model also has to change.
Pat Gelsinger: Yes, all of that, for me, is the next batch of major bottlenecks. How do I improve each of these one by one? If I do, then we truly enter an exciting new era. Thanks to those great AI tools, but even with all these tools deployed, the challenges I face remain enormous.
Raghu Raghuram: We are indeed seeing a kind of renaissance in the number of chips and chip architectures.
Pat Gelsinger: Yes.
Raghu Raghuram: All thanks to these workloads.
Pat Gelsinger: There are probably about 100 AI inference accelerator chips now, and I definitely don't know all of them.
Raghu Raghuram: Yes, indeed. So...
Pat Gelsinger: Why are you investing in so many of these projects?
Chapter 6: Will 100 AI chips converge to a few?
Raghu Raghuram: You've funded quite a few yourself. Yes, so, you know, they're all good at making different design assumptions about the components you mentioned, or about some part of the Pareto curve they target. So do you think all these chips have room to exist? Or is it that, partly because of the AI toolchain developments Guido mentioned, we can now build more kinds of chips? And there's plenty of time. However you look at it...
Guido Appenzeller: More provocatively, historically, no industry has ever had 100 competing processor vendors, right? So is this a more heterogeneous future? Will we see—is this just temporary? Will it eventually converge back to a few?
Pat Gelsinger: I think it's more of a temporary phenomenon. I expect it will eventually converge. I see about three reasons for this. One is the emerging heterogeneity you see in computing environments—okay, now I have different chips specifically for prefill and decode. Then people start saying we really need a layer specifically for the intermediate phase, not just the prefill phase. So now I have early prefill and middle fill, and they're different.
Guido Appenzeller: Speculation and verification. Listen to what OpenAI says.
Pat Gelsinger: Yes, you know, suddenly the whole compute cluster is becoming more fine-grained across different workloads. I think that has historically been unsustainable at any time.
Pat Gelsinger: Then suddenly, when you see this, people start saying again that when I move to reasoning models, I want it to look more like a CPU. So suddenly I don't need that layer—that layer didn't become the dominant part of my compute capability, and instead I need more resources over here, right? Plus model workflows and so on. So I really don't like—I would say, large-scale specialization, right? Because I think workloads will continue to adjust and migrate substantially. That's the first point. The second point is that I think today's models are also about to undergo some evolutionary breakthroughs.
Pat Gelsinger: I think large language models (LLMs) are hitting certain limits. Now people are thinking about how to truly model three-dimensional structures—when you're dealing with molecules, chemicals and imaging, flat LLMs aren't very good. So I think we'll see limitations there, which will trigger different shifts in the algorithmic field. I also think some of the most interesting workloads are those that bring HPC-like things back into AI—for example, suddenly 64-bit precision becomes important again.
Raghu Raghuram: Can you give an example?
Pat Gelsinger: Okay, for example, suppose my AI model is now deriving the five most valuable areas for a chemical analysis. Those chemistry algorithms run at high precision—the most computationally intensive part is whether to find the five algorithms I need to run, or to run those five algorithms on different chemical or biological systems? So I do think that as we optimize AI systems, it will bring us back to some more traditional high-performance computing areas. So I also see this workload trend. Taken together, this extreme specialization, from my perspective as a computing architect, I don't like it much, because I don't think that's what workloads will look like 2, 3, 4 years from now. That's one reason I'm not bullish on these AI chips that become increasingly focused on one part of the compute workload.
Pat Gelsinger: The second point is that there are now about 100 such companies, right?
Pat Gelsinger: There are several doing optics, several doing probabilistic computing, and they can't all win—they can't all win because, in the end, these things need scale. And scale requires you to win the market, obtain capital, and get workloads running on it. So I think saying you'll have 100 such companies is itself against logic. Therefore, I think due to factors like capital and market share, the industry will return to—who is the winning team, the winning design, the winning architecture. I think workloads will drive that.
Pat Gelsinger: I foresee natural industry consolidation happening. At the same time, I also expect winners to pick some winners—for example, OpenAI, Nvidia, Anthropic will say, I like this one, because it's not just hardware, it's how hardware and software co-evolve. Right, moving forward. I do expect there are 10 I can choose from, but I like that one—and it really requires investment in software and workload evolution to fully exploit those platforms. For these three reasons, we'll see this field narrow.
Raghu Raghuram: And there's the issue of deployment complexity. Sorry, please continue.
Guido Appenzeller: That makes complete sense. I largely agree. But let me stir things up a bit and create some controversy.
Pat Gelsinger: You changed that, right?
Guido Appenzeller: So 100 companies may exist, or they may consolidate, right? I completely agree with your argument—in the end, large consumers will select winners to some extent.
Guido Appenzeller: I think to be a pioneer in this field today, to some extent whether you're in it depends on how much traction you get with the really big players. There's also an argument that during any transition period, you may see an explosion of more different approaches, which then gets eliminated over time. That said, I think there's an argument that historically we standardized on one or two architectures because different instruction sets were hard to handle—we needed to build multiple compilers, and basically software complexity dictated that you could only have a few hardware platforms. That part seems to be changing with AI. But today, writing optimized kernels for a chip—I think Kleina is a good example—as I said, humans can no longer really program this thing, but an agent can do it overnight. That's what we need, right? Basically you can say that if you can build a chip with a certain capability, it's now much easier to build the software layer on top that fully exploits it. Could the result be that we have a bit more heterogeneity simply because programming for that heterogeneity has become easier?
Pat Gelsinger: I do think that changes where the bottleneck is. The first bottleneck I'd point out—we've talked about it—is that building these things still takes a year and a half, right? The second bottleneck is that not only does it take a year and a half to get these things into racks and to scale, but to make them truly useful, can you borrow $50 billion?
Pat Gelsinger: I mean, we're not talking about small capital expenditure, but very large capital expenditure—data centers, power, commitments and so on. So I have to build these things at large scale, which is related. That will break a lot of heterogeneity.
Pat Gelsinger: Of course, the big players—we've seen what Nvidia did with Groq—will have some heterogeneity to some degree, right?
Pat Gelsinger: Two years from now, you won't know what architecture is running inside Nvidia; it will be like something else, right? So you'll see these architectures, and heterogeneity will be hidden inside, which is completely fine with me. That's what layering and abstraction have always done. But I think those forces are powerful and will expose a lot of heterogeneity—no, I think it will be abstracted beneath these things. But in the end, building these things at large scale is still hard.
Chapter 7: Why HBM is a hideous memory
Raghu Raghuram: Okay, you mentioned that HBM is a hideous memory.
Pat Gelsinger: But it's the best memory we currently have.
Raghu Raghuram: Indeed. But the frustrating thing about memory is that I really can't remember the last memory innovation, maybe eDRAM (embedded dynamic random access memory). I think you may have been involved in that too.
Pat Gelsinger: Yes, that project died, so I count that as a failure too.
Raghu Raghuram: We killed it. So do you think memory innovation is coming? Or what needs to happen? Obviously, we understand the scale problem, because the scale problem is bigger there. But if we set scale aside, even from a purely technical perspective...
Pat Gelsinger: Yes, if you look at that history, I personally may have been involved in at least five different new memory architectures, and eDRAM was only one of them; none of them saw the light of day. I probably also know about nearly 100 similar cases in the industry over the past 30 years, and how many major new memories have we really had in the past 30 years? Right, you know, okay—DRAM, SRAM, Flash, okay, what else? DRAM, SRAM, Flash, yes. So I do think...
Raghu Raghuram: Flash stacking counts a little.
Pat Gelsinger: Barely, yes, right, you know, anyway, yes, so in that respect it's really disappointing that we haven't created the next physical technology that can deliver low-cost, high-performance, reliable and manufacturable memory. So I do think that over the past 30 or 40 years, memory has been a terrible industry, because you might be profitable for one year and then lose money for the next four. So doing R&D in an industry that is profitable only one year out of five is extremely hard, because it has always been driven by such severe commoditization cycles.
Guido Appenzeller: But that has changed to some extent.
Pat Gelsinger: In the past five years.
Guido Appenzeller: All three major memory companies are among the top 20 most valuable companies in the world. That's a crazy shift.
Pat Gelsinger: Truly astonishing, right? Who would have thought? Anyone paying attention to AI workloads knows it's easy to predict the past and hard to predict the future, right? And when you look at all this, you'd say AI is a memory world, and indeed it is, right there in the middle of it. So I think the combination of capital and technological demand will drive memory innovation.
Pat Gelsinger: There are two areas that excite me now. One is that I do think HBM's bandwidth limitation is a fundamental problem. So based on that, the idea... obviously we have a company, D-Matrix, doing this, but I think other companies will join in too, whether it's Cerebras and its technology, or others.
Chapter 8: How high can chips stack?
Guido Appenzeller: The future is stacked.
Pat Gelsinger: You have to find a way to integrate memory.
Raghu Raghuram: And you have to consider it from multiple dimensions.
Pat Gelsinger: You have to integrate memory and compute together, right? There has to be an entirely new foundational architecture. I do think that, given the huge capital already invested in AI to address these workloads, we can actually break through some of these physical challenges. So I believe, you know, I just...
Raghu Raghuram: Invested in something, right?
Pat Gelsinger: You know, I just invested in a new memory company too. It's still in stealth, but you know what?
Pat Gelsinger: I'm very excited about some of the innovations about to happen in the new memory space. You'll see new materials emerge, such as ferroelectrics for memory, searching for non-capacitive, high-density memory, searching for memory that can achieve high performance, high density and stackability, right? But, you know, some people are looking at using low-cost flash to achieve this. Personally, I'm not particularly enthusiastic about that idea. I think flash or MRAM structures have some issues at the underlying cell level in terms of speed and performance. Yes, but I do think memory is about to see its first innovation in 30 years. I think it will be an exciting field. I believe several companies will emerge there.
Pat Gelsinger: Okay, you know, we'll find some things to do here. But I think the first memory innovation in 30 years is right in front of us.
Pat Gelsinger: You know, I think the memory industry has added $2.5 trillion in market value over the past four years. I think that number is roughly right.
Guido Appenzeller: We could have just invested there and not done anything differentiated.
Raghu Raghuram: Yes, exactly. Just use your money. Hmm, what do you think? Sorry, go ahead.
Guido Appenzeller: Let me expand a bit on the stacking part. If you were to predict how fast chips will "grow taller," would it be like... you know, will we get to 2 layers and then stop there for a while, or will we have, say, 16-layer sandwiches? You know, logic layers and memory layers happily mixed together?
Raghu Raghuram: HBM is pulling that down.
Guido Appenzeller: HBM is moving downward on Nvidia, but how do you see this evolution?
Pat Gelsinger: Yes, you know, I think the problem with stacking is that as the stacking dimension increases, the value of the entire stack also increases, which means the yield of each layer and the stacking process itself must continuously improve. Basically, it can't improve linearly; it has to improve exponentially. When you stack to 8 or 16 layers high, improving manufacturing and design characteristics exponentially becomes extremely difficult, and when you do the math, your manufacturing process has to be nearly perfect.
Guido Appenzeller: They may see designs that take into account a particular class of models, just as we saw in the 2D world for a long time.
Pat Gelsinger: Of course. But remember, that's the nature of DRAM. You always have bad-block replacement mechanisms, like spare rows and so on. So I'm not minimizing this in any sense, but suppose you've done all that, you still need very high yields. You know, if I have a cracked chip, no matter how much redundancy you build in, you can't fix it—it's just a bad chip, right? So because of that, I'm not the kind of person who crazily pursues 16-layer or 32-layer stacking. I think 3, 4, 5-layer stacking will emerge, and that will ultimately become the optimal balance point.
Guido Appenzeller: Mid-rise buildings. Yes.
Pat Gelsinger: Yes, you know, partly because of manufacturing, partly because of the associated complexity, and where the yield sweet spot is, so that's where I think most things will land. But even so, right? For example, suppose I now have a great compute chip—how big does that chip need to be? Now everything is limited by the lithography field, right? Can I use more chiplets to get better yield characteristics and better heat dissipation, and then you'd say I want a memory stack on top.
Pat Gelsinger: Is it a two-layer memory stack or a four-layer memory stack? I don't think it will become an eight-layer memory stack. So probably 2 or 4 layers, right? But power delivery also needs to be integrated, which becomes another element. I want it in the Z-axis direction, because I want my power array to come in from that direction too, right? Then I need fairly intelligent RDL layers to distribute signals across the entire platform. So to some extent, you know, when you add all these parts together, my standard structure is already about 8 layers thick, and I think that's roughly the limit, at least given our current understanding of manufacturing processes.
Pat Gelsinger: Of course, we expect optics to join in as well. So there will be more III-V material structures closer to this complex, because I want my IO to be deeply integrated into this architecture. So my silicon-centric part needs to be complemented by these other III-V materials. How do I integrate optical interconnect? How do I integrate power rails?
Pat Gelsinger: You know, it's an engineering feat and a manufacturing nightmare to bring all these parts together, but that's where physics is taking us over time. So a 16-layer-high stack is crazy to me. A 2-layer or 4-layer-high memory stack, I think that's roughly where the optimal balance point is.
Raghu Raghuram: You could also talk about thermal issues and the new materials needed for that.
Pat Gelsinger: Obviously, first there will be many challenges—you simply cannot allow the physical hotspots that exist today. So you have to use heat more intelligently and spread it out. But you know, then you say, okay, how do I both input power and dissipate heat in these environments? By the way, if I can input power more efficiently at the voltage I want, while having better power response and voltage regulation, then I'll generate less heat, but I do need to think about, okay, how do I cool these things?
Pat Gelsinger: New materials, such as diamond and other proposed things, can play a role in improving heat dissipation and cooling. How do I adopt better liquid cooling technology? You know, attached directly on top, with good thermal conductivity. So, you know, we engineers are becoming plumbers, right? Indeed.
Raghu Raghuram: Yes. You're worried about chip-level electromagnetic compatibility. Yes, let me ask you about the opposite phenomenon. Some people say, hey, let's take memory off the package and connect compute and memory entirely with optics. That can give you more degrees of freedom and more ways to build systems. Compared with stacking, how do you view that approach? That's the other extreme.
Pat Gelsinger: Yes, you know, I think the more OEO—right, interfaces, meaning optical-electrical-optical conversions—the more of those conversions you need in your core compute complex, the harder it is for me, in that sense. And you always have losses associated with it.
Pat Gelsinger: You know, people are trying to come up with smarter techniques to reduce cost and power, right? How much power am I wasting going through DAC and various conversions? So I don't like this approach, because I think that while it lets you access large memory pools, it consumes a lot of power to get there. Yes, right, remember, basically the femtojoule per computation is now quite astonishingly low, right? But the femtojoule per bit of communication is 1,000 times worse by comparison. If I have to spend that much energy to access large memory pools, I don't like these methods; I think they go against the underlying physics of the solution to some extent.
Guido Appenzeller: I mean, there's another approach, which is, can we actually integrate compute and memory directly, right? Move some of the computation into memory?
Pat Gelsinger: That's some of the PIM (processing-in-memory) approaches.
Guido Appenzeller: Exactly. So what's your view on that?
Pat Gelsinger: I don't like those either, right? Because I think workloads—you know, because now I have to take my workload and push these little blocks of computation into memory, right? That way I don't need as much memory bandwidth flowing into my real compute units. But that imposes constraints on the workload. Again, I may be wrong, especially given some of the design flexibility and intelligence we can now bring through design, but most PIM ideas have existed for 25 to 30 years, and I think in another 25 to 30 years they'll still be there, in that sense. You know, give me high-density, high-performance, large memory bandwidth, and put it close to the real compute engine. I expect that will be a more general-purpose solution for many evolving workloads.
Guido Appenzeller: So we're still stuck with increasingly dense chips, stacking, more logic, more power.
Pat Gelsinger: But it will be much better. I am a huge fan of optics, you know, moving to optics for all IO functions. Yes, right? You know, I declared the death of copper about 20 to 25 years ago, and eventually I'll be right.
Guido Appenzeller: You're still right.
Raghu Raghuram: One day.
Pat Gelsinger: Soon it will be right. But I think, for me, all IO should move to optics, right? But the core compute-memory complex, no. Right?
Raghu Raghuram: So when you take that first step, our article is about scale-up even more than scale-out. What's your view on that?
Pat Gelsinger: Yes, you know, I think we now...
Raghu Raghuram: If you could, what would be the application scenarios driving the workload?
Pat Gelsinger: Well, I think what drives the workload is the one right in front of us, right?
Pat Gelsinger: I mean, you know, today, in our scale-up environments, we lay down a lot of copper cables, right? You know, copper is essentially becoming a waveguide, and increasingly shorter. Suddenly, you know, making copper work over 5 meters costs more than making fiber work over 100 meters, right? So, you know, I think physics is pushing us down this path. So, Pat, if that's the case, why haven't we migrated yet?
Pat Gelsinger: It's a complex supply chain, you know, never tested at scale. I think right now everyone is considering some form of NPO, CPO, etc., with a timing around '28, '29, because without making that shift, I can't scale the scale-up environment to large-base compute clusters. So clearly Nvidia has indicated this. I think everyone else in the industry expects that to be the turning point. You know, honestly, I shouldn't have built NVL 72. It's an engineering miracle, but also a manufacturing nightmare. You know, the whole industry spent an extra year and a half digesting this monster, right?
Pat Gelsinger: Truly amazing, you know, that we could scale it to that size. But, you know, it does prove that if we had had a more mature optical supply chain at the time, we would be more advanced today. We could build better scale-up environments today, with better power performance, lower cost, better energy efficiency, but we didn't have a sufficiently mature optical supply chain to support all of that.
Pat Gelsinger: These problems are not simple to solve, and there's still a lot of work to do, but I think '28 and '29 are the years when capacity will explode.
Guido Appenzeller: We may have accumulated a lot of experience in how packaged optics apply to networking, right? I think a lot of that experience can transfer well to GPUs.
Pat Gelsinger: Yes, I do think so. But suddenly, you know, it's like, okay, these things have to operate at the scale of 10 million GPUs, right? It has to be integrated into the package, and there are many other issues. You know, many forms of optics don't like harsh thermal environments. But guess what? Harsh thermal environments are exactly the norm in IT. So we have thermal management issues to deal with, package integration issues to deal with. There's also not enough laser capacity in the industry, so there are a bunch of issues, but I think there are no fundamental obstacles here, just a lot of hard work integrating all these parts together.
Guido Appenzeller: On network architecture, there are also some subtractions, like switching, and what else, more circuit-switched networks, simply because you can't scale these very fast optical networks indefinitely.
Pat Gelsinger: Well, you know, I like you, Nick McKeown, or another example, the professor here, also a good friend of mine, you know, I mean, you know a little about him too, Guido.
Guido Appenzeller: He was my thesis advisor, I...
Pat Gelsinger: So, but you know, he described it well—you know, we build networks to handle any packet going anywhere, without needing to know where it might go. When you think about AI, it's almost exactly the opposite, right? We're dealing with massive and predictable traffic flows. So just from that description, it looks more like circuit switching than packet switching. So I do think the convergence of migrating to optical networks and having a more predictable, manageable, high-volume workload—for me, I do see that. Yes, we'll go all the way to bringing optics into the compute complex, and we'll go all the way to OCS or some form of optical switching, you know, paired with more predictable traffic flows—that's the right architecture.
Pat Gelsinger: Then you ask, is this scale-up or scale-out? In some ways...
Raghu Raghuram: It blurs that boundary to some extent.
Pat Gelsinger: Yes, right? If my scale-up is large enough, then you just have to ask, what is the radius of my scale-up environment? One day I want to connect it to other really large clusters. For me, that will ultimately be a pretty elegant answer, you know, future-proof. But I'm looking at someone here who is more expert in this than I am.
Guido Appenzeller: Now, look, I've always thought that scale-up versus scale-out is somewhat of an artificial distinction. You can understand where it comes from, right? Because we use different protocols, one for the front-end and one for the back-end network, and so on. So now, in practice, it really is quite different, but if you step back and look from a system architecture perspective, they achieve the same goal, just with slightly different characteristics and slightly different protocols, and I think over time these will converge.
Raghu Raghuram: Yes, I mean, especially for training workloads.
Guido Appenzeller: Yes, although I mean training and inference look quite similar.
Pat Gelsinger: Yes, now I'd say, you know, right? Does your brain stop learning? You know, by the way, that's another feature of workload evolution. I do think the separation between training and inference will become less and less common in the future, not more, right? That's also a small argument against more heterogeneous architectures, I think: continual learning. I think there will be some algorithmic areas that say, yes, continual learning becomes the essence of the algorithm, where, yes, I do want my inference environment to update my model weights, because, you know, that gives me more and more granularity, evolving as workloads migrate into those less formal training environments, or in environments with smaller data sets. So I'm not sure that's...
Raghu Raghuram: But the number of nodes you need to be useful—the diameter of inference is smaller than training.
Guido Appenzeller: Today, yes.
Guido Appenzeller: I still think from another angle that heterogeneity may play a role, which is that different types of AI workloads require very different underlying support; some are more compute-bound, some are more memory-bound. We also have small diffusion model workloads, right? Their optimal chip looks very different from that top-tier LLM with, sorry, 50 trillion parameters, right? The latter requires creating massive clusters. I think this may drive heterogenization.
Pat Gelsinger: Yes, I think there's some room for that. Yes, but I also want to say, okay, now I'm talking about getting a million GPUs to work together in my data center, right? You'd say, okay, for that particular diffusion model workload, I really should put in some specially optimized thing, right? So I won't start putting in, you know, 1,000 of those, right? And for that workload, I need, you know, 2,000 of these, and then for some... suddenly, you know, you lose some of the flexibility that scale brings, right? You'd say, you know, how much will management and upgrades, connectivity, failure domains, all these things cost me?
Guido Appenzeller: Agents will do all the upgrade work.
Pat Gelsinger: Yes. Yes, thank you. Everything becomes simple, great, I don't care anymore.
Chapter 9: Energy capacity equals economic capacity
Raghu Raghuram: Okay, so fleet management is done. We discussed compute, memory, networking, and next is power.
Pat Gelsinger: Yes, yes, yes.
Raghu Raghuram: You mentioned vertical power delivery inside the chip, on the chip. But I mean, there's also 800-volt DC, data center power transmission, and the whole ecosystem extending all the way to generation, plus the related sociopolitical dimensions, and so on.
Pat Gelsinger: Yes, yes.
Pat Gelsinger: If you start from basic principles, our country has done a very poor job on energy capacity. Basically, we went through 10 to 15 years during which the pace at which I retired coal roughly matched the pace at which I added renewables.
Pat Gelsinger: The country's overall energy capacity has essentially stagnated, which is very bad, because in the AI digital era, energy capacity equals economic capacity. So basically, as a country, my economic capacity has stagnated for 15 years. If you accept that—and I think it's very demonstrable over the past five years—we've seen a huge influx just to give me more capacity. But I think even so, our country's energy capacity is growing only about 4% a year, from basically 0 growth to 1, right? Wow, a 4x increase, or 4%. So for me, this is a bad situation, and we really do need more energy capacity.
Pat Gelsinger: First, obviously we have some companies in this area.
Pat Gelsinger: We talked about nuclear, one of which operates with Alva. I think nuclear is very good as baseload, and we want to continue building more. Unfortunately, all our renewables are deeply dependent on China, which is very regrettable. You know, gas turbine lead times are currently eight years, right? So we really face a difficult environment for rapid expansion, but first, we simply need more capacity. This fundamentally constrains how far we can go in AI, because if I can't power it, why build new data centers and buy millions of GPUs? I think you'll see more and more data center projects default, with energy simply not available.
Raghu Raghuram: So you think this will be a major headwind?
Pat Gelsinger: I think it will be a headwind, and I think you're already starting to see some early signs. That Oracle comment, I think, is just the first of many cases you'll see. Now, that doesn't mean we'll slow down because of it, but I think people will start asking: do I really have enough energy to support the capital commitments I've made for cement, data center racks and GPUs? So do you think this will become a ceiling? We need more energy capacity, more innovation in this area, and more supply chains built in the United States. Nuclear is one of my favorites. When was the last time a U.S. nuclear reactor came online?
Pat Gelsinger: It was 20 years ago, right? We just stopped a very capable industry like that. So this is one of the areas where we need more. We need to become more efficient in power transmission. As you said, moving data centers to 800-volt DC is absolutely necessary, to be able to eliminate many conversion steps along the way that exist for historical and standardization reasons. Well, we have to solve this. 800-volt DC, Edison's revenge is coming, as I like to say. I do think we need better power conversion. We have several companies, one of which does vertical GaN, and I think that will be a killer technology, enabling 800-to-48, 800-to-12, or even 800-to-5 conversion in a single conversion step, which gives you more energy. We mentioned
Raghu Raghuram: Solid-state transformers, you know.
Pat Gelsinger: Yes, all of that—solid-state transformers, switches and so on—rebuilding the entire power transmission network. I think there will be a lot of innovation there, as if, okay, power is cool again, right? And then you need all of this in cooling systems too, right? So there will be a series of innovations—turbines, refrigerants—and suddenly they become exciting new technologies again. So for all of us hardware people, it's like a renaissance right in front of us, and it will...
Raghu Raghuram: After 25 years, I have to go back and relearn the Carnot cycle.
Pat Gelsinger: Carnot efficiency, oh yes, is that good? How the first law of thermodynamics fights back.
Guido Appenzeller: After all, there are still some things here.
Pat Gelsinger: So this will be an exciting era. But we also say we need to change the laws of physics, because over the past five generations of Nvidia chips, CMOS has shown no improvement in basic power consumption or power per teraflop—completely stagnant, and that's unacceptable. That's why we have the superconducting company Snow Cap, which offers the opportunity to fundamentally change the laws of physics, with a thousandfold improvement in power performance. So I do think these innovations will also arrive in the near future.
Chapter 10: VMware for agents
Raghu Raghuram: Over the past hour, my time management has been terrible. There are too many topics to discuss and we haven't covered them, but before we finish, let's talk about at least one more.
Pat Gelsinger: Okay, you lead.
Raghu Raghuram: Okay, here's what I think. Now virtual machines are back. How do you view this, and where is VMware heading?
Pat Gelsinger: Hmm, you know...
Guido Appenzeller: I feel left out of this conversation.
Raghu Raghuram: Let's talk about networking too.
Pat Gelsinger: Yes, I think every innovation leads to the next layer of abstraction, right? I think a lot of these ideas... you know, I like to say, what's new in our world today? Data structures, algorithms, abstractions... we trace them back to those foundational things. I think the virtual machine abstraction, whether implemented at the infrastructure layer or the application layer, remains a foundational abstraction model and always has a place in the computing hierarchy. So, what do you think? Are any of us going back to run VMware again?
Raghu Raghuram: Right? Back to what year?
Pat Gelsinger: Yeah. You know, I do think these abstractions... but you also have to think about it. Like what we did at VMware, we basically managed every aspect of the computing hierarchy—managing workloads, managing networks, managing storage systems. How do you abstract these? In the context of AI, when we think about the shape of these agents—who manages these agents? Who creates secure configurations for all agents? Who manages agent performance? Essentially, every basic element of virtualization and management needs to be rebuilt at this next computing layer. There will be many great companies making breakthroughs in this area to make these things happen.
Raghu Raghuram: Do you want to end on the networking topic?
Guido Appenzeller: I want to stay on the virtual machine topic a bit longer. What's most interesting to me is that we're now building virtual machines for agents to use, not for humans. What we've seen from other companies so far is that this changes a lot—the product shape changes, your go-to-market changes, for example whether you want agents to choose your virtual machine product, that changes the level of complexity you can afford, and it changes the requirements for startup speed. Humans are much more patient than agents. So if you were rebuilding VMware today, what would be different?
Pat Gelsinger: Yes, I do think...
Guido Appenzeller: VMware for agents.
Pat Gelsinger: Yes. I also do think... you know, do you need both a VMware for agents and a VMware for humans?
Guido Appenzeller: I see what you mean. Yes.
Pat Gelsinger: Right. You know, in a sense, I still need humans involved.
Guido Appenzeller: Let's focus on agents. They're much more interesting than humans now.
Pat Gelsinger: Yes. But, you know, I think this is an important point, because I do feel the need... let me borrow Dario's concept of a "constitution," meaning guardrails, or being able to set policy. Then you can really say, okay, every aspect of the virtual machine, as we know and love it today, is now serving agents, rather than serving hardware and humans. So I think that becomes your basic design constraint.
Pat Gelsinger: On this side, you have to think: how do I make agents better? How do I make them safer? How do I make them more performant? How do I abstract them? How do I migrate them? All these questions will have new implementation forms—new forms for agents. But at the same time I also need to think: how do humans set policy, write the "constitution," get dashboards and so on? I think that's where it gets interesting, because you need to do both, not just one. I'd say, just like those core design challenges we faced for virtual machines back then—abstracting hardware for applications or operating systems in a high-performance way—here, abstracting hardware and operations for agents is the fundamental thing you're really serving.
Raghu Raghuram: This discussion has been excellent. Thank you very much for coming. Let's talk again next time.
Pat Gelsinger: You two are my favorite people. I'm available anytime.
Raghu Raghuram: Okay.
Guido Appenzeller: Great, thank you.
Pat Gelsinger: Thank you.