0:00
/

🎙️ NEW EPISODE: OpenAI’s Jalapeño! Feeling Hot Hot Hot!

OpenAI's Jalapeño inference chip, a 9-month RTL-to-tapeout design cycle, NUMA-style local HBM slices, ESUN scale-up networking, and more

Austin and Vik dive into OpenAI’s Hot Chips announcement: the Jalapeño inference accelerator. They break down the chip’s key architectural decisions, from its NUMA-style HBM slices to its balanced design philosophy of “dark silicon is cheaper than idle accelerators.” The hosts also explore the implications of a nine-month design cycle and what it means for the competitive landscape.

Things we cover:

  • The Jalapeño chip’s design philosophy: user experience and energy per request

  • A nine-month design cycle using AI tools

  • NUMA-style architecture with local HBM slices

  • Scale-up networking with Broadcom and ESUN

  • The “dark silicon is cheaper than idle accelerators” principle

  • Jalapeño’s performance on the Inference X benchmark

This podcast is lightly edited for clarity.

A Hot Chip for Hot Chips

Vik: We’re recording this less than 12 hours since OpenAI actually announced Jalapeño at Hot Chips. And I have got a flight to catch in about four hours from now to travel across the globe. But this couldn’t wait, so we wanted to talk about OpenAI’s Jalapeño inference chip as soon as possible and break it down to the level of understanding we have and from the conversations we had with people at Hot Chips. So let’s get into it.

Austin: Hello listeners. We’ve got another Semi Doped coming for you today. So like Vik said, Jalapeño and OpenAI’s talk was so awesome. We just wanted to get straight into it as soon as we could. Obviously lots to digest. This won’t be comprehensive, but I think we both have a lot of interesting takeaways as we watched it. And I’ll just say that I personally thought it was the best talk at Hot Chips that I saw. I’m watching it remotely.

And, you know, the first thing that came to mind was like, did they name this chip Jalapeño because they knew they were going to launch it at Hot Chips?

Vik: Yeah, that is so funny because everybody in the volunteers and the organizers there were wearing these yellow shirts with like red jalapeños on them and I’m like, wait, is this like a self-organized advertisement for OpenAI Jalapeño? Like everywhere, like this subliminal messaging across the whole conference was these chili peppers. And I’m like, okay, I don’t know if it’s a coincidence or planned, but it’s pretty cool.

Austin: Totally. Yes, 4D chess if that was the plan, that’s genius.

Designing for User Experience, Not TCO

Austin: But okay, so let’s get into it. So, you know, I just took some of their slides and a few screenshots from SemiAnalysis, who had an amazing article, so go read it if you haven’t. Props to SemiAnalysis and OpenAI for benchmarking this chip together in advance so that SemiAnalysis had a great article to drop.

But let’s start talking through some of the slides from the talk and pull out our insights. And I thought what would be really interesting to start with is sort of the vision that was set by the OpenAI team was Richard Ho and Ravi and Chris. So Richard was previously with Google working on TPUs and has like a long and interesting legacy. And then Ravi was the chip architect and Chris was kind of like the software co-design guy. All of them were very interesting. And, early on they said that really the two metrics that they were designing for, which I put on this slide, are the user experience and which they defined not as the time to first token, but really like the time to last token, the end-to-end latency. Ultimately, it’s not like how quickly does this start processing, it’s just how quickly does it get done, right? And then also the energy per request. So how much energy did it take to do that?

And what of course I thought also was interesting was they basically said, look, those two are at odds with each other. You can always go faster if you use more energy. Therefore we ought not to report stand-alone numbers, but we should always try to show Pareto frontier curves whenever possible to show you as we dial one knob, how it impacts the other. And so I thought I kind of liked that because right away they’re saying like, we’re not trying to game this, we’re going to show you the full curves. But what the other thing that stood out to me right away was like, hey, wait a minute, this is really interesting. This is a model lab whose customers are AI users designing chips and they are designing first and foremost for the user, the end AI user experience and not cost. And I thought that was interesting. I’ve seen a lot of presentations from merchant silicon vendors and they are selling to the AI labs or to the hyperscalers. But they’re ultimately, they have a different set of customers, so to speak. And that can impact even design decisions.

And so, you know, traditionally, you hear a lot of hyperscalers or sorry, merchant silicon vendors right away talking about designing for TCO, which makes a lot of sense. But I just thought it was interesting here that we’ve got a different person design, a different sort of team designing the chips that is closer to the end user. And then therefore they were actually thinking about you and I using AI more than they were thinking about the buyer of the silicon. So I just wanted to pause there and that I thought now that we’ve got a model lab kind of vertically integrating and designing their own chips, it’s interesting to see how they might make design decisions differently than a silicon vendor designing the chips.

Vik: So, Nvidia and Jensen actually has said this before, if you’re choosing an architecture for a chip that is based on making the chip cheaper to make or something like that, it’s a bad design decision. You’ve got to make the best performance decisions period because we are not in the era of trying to save money in making of a chip, but like the running of a chip, the tokens per joule, which is a nicer way of saying tokens per second per watt, which is like too many division symbols there, but yeah, watts and seconds can be mangled to form joules. So tokens per joule, joule is a unit of energy. Watt is a unit of power, right? So, the tokens per joule is a very important metric. So that is the ultimate cost of ownership because you are going to put in energy and you’re going to get out tokens, right? You’re going to put in joules, you’re going to get out tokens.

So the energy per request is one important metric. But I don’t think the concept is new. People have been putting tokens per second per megawatt on these interactivity curves that you see all the time and I’m sure it’s there in this presentation too. All the time. So it’s always like this is the Y-axis in most curves. And then some form of latency or how the user is served. So on these charts we will see, it’s usually tokens per second per user, right? That’s the X-axis in most of these interactivity curves you see. But yeah, so their idea is just essentially to bring on the TPU team and vibe code a chip in nine months. Okay, that’s their whole play.

A General-Purpose Inference Chip?

Austin: Yes, totally. And I think, yeah, maybe last thought here, there was a quote, I think from Ravi as he was talking about that said that there’s a market for tokens, it wants them fast, it wants them cheap, and making things go fast gives us joy. So I thought, oh cool, what a cool team to work on too.

So then they had a slide where the OpenAI team gave major props to SemiAnalysis’s Inference X and said, hey, we benchmarked against Inference X. We think it’s a fair public power normalized comparison. So again, saying we want to be as transparent about how our system performs as possible. And a couple things they pointed out is instead of picking a particular model that fit well, what they liked about Inference X was there’s multiple open source models. So they can be tested on something that’s small, something that’s big, something in the middle. And they also liked that at the end of the day, all that really matters is that user experience from the entire system. So they thought this was a fair way to capture and compare end-to-end latency and how the whole system does. So tune your hardware, tune your software, tune your networking, whatever you need to tune to make the best experience possible for users.

I will point out, you know, some of the feedback that in the conversations that have happened since were people pointed out that this was Jalapeño was tested with a pretty small input and output context length. And so, you know, I think there’s probably questions on, okay, that’s cool, that’s awesome, but would love to see how it performs with like, you know, a million input context length. Did you talk to anyone on that topic?

Vik: No, actually. People were generally quite happy with this, but not amazed. I was more towards the amazed side than, you know, I mean, I was obviously happy about seeing the chip come out or whatever. But to me it’s just like the performance, I read the SemiAnalysis article at least halfway even before I attended the talk. And it’s a very performant chip. To be clear, I think even SemiAnalysis says that they have not run any agentic workloads on it. So they also have the AgentX platform that it has to be benchmarks against at some point. I’m sure they’ll get there, but Inference X is a good start.

Throughout the conversation, I was trying to figure out what are the parameters here that would break the Jalapeño chip. Like, are they cherry picking, let’s say, I don’t know, the input sequence length and output sequence length here that is like 8K and 1K is one metric. But I’m just always trying to think of like, is this a cherry pick use case? You know, so I’m trying to be critical the whole time, but it’s really hard to tell because if you think that this is an easy case, then you’re like, oh, wait, what what’s the number in the other case? So it’s hard to discern, right, figure out the whole thing.

Austin: Yes, yes.

Vik: But one thing is very useful is this slide. Yeah. Actually, you should just go there. Yeah, because if I mean, if anybody’s listening to this, we are actually talking through the slides on YouTube. So if you want to like look at the slides and see our annotations on the slides, we have them all marked up. So, you know, we could obviously, try to talk through some of the stuff we’re seeing. So if you’re watching this or listening to this while driving, shouldn’t be watching while you’re driving, but yeah. We try to explain, we try to explain some of these things.

One thing that people immediately thought when Jalapeño came out was like, oh, it’s optimized for the OpenAI models. But they actually showed that it’s not so, and they actually ran GPT OSS, which is a narrow small model, all the way to a Kimi K2.5 model, which is a much different shaped model. So it’s not designed for that, and SemiAnalysis points out that in fact, they were able to play Doom on it, you know. I guess it could do other things than inference, but I’m not sure how Doom plays into this. I guess, yeah, cool, it can do other things is all I take away.

Austin: Yeah, yeah. No, okay, so you make very interesting points, which is like, we want to ask what all is this chip good for and where does it kind of break and are they cherry picking something here? And I think the point of this slide was the OpenAI team saying, look, we made a sort of generalized inference chip. So it is an inference chip. It is focused on inference, not training. It is focused on LLMs, not convolutional neural network inference, right? But it’s general enough that it can run a small model, it can run a medium-sized model, it can run a huge model. And so I think, you know, as you’re trying to think about all the different corner cases that it can and can’t handle, I think probably the only knob that we would like to see more is really that context length, because I think right here they’re showing, you know, hey, we can run these different models. And to your point, it’s not just specifically the OpenAI models, but it’s more broadly, it can run any open source model.

Now, of course, right away you have to ask, oh, interesting. They’re obviously as engineers trying to show that it’s flexible enough to handle innovations that come down the pipe in the future, whether it’s algorithmic innovations or harness innovations or what have you. But right away, if you’re thinking from a business perspective, you’re like, wait a minute, this could run other models. That’s super interesting. What is the implication for Nvidia’s or AMD’s general purpose GPUs, which kind of make the argument of like, hey, they can run all sorts of different shapes of models. It’s not set in stone and then when the ratio of this to that changes, oh no, the hardware is dead on arrival. And then even the implications for Google TPUs, which Google TPUs used to counter position against GPUs to say like, hey, we are more specific for, you know, matrix multiplication and data flow with a data flow architecture, but we aren’t as narrow as like it only works on a particular model. And so here is OpenAI coming in and sort of saying like, hey, we’re fairly generic too, but of course, still really focused on large language model inference.

Vik: Yeah, do you think just like Google TPUs, they weren’t selling these initially to the general market, they were doing it for themselves. Do you think, you know, OpenAI is going to start selling these chips? Clearly, it runs any models. So do you think they’re going to start selling it to everybody?

Austin: Yeah, that’s a good question. On the one hand, you could say, no, they shouldn’t. They should use their vertical integration as a means to compete against Anthropic, against Google, have better margins, move faster, control their costs, make a better user experience and so on. Of course, what gets complicated is if you’re like on the finance team, you’re like, okay, that’s great, except this is so expensive to, you know, build and run these things at scale. Wouldn’t it be nice to offset these costs by being able to rent them to others or sell them to others?

So, for example, I kind of think of the analogy of foundries and IDMs where it’s like, hey, for the longest time, of course, Intel wanted to use their foundry captive, they wanted to keep it to themselves and take all the advantages for their product team saying like, hey, we’re on the new node sooner than everyone else and our product team can benefit that and that’ll ultimately benefit our business. Why would we share this node with AMD or a competitor? But then of course, the cost of a fab just keeps doubling and doubling and at some point, you know, you have to like amortize those costs. And so, now, they made this chip with a small team and they did it quickly, but obviously still the cost to tape it out and to ramp it up were still, you know, talking $100 million or whatever. And and so that would be the question is like, is there a reason where they would want to amortize those costs and rent it or sell it to others, but do it in a way that’s maybe sort of doesn’t lose a competitive advantage.

So, what I think, the first thing that I think is like, what about enterprises? Like, don’t sell them to or rent them to Anthropic, but what about enterprises who want to run on prem or run cheaper? Could you partner, could you sell them to a neocloud, have the neocloud, you know, rent it to enterprises? And of course, it doesn’t have to be bare metal rentals. It could still just be like, I want to run OpenAI inference as a service. I just want the cheap version. This is what AWS is doing with Tranium. It’s like, hey, you know, we can give you, we can run our models, whatever models at 30% cheaper or whatever because it’s running on our XPUs and not on GPUs with the various middleman margins. And so, yeah, I’m just thinking out loud here, but I think there could be ways that OpenAI could pursue this line of business if they wanted to. Time will tell.

Vik: I think I’m going to go at it a little differently. I think that OpenAI’s benefit or unique advantage here would be to make a chip specific to their architecture. I’m not exactly sure why they are showing that this works with every other model, unless they want to go and sell this to other people. Like, otherwise, in the spirit of extreme co-design, you want to get, let’s say there is a way, and we’ll get to some numbers here, but let’s say there is a way to make this chip perform like a Cerebras class chip because you can essentially finally tune the hardware for the actual software you’re going to run on it and the actual workloads that you’re going to run on it. Because this is the information that Nvidia does not have. Nvidia does not sell these tokens or anything like that. Anthropic and OpenAI have these chips that are, you know, can be finally tuned to what their models are supposed to run, you know.

So, if you can make extreme performance just by co-designing with it, they should do it because they even have all kinds of statistics and data and things like that that nobody else has about how their tokens are being like, you know, what kind of workloads and how their models are being used. So, they should use that to their unique advantage. I mean, like, I don’t see why it’s a bad idea to design Jalapeño only for OpenAI models and finally get performance up to Cerebras levels.

The Regret Factor and a Roadmap

Austin: Okay, Vic, so to the question of should they design even more specifically for their own models or should they keep it general? Because to your point, they know their models and they know it’s coming down the pipe better and therefore they would have that advantage of we are vertically integrated, we can co-design with our model team and run out ahead and that could be our secret sauce is just knowing what the model team’s doing and designing to that. I thought I included this slide. There was a fairly related point made by Ravi, I believe, during the talk where he said, at design time, there’s two costs that we have to consider. Marginal costs, which is, hey, how much will it cost to include yet another feature? We’ve got a laundry list of features, how many transistors or how much floor plan area would it take to include that? And opportunity costs, which is if there’s user demand that you can’t supply because you didn’t have the feature, then too bad, you’re out of luck.

And he said, you know, the opportunity cost is generally much more than the marginal cost, which he said there’s like a regret factor there. So ultimately, they have a laundry list of features, they want to include them all, and you have to think about like, ah, which ones are we going to regret if we don’t do that. Now, of course, you take all these features and you try to fit them into your floor plan and then it’s like, oh, wow, we have way too many features and we actually there’s no way we’re going to make this fit in the chip. So you just have to like start cutting stuff and eventually you just need to tape out the chip. So that’s where the road map comes in. You say, ah, maybe we... maybe that you aren’t going to talk about, but that could take advantage.

Vik: Yeah, so this is good because I think it by keeping it relatively general, I mean, they don’t have to run Kimi on it, but having generality will help them develop their own future models on it. So I can kind of argue against my own argument on that front. The second thing that is that somebody in the audience actually asked, hey, what was the hardest decisions you had to make on this chip? You know, I like that question because it really opens up the door to a lot of answers. Of course, Ravi was very measured in answering and he said, the hardest decisions were the ones that we had to decide what to cut from the chip. Like we wanted to make all these cool things, but you can’t, you know. So how do we decide what to cut? So that’s the hardest decision.

Austin: Exactly. Yeah, which I thought that was an insightful question. And then I think Richard also pointed out with this next slide, like, hey, but guess what? We have a road map of chips. We’ve got I’m holding Jalapeño in my hand. Our Gen 2 is already approaching tape out and we’re thinking about Gen 3. So I’m sure it’s a little bit of kind of just scheduling and prioritization of like, we really want this feature, but these things are not going to make Gen 1, but how can we get them in Gen 2?

Vik: I like the approach that this is our first take, and they made this chip in a really short span of time. It is about nine months from first RTL to tape out, which is an incredibly short period of time for anybody who’s worked on chips would know that this is ridiculously short. Usually these things used to take two to three years at the very minimum, and it’s it involved a lot of manual work. It seems like AI has really helped them. And I was like, wait, is this a way of saying that you should use more AI because it kind of is a proof of concept that more AI is better and it helps OpenAI’s business by showing how the sausage is made, you know, to use AI to make chips and then chips to make AI, you know, it’s so I was like, yeah, it’s cool. I don’t know. It’s cool if people use AI to make a chip. Yeah.

Austin: Totally, totally. And, yeah, SemiAnalysis, I believe, made the point of like, hey, most people their first chip, it’s like almost throw away. It’s like proof that they can do it and they learn a lot through it, but where they want to compete is on chip two or three or four. So it’s very interesting that they came out swinging with chip one.

Why Aren’t We Using All the HBM?

Austin: Let me actually go back a couple slides though because, you know, okay, here’s a slide screenshot from SemiAnalysis. This is they’ve got Jalapeño, which is the purple line here that’s out on the Pareto has set a new Pareto frontier. It’s with deep seek R1, 671 billion parameters. So not the biggest model, but a fairly decent sized model. And you can see that, you know, at given ISO interactivity, it’s much higher throughput and can it can also reach much further out in interactivity than a GPU. Now, I think to be fair, some of the pushback was the Jalapeño will have HBM4 and a lot of these chips that it was comparing to, the Blackwell, is on HBM3E, I believe. And so the and this, you know, MI 355 as well. And so the fair comparison is going to be Jalapeño against Vera Rubin and Jalapeño against Helios. But I just wanted to make the point that this is their first chip and it’s obviously instantly competitive. And so I think that’s also what’s impressive. Not only was it fast, but it is competitive as well.

So that was that kind of medium model when they when they showed like, look, we can do models for, you know, from left to right on the spectrum. This was that point in the middle. They also showed when they run a small model, GPT OSS 120 billion parameters, probably what really stood out and has very interesting implications to sit and think about for a while is, hey, look, Jalapeño can reach out to above 1,000 tokens per second per user. And that’s the SRAM territory. And so, you know, GPUs just can’t get there. And so then that’s where GPUs said, okay, fine, we will you have Dynamo, we’ll split the workload to pre-fill and decode, and we’ll run the decode on these SRAM chips, Groq, or it could be Cerebras. And now of a sudden, Jalapeño’s coming in and saying, wait a minute, which this is the same argument, by the way, that AI ASIC startups are making, which is like, hey, actually, what if you can have one chip that can do both high throughput at a lower decode speed, but also reach those, you know, thousands of tokens per second speed.

Vik: This is compared to Blackwell, of course. So, if you put the Rubin system with the LPU, you know, basically the Groq chip, you can get this Pareto curve too. So it’s not like there isn’t hardware out there that can’t do this. It exists in different forms. But it’s nice to see that this chip in a small model can do over 1,000 tokens per second per user.

Austin: Yes, which again is interesting, you know, can it would be interesting to see how far Vera Rubin can get by itself on this because then you have to ask yourself like, oh, okay, well, you can get it with Groq, that’s nice, but remember how many racks of Groq it takes, you know, of LPUs. So then you have to ask is it is the cost worth it and probably the power consumption worth it to have a NVL 72 rack and nine Groq racks or whatever it is, or could you just have a rack or two of Jalapeños?

Vik: So Jalapeño’s benefit really seems to be the token throughput per watt. It is actually very energy efficient in generating this kind of performance. That’s one of the key takeaways, especially if you normalize it per watt. It’s a 700 watt TDP chip, which compared to the Blackwell, you know, the Blackwell GB200 is over 1,000, I think it’s 1,200 watts. So that’s a big difference in the power consumption. So it’s a very energy efficient chip as well.

Austin: Yes, totally. Which matters. And they another that was another point why they liked SemiAnalysis Inference X was it is power normalized. So it’s not like, oh, just use more power and you get more tokens or higher interactivity. Well, that’s kind of cheating. So it’s you know, power normalized.

Okay, one other thing. So back to this slide that had their road map. At the bottom of the slide, they called out, you know, and hey, thanks to our incredible partners, especially Broadcom and Celestica. So I thought it’d be worth talking through really quick, who are some companies that are helping and participating and if you’re thinking about like what happens if this goes really well for OpenAI and as literally one of the two biggest consumers of compute, if their own chips seem to fit their needs very well and they eventually spend more and more of their gigawatts or megawatts on their own chips, who might be some of the other supply chain vendors who will benefit? Obviously Broadcom and Celestica and then there’s some other ones as well. Do you want to talk through these quick?

Vik: Uh, yeah, so the CPU is an X86 CPU. People believe it’s a Turin class CPU, which is a which is a good CPU. I’m sure the Venice will do better if they want to do AMD. There are Intel alternatives. So there’s a lot of CPUs out there, but this is a Turin class CPU, which is a good one. I mean, it’s a good CPU for sure. It’s built on TSMC N3 node, and I think they have the N3P versus N3E, which are slight variations of the same process node, but for different power optimization and speed. And I think the biggest speculative thing is that this uses HBM from Samsung, HBM4 from Samsung, and seemingly has a slightly better speed per pin compared to, I don’t know, like SK Hynix, perhaps. So that’s maybe giving some more boost to the performance, but that’s entirely speculative.

Austin: Yes, and thanks for SemiAnalysis, you know, they talk about this. The it was SemiAnalysis speculated that Celestica was the partner here at the system level design. So, you know, I think you kind of think of them as like the OEM.

Vik: That’s a very interesting calculation essentially that the fact that, hey, yes, if you have the total aggregate bandwidth, if you take all the HBM4 chips across 128 chips, which I believe is what they goes into their rack, you get like one petabit per petabyte per second. And then at a four-bit, you know, FP4, if you have a one trillion parameter model, and you have, you know, let’s say 0.5 terabytes of data, you should be getting about 2,000 tokens per second. But even just from HBM. But like, we’re not getting 2,000 tokens per second from HBM because that’s the whole thing about why Cerebras is using SRAM and LPUs are using SRAM and all that. So that you can get it faster, you know, above 2,000. I mean, Cerebras actually in in a separate discussion we can probably have promises like, you know, 4,000 tokens per second because of SRAM, but the complexity of the system is enormous.

This is a good argument saying like, hey, why are we not even using HBM to the max? And you can see people like Nvidia are pushing the future per pin lane rate, per pin rates from, you know, 10 gigabit per second to like 16 gigabit per second. And I saw some charts in the memory talks at Hot Chips that HBM5 will reach more like 23, 24, you know, gigabit per second. But are you actually using all that bandwidth and does it show up in the token count? OpenAI says, no, it’s not. We are not using the whole capability of even HBM. So why don’t we do that? So, I like that approach very much.

Austin: Totally. Which, you know, has all sorts of interesting implications then when you’re thinking through memory companies, which is like, well, how important, like, exactly what you’re saying, which is like, okay, today, we’re not even fully utilizing HBM, but our solution is just like, make it go faster, make it go faster, make it go faster. But it’s like, whoa, whoa, do we need to make it go faster every single year, or should we slow down and focus on taking advantage of what we have before we go faster?

Vik: Yeah, and I like the KV cache locality idea as well, because it’s like, you know, KV cache is the biggest problem in decoding because you need to move all this data back and forth, and you have to do it token by token. So, they argue that this is not the right idea. This is long-term, don’t move KV cache around. If you don’t have to move it around, then it solves so many problems. Keep it local somehow. And they have another statement later that we will get to that, you know, dark silicon, like if you can turn off some stuff, it is still better than not using GPUs. Like dark silicon is better than unused GPUs, but we’ll get to that.

Austin: Yes, yes, we’ll get there. Okay, so as Chris was talking about like, hey, so why aren’t we using the full bandwidth of HBM? It one of the things he pointed out was like, well, guess what? When we’re trying to do some operation, some matmal, the data is not there when we need it. Like that’s a really big problem. And so he, you know, he called it the operands arrive late. So the data is not in the registers when we need. And why? And so they talked about like, oh, we’ve got these unified memory subsystems and you’ve got all this contention. So you’ve got all this HBM, and you’re sharing HBM with all your neighbors on the scale-up network, and there’s copies of things in various places and there’s all this contention. So yes, you can get the data off quickly from the HBM, but then getting it to the right place at the right time is a problem. And so his point was that utilization falls. So sure, we have all this HBM4 bandwidth, but the utilization falls if compute or memory is blocked just waiting on data in transit.

And so that’s where this local HBM slice thing comes in. And I’m going to show the the diagram of it, but maybe some quick quotes again. I probably won’t read these whole things, but he said, okay, so how do we stop this contention? What they came up with was, what if you have little local HBM slices for every accelerator, and they have dedicated buses that have high efficiency and low latency. So how can you use up all of those flops? And he then there’s like this one nice quotable sound bite where he said, because ultimately the show flops don’t matter. It’s how many flops you actually deliver. Meaning like, who cares if you have a bunch of transistors and you do you show on napkin math that you could achieve this many flops. If the data is not there in time, it doesn’t matter. All that matters is the user experience, which is kind of back to the argument I was trying to make of like, they’re really thinking about the user experience. Okay, how do we just make the user experience as good as possible? Well, we have to actually utilize our memory bandwidth. We have to actually utilize our flops. How they came up with it was this HBM slice with a local low latency view, which he he called a NUMA style architecture. So here’s the screenshot of that. I’ll let you jump in here and take the first stab at it.

Vik: Yeah, so instead of so NUMA, NUMA, for those who haven’t heard the term NUMA, it stands for non-uniform memory architecture, and it’s often used in like CPUs because when you have a multi, you know, core CPU, how does each core have memory shared along with it? So what you can do is you can, you know, have parts of the memory cordoned off for each core, very broadly speaking. I’m not going to get into too much. I’ve actually written about NUMA architectures in the CPU article that I have on the substack. So, the same way here, the idea is that instead of having the HBM as a resource that is contended for, and then that way there is no contention for HBM memory and everybody has their own lane they can keep operating on. So I like the idea. It’s a simple fast approach to minimize contention, right?

Austin: Yes, yes. And now, and they made the point on in the presentation, which was like, this we think is the best design. Every HBM slice or every accelerator has its own local HBM slice, and this starts to lead into the networking. And then therefore, we have like, they can talk very quickly to the HBM, then we have this next level of communication, this collective network, which is high bandwidth and low latency. And then to kind of talk further out, we’ve got this general knock that’s not as fast, but it gives us the flexibility to talk between them.

But the point that was raised was like, hey, okay, if this is how the computer designed is designed, yes, that could make things a little bit more complicated because now you have to ask, okay, how do we make sure that the right data is in the right place and how do we think about taking the workload and mapping it to this architecture? But ultimately, that complexity is was worth them solving because this results in a better user experience and better performance. They just had to solve the problem of like, okay, we need to think about this differently. What happens when you have all these NUMA style HBM slices? What should we put where ultimately?

Scaling Up with ESUN

Austin: So with that, you see on the slide, there’s mention of the scale up Ethernet bridge. And so I think I have the next slide. Yes, okay, so the next slide is they said, okay, that was the architecture or like thinking about, you know, each chip, but let’s think let’s talk about the system. And they called out a large scale up domain for the 128 Jalapeños can talk on this large scale up domain. There’s Broadcom, Tomahawk 6 switches that communicate at 600 gigabits per second per chip. And then there is a broader scale up domain that scales all the way to 2,048 Jalapeños and that’s at 200 gigabits per second interconnect. And this is communication using ESUN, which is the scale up networking sort of protocol or group that Broadcom ultimately spearheaded. And it’s different than UALink, which is a different one. But yeah, take it away on this slide.

Vik: Um, yeah, so this is a ESUN based scale up network. So they have 200 gigabit per lane connections here. And so, I so this picture itself doesn’t explain all that much to me, but yes, so it’s a it’s a bunch of chips connected with networking and Tomahawk switches. And people were trying to count how many switches and how many Tomahawk switches obviously, but yeah, it’s a networking setup. You need to network a bunch of these chips together.

Austin: Yeah, yeah. I think probably the maybe the interesting point is this is all scale up and there’s sort of like a two-tier scale up network here. So there’s probably like within the rack are those 128 are probably all within one rack and then the 2048, I think was 16 racks. So their scale up network is kind of two-tier and it actually can spread out across 16 racks.

Vik: Yeah.

Austin: And and with that, I guess I’ll point out a lot of times you’ll hear people say, oh, scale up is within the rack and scale out is like rack to rack and this that this continues to show that that’s not the case. That was a case for a point in time, but obviously really scale up is the accelerators that are sharing memory.

Vik: Yeah, they they say that this is what is called a half flattened two-level Clos topology. I was like, okay, I really need to think about this. Like, I because I don’t understand what that exactly means at this moment because it I’ve heard it for like, you know, less than a day ago, but I really want to think about what that means for networking. But that’s an interesting point there.

Dark Silicon is Cheaper Than Idle Accelerators

Austin: Indeed. Indeed. So, on the next slide, here’s this quote that you mentioned, the dark silicon is cheaper than idle accelerators. So I thought this this slide was really interesting. Again, zooming out, the OpenAI team was saying, hey, we’re not designing to one specific algorithm implementation as it stands today because in this chart in the top left, they they say like, depending on the type of workloads, I’m sure they’re running all sorts of different models, right? Like they have their big models and their medium models and small models. There’s a different ratio of pre-fill to drafting and speculative decoding to verifying and it’s always changing and we don’t want to be locked in forever on one particular ratio. One way that GPUs solve this is say, okay, have some GPUs that are dedicated to pre-fill, have some that are dedicated to the decode side of things.

And OpenAI said, why? So the only problem with that is then you literally could have like GPUs or racks of GPUs sitting idle, if you, you know, you’re focused on pre-fill, you’re focused on decode. Oh, they said, why not just have a single balanced chip where every chip has enough compute and enough memory bandwidth and enough IO that it they could handle different parts of the workload and then we can just gate and not power on the sections that aren’t needed. So yeah, did you want to say more on this topic here?

Vik: Yeah, so I wanted to explain what draft and verify means in a workload because everybody may not be aware of what that means or what speculative decoding is. So the idea of speculative decoding has like two two two parts of it, right? The drafting and the verification process. So what the drafting is is that, you know, instead of trying to just generate the right token every time, what this speculative decoding is is it tries to make a guess. So a smaller model often called a draft model will generate a bunch of tokens, like let’s say it generates eight tokens, all at once. And now we want to make sure that like one of those tokens is correct. Because it’s a small model, it’s not very smart. Let’s just call this like a not a very smart model. So it could make mistakes. So out of the eight tokens it has given you, maybe only one token is correct, right?

And the idea is now you need to verify those eight tokens and find out which is the right one. So the drafting and verification process, it it is really a dynamic thing. Like if you have a smarter draft model, you could probably only generate two tokens and then have it the verification process decide which of those two tokens is the correct one. So this is like a faster way to improve token rate because you don’t have to make the large model generate the right token. Instead, you just just eyeball it and then just decide which is the right one, right? But that changes, like if you want to use a smaller draft model, maybe the approach is just you spray a lot of tokens and then the verification workload takes over because now you have to find out which of these is the right token. Now you have eight of them to verify. So, none of these are like set in stone. So the draft, the speculative decoding has these knobs you can turn and each of these, depending on the workload, can generate different token throughputs. So, the the point here is that, yeah, there could be a different kinds of mix of draft and verify models that really require the chip to operate differently based on the need of the workload.

Austin: Yes, yes. And so then the argument is back to, hey, these one chip can handle different workloads, different parts of the workload. And actually there’s probably an argument there for the useful life of the chip. When I think about a Vera Rubin paired with a Groq LPU, you have to ask like what’s the useful life of that Groq LPU? It’s really focused on decode in a particular format and if that should ever really change, either it’s going to be potentially like less efficient on a new workload or you could say like, oh, is all that silicon just have a shorter useful life or it’s only fixed to run a much smaller set of workloads. Where here they’re trying to say like, no, the way we are designing it and thinking about it is it could run all sorts of different workloads, probably have a very long useful life, even beyond when it’s not when they have version two of the chip out.

What Does a 9-Month Design Cycle Mean?

Austin: Okay, so this was just the last slide that I had and it kind of had all their specs listed again and a flow.

Vik: Yeah, it’s a pretty cool chip. Overall, it’s a pretty cool chip. Then they made it in a really short amount of time. It’s by no means, you know, optimized or anything like that, but, you know, they their next version of the chip and the one that follows should all be very interesting. So, it’s a great start for sure, and it’s very exciting that, you know, they have all these TPU guys. So, I was like, how did they design a chip so quickly? I think it’s a combination of AI and having the right people on the team. That’s my takeaway from talking with people.

Austin: Yes. AI, having the right people on the team. I think Richard specifically said or somewhere that also just starting with a blank sheet is an opportunity where, you know, you’re they’re not thinking about like, how do we make legacy software fit into this or anything like that. But obviously they have the right team, they have AI. We haven’t even touched on AI. It probably could use its own podcast, but, a lot of the on their slides, a lot of things they did were very interesting on how they used AI to go a lot faster. If you haven’t watched the talk, definitely just watch it, but maybe if people are interested, you know, we can keep going into that because then again, I think there’s all sorts of questions about like, what did OpenAI do to help them write RTL, and is that apples and oranges versus what Cadence and Synopsys are doing, or is it apples to apples and just starting to think about the AI for EDA space and where OpenAI fits in there?

Vik: Definitely a lot of people are, you know, using EDA already for like chip design and AI enabled EDA on top of that. This is kind of proof, public proof that, hey, this really is worthwhile and it really works and it accelerated the design, pretty quickly and had a OpenAI develop a chip in nine months. So, that’s a good story to tell, but it also is a wake-up call to a lot of people building inference accelerators now because if somebody can come in and generate a, chip that’s better than Blackwell in in under a year, the question is, where does acceleration take us now? Like, how much better of a chip will a company with the right AI tools and the right people be able to design given two years? You know, and that will dictate like how money should flow, like who should be funded, who is capable of doing this, right? And what kind of hardware will emerge because of AI building AI, in a sense. And what does that mean for people like Nvidia? Like, is there really, you know, something to worry about for the hardware that, you know, because if in one year somebody can come and beat a Blackwell, I understand there are like HBM4, HBM3 discussions, but those are not not the broad view. The broad view is that some, you know, I think the SemiAnalysis article has some comparisons to Rubin as well. So, it’s it’s much more of a Rubin class chip than it is a Blackwell class chip because of HBM4. But regardless, you know, a company with the right people and the right tools could make a Rubin class chip in a year. That’s the takeaway.

Austin: Yes, yes. And actually, what’s really interesting, so there’s you called out, you’ve got to have the right team and they used AI. And so also if you compare to other startups, so there’s startups that had teams that, for example, came from Google, like I think of d-Matrix and Reiner Pope, and, they they had folks who came from Google and so they started with people and the know-how, but they didn’t necessarily have the AI tools right away. And so you can kind of compare, merchant silicon vendors, GPU vendors versus AI ASIC startups who had people and know-how, now versus have the know-how and have the AI tools and, you know, compare and contrast how quickly these folks move.

Vik: Actually, during the talk, one of the speakers mentioned that the initial RTL was developed with a GPT-3 class model. And of course, as they kept developing, their models got better and they could use newer models, but it’s not like any of these companies don’t already have a GPT-3 class model available. I’m not saying that it’s all because of AI that they could do this, but I think it’s a part of it. And it’s something to pay attention to. To me, it’s a, as I mentioned at the start of this episode, it’s kind of amazing to me that this could be done in under a year and you can generate Rubin class performance out of a chip.

Austin: Yes. That’s pretty cool. Yes. That is very cool. Obviously, you know, Nvidia can totally push back on lots of this to say like, okay, let’s see you manufacture it at scale, let’s see you ship it at scale, how reliable is it at scale, right? And so, so obviously the, the market leaders that have all the experience and all the production hours, are doing fine, but still very interesting implications. Like, honestly, one thought I had was, I wonder if Nvidia is already or sees this and will have a little skunkworks team go off to the side and say, white sheet of, design and do the same. Let’s stand up a chip in a year and maybe don’t worry so much about CUDA and legacy stuff, but like, hey, what if we internally built our own XPU? What design decisions would we make and how quickly could we make it? Yeah, yeah. That would be fun to see.

Yes. And then I also thought, yeah, totally. Oh, OpenAI hasn’t gone public yet. Maybe Nvidia could just buy them, but I’m just kidding.

Vik: Stop it. No more. No more $20 billion, okay? Yeah.

Austin: Okay, okay. We’ll end here.

Vik: There’s so many stuff to buy. I mean, yeah, there’s there’s so many like what’s AMD’s path now? What’s forget about AMD. What about Anthropic? What’s Anthropic going to do, right? I mean, where’s where’s Anthropic’s chip?

Austin: Yes. And and Anthropic, correct. They have a team, and so, you know, Anthropic, we’d love to have you on to talk more about it, but of course, OpenAI, we want to have you on because you just shipped at Hot Chips and you have, a chip in hand, of course. So, there’s so much more to talk about. I think even the the chip and system itself, we didn’t even touch on everything, but obviously the AI angle is just kind of mind-blowing and to your point, Vic, just, this is ultimately good for EDA in that it just goes to show that small teams can use AI and use their know-how and bring competitive chips to market quickly. So, you know, there’s going to you could argue that the barrier to entry may be kind of lowers in some sense and therefore more EDAs is going to be used. We would love to talk to EDA companies as well to get their two cents. A lot of the 2027 is going to be very interesting.

Vik: Totally.

Austin: So with that, we’ll cut it here. Thanks everyone for listening to Semi Doped. Thank you for sharing us on X. Thanks for all of your comments on YouTube. Um, share it, tell a friend, send Go ahead.

Vik: Oh, no, I wanted to add like, I met so many listeners on, you know, in person actually at Hot Chips and it was like amazing. Like, I was so humbled to hear like so many of you enjoy the podcast. So, so everybody who like actually spoke to me in person, thank you.

Austin: I love it. And with that, you know, send us your feedback and we’ll talk to you next time.

Ready for more?