Astera Labs, a company few knew before its March 2024 IPO, is now a Nasdaq 100 component with staggering growth. Austin and Vik explore how Astera’s core product, the retimer, became an essential part of every AI server by solving the fundamental problems of high-speed signaling over copper. They also discuss how the company is leveraging this success to challenge incumbents in the switch market.
Things we cover:
The physics of high-speed signaling over copper
Retimers vs. redrivers
How Astera Labs won the Nvidia H100 socket
Astera’s expansion into switches (Scorpio) and CXL (Leo)
The battle between UALink and Ethernet
The role of signal integrity and eye diagrams
This podcast is lightly edited for clarity.
Introduction: The Copper Problem Inside the Server
Austin: In March 2024, a company almost nobody outside of the industry knew IPOed at $36 a share. This June, it hit $499, joined the Nasdaq 100 and printed 93% year-over-year growth at 76% gross margins. Astera Labs doesn’t even make GPUs. They don’t make memory, they don’t make optics. Astera makes the chips that fix copper, the little devices that catch a dying high-speed signal halfway down the board and relaunch it clean.
Unglamorous, invisible, and in every single AI server Nvidia’s partnerships. Today’s episode is the story of how a repair part became a franchise and how that franchise is now going after the biggest socket in the AI rack. Hello everyone, welcome to Semi Doped. I’m Austin Lyons of Chipstrat and with me is Vik Sekar of Vik’s Newsletter. So Vik, let’s talk Astera Labs, but let’s talk retimers. When I first heard about retimers, I was just like, what does that even do? What does that mean? So what do you say? Should we get into that today?
Vik: Yeah, and let’s talk about the problems of copper connectivity, which we have spoken about a lot before. But I think retimers is an interesting concept because we always talk about stuff that is outside, let’s say a server tray. We talk about how GPUs are connected together in scale up. We talk about how to connect racks together in scale out and scale across, all of that stuff. But what we’ve never covered is how does a GPU connect to a CPU within a server tray or how does a GPU connect to the network interface card or the NIC? We’ll keep referring to this as the NIC, NIC, network interface card. So how does that connection happen? It may not seem like it’s much of a connectivity problem because come on, how far are these things, right? The server tray is maybe just a big with about two feet wide by four feet deep. But that is still a problem because the way you connect a GPU to a CPU or to a NIC is through the PCB that is inside the server rack. It’s a host board that is holding all these chips. And so those are not like excellent conducting mediums to carry high-speed signals. And so that is the problem we’re going to talk about today and how Astera Labs solved that problem and became pretty much an integral part of what is inside the server rack, server tray today.
Austin: Awesome, interesting. Let’s get into it. Yes, because already you got my wheel spinning because you’re talking about copper PCB traces and high-speed signals. And we’ve been talking about with Credo and active electrical cables, the copper problem. So yeah, let’s get into us. When a GPU talks to a NIC and it talks over this PCB trace, what are we talking about?
The Physics of Signal Degradation
Vik: Okay, so let’s go back from the beginning and set the stage quickly, but also in a complete fashion so everybody understands what’s happening in this part of the AI stack. Fundamentally, what happens outside, we always talk about NVLink for scale up or Ethernet or InfiniBand, all these things. These are all outside the server tray. What happens inside is something that has been around very long. This is in every PC that you build out, it’s been in gaming PCs forever. It’s called PCI Express or PCIe. So PCIe is how the GPU communicates with the CPU or the NIC, for example. It’s just a protocol that’s been around for a very long time. So the whole idea is that these don’t have to run very fast on a single lane basis. However, they have many parallel lanes. So a PCIe interface can have 16 lanes and so it’s called a PCIe X16.
Now, the way this is dealt with or spoken about is that you have so many giga transfers per second. So you can have let’s say a Gen 5 PCIe has essentially, let’s say 32 giga transfers per second. This was actually used in the Hopper era anyway for connectivity. So what that means is that why are we saying transfers and not bits? Transfer, we are talking about transfers because we are talking about transferring a symbol. A symbol can be a collection of bits. It can be one bit is one symbol, which is basically this NRZ modulation, zero or one, very simple. Or it could be a PAM4 modulation, which is like four levels. So you could have different kinds, so when you do PAM4, you have two bits to a symbol. So giga transfers means you’re transferring 32 billion symbols of either one or two bits every second. So if you look at that in the time domain, let’s see how many symbols every second? So every 31.25 picoseconds, you have a symbol going by. So you think of a string with a lot of beads on it, and each bead is a symbol and the beads are moving from left to right. That’s how it looks if you kind of imagine bits moving on a wire. That’s exactly what’s happening between the GPU and the CPU.
Austin: Did you say picoseconds, like 10 to the minus 12?
Vik: Yes, yes, it’s very less, it’s a very small amount of time and that’s what it takes to transfer a billion symbols per second, it takes picoseconds, it’s very, very small period of time. And so the thing is that it’s not easy to do this. It’s not easy to move bits across from one place to another because it faces several problems. Copper is a problematic medium and when you try to push these bits from one place to another, things go bad. So one thing is loss, you put a signal of high strength in one end and it loses all its strength by the time it reaches the other end of the wire. That’s one problem. Then the PCB itself has problems and all that. And then these things are like high-speed signals, so you can think of them like sometimes they ping-pong, so they reflect from that end. So you can think of the beads trying to come back on the wire, you don’t want that. Those are reflections. Then crosstalk, if you put two sets of beads next to each other, imagine if one bead on this string is affecting the bead on the next string, that’s a terrible outcome. But these things happen in copper. So all of these things, they basically need correction. They need some way to fix this.
But that’s just the copper thing. Actually, I should mention that there are two more problems we should talk about. And that is inter-symbol interference is one of them, which means the beads on a string don’t actually stay like one bead on the string. Imagine that they’re made of sugar or something and they try to stick to each other and try to become one larger ball of sugar instead of staying distinct sugars. As they travel really fast, you can imagine they heat up or something and then try to stick and become one and a blob. That actually happens all the time. So when one symbol travels, the next symbol kind of starts to overlap with it and it starts to spread out and that’s terrible. You don’t want inter-symbol interference.
Austin: Yes, yes, let me jump in there for listeners. You can think of if you’re familiar with square waves or sine waves, if you have a square wave, you can imagine that as it travels down this medium, it starts to spread out instead of being a square, maybe it’s more of a hill. And so you can imagine if you’re sending these different pulses in the simplest sense that as they start to spread out, they could start to overlap each other to Vik’s point of the sugar sticking together. So now all of a sudden instead of a steep up to one, steep down to zero, some space and then another steep up to one, you start to get a ramp up to one and a ramp down and then maybe it overlaps another one to the point where what you receive is like a 0.5 and it’s slowly transitioning and it’s a little bit confusing is this a one or is it a zero? And then of course, to your point, as you mentioned earlier, there’s this attenuation. So by the time it gets all the way down the copper, it’s just lost the peak amplitude as well. But back to the inter-symbol sort of overlap here, yeah, there’s this smearing where you tried to send certain symbols and they’re kind of overlapping each other and somehow we have to figure out what was said at the end. Yep, carry back on. What’s the next one that you were going to mention?
Vik: The other one is jitter, which means that the pulse is not where you think it is. The pulse you think you sent it some time ago, but by the time it reaches a given location, it isn’t there. There are random variations in the system that cause the pulse to be further than you think it is or earlier than you think it is. So when you go in be like, “Okay, at this moment, I’m going to look and see if this is a one or a zero and I’m going to lock it down.” That’s what you’re saying, let’s say that the 0.5 question isn’t a thing, we are constantly monitoring this particular point on the wire and we’re going to be like, “Okay, when the thing comes, I’m going to call it a one or I’m going to call it a zero.” But the thing doesn’t come at all or it comes late and you’re like, “What happened?” So now you can’t detect the bit because the thing wasn’t there when you expected it to be there. That thing is called jitter and that also causes problems in communication. All of these are big issues.
Austin: Gotcha. Okay, yeah, so you’re saying with jitter, it’s like if I’m thinking about, I’m on a clock, I’m at the end of the wire, I’m checking and it’s like, what am I reading? What am I reading? What am I reading? And it’s like, we can figure out, oh, this is copper and we know how a high-frequency symbol behaves and how much of that is just like random noise or random something, randomness.
Vik: I think that inter-symbol interference is somewhat deterministic and you can kind of figure it out. And there is a way to do this and I’ll get to that. Jitter is more random. You really can’t do much about it. And all this affects how we will ultimately detect if it’s a zero or a one.
See, the whole problem here is to tell whether at any particular instant you have a zero or a one. If it’s a two-level modulation, if it’s a four-level modulation like PAM4, you have to tell which of the four levels it is. And typically they talk about it in PAM4 as 0, 1/3, 2/3 and one. Instead of trying to use something else, they just break up the one into four levels. So it’s up to you to tell what the level is. That’s the whole problem.
Austin: Ah, yes. Okay. Yeah, I could see how that’s challenging, especially for PAM4. Like if you’re trying to read 0 volts is 00, a third volt is 01, 2/3 volts is 10, and a volt is 11 or something like that. How now all of a sudden when there’s that smearing and stuff, it’s not like, oh, is it a zero or one and we’re at 0.5 and we don’t know. Now it’s really crazy. It’s like, oh, it’s supposed to be 0 or 0.33, but I got 0.28 volts or 0.16 volts or, it’s tighter. There’s less margin.
Vik: Yeah, yeah. There is a nice way to visualize this and you will see this mentioned a lot in this interconnect world. And those are what are called eye diagrams. And I’m just going to explain quickly what that is. So imagine at this one particular time, called a unit interval. That is basically the 31 picoseconds that I mentioned. Because you know that that is the symbol period. That’s the period. So that thing is called a unit interval and is usually on these charts you’ll see it as UI. And so you’re looking in that unit interval and you’re seeing all the bits that go by. You’re looking at just one bead on the string, but that bead is moving, right? You’re continuously looking at the string.
Think of it that way. And what you’re looking at is an overlapping bunch of zeros and ones, for example. And always you want to have some kind of a separation between zero and one, so that you can clearly say anything above 0.75 is a one or anything below 0.25 is a zero. We just round up that round down that way. And so, but nothing should be there between 0.75 and 0.25 because you don’t know where it goes. The uncertainty is too much. So that gap is usually referred to as the eye opening. The eye of this signal must be open because if the eye closes, which means there is no space between the voltage levels, you can’t tell what it is. So the eye opening determines everything.
Austin: Okay, okay. So you basically take a bunch of snapshots at this unit interval and then plot them all on top of each other, right? So you take a thousand of them. And in the ideal sense, you would have a rectangle. You have a bunch of ones and a bunch of zeros and that’s it. But of course, it takes time to get up to one and time to get back down. And so that’s why you kind of end up getting this eye shape because it ramps up and ramps back down and same and but your point is, when you throw thousands of these measurements on top of each other, there should be some sort of gap where it’s very clear, hey, there’s nothing in here, so it’s clearly a zero or one. But if they start to, if that gap closes, that’s bad news because, hey, some of these measurements that we took are in the middle and we won’t know what to call it.
How to Fix a Broken Signal: Retimers vs. Redrivers
Vik: Exactly. So now, the whole problem of why Astera Labs solved this problem is that they needed to find a way to keep this eye open when the communication is going between a GPU and a NIC or a GPU and a CPU. And the distance is in the range of about 30 centimeters. It’s not that much, but it’s still required. Because at the speeds that we’re talking about, because Gen 5 is 32 giga transfers per second, Gen 6 is 64 giga transfers per second and it goes to PAM4. So that makes it harder because you’re adding more levels and you’re going faster. So it really requires retiming and then in the future, we will go to Gen 7, which is at 128 giga transfers per second with PAM4. So Gen 6 is already in, I think in the Blackwell era of chips, it’s being used. So, yeah, we’ll get to all this. But so Astera Labs is basically their job is to keep the eye open within a compute tray. That’s their whole purpose with retimers.
So what are retimers? Let’s get to that portion because we need to talk about, okay, now we have inter-symbol interference. And we have jitter, whose main effect is to close the eye. Now we have to find a way to fix it. And there are a few tools to do this. I think we should first introduce the idea of equalization. It’s a fancy way of saying, look, I’m going to lose half the signal down the line. So what if I just boost the signal twice now, so that by the time it reaches that point, it’ll still be what it was here? Because I amplified it beforehand and so that when it degrades later, I can keep the voltage level constant. So that is equalization. You’re basically applying the inverse of what will happen to the actual signal.
Austin: So, okay, let me interject really quick. So, back to the eye diagram, so you’re saying the problem is the eye starts to collapse, especially if we just don’t do anything, but we’re increasing the speeds. Of course, it’s copper and so the distance at which you can effectively keep the eye open is shrinking. And so the question is, how do we keep the eye open? And you’re saying one of the simplest ways is sort of, hey, if we know that it’s going to attenuate or degrade as it goes through, can we just amplify it? And it’s kind of like, if you and I were in a hallway and I was talking to you and then you went a lot farther away, I would just shout so that by the time it got to you, it would sound normal.
Vik: Yeah, exactly. That’s a good analogy because you can kind of see how far away that person is and there’s no point talking softly because you know they’re not going to hear it. So you will adjust your volume to be good enough so that you have a good chance of them hearing it at the other end. Yeah, exactly. It’s a good analogy. I like it. That’s on the transmitter side though. That’s what you can do by shouting louder.
But what can the receiver do? What can the person on the other end do? That’s what you will often see as CTLE is. It’s called continuous time linear equalizer. Okay, so this is a long acronym, but in concept, it’s relatively simple. It is just that you want to amplify the high voltages more compared to the low voltages. So that you can kind of stretch the eye open. Imagine that you have, remember I told you that, 0.75 or above is a one. Now, what if you can just amplify the 0.75 to make it 0.9, so that the eye is even more open? And what if you can make the 0.25, 0.1? So now your eye is much nicer and much wider open. So you can amplify the highs while keeping the lows low.
Okay. So that’s one way the receiver can deal with this problem of all these channel imperfections that happen. This is still analog stuff. There’s no digital work going on. This is just a simple amplifier. But the problem with this is that it amplifies all the high-frequency noise stuff. If there’s any noise and all that, everything gets amplified. Because it’s just a dumb amplifier. It amplifies high signals and keeping the low signals low. That’s it. That’s not a good thing. And if there is any interference from the neighboring bits moving up and down, it’ll amplify that as well. So it’s not entirely very clever, but it’s a necessary part of this system.
Austin: Yeah, it reminds me sort of of naively, if I’m the receiver, if you’re at the other end and we got far away and you shouted, and then I can kind of hear you, but instead I’d put headphones on and I just hold a microphone and then the microphone sort of does that. It hears you and I turn up the gain or something and it’s like, oh, okay. What you said at one volt was down at 0.75 by the time, but then this microphone just amplified it back up to one volt and I can hear it fine.
Vik: This is called a decision feedback equalizer. So it just knows what’s going to happen in intersymbol interference and it cancels it out.
Austin: Gotcha. Nice. That is pretty clever. So does it look at like it depends on what I just received, like, oh, I know that I just received 1101 and therefore I can calculate sort of in real time what kind of noise or smearing or interference I would get from that particular pattern based on this particular channel.
Vik: Mhm. Yeah. So it has some kind of a characterization and then it reconstructs it based on how much it knows the spreading or the smearing has happened. Very clever stuff, very complicated stuff. And it is because of this that even Gen 5 PCIe works. It’s a very essential component. These DFEs are very essential components.
Austin: Does that really quick, does that have to get tuned in manufacturing like per board or something?
Vik: I think it does. Yeah. I think it does.
Austin: Interesting.
Vik: So it’s not all that straightforward to do, but it is a system that’s been around a while, so it’s not like ultra new. But these are all very clever chip design techniques. I have some friends who actually do this for a living and they’re really good at it. They’re these analog guys, they’re awesome. And when you go to these conferences like ISSCC and all that, you’ll see all these designs of how they do this. It’s quite complicated. Let’s just say. We’re just trying to cover it at an extremely superficial level, but the reality is very, very hard.
Finally, I think it also extracts what is called a clock and data recovery. So what that means is whenever, let’s say, I’ll be like, Austin, whenever you see a symbol go by, just clap. Okay, just clap for me. And I’m going to sit here and kind of write down all the times that you clapped. And I’ll be like, okay, I know the period of the symbol. I know the clock now. So now I can set like a metronome. You know what a metronome is? It’s that music thing that goes click, click, click, click. So I can set the speed on my metronome and that should work out well because now I know when the bits are coming just by hearing the sound. So I have recovered the clock this way.
Austin: Yeah. That’s pretty clever. And that kind of helps with that problem that we talked about where it’s like if I’m on a particular clock and stuff isn’t flying by right when I expect, I’m like, man, this isn’t lining up. But you’re saying, no, no, do the opposite. Just watch what’s happening and then figure out the clock from that.
Vik: Yeah, figure out the clock. So this is all these are all the techniques that people used to correct it. And pretty much this is all there is to it. So you can do the just to quickly summarize, you can do the equalization at the transmitter thing. Remember you said like just shout louder, or you can do at the receiver side, you can do the equalizer on the receiver side, which is like just put a little amplifier or some microphone and turn it up. Or you can do the decision feedback equalizer, which is you know what the smearing is happening. I don’t think we have a very good analogy for this, but you can reconstruct based on the intersymbol interference. And finally, you have to recover the clock too by basically clapping out and fitting a metronome to it or something, and you figure out the clock.
With all this information, you can kind of tell to a high degree of accuracy what the received signal is going to be doing. And you can correct for errors with these methods. And that is essentially with all these build up a full PCIe channel in the transmitter and receiver, and you with this chain, you can build what is called a retimer. So and I wanted to also mention the clock and data clock recovery process because it retimes the signal. And I’ll tell you what that does in just a minute. So now I’m going to just explain what is a redriver and what is a retimer because this is a very there is a difference here. So redriver has it does everything that I just said, but it does not do the clock and data recovery part. And it also does not do the decision feedback equalizing because I think it needs some amount of sophistication to do that. So imagine you’re just shouting louder and you’re using an amplifier on the other end or a microphone and picking up the signals better. That is called a redriver. It’s just an analog megaphone. That’s it.
Austin: Yeah, yeah, you’re just driving the signal on one under the other.
Vik: Yeah. So what are the benefits of this? Yeah, it uses lesser power and energy. It’s less complicated. A redriver is simpler and things are good. That’s it, and you don’t have much latency, you can do this quickly, low energy, all that stuff. But the downside is that when you go to faster and faster signals, you got have more and more crazy things happening to the signal, you can’t recover it. So now what do you do? You have to go for the whole enchilada and go for full retiming where you get the decision feedback equalizing to remove intersymbol interference and you recover the clock. So it’s not just that you amplified it, but it’s the retiming process basically accumulates all these bits and sorts them nicely into the windows they are supposed to be in. Because you remember you have the metronome clock, so you know where the bit should lie because you’re the clock that you recovered tells you where they should lie. So you can take these bits that arrived late or that got messed up somehow and slot them into the correct time slots. And now you have a beautifully recovered signal because you took the bits you received and you put it on where the clock should be. So you retimed the received signal based on your recovered clock.
Austin: Yes. Okay, okay. So to sort of simplify it, like back in the analogy of when you and I spread out, the further out we go, I shout louder, you have a little microphone and headphones, but eventually it just doesn’t work. And so because whatever, in this analogy, we’re just too far apart. And so this essentially with the retimer, are we ultimately just putting someone in between us and I’m shouting and they’re listening to what it what I was trying to say and then they’re sort of cleanly repeating what I was saying down to you?
Vik: I think it’s a little bit different. It that sounds like it’s more like a repeater to me, a repeater amplifier to me. This is not really that. This is like somebody who knows what is going to be said on the other end to some extent. Like I’m going to talk to you about retimers and you’re an expert on retimers. And so even if you hear it faintly, you kind of have an idea of what it is that is being said. When you hear a garbled word, you’re going to be like, oh, he’s talking about retimers. So he probably meant clock and data recovery because a general person who’s not familiar with the domain would not understand that word. But because you heard it halfway and you know the context, you can kind of reconstruct it better than the normal person would.
Austin: Nice. Nice. Very good. Okay, okay. I’m tracking.
Astera’s Big Break: The Nvidia Socket
Vik: Okay. I hope everyone else is too. But yeah, so this is very important. So these retimers, now that we got to the point of retiming, these retimers are what make sure that bits from a GPU make it to the NIC so that you can then go do scale out. Or the bits from the GPU make it to the CPU and then you can do all this agentic fancy stuff. The bits have to make it from point A to point B, otherwise you’re not going to have any communication. So this was this is the problem that essentially Astera Labs solved for Nvidia. And in their HGX baseboard, I think the H100, which is just it had like eight GPUs in it. They needed to make sure that things can communicate with each other on the board, on the board. So that’s what this this retimer chip did. And it really got off Astera Labs to a running start for sure because you need about a
Austin: Yeah, yeah. Yes, yes. Okay, let’s get into Astera. So what I’m hearing you say is that when Nvidia exploded and things took off with the Hopper era, for every single GPU that was ever sold, there needed to be a retimer so that the GPU could talk to the NIC, for example. And my question is, how did a startup own that socket? Where did Astera Labs come from and how did they get into presumably the H100 reference design so that they would get designed in? Like that’s amazing. That feels hard to do, but it’s incredible.
Vik: That’s a good question, but these guys I was looking at the company earlier. All the founders are actually ex-Texas Instruments guys. And they were basically high-speed interface folks. And their ultimate bet was that in 2017, they left TI, I believe. And their bet was that that PCIe 5 would be required on every server board. That bet was turned out to be right.
Austin: Gotcha. So they saw they saw this coming. They said, oh, every server board, even presumably CPUs, would need PCIe Gen 5 and at that speed, we think copper is going to fundamentally have a problem and that these retimers will fundamentally be required.
Vik: Yeah, they they they foresaw that coming and that as soon as servers go faster, that they’re going to have to use these chips. And so they they made this whole company. And then it was it was great when the H100 and the ChatGPT revolution happened. It was great for Labs. But the question was when I was researching this, I was like, why not other companies? PCIe has been around a long time, right? But what’s the what’s the big deal? Like why Labs, why not Broadcom? Why not even TI? I think one of the smart things they did was they have their I think they were first to PCIe 5 market. So that is one thing. And the other thing I believe is that they have this monitoring platform they built pretty early on called Cosmos, which not only tells you when the link is going down or something is going bad, but it you know, all this channel impairments, all this data that we all this technical stuff we spoke about. They could capture everything on that platform and hyperscalers love that stuff too. So I think their software platform had something to do with it as well. And then when they get qualified, it is a major win because now you’re locked in. It’s a very sticky slot. Nobody wants to change this stuff if it works, right?
Austin: Totally, totally. Interesting. So you’re saying when they were seeing that this socket was needed and retimers would be needed, presumably others saw it as well. People who work on the Gen 5 spec, for example, would have known, well, on the dance, hey, we’re trying to qualify 32 gigatransfers per second. This is going to be a problem. But Texas Instruments has a huge catalog, for example, and so do other companies. And so they might have just thought just about the retimer as a component in their catalog. But what you’re saying is maybe Astera thought, no, it’s more than just a retimer. It’s also a sensor that can give you health monitoring information about your entire fleet. So we’ll capture telemetry, we’ll build software around it, give you a user interface, and it’s not just a little black box component on your baseboard, but it’s actually something that can give you all sorts of insights. So that, for example, if you imagined that the world of tens of thousands, hundreds of thousands, millions of GPUs in a training cluster, imagine if they can all be telling you about the health of their link, for example. I can definitely see how obviously that’s supremely valuable for the end customer, but also that it might take a company oriented around that mindset of this is more than just a little component on a board, but obviously it’s very customer-centric. And then to your point, of course, sometimes when there’s I should say a good time for someone to try to take some market share is to obviously be first to that next generation thing, or create a market. And of course, if they were there first and they got designed in, then all of a sudden now it’s their socket to lose. Even though they’re a startup. And again, this is the beauty of fabless companies, which is, your fab partner can manufacture at scale. So if you can get designed in, you can figure out how to build these things at scale and now it’s sort of your game to lose.
Vik: Yeah, this is the benefit of startups, right? They can pick a problem and just go after it. Bigger companies are not as nimble and that’s one of the nice things about this industry that somebody can come in and disrupt something amazingly different. And which is why we would love to also talk to startups on the podcast and stuff. This is a very interesting industry because Broadcom did come in later as a second source. They did come up with a similar retimer chip doing Gen 5 timers in 2024, but by the time when you’re designed in, you’re designed in. It’s hard to move because the Nvidia will have to requalify everything and they have to make sure that the channel link budgets are correct. It’s a big problem. Once you’re designed in and it works, you’re done.
So the question is, did keep the slot when Hopper went to Blackwell? Yeah, they did actually. They did keep it in the B200 boards as well. And I believe it was the Aries 6, which is the Gen 6 part of the same family, which was doing 64 giga transfers with PAM4. So they had a fancier chip, of course, this chip has more ASP, good for money. Yeah, yeah. So all through, the money every time they sell these chips, the revenue keeps going up and up and up. So you can see all these charts probably elsewhere that people are charting all these revenue tramps for Labs. So it’s all public information. But yeah, but one interesting thing that happened in the Blackwell era was people got a little spooked because Nvidia said like, okay, this is not happening, guys, the distance between the NIC and the GPU is too far. So we’re going to move it close by. And now everybody freaked out because what does that mean for Labs? Now, if you put them close up close together, why do you need like Gen 6, I don’t know, some fancy retimer chip or maybe you don’t even need retiming. If it’s close enough, maybe you just go with a redriver, cheaper chip, lower power, all well and good. So that was a bit of a scare. But as reality would have it, it turned out differently. Do you want to know how?
Austin: Really? Yeah, yeah.
Vik: Yeah. So what happened was, it turns out that not everybody deployed racks this way. The Grace Blackwell platform had variations. Like there were a lot of customized deployments going on. For whatever reason, I don’t fully understand why there were customizations, but everybody didn’t use Nvidia’s so-called reference design where things were placed in a certain way. There were customizations. As soon as customizations started to happen, it wasn’t it the reach went up again. And so the retimers were required. And then the second thing happened, which is the birth of like custom accelerators, like XPU, the Tranium thing, all of that came in. And once all that came in, you’re no longer tied only to Nvidia. Like the ASIC world is wide open for you. You can go and do put retimer chips in all of those accelerators too.
Austin: Yeah, absolutely. Totally. There’s a new tamp for you. And I could see, once people got used to having the retimers and the telemetry information, it might also be a little scary to take that information away. Even if Nvidia said, oh, if you put everything close together in our NVL72 configuration, our reference design, you won’t need it. It’ll be great. But I could see being like, nah, I think we do want to pay to have that telemetry and to make sure those signals are clean.
Vik: Yeah, yeah, yeah. Exactly. So I have some numbers actually. So in 2023, their revenue was about 115 million.
Life Beyond Retimers: The Scorpio Switch
Austin: That’s amazing. So, of course, one must ask, is it just retimers or is there life beyond retimers?
Vik: There is life beyond retimers because this is where their next, big, big move is really, the big socket that you mentioned in the cold open basically is that they’re like, why stay within, why stay within the tray? We know how to condition signals. Why don’t we go up and make a full switch? We’ll make a switch. And this is competing with basically the NV switch or it’s competing with the Broadcom Tomahawk switches. So, they have their Scorpio, I believe P and the Scorpio X switches. The P is basically for scale out. So, this is going to be pretty much a Tomahawk replacement switch. And they have Scorpio P series switch with up to like 320 lanes. It’s a very high radix switch. And then the X series that they have is basically for scale up. It’s like the NV switch replacement that anybody else could do with like merchant silicon. So, NV switch is basically NVLink Nvidia’s switch, right? But if you wanted a similar performing thing, you could do a Scorpio X. So, there’s that.
Austin: Yeah, totally. So, you’re saying Astera said, hey, what are our core competencies? It’s networking, but it’s signal conditioning and we’re doing that really well with retimers, but ultimately, let’s move into switches. Is that because conceptually a switch is like signal conditioning of every lane plus routing or something. So, like is it pretty conceptually similar?
Vik: Yes, the signal conditioning part is, but then the actual switching process in and packets and stuff isn’t very simple that straightforward. Because so far we’re just talking about signal conditioning, but how the packet is actually switched within a Scorpio switch isn’t that straightforward. It’s like it literally like how to build switches. And this complexity because it’s not incremental is exactly what makes this a very high ASP part. Like they expect a lot of revenue from the Scorpio switches because I believe the Tranium 3 uses this Scorpio X as well in scale up. So, yeah, it’s probably going to be quite a bit of revenue from the Scorpio parts going forward.
Austin: Speaking of Tranium, does Amazon have warrants in Astera Labs? I feel like a lot of times, I’ll have to look it up. I think they do. I think a lot of times when a company like Amazon works with a very early stage company like Astera Labs, they also to make sure that all their incentives are aligned, they end up getting some warrants or some ownership of the company. And that obviously can help Astera get the opportunity to build out their port silicon portfolio and have the get the right to play in future Tranium. And then of course, it may be incentivizes Amazon to put Astera Silicon in their silicon when they can instead of, for example, Broadcom.
The Interconnect Standards Battle
Vik: Yeah, yeah, yeah. Oh, I mean, that was that there’s another big big story that which I’ll briefly mention because it’s really interesting. Because, so AMD actually wanted to put as because they are all about UALink, right? Like the open standard of hooking stuff up. AMD actually wanted to put a UALink switch, but there was no UALink switch in the market. So, what Broadcom did was they were like, okay, this is the only thing that was available was a Broadcom Tomahawk switch outside of like Nvidia’s NV switch or whatever. So, there’s no option. AMD had to go with Broadcom. And so what Broadcom did, this is what this is ruthless business, right? What Broadcom did was they immediately exited the UALink consortium and they’re like, Ethernet is it. They had like scale up Ethernet SUE and then they merged it with like OCP and called it Eson. And they were like, that’s it, we’re not doing UALink. We’re in. If you are designed in with us, we’re going to go Ethernet all the way. It’s basically puts the nail in the coffin for all these other people who are trying to do UALink. And they’re like, no, we have the only switch in the market and we’re not doing UALink. Screw that. We’re going Ethernet. That’s it. No more. And remember the Tranium thing with the Scorpio switch right now is a PCIe switch. It’s not Ethernet. Right? PCIe could be used for all this as well, by the way. And Amazon Tranium is running on PCIe. But the future of Scorpio switch is actually has UALink in there. And it’s probably a 2027, I think story. Along with Marvell who’s also building something like this for UALink. Now, the risk here is like, will AMD and the others go away from using Ethernet, Eson and go jump on UALink and buy all these people’s switches or are they going to say, no, forget it. Ethernet is it. We have been designed in. Remember Astera Labs’ own story, right? When they got the Aries slot on the Hopper series, Broadcom couldn’t nudge them out. Now what has happened, Broadcom took the AMD slot because there was nothing else in the market. And now they’ve cornered the Ethernet way and exited the UALink consortium. And now can they get back in? That’s very interesting to see.
Austin: Yeah, well, we yeah, we should do a deep dive on this because I know that with the Helios platform, AMD is doing UALOE. UAL, UALink over Ethernet. And so they’re using the UALink protocol, but they’re doing it over an Ethernet switch. And so I do think they’re planning to take a step in that next gen of Helios to moving to UALink. But to your point, Astera Labs in the market didn’t have the UALink switch ready in time. So this is like an intermediate step. So I think that, you know, Helios 500 and beyond will be interesting to help tell us what direction things go.
The Rest of the Portfolio and Closing
Vik: Yeah, for sure. UALink is supposed to be the more purer networking, you can, GPUs, you can access each other’s memory and all that nicely with the UALink protocol. It’s built for this kind of stuff. It’s very low overhead, very lightweight protocol. Ethernet and Eson and all these are kind of derived from the previous era of Ethernet and is not optimized for what AI needs today. UALink is more so. So we’ll have to talk about these standards another day because that’s a whole different subject. But I just wanted to like mention the two other things that it’s only a mention. I think Astera Labs has their Taurus product line, which is essentially basically the same signal conditioning product, but you put it inside cable and you make an active electrical cable out of it, like AECs. So they have these chips that they can put in AEC cables. This is not like a high margin business and they are very clear that look, our gross margins may go down because this is not the high margin business, but we want to do it anyway. Like if there’s a way we can use our existing technology, we will. So they’re like, we’ll do it. And then their newest line, I think, is the Leo product line, which is basically a CXL controller, which is used to maybe pull memory and do that kind of stuff. So we it’s still that’s still new. So there is some promise. There are apparently some design wins already. So we have to keep a tab on that. But CXL is a new topic, a new episode for another day.
Austin: Totally.
Austin: Totally, totally. CXL is old, but it’s new and it’s coming back maybe. We’ll talk about it another time. But with that, listeners, we hope that you learned a lot about retimers. We hope that you learned a lot about Astera Labs and you had fun with it like we did. So thanks for listening. Check us out on YouTube, leave us comments, share this with your friends, check us out on Spotify, Apple podcast, wherever you listen. And of course, check us out on X. We’re trying to post this, of course, on X, as well as our daily takes. If you like our daily takes, check out our newsletter. Just go to semidoped.com and you’ll find it under the daily tab. And then, yes, also follow us on X. We’re trying to share some clips there as well, so that if you missed some old stuff, you’ll get to see it again. So thank you for your support and until next time.


