Austin and Vik break down the single biggest bottleneck in modern AI: data movement. They explain the three tiers of datacenter networking, the intense debate between copper and optics for different applications, and why the future of high-speed interconnects might be co-packaged with the chip itself. 🤖
Things we cover:
* The three types of datacenter networks: scale up, scale out, and scale across
* Copper vs. optics for different reaches
* Pluggable optics and the move toward co-packaged optics (CPO)
* Nvidia’s NVLink vs. open standards like UALink
* The role of DSPs and transceivers in optical links
* Networking speeds from 400G to 1.6T
This podcast is lightly edited for clarity.
The Biggest Problem in Computing
Vik: The biggest computers in the world today are basically not in single boxes, right? They are in entire data centers. And the whole point about this is that all these chips are intended to work as one big computer. But the whole problem with this approach is that the compute capability has grown much faster than how you can hook up all these chips in a data center together. So today’s compute capability is more restricted by how the CPUs and GPUs are interconnected in a data center, rather than the performance of the silicon by itself.
So, the biggest problem facing data centers today is how to hook up all these chips to work as fast as possible. Networking is all the rage today, and it goes from a wide spectrum of whether to use copper, whether to use optics, and where to put all this stuff together. So, by the end of this episode, you should be able to understand why we need networking to begin with, what are the three kinds of networks that all data centers are built up from, and why perhaps co-packaged optics is coming. That’s what we hear. We don’t know.
Austin: Hello everyone, and welcome to another Semi Doped podcast. I’m your host Austin Lyons from Chipstrat, and with me is Vik Sekar of Vik’s Newsletters. Hey, Vik, what do you say? Should we talk about data movement interconnects?
Vik: Yeah, we should do this because it’s the biggest problem that we are facing in the history of computing. Because in the past, we’ve always had the discussion of what is the per-core performance of a CPU, or then we went to the multi-core era, like, hey, we have 16 cores in this chip, or even server-grade chips have like 256 cores. And then the GPU era came and then you could put in the accelerator cards in your gaming PC, and then you got you could get graphics acceleration.
And if you remember, there was this Nvidia SLI, which is basically that you could put two Nvidia GPUs together, and you could put this local interconnect bridge. I forgot what SLI is actually called, but it’s basically think about it like a bridge that is like an interconnect between two GPUs so that they work as one. So that’s like the simplest form of what we’re talking about here. Nvidia SLI was a way to make two GPUs work as one. But now, we need to make 100,000 GPUs work as one for LLM reasons. So, networking is the biggest discussion today.
Austin: Totally. And obviously for training, it’s all about, yeah, how can we have a data center-sized brain? But even for inference, this is why we’re going to scale-up domains of 72 GPUs and eventually beyond. Even with large, frontier models that have 10 trillion parameters in their mixture of experts, we are getting to the point where pre-fill and especially decode need as much memory bandwidth as possible. There’s so much data movement that it’s all about how can we actually put together a bunch of chips on this scale-up fabric? And we’ll talk through all of these—scale up, scale out, scale across, scale in, all that stuff here.
But it’s not even just about training, but even inference is demanding lots of data movement to be able to do frontier models at high throughput, but also obviously large batch sizes. And so, don’t let people think that this is just about training, but this is also an inference story too.
Vik: Yeah, for sure. So, whenever we talk about data movement, it involves data movement at all levels. That’s what is interesting about this. Data movement could be restricted to information moving between memory, like HBM, to the GPU. So, that is like a really short distance over which you need to move a lot of data really quickly. And that’s just because of the way LLMs work. You need to read a lot from memory all the time, and the faster you can do it, the more performant your LLM is in inference, for example. In training, it would mean that faster network bandwidth between memory and the GPU, or overall bandwidth being high implies that your training run finishes faster. And that is everything because everybody wants to be in the next biggest model, everybody wants to train it before the next guy gets there. So, that’s the whole idea. Everything gets faster.
Austin: Yeah, yeah. So, should we share some slides in this talk so it’s a little bit more visual? We’ve heard from lots of our YouTubers, who are also, thank you for watching us, and they’ve said, we want more visuals. So, if you’re listening, of course, we’ll try to describe it so that you can understand it as you’re lifting weights or something. But for those of you who are watching, we’re going to pull up some slides here and we’ll talk through some of them as we go.
The Three Tiers of Datacenter Networking
Vik: So, the whole data center essentially has networking on three different layers. So, the first one is scale up, which is literally what you think it looks like. You go up a rack. So, you have a rack full of GPUs, and then you hook them all up going from bottom to top. So, that’s scale up.
And then the other dimension of scaling is the scale out network, which means that you can have different racks, each of which has its own scale up network, but now you hook them all up together. So, that is scale out. And just between these two, you can already start to tell that scale out networks are kind of longer reach networks than scale up, right? Because scale up only means that you have to go within a rack. And if you have never stood next to a data center rack, you can imagine that it’s about the height of a tall adult human or a little bit higher. I think it’ll be like maybe seven foot, six to seven feet high racks. That’s about the whole length of the rack that scale up network has to cover. Scale out, it depends on how many racks you’re going to put. So, if you put one of these super pods or something, you could travel a dozen racks across from east to west by the time you hook them all up.
And then all of the racks in a data center are all scale out, but if you decide that you want to now hook up an entire second campus, maybe located 20 miles away, 50 miles away, you need a whole new network to do that, and that is called scale across. You can even hook up data centers to act as one big chip. So, it’s not just one data center that acts as a big chip, you could hook up five of them or two of them, whatever. That’s called the scale across network. And this scale across network has the largest reach of the three because it spans obviously tens of miles. So the technologies used in each of these layers are substantially different.
Austin: So, let’s talk about these because someone listening might say, “Well, why does scale up, why do we call it something different than scale out, for example?” And you hinted at it that there’s different technologies that are used, whether you’re staying within the rack versus rack to rack, versus building to building.
Now, a couple things I wanted to mention. So, scaling up, the problem that we’re trying to solve is in a dream world, we would just have the biggest compute with the most memory possible, and that would be your rack. It’s just like Cerebras, for example. They’ve just got all these chips on one wafer, and they got a ton of compute, and they can all communicate really quickly on that wafer within each to each other. The problem is the way that the industry works today, GPUs are packaged—maybe it started out one die, then eventually made its way to two die—but it’s packaged with a discrete amount of high bandwidth memory on a particular chip. And in the future, some of these AI ASICs and startups are choosing a certain amount of HBM and maybe a certain amount of SRAM or making different memory decisions that we’ve talked about in other episodes. But at the end of the day, each GPU only has access to so much HBM. And what we want is more HBM than what’s possible to physically package on it. So, we want to make two different GPUs, let’s just say on the same board, for example, or in the same server node—forget the whole rack, even within the same server node—we want neighbors, GPU A and GPU B to be able to access each other’s memory as if it was their own.
So, scale up is really defined as how many other GPUs can I talk to at low enough latency that it feels like we’re sharing memory? And this is reaching out and talking to other accelerators’ memory. That’s what this is all about. And so, when you ask yourself, “Well, how big could my scale up network be?” one way to think about it conceptually is, well, it can be as big as when I talk to every other GPU, it feels like I’m actually accessing my own memory because the latency is so low, using NVLink or something, for example, some high bandwidth, fast network. So you can actually have a scale up domain that spans two racks. It’s technically possible so long as GPUs between the neighboring racks are still so close and connected in such a way that it feels like you’re accessing each other’s memory as if it’s your own. So that is scale up.
And then scale out is of course, now when I’m talking my rack and that rack 30 meters down, obviously, there’s going to be enough latency, enough hopping through switches and things like that that it won’t feel like I’m accessing their memory. So I won’t try that. We’ll communicate in a different way and send information differently. I also want to just mention one other thing. So East-West, like everything on the same sort of high bandwidth GPU to GPU communication fabric, my understanding is that’s East-West. And technically when people use North-South, usually I think they’re talking about talking to the front-end network, like GPUs talking to external CPUs. Like the chat, I’m chatting on a chatbot and it gets sent in. That’s like the North-South network.
Okay, so we’ve got this slide from Nvidia that starts to show that scale up, scale out, scale across, like we’ve hit on it at a high level. And of course, now there’s another new term because why stop there? On the far right side, we’ve got scale above. So, hey, what if you have data centers that aren’t physically connected via fiber optics, maybe 50 miles apart, but are actually literally in space and there’s communication, I guess, through satellite frequencies instead of fiber optics. And then on the far other end, we’ve also started to hear about people talking about scale in. Do you want to say anything there on that or do we plan on covering that today?
Vik: Not too much, but there is a whole networking effort to make the connections within a compute tray—the thing that slides into a rack that holds maybe the two GPUs or whatever. You could talk about networking even within that, like how do you make that faster? We won’t cover it too much today, but we’ll briefly touch on it as we progress along. So, these two are really outside the norm of what normally is in the conversation of data center networking, that is scale in and scale above, but they do exist.
Scale Up: Copper vs. Optics
Austin: Yes, totally. All right, so let’s keep going on scale up. So if the goal is to make a group of accelerators behave like one large accelerator, they can access each other’s memory. You obviously want the highest bandwidth possible, you want the lowest latency possible. These are going to be short reach. So we’re talking potentially next to each other in the same board or within the same server node or within the same rack. So like a few meters at most. So tell us, Vik, this is the question that everyone talked about a lot and people are still talking about, is this copper? Is this optics?
Vik: Yeah. What is scale up? It’s a big question, right? At this point, it all comes down to what is the speed it can handle. The whole motive, the whole mantra of this networking at this level is: copper when you can and optics when you must. So, copper is cheap, it is low power. It has always worked for decades. There’s no need to do anything fancy if we can avoid it.
However, there is another problem is that when you do go to optics, if you go to optics, it is very, very power hungry because the number of cables in a scale up network is enormous. And we have some pictures coming up, we’ll show you. But it is a lot of cables and it’s a lot of transceivers on either side. So, you’ve got to make conversions into optics from electrical, and then go from one point to another and convert back into electrical. Those conversions are handled by basically transmitter-receivers or transceivers. And there are way too many of them. So, nobody wants to go to something like optics, which is more power hungry because of these conversions from electrical to optical and back, if they can avoid it.
And in this picture, you see this as basically the back-end network. The back-end network is this really high-speed fabric. Think about InfiniBand. It could also be Ethernet scale up. So, these are like really fast connections, which bring down the latency between these GPUs to be as low as possible so that they can act as one domain.
Austin: Yes. So, scale up networking, historically it’s been NVLink from Nvidia with NV switches. There’s also UA-Link, which is an industry standard that AMD is spearheading, but there’s lots of other people coming along to try to make an open alternative for fast scale up networking. And then InfiniBand has historically been scale out networking, and also Ethernet. So, Spectrum-X Ethernet from Nvidia and Ethernet from others for scale out. But back to scale up, Broadcom has also been pioneering using Ethernet for scale up. They called it SUEI, I believe, scale up Ethernet.
So, there’s definitely appetite as other merchant GPU vendors like AMD or AI ASIC companies are also trying to connect all of their accelerators together using copper, talking very high bandwidth. They need protocols so that you could buy switches from Broadcom or use a Marvell switch or whatever. And so there’s some industry standards that are still being developed. So if you hear of UA-Link or if you hear of I think it’s called E-SUN now, you’ll know it’s scale up Ethernet, Ethernet for scale up networking, I think that’s what E-SUN stands for.
Vik: What about the front-end network? I think it’s worth a brief mention.
Austin: Okay, so what is the front-end network here? The front-end network here is all of the information coming into the data center from the outside world. So this is connecting, yeah, obviously you’ve got all of your GPUs, but when I type when I use my cloud code locally and it needs to send all this context in, it’s coming in over the front-end network. And then the GPUs are all talking to each other doing their thing and then it comes back out over the front-end network.
Vik: Yeah. And then you have data center interconnect, which is like the scale across thing that happens outside of the data center. So that’s what it is.
Cabling the Datacenter
Austin: Yeah. On this slide that I’ll pull up, again, so scaling up, conceptually you can think of it as adding more and more accelerators. They’re added in three dimensions. So you have some within a board and then some within the rack and you might even have rack-to-rack connections. But ultimately, they all have to talk to each other. And so there’s a lot of thinking that goes into how do you design the topology of the network. You can have leaf network switches, spine network switches.
But obviously you can see the tradeoffs even when you look at a picture like this. If you in the far if you got the far left bottom GPU trying to talk to the far right bottom GPU, there’s going to be hops through the network switches. And so what’s actually really interesting here is if you go look at how long does it take for even an NVLink switch to send information, you can, let’s say it’s like four milliseconds or something. And then you start to look at for an LLM like a mixture of experts LLM, how many times does information need to be sent back and forth? You can start to calculate the minimum amount of time for inference to happen. So, if you have a certain number of milliseconds and you have to make a certain number of sending information back and forth, you can multiply that and you can get a number of milliseconds, which then can actually tell you what your throughput will be. So, you know, even for one user, we might only be able to get 400 tokens per second. And that is the limit that maybe a GPU could do in a particular architecture, and it’s not even compute limited, but it’s about data movement limited.
I just wanted to share that because I thought it’s interesting that inference is just another use case of showing that inference can be limited by how quickly you can move information and how many hops you have to make and how each piece in the chain, how long it takes to respond. It can actually cap the interactivity, which is ultimately the user experience. And so this is why again, you’ve got these SRAM-based systems, the Groqs and the Cerebras, where because maybe they don’t have it, in Cerebras’s case, everything is just on one wafer right next to each other. If you can fit all the model weights in and you can keep it all onto one wafer or a few wafers, you can skip a lot of this communication hops and ultimately get a much higher interactivity.
Vik: Awesome. Yeah, there’s a lot of networking decisions that goes into how well tokens work out for you when using it at scale. Oh, this is one of the pictures that always fascinates, right? Because you can see this rack there with all these cables going up the rack, hooking up all the GPUs within it. It’s just a nice way of visualizing what kind of cabling goes on within a single rack.
Now, for a 72 GPU rack, because you want to have an all-to-all connection between all the 72 GPUs, which is pretty much 72 squared, you’ll get a little over 5,000 cables that require to go between GPUs in a single rack. And I read somewhere that this adds up to about 2 kilometers of cables within a single rack. And they carry an enormous amount of information that is just staggering. Because if you think about how fast each of these cables are, you are looking at in the Blackwell era, you’re looking at 200 gigabits per lane. So, that’s really fast. You can multiply that up by whatever. You have an incredible amount of data going by everything. In the Rubin era, that’s going to be 200 GHz, but bidirectional. So it’s going to have both directions running on a single cable at that speed. So that just goes to show that the kind of speeds that you have on this stuff is insane.
Because I bet that if you go to CAT 6 cables that run in your home, for example, most people don’t even run a 10 gigabit per second network at their homes. They don’t run 10 gigabit networks, because you need special switches for that and kind of stuff. Nobody even cares about that. So all this stuff that runs in your house, your networks, they’re all very slow, if at all, it’s maybe running at one to five Gbps or something. We can do a network check in your own house later on a wired network, don’t do it wirelessly, on a wired network and see how slow that is. This is incredibly fast, and you got so many cables, and that’s just one rack. And then you have so many racks in a data center, all of which are carrying this much data. Then you hook them all up together. So it’s just amazing how much networking there is.
Austin: Indeed, indeed. We’ll move to the next slide. So I remember when Elon posted this, when they were standing up the XAI Colossus 2 data center in Memphis, I believe. And obviously, these are purple cables. And so, Credo is synonymous with purple cables. And so I remember when everyone saw this, they’re like, “Credo, go Credo, this is amazing.” Now, I think interestingly, now, I think some customers are asking for different colored cables so that people aren’t getting caught up in what color is your cable and should I invest in a particular company? But again, just looking at this beautiful picture of all these well-managed cables is just, you know, still emphasizes how much information is getting sent back and forth, just seeing all these pipes. It’s crazy.
Vik: Yeah, this is Elon actually posted it himself on X. It’s like, look at this beautiful cable arrangement. And it’s so satisfying.
Austin: It is satisfying. It’s also a little terrifying. Like, holy cow, that’s someone’s job to make sure that all of those are connected and that it all works. Man.
Vik: I know. And imagine them getting tangled up, right? Like imagine tangled network cables at data center scale. It’s impossible. You probably have to burn the data center down. You can’t fix that.
Austin: Right, totally. So, what are we looking at here in this one? Another Elon tweet?
Vik: Yeah, this is another one in XAI Colossus where you see all these, I think these are optical cables that go within a data center. These are the cables that hook up racks and stuff.
Austin: Uh-huh. Longer reach ones.
Vik: Longer reach ones, yeah, and that’s very beautifully laid out as you can see in the picture. It requires a lot of reach actually because you just can’t hook them up like you would with the minimal possible thing or whatever. No. To actually do cable management at the data center level, you actually need more reach than you would physically measure the distance between tape between different racks. You can’t just go like, “Oh, 12 racks. Okay, how about we just multiply the each rack is like two feet wide, 12 racks, 24 feet, and that’s the reach.” I mean, that’s just the beginning estimate. Usually you need much, much more because of all this cable management that happens. I just wanted to put this up there just so that these are like GB200 deployment in XAI Colossus. Save us going inside and taking pictures ourselves, right? By the way, which we are totally willing to do. This is the second best thing.
Austin: Yes, Elon, let us know. We’ll be there. So, I guess for people who are listening and not watching, we’re obviously showing pictures, first of connections within a rack and then between racks, and this is long distance ones. You’ll just have to go check it out to see it for yourself. But I’m going to flip back through these slides and make an interesting point, which is if you look at the cabling within even an individual rack, there’s tons and tons of connections, lots and lots of cables. Eventually, there’s some mid-plane stuff that we could talk about.
But the point is when you’ve got on the scale up network, when you have all these GPUs and they’re very close within the same rack and they’re all connected all to all, there’s a ton of connections. Now, even when you go look at the next picture and you start to get some interconnections between racks and between switches, there’s still lots of cables, but as you go farther and farther out, like on the scale out, and eventually on scale across, there’s less and less cables. And so what’s interesting is you can think of that as sort of a proxy for addressable market size.
So these, as we’ll get into later, when the optical companies are saying, “Oh, we want to get into scale up,” they’re very excited about it because in scale out, where they already play, and scale across where they play, you’ve got fewer connections. You can think of it like, there’s highways that run between towns and then you’ve got interstates that run between LA and all the way over to New York City. But obviously within a town, even though it doesn’t feel like it, that’s where you’ve got the most miles of roads because you’ve got house to house to house to house, street to street to street. And so scale up interconnects, there’s going to be tons of connections and huge, huge opportunity and huge TAM. Scale out, not as many connections, scale across, even fewer connections. And you can physically see it when you look at these pictures, so I just thought that was cool.
The 78-Layer Mid-Plane
Vik: Yeah, I wanted to go before we go ahead, I wanted to touch on the the mid-plane PCB you mentioned. Can you go back to the Nvidia slide and I’ll tell you. I think this is the perfect place to mention it. So, the basic thing is that when your speed gets higher and higher, this amount of cable doesn’t reach. So, in the next generation of racks, Nvidia decided that they will instead build a gigantic PCB and then come up with a very funny and novel way of hooking up the compute on one side and the networking on the other side. And then make all the connections within the PCB itself.
That’s not an insane idea because I think people have done these kinds of PCBs before. Like data center PCBs can be pretty big and complex like this. They’re usually in the range of 25 layer PCBs and stuff like that. Problem is that this one, this complex PCB is about 78 layers, I believe. So, that is because you have so much routing that needs to happen. The way I describe this on PCB scale is like think of a flyover, like a clover or something in a big city. You got to make all these transitions from the north highway going east and the south highway going west and nothing should intersect, everything should move smoothly around, go up and above and below. That’s exactly what happens on the PCB as well. All this traffic needs to be routed somehow and you need 78 layers to do it.
So think about the freeway intersection. Now you have 78 different levels doing that. It’s quite complex and it’s very difficult to manufacture too. So it’s quite an engineering challenge going forward. And this is all because they obviously want to avoid going to optics. This is really pushing the copper story. I don’t want to say any more about it. There’s always people are contemplating whether this thing can be made or it can’t be made because it’s too hard. But regardless, I think that’s one nice engineering approach to do it.
Austin: Yeah, that 78 layers, that’s wild, man. Yeah, tell us about large scale AI data center challenges.
Long-Reach Optics and Transceivers
Vik: Yeah, yeah, yeah. So, when you go to multi-building campuses or even between entire campuses, now you have a whole different problem. Like you have to route enormous amount of cable from the data center out and send it through all these underground fiber networks that have already been laid out and all that. I just want to put this in here because it’s like it presents a different challenge. You could have a few kilometers of reach, you could have tens of kilometers of reach. So the technologies tend to become different.
Within a rack, you could use simpler forms of modulation like PAM4, which is one way where you just use four different signaling levels. 00 is one voltage level, 01 is one voltage level and so on. And you can just modulate like that and send information. It’s simple. Now, in data center scale, when you’re trying to go far off, the modulation methods are not as simple like PAM4. You have to end up using a lot more complex schemes like coherent optics where you use both the amplitude and the phase information to send the data over long distances. And this is all well known. This has been done for a long time, using also other techniques like wavelength division multiplexing, which if you’re looking at the slide on YouTube, you’ll see it as WDM. In WDM, the idea is that you will send data on multiple wavelengths all together. So if you use eight wavelengths at the same time, you can send eight times the data and so on. And the example of eight is more of a coarse wavelength division multiplexing. So many times it’s referred to as CWDM. You could have dense wavelength division multiplexing, which is you could even have a hundred wavelengths traveling through a single optical fiber. So there’s a lot of different technology when you go at scale here. And remember, this is all optics. There is simply no way you can send copper over the distances required between buildings or between data centers. At that domain, there’s no argument here. Optics is a must.
This picture, if you’re seeing it on YouTube, I saw this online, so I thought I’ll throw it in here. It’s basically just the intra-building fiber trays in the Meta data center. So you can just see how many cables are going through that and it doesn’t even look very cable managed to me. It just looks like a bunch of optical fiber just clued together. I’m like, how does that even work? Yeah. So it’s just a lot of fiber. It’s a lot of fiber. Anybody looking at this is going to go like, “Oh, who makes all this fiber by the way?” Corning is one. And so they, as optics picks up, I guess there’s going to be more demand for their stuff. But again, it all comes down to like you said, is it going to be within a scale up domain? That’s even more optics. There are a lot of cables, but then scale across also has a lot of optics and a lot of fiber optic cables.
Austin: We’ve talked about scale up, scale out, scale across. We talked a little bit about fibers, but should we talk a little bit more about lasers and how you actually send the information?
Vik: Yeah, yeah. I think it all ties into essentially having these optical transceivers. If you’re looking at this on YouTube, what you’re seeing on the screen right now is an optical transceiver. It’s kind of a beautiful thing to look at because it has all this gold platings and this transparent glass fibers coming in and they’re all arranged beautifully and symmetrically. I’m just describing what I’m seeing basically. And the whole function of this optical transceiver is to convert electrical signals into optical signals and vice versa.
So the way it converts electrical signals into optical signals is that there is a laser on this transceiver. And then the laser is turned on or off based on the maybe the information that’s coming in. Like if it’s a zero, the laser is off. And if it’s a one, the laser is on. This is a simple example. This form of modulation is called essentially intensity modulation. It’s as simple as that. And on the receiving side, you have basically a photodiode, which is a device that converts light back into electrons. So once it senses that the light is on, it says, “Hey, it’s a one.” Once the light is off, it’s like, “Hey, that was a zero.” So you have both the transmitter and the receiver inside this. And then you have the appropriate electronics that goes with this. The laser needs a driver chip and the photodiode on the receiving end needs some amplifiers, because what you’re sensing is basically an analog signal, which should be eventually converted into digital signals and so on. So all of this takes energy, it takes power.
And on top of all of this, if you need to and if the reach of this optical cable is long enough, you also need a digital signal processor, a DSP. The reason you need a DSP is because sometimes bits don’t arrive the way you think they arrive. Like a zero becomes a one, a one becomes a zero. And by using DSP, you can use parity bits. What these are are like error bits. Like you say, “Hey, wait for 10 bits to come and then the 10 bits, if you do certain mathematical operations, should match these next two bits that come.” And if they match up, that means the 10 bits that you sent were correct. And therefore, it’s all good. Otherwise, the DSP has a way of finding out the error bits and then correcting for it. So what ends up happening is that you can actually, even if the whole link screws up and it’s really terrible, you can actually fix it with a DSP. So that’s the whole point of this optical transceiver assembly. And I just want to put it up here so that we understand that this is how the whole optics works. And all of this takes a lot of power. This is primarily the reason that you only want to use this when it’s really required.
Austin: Yes, yes. And I just want to lean into that a little bit more. So obviously, the AI accelerator, everything, all the communication within the chip is happening in the electrical domain. It’s electrical signals traveling around, currents, voltages and so on. Ideally, if you just communicate over copper, you’re still sending electrical signals, no big deal. Anywhere you’re converting to optics, to what we’ve been talking about here, you have to do the electrical to optical conversion, then you have to convert back from optical to electrical. And to your point, sometimes you might have to have a DSP, you might have to send parity bits, do this all this error correction. There’s power, there’s latency, there’s time. And I know you said all this and I’m just emphasizing that now when we start to scale up with optics, you have to ask yourself, “Oh man, is every GPU going to talk to every other GPU and have to have all of these transceivers on every lane?” And so what’s your reaction to the idea of having to have so many transceivers for even for scale up?
Vik: Yeah, that’s what Nvidia has been pretty vocal about it in all the conferences that I’ve been to that they don’t want to do this. They don’t want to use pluggables to scale up a network because each module you’re looking at here is about 30 watts, I’d say. That’s a lot, considering that you need two of these on either end of the cable and then you got like 5,000 cables, you can add it up. It becomes quite a lot. And they don’t want to add to the power that is already being sucked up by the compute trays in a rack because those GPUs are massively power hungry. And nobody wants to add another 10, 20% to it because modern data center racks are already 100, 200 kilowatt racks. And we are not even talking about the Kyber era of racks which are like 600 kilowatt supposedly. And in the future, we want to make one megawatt racks. So, these things are already pretty power dense and they don’t want to add to the power usage if they can avoid it. So they’re like, “No, we’re not going to do this. So we’re going to stay scale up in copper as long as we possibly can. If we have to go optics, we’ll do something else.” We’ll talk about what that is.
The Future is Co-Packaged Optics (CPO)
Austin: Nice. Very good.
Vik: Yeah, this is just the transceiver looking at it. We already spoke about this, but essentially the TOSA in this picture refers to the transmitter optical sub-assembly. It’s just a fancy way of saying all the electronics that goes into the transmitter side. So the ROSA here refers to the receiver optical sub-assembly, which is another fancy way of saying all the receiver components in one place. And you can show this picture shows essentially how the optical signal goes in one side and the electrical signal comes out the other side and there’s these heat spreaders and the whole housing. This whole thing is called in this picture showing a QSFP connector, but you have various kinds of connectors. You have OSFP, you have OSFP-XD, which is extra dense, which means you can even put more optical cables into this thing. It has more channels. Or the modern one that is really massive and huge is the CPO form factor, which is basically like eight of these things put together.
So, this is just like the pluggable module. It’s not like these are useless and it’s not like they’re going to go away anytime soon. But they have their uses and disadvantages. And what is happening is as we’re going to more and more faster networks, you can see in this picture that some of these generations of speeds are kind of going down and newer generations are picking back up. For example, if you’re looking at the graph here, most of the networking was 100 and 200 Gbps networking early 2020s, 2021, 2022, and then you slowly 400 Gbps started picking up. And 2023 and 2024 were the time when the 400 gig networking took off. And then you started seeing the entry of 800 gig networking. And that started occupying a sizable portion and 400 gigs is like kind of dying away in ‘25, ‘26, at least according to this graph. And then the projection is that the next generation 1.6 terabit per second networking comes in and takes up an increasing portion of the networking hardware and displaces the previous generation.
Austin: The one thing I want to say here on this chart is that, so this is aggregate bandwidth, 1.6T, that could be like eight lanes of 200 gigs per lane, for example. 1.6T is just beginning. We in earnings calls, you hear people talk about it a lot, but this chart is nice to show you that no, no, there’s still lots of 800 gig out there. There’s still 400 gig out there even. And then I’ll also just remind people that with each industry jump to the next speed, which is usually a doubling of speed, obviously, there’s opportunities for every player to try to get there first, and there’s opportunities to charge more per component, wherever they are, if they’re the transceiver or in the supply chain. And so usually, the exciting thing is that the addressable market can end up being bigger and bigger because it’s more and more viable because it’s frontier, and obviously the hyperscaler is going to willing to pay more for 1.6T than maybe they are for 800T and so on. So, these are just the continued cycles that companies are racing for and trying to capture outsize value. But even if you’re not there first, there’s still money to be made in the other speeds that you were competing in.
Vik: Yeah, the next generation is interesting because you would imagine it’s 3.2T, right? But there is some talk in the industry generally that there is a possibility that we would go to 300 gigs per lane and have a 2.4T generation before we go to 3.2T. That’s interesting. I think it’s primarily pioneered by by Google, because they need all kinds of fancy interconnections for the TPU pods. So I heard that this is what is being driven.
The final thing that is worth mentioning here is, really we’ve all the time all that we’ve spoken about optics so far has been basically the pluggable optics, which is also basically what goes into the face plate of the switch, so you can just go plug it in and that’s your transceiver. But like we mentioned earlier, it’s essentially very power hungry. And the reason it’s power hungry is that between the front panel of the switch and where this actual switch silicon lies, there is a long pathway through a PCB. And that pathway is so long that it’s basically terrible for signal integrity when the speeds get to 200 gigs per lane and or faster or whatever. So it really doesn’t scale. And now the only way to overcome that is like using these DSPs and all that, which is power hungry, nobody wants it.
So one idea is like, okay, why do we put this optical transceiver at the edge of the face plate of the switch, which is basically outside the box? Why are you putting the transceiver outside the box? Why don’t you put it inside the box and put it as close to the switch as possible so that you can stay in optics as long as possible, and then you can just go through the copper interconnect really quickly to the switch. And then that’s it. Like now you don’t probably need all those DSPs, and life is great because you don’t need to make any corrections and you don’t need to blow all this energy running the DSP and things like that.
So one option to do this is what is called near package optics, which is you put it on the same board, you put the optical engine close to the switch. You don’t put it outside the box, you put it inside the box, close to the switch. And yeah, and then you have a much shorter path and then stuff is working great. But the ideal solution that the world wants to go to is basically co-packaged optics or CPO, which is you basically put the same the GPU and the optical engine right next to each other and co-package them using some advanced packaging technology like TSMC’s CoWoS or Intel’s EMIB or something like that. And then what you have is you have basically a switch with an integrated optical transceiver that is almost an entirely optical link now because there is almost no copper interconnect left, except that very tiny sliver that goes between the GPU and the optical engine.
So this is where the world wants to go to, but this is not easy. The CPO approach has been in the industry, people have spoken about this for a very long time now, could be even a decade. And there are good reasons it hasn’t come to market yet. It’s a challenging problem because if any one of those optical engines die, and they do, because if it has a laser in it, lasers are terrible when it comes to handling temperature, they’re going to die. So when they die, what do you do with the switch now? You can’t just change out an optical engine, you have to throw away the whole silicon switch, made in probably a 2 nanometer process node technology, which is thousands of dollars and it’s a big waste. So nobody wants to do that. So this is where the industry is. So they’re looking at it saying, “Hey, okay, what dies the most? Is it the laser that dies the most? Okay. Why don’t we move the laser outside the rack, but we’ll keep the rest of it inside the rack?” And so that’s one approach. But then, there are still other concerns with co-packaging it and so people said, “Okay, fine, fine, fine, let’s take a step back. Maybe we’ll go to near packaged optics.” It’s not as good as co-packaged optics, but maybe this will just work out just fine.
The beauty of this whole CPO/NPO approach is that it is actually much lesser energy usage than you would do with a regular pluggable transceiver. You could make the energy drop to a third of what you could get from a pluggable transceiver. And so now all of a sudden, optics has become an interesting technology for scale up. And that is where the industry is today because if people realize that if you can go to near packaged optics or co-packaged optics, it’s amazing because now you’re just only limited by what optics can do. And remember, optics can go very, very fast. Optics can go very long distances, like we spoke about in scale across. So the possibilities are endless if you can get co-packaged optics to work and then hook up the entire scale up domain, hundreds of thousands of GPUs, on the scale up and scale out networks, all with optics. That is the holy grail of networking, which the industry is headed towards and the big question remains as to when. Nobody knows that.
Closing
Vik: All right, so, that’s the end of this episode. I think the networking idea is a very fascinating one and there’s really a lot to be said about what happens in each part of this whole supply chain, and each little bit of technology has its own deep dive. So we could talk about just external lasers, and what kind of lasers you need for CPO for like a whole episode. So, we’ll cover those on future episodes, but if you’ve made it this far, thank you for listening and tell your friends if you enjoyed this episode. And you can always find us on YouTube, obviously, which we recommend watching this on because of all the pictures, but also on all the podcast platforms. And if you’re listening on Apple Podcasts, please do give us a five-star review and hope to catch you on the next one.


