0:00
/

🎙️ NEW EPISODE: Grok Bots and How CPUs are used in Agentic AI

Grok bot's agentic AI platform, the GPU-as-genius and CPU-as-assistant analogy, host node vs. agentic CPUs, the Mac Mini craze, and more.

The rise of user-friendly agentic AI platforms will create a massive new demand category for dedicated, high-core-count “agentic CPUs” to execute tasks in parallel, fundamentally reshaping the server CPU market beyond just feeding GPUs.

Austin and Vik explore the evolving role of CPUs in the age of agentic AI, sparked by the release of Grok bot’s new agent platform. They introduce a powerful analogy, casting the GPU as the “genius” and the CPU as the “assistant,” to explain the different types of compute work. From host node CPUs keeping the genius fed to racks of “agentic CPUs” handling spillover tasks, they map out the new landscape of AI hardware.

Things we cover:

  • Grok bot’s agentic AI platform

  • The GPU “genius” and CPU “assistant” analogy

  • Host node CPUs vs. agentic CPUs

  • Why the Mac Mini became the early agent computer

  • Rack-scale solutions for different workloads

  • The need for new orchestration software

This podcast is lightly edited for clarity.

Welcome to Grok bot

Vik: So check this out. This week was actually the release of Grok bot, which is an agentic AI platform. You think about it like an open claw, but extremely easy to use so that everybody can run agents all the time right now, even if your computer is closed. This is a very important development, I think, because it’s not so much for what the Grok bot is as a product, but what it means from an underlying substrate of compute and specifically what it means for CPUs.

Austin: All right, hello listeners. Welcome to another Semi Doped episode. I’m Austin from Chipstrat and this is Vik from Vik’s Newsletter. And if you’re watching this on YouTube or something, you’ll notice Vik is in a different location. So Vik, tell us where are you and then I also saw that you are trying this Grok bot, which I saw on X, but I have not tried. I honestly don’t even know anything about it. So, I want to hear where are you and then tell me about Grok bot because I know you’ve been playing with it.

Vik: Awesome. I am in the Bay Area. I came here for Hot Chips next week, which is very exciting. A lot of people are coming down for this conference, but I decided to show up a week early because I was actually in Taipei attending the Open Compute Project APAC, the Asia Pacific region version of this conference. And I was like, I think it makes just sense to come here a little bit early because I’ve got so many people to meet in the Bay Area. It’s been amazing. I’ve met so many amazing people and so many amazing things have happened that eventually it’ll all come out at some point and we’ll discuss it on the podcast or whatever when the context is right. But it’s been an amazing week and so I am in the Bay Area.

I got my little camera rig with me, but I don’t have my Semi Doped sign, which I’m very bummed about. So I’m going to travel with my neon sign next time. So Grok bot. Grok bot is an interesting thing, because this is an agentic AI software. They call it an agentic harness is the term for it. It’s a piece of software you just download off the website and you install it and you got an interface where pretty much everything works out of the box. So what you do is the first thing that comes up and you go, “Oh, hey, what do you want it to do?” And I said, “Hey, I want you to be a chief agent and you’re the chief of my agents and you’re going to talk to all the other agents.” And then it’s like, “Oh, what other agents do you want?” And I’m like, “Oh, I want an agent that does this. I want an assistant agent. I want a news monitoring agent,” something like that. And it just spins it all up automatically. It’s fancy. It’s really nice. So I think this will open up the doors to a lot of people starting to use agents now because the whole open claw approach and the Hermes approach was a little bit difficult. You got to install it and you got to go into the command line and you’ve got to set up an API key and you’ve got to set up all of this stuff. It is a little bit involved for everybody to use, but Grok bot takes care of everything.

Austin: Nice, amazing. Okay, so like open claw and Hermes, however it’s pronounced, those ran locally on your own machine, right? And this Grok bot, how much of it runs locally on the machine versus in the cloud? You said you downloaded something.

Vik: This is just a piece of software that essentially communicates with a computer in the cloud. This is the nice part of it, which is what I really like about this is it is actually running your own little dedicated computer in some server somewhere. So you have a little virtual machine that’s running somewhere and everything is within that box. All your tools are installed there. All your sign-ins are installed there. You don’t need all these MCP connections and these plugins and those plugins. It’s just think about it as your computer, but somewhere in the cloud. It’s a VM, it’s a virtual machine. The nice thing about this is that all of the security practices and everything that matters for these agents to run in its sandboxed environment is all contained within there. So if your agent has signed in into one tool, it is available easily across all your agents and everything because it’s a local computer you have. That’s the difference compared to open claw and Hermes.

The Mac Mini Craze

Austin: Nice, amazing. That does sound so convenient. So in this, would you say it’s better than just having a bunch of Claude agents running locally on your machine that you have to manage?

Vik: So it’s a give and take. There’s a lot of control when you have your own Claude agents. You can specify a whole lot of stuff to the finest detail like which model you should use for doing what and you can token optimize and be saying, “Hey, for these queries, you can set up OpenRouter as a model provider through the Claude interface.” And you can in OpenRouter, you can say, “Okay, use the auto router or use this hierarchy of models such that these requests are routed here and those requests are routed there.” This takes all that out of the equation. You just don’t know what’s happening. You just have tokens and you use those tokens. That’s it. You don’t know what’s being used.

Austin: I think the easy button is how it’s going to go. Agents will go to the masses where they don’t need to go buy a Mac Mini, they don’t have to install something, they don’t have to know what OpenRouter is, or which maybe is getting acquired by Stripe now. But they don’t have to configure and token max or get token efficient, token min, they can just hit the easy button. So then, does that mean I should buy a bunch of CPU stocks if Grok is going to spin up all these VMs?

Vik: We’ll get to that. We’ll get to the not financial advice part really quickly. But you mentioned Mac Mini, so I just wanted to touch upon that. When agents were rolled out, not so long ago—it was only December last year that we even heard of this thing—it was such a rush to get compute that even Intel was caught unawares of the demand for CPUs. In their earnings call in February, they said, “Oh, we have a lot of CPU demand. We had no idea it was coming.” So the question is, why did all that happen? If people observed keenly enough at the time, you could kind of tell what’s happening because the moment agents showed up, everybody ran and bought a Mac Mini. So what exactly were they buying? Were they buying a Mac Mini for its memory? Were they buying it for its NPU based graphics processing they have in there? Or were they buying it for just compute or a CPU? So that question if you revisit and ask yourself why people bought a Mac Mini, that will tell you something. And the reason is people didn’t want to run open claw on their existing laptops, because you’re giving it access to your whole computer. And that’s dangerous because you don’t know what it’s accessing, you don’t know what it’s going to do with the information on your computer. Does it access sensitive information, medical records? You have no control. So the idea is, what if I could buy a separate machine, a Mac Mini, keep it only for this purpose, install open claw on it, and then have that as my agent computer, and then there I don’t put all the sensitive information. I give it only what it wants. And then it works all the time.

Now, the problem with that setup, and this is familiar to anybody who’s run a home server setup, like I’ve run one for 10 years. But keeping that up and running is a real big pain because you are responsible for power. If there is an outage or if your computer’s PSU blows up, something happens, power goes down. Or if your internet is out, that’s a problem because you don’t have connectivity into your own machine at home. And you have to set up networking, like how are you going to access your PC at home when you’re traveling like I am right now? You’ve got to set up some kind of a VPN network. Like you and I set up Tailscale for all this stuff, all that time. It’s not that simple for everybody to go off and be like, “Hey, if I tell my mom, ‘Hey, mom, all you need to buy is this Mac Mini and then set up Tailscale and set up a UPS so power doesn’t go down and have some failover internet, so nothing happens.’” She’ll be like, “I’m out. I can’t use this stuff. You keep your agentic PC. I’ll just go back to my regular laptop.”

Austin: Yeah, 100%. It is all about the easy button for 99% of people, seriously.

Vik: So that was the problem. So people were buying these computers just to sandbox their environment and have its own area for agents and stuff to run. At that time itself, I thought about it and thought, what is the right way to do this? Do you have to buy a PC and go do this? No, at the time too, what you could do was you could get on a VPS, you could rent out a little virtual machine on a server, and you could install open claw on that one. And then you just go off and—

Austin: Right, right. Yes, I have an agent computer by the way, and I do run a few models locally, but it’s obviously smaller things for the non-frontier tasks, including audio, basically transcription, voice to text. So I’ll run this podcast through on that when I post it to our Semi Doped newsletter to give the transcript to people. That’s the simple kind of thing that can run on a local box, but for most people, when they want to spin up agents and they want to do things, I think that there’s a lot of work that you would want an Opus or a Fable quality frontier model.

The Genius and the Assistant

Vik: Yeah, depends on the workload, you might want a better quality model. So, what is the right way to do this? Now, that’s what Grok bot does all of this, it does all of this for you because you don’t actually have to, when you get Grok bot—this is not a sales pitch and they’re not really a sponsor or anything, but what problem it solves is all I’m trying to articulate. I’m not suggesting anybody go and I’m not shilling this product, okay? So, don’t get it. Actually, it’s $200 a month. Most people probably won’t get it. It’s very expensive. So, not saying you should. But what Grok bot solves is you just get the virtual machine with all the integration tied in. It asks you when you install, “Hey, what integration do you want?” I said I want my Google Calendar, I want my Google Drive, some of these things you can just say add, add, add, and all those things are natively installed and ready to go for you. You don’t have to do anything. There’s Canva, there’s Figma, there’s so many tools out there. You don’t have to set up each one. So, you set it up there into your Grok bot instance and it’s available across all your agents. And then it runs on a virtual machine and you don’t have to be responsible for anything, including the security of that virtual machine because that’s Grok’s problem. I hope they have experts taking care of that. So, I don’t have to become a cyber security expert now.

Austin: Right. I mean, I don’t know. Elon fired all the people at X or slimmed it way down. But anyway, carry on. I’m sure they’ve got their ducks in a row.

Vik: Maybe they don’t. Okay, fine, then deploy Fable to be your security expert.

Austin: There you go.

Vik: You got to be a creative problem solver, Austin. Whatever. All this is, you don’t have to worry about anything. You get a virtual machine that’s sandboxed, security is supposedly ensured. Then you’ve got all your tools installed, you run your agents and it’s amazing. So, now we go to the question of CPUs. So, do we need them? Let’s now go into the question of why do we really need CPUs? Because we never even mentioned GPUs in all of this stuff. Sure for doing the inference, we want to do GPUs, but we really haven’t even touched upon what CPUs have really done for us in the last six months since agents showed up and I thought that’s what we should dig into now.

Austin: Let’s dig into it because from your example, okay, so we talked about running Claude on your laptop, so there’s going to be work done on the CPU there, as well as calls to a GPU in the cloud. We talked about running open claw on a Mac Mini, so now you’ve got another CPU on your desk. It’s not your laptop, it’s your Mac Mini. It’s in its own little contained environment. It’s got a CPU and yet it’s still going to the cloud for the GPU. And then of course, now you’re talking about using Grok bot, which is actually spinning up a VM in the cloud, so a CPU in the cloud. And so it’s a CPU in the cloud talking to a GPU in the cloud and then communicating results back down to you. So there’s actually, and we’ll get into this, but what I’m alluding to is there’s different work happening. In some scenarios, it’s happening on your local CPU, some it’s on your home server CPU, and others it’s on a rented cloud CPU. And so there’s obviously interesting implications. But let’s get into it. Let’s first talk about what do the CPUs even do versus what is the GPU doing?

Vik: Before the show, we were talking about a nice analogy we could come up for this. I’ll let you explain that analogy because this is a very nice framework for which people can think about where GPUs lie, where CPUs lie and how we should think about this going forward.

Austin: Yes. Okay. So, ultimately, I think of a good way of where the value happens is the GPU. The GPU is the genius. Now it’s the genius that has a PhD in everything, which of course, you don’t always need. But the GPU is the genius and then but at the end of the day, once the brain, the genius, thinks of all the work that needs to be done—which I mean, this is what I’m doing with Fable is I’m trying to say, “Give me your very thoughtful approach,” or Opus too. And then once it is decided what work needs to be done, ultimately, most of this work, or it depends on the workload, but a lot of the work can actually be given to assistants, like little doers, little task carry-outers. So ultimately, I think about the GPU as the genius that is doing all the thinking and delegating work and then ultimately the CPU as the assistants that once told what to do, they can go carry out all the work that the genius told them to do.

The Rise of the Agentic CPU

Vik: Yeah, so ideally you don’t want the genius doing all the work. The genius is thinking about it. A genius is paid a lot of money per hour to come up with these brilliant ideas. And you don’t want the genius going off and, I don’t know, putting paperwork in the drawers or you don’t want the genius looking up some information from a file that’s in the drawers or you don’t even want the genius going out to fetch the mail. Ideally in an ideal world, you want to have the absolute genius, aka the GPU, just doing what it does best. Actually, we can break this analogy down into a little bit more detail. So, if all the workers are like CPUs, CPU cores doing their thing, for example, I think we can break down what kind of worker that is too. Because there is something called the head node CPU. And the analogy to the head node CPU is an assistant to the genius. The assistant to the genius, which is the GPU, the job of this assistant is to keep handing the genius work. If the genius needs coffee, you give the genius coffee. If the genius needs water, you give the genius water. The genius doesn’t stop working because it’s a loss, you’re losing money if the genius stops working. So, you want to make sure that once the genius is done with this task, the next folder of materials is ready for the genius’s work to be done and the host node CPU will be like, “Here you go, sir. This is what you got to do next.” The genius is like, “Let’s go. Let’s do this.” So, that’s the host node. And it has to have a, I’d say it has to have a pretty fast per-core performance because it’s constantly monitoring what the GPU is doing. Are you done? Are you done? No, okay, let’s get this ready. Let’s keep it ready. As soon as the GPU is done, this has to hand it in. No delays, no wasting time, don’t waste the genius’s time. So, that’s the host node CPU.

Austin: Yes, I love it. It shows that you want a high clock speed or a very responsive CPU, but also obviously you want the communication between the CPU and the GPU, between the assistant and the genius to be as fast as possible, right? So that there’s no, for example, you ideally you wouldn’t have the assistant be in a different building and then they have to come over and check in and then walk back to their building and then come back and check in. You want them in the same room talking with each other. And so maybe to take it a layer deeper, there’s differences between a coherent host and a standard host with these host nodes. So a coherent host would be like Grace Blackwell where you’ve got the GPU—

Vik: The chip-to-chip, the C2C protocol that the GPU and the CPU uses, the coherent protocol, the coherent CPU would be somebody who’s standing right next to the genius and looking over and like, “Oh, here you go. Here you go.”

Austin: Right. Right there, right? And if they have memory coherency, then it’s like I can see the—I’m the assistant and I can see the genius’s notes and I could jot down on his notes too or something.

Vik: I was thinking about what HBM is in this analogy, and this is where we push it too far and we shouldn’t, but we’ll do it anyway. But I think of HBM as the stack of papers next to the genius’s desk. The genius is not going to get up off the chair and go and get it from the cupboard or anything. HBM will be the stack of papers on the desk, and the genius can just take it and keep working on it and putting it back on or whatever. The memory coherency comes from the assistant having access to the same pile of papers on the desk because they’re in the same room. So if the assistant wants, can also access the memory directly.

Austin: Exactly, exactly. When you zoom out to your point, this CPU assistant, this host node, it has one job and it’s keep the GPU fed. And that’s ultimately what the host CPU needs to do, keep the GPU fed. So, then it starts to raise the interesting question. That made a lot of sense in the ChatGPT era where everything’s just a chatbot and it’s just like, “Yo, I’m asking questions and it’s just like, just keep the genius fed so he can respond to questions.” Now, all of a sudden, when there’s all this extra work to be done, like the genius is spewing out code and you need to compile the code, see if it compiles. Maybe you need to run the code. Maybe the genius is like, “I need information from 50 different sources. Go hit SEC filings and go search the web, and go do all this stuff.” Now, all of a sudden, the question is, can that CPU assistant go fetch all these filings? Can it take all this data? Does it have the memory capacity and bandwidth to analyze all this data and keep the GPU fed? Or is this too much work for that assistant standing right next to him and do we need an army of other CPUs to help?

Vik: This brings us to the concept of what is it? I don’t know. Is there a thing called an agentic CPU? It’s a marketing term probably, but I think it makes sense in this context.

Designing the CPU Workforce

Austin: I’ll give you my take on the agentic CPU and actually I just recorded a podcast recently with AMD and they were aligned and thinking very similarly, which is, if you zoom way out, sort of breaking this analogy, just going back to how things used to work. Back in the cloud era, obviously we had all these general purpose CPUs and they back all of the SaaS products that we use in the cloud. And so, now, as we’re standing up all of these GPUs so that we can ask them our little chat questions, of course, they’re not necessarily going to go communicate with a CPU that’s already running an API server or running my database for my company or something. That already has a job and that CPU is kind of already shaped to fit the workloads of, “I run big databases quickly” or “I run tons of small little API servers quickly” or whatever.

So then the question is, okay, now, fast forward, we’ve got all these general purpose GPUs that already have jobs, they’re already running SaaS products or web servers or databases or whatever. And then now we’ve got these new geniuses that are standing up and they have a little assistant CPU next to them trying to keep them fed, but if now there’s work that’s going to spill out because it’s too much for the host node CPU, then the question is, what CPU should those fall on and where should they live? And so, there is this term of art that has come up lately, agentic CPU, which is more about, “Hey, these are CPUs that are dedicated to doing this agentic work, all this spillover work that the host CPU could do, but it’s too busy keeping the genius fed.” So, there should be racks of CPUs that are dedicated to this task. But of course, it then raises the question, well, what should those CPUs look like? Should they look like the host? Should they look like the assistant? Or should they look more like the general purpose or do they need to have their own shape? So that’s kind of the framing for this new kind of middle ground agentic CPU.

Vik: Agentic CPU. So you don’t want to use the host CPU, like you said, to do any of this work that the genius is asking, like, “Hey, go search the internet.” No, because once the host node goes off and does something like this, it stops feeding the GPU and the genius and then the whole thing goes down. What’s the point? Now the genius is idle, which is the worst case scenario. So you need a different kind of CPU that does this stuff. It doesn’t even have to be a different kind of CPU. In function, it’s a different CPU. You could use the same CPU, like the host node CPU. It may not need the coherency that we spoke about because it’s not talking to the genius all the time. But otherwise, it could be the same CPU. But ideally, what you can think of this is, let’s take a CPU with, let’s say 128 cores or something. So what you can do is you can think of all the cores as a floor plan of an office building, and then you can kind of draw little boxes and say, “Okay, these four cores are the finance department. These four cores are my research group. These four cores are my facilities team,” whatever it is. And so what you do is then somehow you’ve got yourself a little company and the genius is saying, “Hey, I need this stuff to be done.” And the host CPU is going to yell across the room and said, “Hey, finance team, I need you to go and scrape up the SEC filings from last night, go and do that.” And those guys will be like, “Okay, cool, we’ll do that” and the team of four cores will go off and be like, “Okay, I’m going to do this stuff” and they’ll report back to the host node CPU, the assistant, who will then feed it to the genius, the GPU.

So, in this scenario, a multi-core CPU that is either one of AMD’s 256 core or Intel’s more recent 288 core CPU. These are great because you can assign a lot of little departments across doing different functions, or you could have basically bigger teams of people. So if you have more cores, you can put like eight people in the finance team versus four people in the finance team. So they kind of get the job done faster. But also you want each person to be competent. So each core should be fast. It should not be ultra slow. You don’t want a team of eight interns versus a team of eight experienced people. Single core performance matters because if you get very slow single cores, it’s like you’re putting eight interns on the job, which is fine. Depends on the task. I’m not against interns, but it really is okay, but it depends on what the task is.

Austin: Totally, totally. And this analogy is good because for example, your host CPU, the one that’s feeding the genius, it obviously it may have something like 88 cores, but it may be super fast and those cores may be dedicated to keeping the genius fed. Now all of a sudden if you take that same CPU and you move it into this agentic situation, you might stop and ask yourself, “Oh, do I want on my little floor plan with all my cubicles, do I want 88 cubicles of really fast thinkers? Or sometimes would it actually be better to have 128 or 256 or 288 or 512?” And it reminds me of working at lots of other companies where you look around and of course, you’ve got some management tier and senior architectures and stuff, but then you also have a lot of fresh out of college hires that are maybe cheaper and maybe they work a little bit slower, but guess what? Some of their tasks may be like, “Go fetch this from the web and it’s going to take two or three seconds to respond and you’re just going to sit there anyway.” So maybe it doesn’t matter, and then you go process and you process a little bit slower.

Vik: It’s about the cost per employee now. How many cores can you get and how expensive is each employee in that floor plan? That’s important. So you want to have the most capable employee and many of them at the lowest possible cost of acquiring them.

Austin: Yes.

Vik: And that has to be suited to your workload. You don’t want to put really, really inexperienced people on a very complex task or very experienced people on a boring task. So when people many times ask what’s the best agentic CPU or what’s the best host node? I think host node CPU you can kind of tell what it is. And we’ve both written about this in quite some detail on our Substack. Host node CPU you can actually tell. You need coherency, you need speed, cores are not all that important, but the single core performance and getting stuff to the GPU is very important. So I think you can identify a host node CPU when you see one. However, like you mentioned, the whole floor plan of employees that the CPU cores are, and you’re going to fill racks of them in a data center, there is no reason you can’t fill several racks with different kinds of chips. You’ve got some high core, but low core speed chips in one rack, when you’ve got really fast chips, but not as many cores in another rack. And the whole problem now comes down to how do you architect your workflow to the workload that you are going to be using it for. I think that’s very important. So there’s no such thing as the right CPU. I think they’re all right CPUs. It’s all about the cost that you get per core, the total cost of ownership of the chip and how you use it together. So that co-design and co-optimization is very important and I spoke to an Intel speaker also at Taiwan and that was a very good talk and he explained that they actually, if you go to Intel, they actually do explain to you what the right CPU configuration should be for a given workload. So they have recommendations that they provide for these kinds of things. So it’s very important.

Rack Scale and Orchestration

Austin: Yes, yes. And this is why Intel had recently launched this past summer a P-rack and an E-rack for agentic AI, trying to make the point that it’s not necessarily one size fits all and there may be particular workloads that you can map to needing more performance even for agentic AI tasks or just wanting as much efficiency and as many workers as possible. There’s a lot of analogies with org design and I’m thinking back to companies where it’s like, “Oh, budgets were bad and the year was bad, budget was tight, so they had to lay people off.” And so then the question is, do you lay off the junior workers where they’re all really cheap, so you’d have to lay off a lot of them, or do you lay off middle management where they’re expensive? And a lot of times it’s like lay off the middle managers. And this kind of reminds me of saying, “I don’t need the host node over here. Actually, I just want a bunch of maybe junior employees or early career employees that are essentially cheaper but can still get the work done.”

Vik: And then you’ve got the general purpose employee, general purpose, general purpose CPU. I even lost the analogy now. Sorry, I’m in office mode now. I’ve even forgotten we’re speaking about CPUs. But the general purpose CPU will be like, I don’t know, it’s just good for everything. Sometimes you’ve got these people who just have the skill to do quite a lot of different things. Those are also can be valuable. So it’s the agentic CPU is one thing, the host node CPU is another, and then you’ve got general purpose CPUs. These are just people you need.

Austin: Right, right. Which the receptionists or the people who maintain the building, all important tasks, but you can’t do without them. Totally. I will say at hyperscaler scale, they tend to still have particular workloads in mind and therefore they will buy a general purpose cloud SKU, but still with a particular shape, like we know this is memory optimized because we’re going to have in-memory databases running here. And so what will be interesting, whereas if you’re just an enterprise, you might buy more sort of generically shaped like, “Yeah, it’s kind of fast enough and it has enough memory and it has enough compute that we think it can do a broad set of tasks.” So it’ll be interesting to see also how agentic tasks and agentic CPUs, if there’s any difference at the enterprise level, you know, can I buy a one-size-fits-all agentic CPU rack versus obviously hyperscale. But now that I won’t even go there because then you start to ask, are enterprises really going to be buying additional racks of CPUs or are they going to be still leaning on the cloud here? How’s this going to play out?

Vik: Rack scale ideas is quite interesting too. I was at—this is a different kind of a—I realized this is a slight tangent, but I just wanted to mention it because yesterday, considering where we are recording and when we’re recording this, yesterday was the Cerebras announcement of their new rack scale solution. And so we always thought of Cerebras as a wafer scale chip, right? A wafer scale chip. But they’re saying, “No, the next unit of compute could be putting them into a rack.” So you think of this as a rack scale solution. So even CPUs could be filling into racks and that’s nothing new. It’s been part of the cloud data center for for decades now. That’s how CPUs used to be filled into racks and they used to do the compute. So now you’ve got a rack of GPUs, you’ve got a rack of Cerebras GPUs, you’ve got a rack of P-core, like Intel calls it, or a rack of E-core. All of these have different capabilities, like fast latency, like this genius’s specialty, the Cerebras genius’s specialty is just speed. It is only speed. This thing, this genius can’t remember. This kind of a forgetful genius, but it’s very good and very fast at doing stuff. Then you’ve got the other kinds like the large HBM based accelerators. Those geniuses have a little bit more context. They’re not like ultra fast. They’re very smart, but they kind of have larger context. They’re a little bit more general purpose genius, not just like genius. And then perhaps you could have the slow genius, the slow thinker. I’m just pushing this analogy because what if you really don’t need that token speed in the genius. You don’t need this fastness of the genius. You just want the genius to think for a long time and come back with whenever the genius has an answer. This is definitely a workload for science and medical problems or solving cancer or something. It’s not like you want the answer tomorrow. We would all like it, but it would be much better if this genius could think for a very long time and come back with a nice answer that we could all work with. And without blowing the budget because you can’t say, “I will put Cerebras on solving cancer tomorrow,” that Cerebras kind of high token speed. Maybe it will. Maybe that’s what it takes because it’s a hard problem. But maybe sometimes you’re like, “I just want to study weather patterns and the inferencing can go really slow. I don’t mind. Weather is slow anyway.”

Austin: And when there’s like overnight jobs that could be like, “Hey, go look at all my transactions from today and summarize them and write them in some log or something.” You might want some intelligence where you can’t write deterministic software. I mean, that use case you probably could write deterministic software, but you might want to extract some insights first from all of those transactions and write that to the log as well. You would need some intelligence. Frankly, that could be an LLM that runs on a CPU too. If you’ve got a cluster of P-cores, which then it makes me ask the question about orchestration, which is like, do we have the right orchestration software to schedule across super fast geniuses and regular geniuses and even CPUs, or is there actually opportunity?

Vik: We’ve obviously got Nvidia Dynamo and stuff like that, but even a layer higher. The Nvidia Dynamo thing is all about how to make the genius work better, but that’s not the orchestration layer.

Austin: Yes.

Vik: So, it’s yes, so we do need a solution. I think Modular is one such orchestration layer, if I get this right. But maybe more so for the genius still. I don’t think it’s going to do it over everything or I’m not entirely sure, but—

Austin: Well, we should talk to the Modular people. You can write with Mojo and it can run on CPUs or GPUs all with one programming language, which is pretty sweet. But I don’t know a ton of details yet about their orchestration level capabilities.

Vik: But the one thing I understand is that they can talk to disaggregated hardware or different kinds. So you can mix and match various pieces of hardware for inference. We’re still talking about the genius. But even that, you could have orchestration like do a little bit of tokens here, deal with a different architecture here, like can you mix an AMD rack and an Nvidia rack all together in a data center and have a software platform that does with all this inferencing. But then yeah, ultimately then you’ve got to have a software layer above all of that. Maybe it does exist and we are not entirely software guys. But that platform will orchestrate how the CPUs and the GPUs interact with each other and how the data moves between all of them, a very, very complicated problem.

Austin: Totally, totally. And I know Gimlet Labs, who I talked to on this podcast, I don’t know, back in May, I believe, they were kind of working in this same space, which is as a neocloud, could you ultimately have lower costs by having different hardware and being very good at scheduling it across the correct so that the correct slice of work gets done on the correct hardware.

The Coming CPU Demand

Austin: But, you know, all these are topics for another time. I think this is probably a good place to call it quits. So we talked about CPUs, we talked about agentic AI, where should it run? Oh, let’s circle all the way back to Grok bot. So, where do you think your Grok bot VM lives?

Vik: It lives in some a general purpose CPU somewhere, I think, because this general purpose CPU is just running an OS somewhere and it has a little computer that’s made only for me. And I go in there and it’s the only job of it is to do what an operating system does. It’s like, “Hey, go access the memory and here you go, I’ll get data from the internet” and all of this stuff. This is not entirely agentic per se, but it’s just a computer. It’s a computer like the one you’re on watching this on your phone or your laptop or whatever. It’s just a computer. And that computer accesses a whole lot of other hardware. That mini virtual computer you have has access to the genius. And it has access to all these host node CPUs, which the assistant that’s feeding the GPU. That entire office building of stuff is accessible to my little virtual machine. And I can spawn as many agents as possible from my little VM and send them off to do various tasks. So it’s a nice approach I think. And it’s going to cause more demand for CPUs.

Now, because you’re opening up not only the layer that was previously inaccessible to people because not everybody could actually install open claw. Now, if more people can use tools like Grok bot—I’m not again trying to say this is one particular product. I think more will come out like this. It opens up the AI world to a lot more people, the agentic AI world, which means that if more and more people start—even regular people, not the high-tech AI using token maxing crowd, if the regular old people who never really wanted to use AI, but now find use in it, start using this, remember how many VMs are going to be apportioned for each of these people who use this tool or tools like these. Then you’ve got those little VMs spawning off so many calls to CPUs and GPUs. Imagine, each one can run 10, 50, 100 agents. There’s a lot of hardware demand. What can I say? There’s a lot of hardware demand.

Austin: Yes. Absolutely. I think 10 million people doing this is not crazy. And that could be 10 million VMs running on 10 million cores in the cloud. And then if people are kicking up, let’s just say whatever, 100 agents per VM, all of a sudden you’re at a billion cores that are needed. It’s not that crazy to imagine that world.

And that’s coming quickly. And so, and maybe I guess to contrast that with the Mac Mini craze, yes, Mac Mini was sending a lot of new requests for tokens from GPUs, but also and so that was also selling lots of host node CPUs, but ultimately a lot of that agentic work or whatever is running on the Mac Mini. But this, the easy button, the fast path as you push that button and it spins up a VM in the cloud. And to your point, I do think that’s where really we’re going to see a lot of proliferation of agentic AI to the normal crowd, to normies. And therefore, it’s obviously going to be good for server CPUs.

Vik: Agree.

Austin: Totally. All right, let’s wrap there. We’ll check back in the future. Anyone who has any interesting thoughts or corrections or interesting orchestration stuff that we should be aware of, send us an email, put it in the comments, whatever. Elon and Grok team, if you’re listening, we’d love you to come on and explain how it works for us. And I think one maybe last little point is, the question is, how much of this stuff is actually running on agentic CPUs versus VMs on general purpose machines? And because these are newly marketed things, I would love to hear from Anthropic or OpenAI, how much of if they would be willing to share, how much work are they actually putting on agentic CPUs versus just still regular general purpose CPUs and what are the pros and cons? But so much that we would love to learn. So anyone who’s in this space that’s listening, let us know. We would love to talk to you. With that, we’re going to wrap here. Thanks, everyone.

Ready for more?