LALatent SpaceSep 2, 2026· 44:03

The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO

A comedic episode parodying AI hardware marketing, featuring Sean Lie discussing fictional Cerebras CS4 chips with absurd specs like 30 billion parameter models running at 4,400 TPS with wafer-scale integration and jalapeño-flavored SRAM designs.

Transcript

Intro0:00

Host0:01

What used to be considered fast at like 100 or 200 tokens per second is quickly becoming the new batch mode. And so I think that, you know, there's a place for that. It's great for prompt processing, it's great for very, very parallel workloads.

Speaker 20:19

But my evals take 20 hours. I can hit a button and run it in 2 hours.

Host0:22

Exactly.

Speaker 20:23

Like, come on.

Host0:24

Exactly,right? But quickly, you know, being able to do prompt processing isn't quite enough,right? That's a big part of the market, but that's not quite enough.

Speaker 20:34

Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content.

We've been approached by sponsors on an almost daily basis, but fortunately, enough of you actually subscribe to us to keep all this sustainable without ads, and we want to keep it that way. But I just have one favor to ask all of you.

The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you, and it means absolutely everything to me and my team that works so hard to bring latent space to you each and every week.

If you do it, I promise you we'll never stop working to make the show even better. Now let's get into it.

Hot Chips1:21

Speaker 21:22

Okay, we're here at Cerebras HQ with CTO Sean Lie. Welcome.

Host1:27

Thank you. Thank you for having me.

Speaker 21:28

Thanks for having us.

Host1:29

Yeah, no, thank you for having me.

Speaker 21:31

And it is the day after Hot Chips, lots of things launching, even like Apple launched their M6 stuff. But you guys obviously had, we were at Supernova, you talked about CS4, we're going to talk a little bit about CS5, OpenAI talked about Jalapeño, all the hot stuff in Hot Chips.

What's your take this year, having been in this industry?

Host1:53

I think, I think, well, first of all, thank you for having me. And you're absolutelyright, it's a super exciting timeright now. Hot Chips is a really good, you know, time when the whole community comes together. And, you know, what used to be, you know, I still remember going to Hot Chips 10, 20 years ago when it's a bunch of like, you know, computer architect, you know, geeks, you know, geeking out on, you know, speeds and feeds.

And now it's like, this is where, you know, industry-changing hardware is being revealed. And so, you know, it's come a long way. It's a really exciting, you know, community. And I think one of the main takeaways that I had after, you know, after leaving Hot Chips yesterday is like, what an exciting time it is for the chip design and the hardware industry.

It's, you know, this is really the golden age of hardwareright now. And, you know, like you said, I've been in this industry for some time. This is a very, very unique time. Not just because, you know, AI is taking off and, you know, and there's, you know, obviously a lot of people doing their own hardware, but the amount of innovation that's happening now across multiple fronts, we've never seen in the history of the semiconductor industry this level of innovation across all the levels,right?

In the chip, the interconnect, the system design, the software, the optics, the, you know, the methodology, the tooling, people are pushing across the board. And so like, you know, as a technologist and, you know, a computer architecture geek at heart, it's like really, really fun to see this community and this industry at this stageright now.

And when we felt that, like at Hot Chips and everybody I talked to, it's like that buzz is there. It's amazing.

Speaker 23:41

We want to go through the chips. Let's, you know, talk about your chip first. What's most exciting? CS4 came out, some crazy numbers here, 30x faster than GPUs, but, you know, walk us through what's new with the new CS4.

Host3:54

Sure. So, you know, we designed our next generation CS4 architecture with a new, brand new system platform with the goal really to make wafer scale mainstream, to make it, to bring it to hyperscale and to solve a lot of the high-density data center problems that, you know, we're facing every single day.

CS43:54

Host4:23

So we've designed this modular platform that provides twice the amount of power to the wafer than we have in our previous generation, twice the amount of interconnect bandwidth, half the latency. And all of it is done at the system architecture level.

And what you see here is the result, is that by being able to provide significantly more power to the wafers, by, you know, packing more density into the rack, we're able to drive up the performance even further. You know, we like to say that Cerebras is, with our current product, already the undisputed leader in, you know,ultra-fast inference.

And what's amazing is this CS4 product will actually push that frontier even further by taking it another two times faster. And this is a perfect example,right? We're here in this demo that we gave at Hot Chips, we're showing GPT-OSS running at over 4,400 TPS, which is just mind-blowing.

It's like, it almost feels like it's fake,right? And we think that this is going to really revolutionize the entire, you know, industry because, you know, not only are we able to, you know, continue to push the frontier on how fast these models can run, it will enable all sorts of new applications, you know, user experiences start to become extremely different,right?

What used to be batch and offline applications start to become real-time. And, you know, all of this, you know, growth and new agentic flows and agentic frameworks now also mean that, you know, you're sitting there waiting for these agents to go over and over and over in their agentic loop.

And all of a sudden, if you're running your model at over 4,000 tokens per second, now you can do, you know, more agentic loops, you can do more reasoning. Ultimately, you get significantly more capable, more intelligent agents. And so this is what really excites us about this, not just the fact that you showed these really, really big numbers, which is also really cool, but the fact that you can, you know, really start to do things that you can't really do otherwise and start to enable brand new capabilities.

That's what makes me so excited about this reveal that we had yesterday.

Speaker 26:51

Yeah, I think it's also very rewarding that you guys started on this journey, the big chip journey, wafer scale, everything. Everything that you've always said, it's just like people didn't take it that seriously until they had to. And then now they're really, really taking it super seriously,right?

Host7:07

Yeah.

Speaker 27:08

I do feel like as, you know, co-founder, CTO, the architect of this whole thing, how are you guys approaching the new generations post, like, AI boom? Like, I imagine CS1, 2, 3 was a bit more sort of calm.

Now, like, literally OpenAI is launching theultra-fast mode with you guys. And like, it is a matter of like, I don't know, like you're co-designing your model with the chip almost.

Host7:36

No, absolutely. I feel like the main shift that has happened over the last few years for us is that, you know, the first generation chip primarily was a technology demonstration,right? We had to demonstrate to the world, to ourselves, that, you know, you can actually build such a thing.

And, you know, we told a lot of people that this is the future and most of them were just like, you guys are insane, this is not possible. And so, you know, at the beginning, it was really about, you know, figuring out the foundational technologies to make it possible.

And what we're now seeing is now that we have, in fact, demonstrated that this can be built and it can have the kind of performance that, you know, that this type of architecture can have, what are all the things that it can enable,right?

And, you know, you alluded to, you know, our recentultra-fast launch with OpenAI. That's a perfect example,right? We're running,

you know, frontier level, one of the most intelligent models,right? OpenAI's largest, most capable, most intelligent modelright now at 14 times faster than their normal, you know, GPU speeds. And it's becoming this completely transformational thing when you can combine the most intelligent models with the speed that the hardware brings.

And this transition from technology demonstration to, you know, solving real problems, enabling new, real capabilities is this transition that we as a company are going through. And how we're coping with thatright now is we're really shifting our focus to ensuring that we can make this scale, we can bring the capacity that we need, we can run the models that, you know, the users ultimately, you know, care about.

And then we continue this flywheel by making sure that we continue to put out faster hardware and keep pushing the frontier. And that's kind of, you know, how things have shifted for us. And it's, I would say, like, you know, unfolding in a way that is, you know, almost better than we could have imagined,right?

It's this perfect confluence of like, you know, having the hardware with the capabilities demonstrated, now matching with the capabilities of the models and the applications. And so our goalright now is just, you know, scale this thing as fast as we can.

Speaker 210:11

And so maybe we'll talk about yesterday.

Host10:13

Yeah, sure.

CS510:14

Speaker 210:14

Previewing CS5. Can you recap for people who are maybe not yet caught up and, you know, maybe are new to the story for the first time as well?

Host10:22

Yeah, no problem. So I mentioned that our CS4 system that we just launched is based off of a brand new rack scale platform that we call the Nexus platform. And as I mentioned, what's special about this platform is the modularity.

It has a modular power supply in the front. It has a modular server in the back that we call the backpack. And this is a platform that enables that 2X performance that we just all saw the numbers for.

But it's also the foundation for multiple generations of products. We've designed it from the ground up to support multiple generations of our products. And in particular, we designed it together with our next generation wafer chip that will be coming next year.

And with that chip, we'll be pushing the performance even further. So we just saw a 2X improvement this year with CS4. We're going to push it even further with another 2X improvement in performance. And what this ultimately means is you'll be able to run, you know, medium-sized models like GPT-OSS or Gemma at speeds up to 10,000 TPS.

And even frontier-level models like Kimi or DeepSeek and GPT-5.6.0, up to 5,000 TPS. Completely game-changing. All again, enabled by this brand new platform that we've designed for multiple generations. And so, you know, this is the trajectory that we're onright now.

And, you know, now that the world sees the value ofultra-fast inference, our bet is that even this is just the beginning.

Speaker 212:06

What's the rollout process like? So you have a deal with OpenAI over the next few years.

OpenAI rollout12:06

Host12:11

Yeah.

Speaker 212:12

When do we get access, you know? So everyone wantsultra-fast. Right now, sure, there's Gemma OSS. When do we get the big Kimis, the DeepSeeks? When can we start to, you know, actually use the bigger models?

Host12:25

Soright now, we are basically, you know, sold out of everything that we're building,right? And we are very strategically making sure that we're deploying every single megawatt in the most strategic way possible. And in particular, quite a bit of it is going to OpenAI,right?

And we've been very public about this. You know, they're our biggest partner, not just because of the commercial arrangement, but also because of the fact that we have this co-design, you know, spirit, which you guys have heard OpenAI talk about this a lot as well,right?

Which we believe will basically allow us to kind of continue this flywheel of not just making models, you know, faster, but using that speed to make them more intelligent and so on. So today, a lot of that capacity is going into OpenAI.

And within OpenAI, they have very strategically decided to use quite a bit of it for themselves. So they're using itright now.

Speaker 213:28

Like internal usage.

Host13:28

That's correct,right? Soright now, internally, they're using it for a lot of really critical use cases where the speed really, really matters. Like they're using it in like their incident response teams,right? When there's an outage in their service, for example, every single second, every single minute matters.

And so they're getting a tremendous value out of that. They're using it in some of the most critical research applications where the extra reasoning is enabling significantly, you know, more intelligent responses. And soright now, that's quite a bit of the capacity is going there.

And then as we mentioned in our launch with OpenAI, we are now also making it open to, you know, to enterprise customers who are, you know, who are able to, again, use it for some of the most demanding and, you know, the most high-value applications.

Over time, OpenAI and Cerebras, we have committed to, you know, to bringing enough capacity to makeultra-fast inference available to a much, much wider audience. And, you know, and our CS4 announcement, our CS5 announcement, these are very much part of that commitment to continue to drive more faster tokens and, you know, basically more throughput to be able to satisfy, you know, all the various use cases out there.

Speaker 214:54

By the way, you know, the way that you framed it, I realized that they didn't have to expose it to enterprise customers.

Host15:00

Yeah.

Speaker 215:00

They could have just kept it for themselves.

Host15:01

Yeah.

Speaker 215:02

Nobody knew. That's an interesting, like, business decision almost. I don't know if you have any weight on that. That's more like a business analyst point of view.

Host15:09

Well, I would say that, you know, obviously I can't, you know, explicitly comment about their thinking,right? Ultimately, it's OpenAI's decision how they want to use this. But I think if you look at it from the outside, it kind of makes sense,right?

I mean, their mission very much is to continue to, you know, push what's possible. And by using it internally, that helps them do that. But they're also a business now,right? And there's, you know, talks about them going IPO and all this.

And so very much, even externally, you can see they're balancing both of these,right? And I think we very much see that playing out in our, you know, in theultra-fast space as well.

Speaker 215:47

I mean, they're balancing a lot, you know, most recently their hot job.

Host15:50

We got to talk about this as well.

Speaker 215:52

We got to go into them.

Host15:53

Yeah, absolutely.

Speaker 215:54

What are we thinking? Prefill, jalapeño, decode, Cerebras? What's the, what do you think?

Jalapeño15:54

Host15:59

I think that's a very, you know, rational, you know, conclusion. I think of all of the hot chips announcements, probably jalapeño was the most exciting, but to me, but not, maybe not for the same reason as everyone else.

Speaker 216:21

I guess to double click on that, you're the expert in chips,right? There's people see token per second. What do you see as interesting when you see that?

Host16:29

Yeah, so this is what I was going to say. I think that, like, they, you know, they pushed a lot on the performance and the fact that they're significantly better than, you know, better performance than the GPU, than NVIDIA.

But what I see is that they've built a significantly better GPU. And that in its ownright is very, is a big achievement. NVIDIA knows what they're doing,right? They own the market for a reason. They're not dopes. And so to be able to come out of the gate and build a significantly better GPU is a big achievement.

But to me, the reason why jalapeño is so exciting isn't even these, all these Paretos. It's really the design methodology behind it,right? They very clearly took a very drastically different approach to building this chip,right? Having an AI-first methodology enabled them to, you know, build the chip faster and achieve some of these very impressive results,right?

And that is 100% the future of our industry. And it's not surprising to see that OpenAI is kind of leading the way here because this is, you know, very much their MO. But coming back to, you know, what you said about, you know, how we're going to use this,right?

I think that, you know, jalapeño is pushing the boundaries for both, boundaries of what's possible for both throughput as well as latency. And for me, even if I take the Cerebras hat off, I think that's awesome for the industry,right?

This is going to lift all the tides. Everyone is going to benefit. You know, they've also now been able to, you know, push the latency into regimes that traditional GPUs can't hit. And that's also great because, again, we believe in speed.

There's a lot of value there. And what's going to be really interesting and what I'm super excited about,right? And, you know, OpenAI is our biggest customer. And so this is one of the things that, you know, I think is very strategically important to us is that when jalapeño is available next year, when our next generation CS5 is available together,right?

We will enable a full-fast inference portfolio,right? That is substantially different and better than what's already available today, which is already substantially different than your baseline GPU. And then on top of that, there's opportunity to integrate even further, like prefill, decode, disaggregation, for example, or other forms of disaggregation.

Speaker 219:07

Which they didn't actually specialize jalapeño for,right?

Host19:11

They did not specialize jalapeño for that, but they specialized it for throughput,right? And so they get a tremendous amount of throughput. And so just as a computer architect, there's like so many different things you can start to do with that,right?

And then we have, you know, the, you know, insane latency,right? Like I mentioned, up to 10,000 TPS in CS5. You can start to imagine some really, really cool products that we can build together,right? And that's what really excites me about having them as a partner.

And we're also excited about collaborations on the AI tooling front because much of the benefits that, you know, that they're seeing from the AI tooling infrastructure, everybody probably says this now, but like, we're obviously doing, you know, a lot with AI, but we're also collaborating very closely with OpenAI,right?

To use their tools to help us also continue to push what's possible in our chip design, in our software, and all that. So both of those together, I feel like, you know, this is a really unbeatable combination.

Speaker 220:12

I think one thing that people are talking about, you know, I'm trying to get to the disagreements of the hot takes now.

Performance per watt20:12

Host20:17

Yeah, sure.

Speaker 220:18

So people are focusing on, you haven't really mentioned like power. And I do think that something that seems to be a consistent theme is, you know, performance per watt rather than tokens per second. Any variation of this theme or what were people sort of talking about offstage that, you know, is more contentious?

Host20:38

Well, I think there's a few things that, you know, in terms of, you know, some of the more contentious things, like I would say one of the themes that came out quite a bit in my discussions at the conference was around the Groq announcement,right?

You know, Groq, obviously not surprising, they have a new chip. In general, I think it's awesome that SRAM designs are becoming, you know, more mainstream now.

Speaker 221:04

We've been here the whole time.

Host21:06

We've been talking about it for a long time. And it's amazing to see, you know, the industry starting to embrace it,right?

Speaker 221:13

Do they feel different post-acquisition? I mean, you've been competing with them for a while.

Host21:19

To first order, no. I think it's really awesome to see that, you know, SRAM architectures are becoming, you know, more accessible. You know, even the biggest of the big guys here, NVIDIA is embracing SRAM design, acknowledging that, you know, that the traditional GPU designs really can't hit theultra-fast, you know, regimes.

A lot of what was being discussed offstage was, I mean, the natural, obvious questions is like, why did they launch on a 30 billion parameter model? And, you know, how come when Jensen spent so much time at GTC talking about attention, FFN, disaggregation, there was no mention of that?

And so I think there's, you know, I think that's pretty telling,right?

Speaker 222:07

Maybe this is separate Rubin, you know, strategy.

Host22:10

Well, so there's Rubin, but, you know, the LPX itself.

Speaker 222:15

That was supposed to be where it was?

Host22:17

It was supposed to be Rubin LPX together,right? If you guys recall, I mean, Jensen spent like.

Speaker 222:24

GTC, yeah.

Host22:25

GTC like half an hour explaining attention runs here and, you know, and the MOEs run here and so on. And I don't know if this is like a hot take per se, but, you know, it's very suspicious that

their product that's in full production, they've only shown performance numbers on a non-disaggregated 31 billion parameter model,right? And I think to me what this just shows is that there's definitely some challenges in running on a non-wafer-scale SRAM design because there's just not enough memory in each of the chips,right?

And I think that's, you know, that's what's happening and we're seeing the evidence of that. And in many ways, I think it's very much validating kind of the design choice that we had,right? If you think about it, to run a frontier-level model like, let's say, a few trillion parameters, you need thousands and thousands of Groq LPUs just to hold the weights,right?

And so when you start to think about it that way, it's like, well, is it surprising that the only performance numbers that they're showing are, you know, on 30B,right?

Speaker 223:38

I mean, so let me see if they're going to gradient distance to this.

Host23:42

Well, I think it's the other way,right? I think what's going to end up happening is they're going to end up focusing on significantly smaller models,right? You know, if you have that limitation in your architecture, then I think that's what ends up happening,right?

Whereas in our case, you know, we're running the world's largest models and in some ways kind of simple because, you know, one of our chips has, you know, order a hundred times more memory than one of their chips,right?

So you got two orders of magnitude difference in scale kind of for free,right? And so I think that's definitely one of the things that was a topic of discussion, again, independent of Cerebras, just kind of odd that, you know, you'd launch a brand new product on performance numbers of such a small model.

But when you peel it back a little bit, it kind of makes sense. I mean, I've been living in this space now for a long time. There's a reason why we needed the wafer-scale integration to be able to aggregate enough SRAM to be able to actually make it useful for large models.

Speaker 224:44

Yeah. And look, the market's large,right? You have a different market than them. And, you know, you clearly are the longest running incumbent now in this space.

Host24:55

No, absolutely. I mean, the market's large. There's a lot of different opportunities for, you know, different hardware to play different roles. In fact, in general, you know, at Cerebras, we believe very, very strongly in a heterogeneous disaggregated, you know, ecosystem,right?

Not just prefilled, decoded, disaggregation, but, you know, I think we're just at the beginning of what's possible in this space. And with the scale of these deployments and of these models,right, inference is no longer just like one workload.

There's many, many kind of sub-workloads within it. And, you know, you really want to use theright, you know, tool for the problem. You really want to use theright hardware for the problem. And so absolutely, I think there's a spot for all the different types of architectures out there.

And, you know, we've chosen to target, you know, the frontier.

Speaker 225:50

Yeah, frontierultra-fast.

Host25:52

Exactly, exactly.

Speaker 225:54

We was going to go into some of the other companies that we're, you know, talking about this year.

Host25:58

Sure, sure, sure.

Speaker 225:58

But it gives me, this gives me an opportunity to follow up on one thing, which is how do you think about the classes of workloads,right? To me, theultra-fast frontier workload is just like very clearly like one of the fastest growing segments of the entire inference market,right?

Like the fact that like I want it and I cannot pay enough for it. It might have been a mistake for OpenAI to offer it even,right? But it's good for Cerebras. Anyway, so just in your experience,right? You say you're strong believers in a heterogeneous inference solution.

Workload buckets26:13

Speaker 226:29

Okay, like what are the buckets and how do you see it from your talking to your customers?

Host26:36

Well, so I think there's probably two different views here. You know, one is the product view and then the other is kind of the technical computer architect view,right? From the product view, it's actually just very simple. It's that bringing more speed opens up significant opportunities, different applications, different use cases, different capabilities, more intelligent models, more intelligent agents.

And so the further you can push that, the more and more that you can enable. In fact, there's probably all sorts of things that we can't even imagine that you can build, you know, when you're even faster than what we are callingultra-fast today,right?

Which is why we keep pushing that. And then the other angle really just comes down to capacity,right? That's, so it's like, yes, I want the fastest, but then you need enough to actually be able to serve your use case,right?

And so it really just is that simple,right? And then finding theright architecture and theright mix of architectures,right, to be able to provide that is really the name of the game. Now, from the kind of computer architect's point of view, in a lot of ways, it's even a little bit simpler,right?

Like I was having a conversation with somebody actually at Hot Chips about this and they were asking me, well, you know, if you're doing disaggregation, like doesn't that mean you have to like partition the data center and you have to deploy a certain amount of this kind of hardware and, you know, different type of a certain amount of another type of hardware?

Isn't that restrictive? And I'm like, I mean, you know, when we design it.

Speaker 228:05

It's modular.

Host28:06

It's modular. And when we design a chip every single day, we're deciding like, am I going to use the silicon real estate for memory or if I'm going to use it for computer or if I'm going to use it for I/O?

And so from a computer architect standpoint, what I see,right, is there's significantly different parts of the workload,right? The simplest ones are, okay, there's prefill, you know, decode, but then you go one level deeper. It's like, well, what's actually happening during these phases?

Oh, well, there's, you know.

Speaker 228:34

Loading.

Host28:35

There's loading the actual KV cache. There's doing the actual, you know, attention. There's spreading the experts across the various different hardwares and balancing them. All of these things can be now addressed with kind of more tailor-made, you know, hardware solutions and architectures.

And at small scale, it doesn't really make sense to bring together a bunch of different hardwares just to solve like one problem. But at the scale we're all talking about, the hundreds of megawatts, the gigawatts, the multi-gigawatt scale, it easily pays off.

And then at that point, it's as if you're like thinking about the entire data center as if it's like one computer. It's like one chip that you're trying to figure out, okay, well, I want this amount of this capability so I can run, you know, attention fast.

I want this amount of capability so I can run prefill really, really efficiently. And being able to piece all these things together is like the computer architect's like dream to have all of these tools in our toolbox,right? So that's how I think, you know, where you get the value from this heterogeneous disaggregated ecosystem.

Co-design29:42

Speaker 229:42

How much of that plays into model design, model architecture? So working with OpenAI, you get to hopefully see what's going on on models. Like there was the whole MOEs were a thing, reasoning models.

Host29:54

It's not enough. There's also voice diffusion. What if there's more interactive models like the thinking machine stuff?

Speaker 230:02

Absolutely. And I think as we start to get into more interactive use cases that we enable throughultra-fast, we're even seeing a lot of this,right? So if you're trying to interactively, for example, do graphic design, now you mentioned diffusion.

Now, you know, image generation or video generation might be kind of now in line in your, you know, creative like inflow interactive use case with the thing. And so there's so many of these different use cases that are coming.

And what I, you know, you mentioned about co-design with OpenAI and so on. What I think is

the most untapped opportunityright now, frankly, for Cerebras, but frankly, for the entire non-NVIDIA environment,right? And the non-environment, sorry, community is that, you know, we're all running models that were designed for NVIDIA GPUs,right? And it's not just NVIDIA GPUs, but like usually they're designed for like one particular NVIDIA GPU,right?

Like, okay, this thing was designed to run on B200, GB200, MBL72,right? And, you know, and so at Cerebras, for example, here we were showing that you're running the model like 14 times faster already, but it was a model that was designed for a completely different architecture.

And so if you start to then open up the possibility of adjusting that model architecture even slightly, you can get massive gains. And if you start to do even more co-design, I think, you know, the opportunity is almost limitless.

And that's really what excites me a lot about, you know, about working closely with customers and partners like OpenAI.

Host31:46

Yeah. And we don't have time for this because we want to move on, but the people on Earth are really sleeping on AI code gen for kernels, which makes actually, which is very good for you guys.

Speaker 231:57

Yeah, no, absolutely. I'm a huge believer in that for sure.

Speaker 332:00

I think before getting into Etched and super spicy chips, any takes on AMD, NVIDIA, what they're doing? Could they be doing anything different? Any hot takes there?

Host32:10

I think that NVIDIA is doing exactly what we all expected NVIDIA to be doing,right? Like I said, they're no dopes. They know exactly what they're doing and they're continuing to push throughput. And in many ways, it's absolutely theright thing,right?

Because we need more tokens, we need them cheaper. That's a reality. However, you know, what used to be considered fast at like 100 or 200 tokens per second is quickly becoming the new batch mode,right? Quickly becoming the new like overnight.

It is true. And so I think that, you know, there's a place for that. It's great for prompt processing. It's great for very, very parallel workloads.

Speaker 232:55

But my evals take 20 hours. I can hit a button and run it in two hours.

Host32:58

Exactly.

Speaker 232:59

Come on.

Host32:59

Exactly,right? But quickly, you know, being able to do prompt processing isn't quite enough,right? That's a big part of the market, but that's not quite enough. And so I think all of the traditional architectures, I feel like are very much all on that treadmill,right?

In a good way,right? Not necessarily a bad way. It's in a good way,right? And so AMD, Trainium, in many ways TPU, like all of these in my mind are all trying to build a better Rubin,right? And there's a huge amount of value in that.

But I think that there's also a lot of value in trying to push, you know, the boundaries in other vectors as well, which is obviously, you know, what we're trying to do here at Cerebras.

Speaker 333:45

Okay. Spicy, spicy chips. There's this page that you have on your website that this company Etched just doesn't have. No numbers, no benchmarks, any takes. That's their one criticism. But what's Etched?

Etched34:01

Host34:01

Well, so they have said very little about what they're doing.

But what they have said has changed also a lot compared to what they originally said.

Speaker 234:12

Which they acknowledge.

Host34:13

Which they acknowledge. And, you know, the environment is changing a lot. And so that makes sense,right? They're going to pivot. They're going to try to find their space. I think my reaction when I see pictures like this is that it's very impressive graphics design, but I also don't see them building anything beyond just or trying to build something better than just, you know, a traditional GPU,right?

You know, they've made claims about having a distributed storage that is somehow faster because it's distributed, which I don't quite understand. You know, it'd be very helpful if they can explain those kinds of things. I think, you know, it's hard to see where the differentiation is that they're trying to push.

I think they had a rack or something at Hot Chips, I think. And so clearly they're building something, but it feels like it's a little bit, you know, far off for now.

Speaker 235:14

Yeah. But at the same time, like $20 billion now, that's like, you know, not a pushover. I think maybe the steel man for this is like what interesting direction would you want to see them pursue,right? That in the best case.

Host35:30

In the best case, I would love to see them or others, frankly, push the boundaries of more innovative solutions outside of just the chip,right? You know, I really believe that the next kind of frontier of taking AI compute to the next level is all about better ways to integrate outside the chip,right?

Speaker 235:58

The rack, the node, the designer.

Host35:59

The rack, the node, the package, the,right?

Speaker 236:03

The interconnect.

Host36:04

The factory.

Speaker 236:06

The factory.

Host36:06

They may be doing something like that. It's hard to tell,right? But I think that this is not just a point about Etched. I think for the entire industry,right? Even if I take my Cerebras hat off for a second,right?

What excites me the most is our ideas and creative thinking outside of the single chip. We all know that the applications, the problems, the models are now so large that you can't just do anything on a single chip.

So everything comes down to the integration. Everything comes down to how do you bring together more compute? How do you bring together more memory? How do you bring together more interconnect,right? And so, you know.

Speaker 236:48

Scheduling, pipelining, all those algorithms.

Host36:50

All of those things,right? And so, you know, of the, you know, the reveals yesterday, I think probably one of the more interesting ones to me, I mean, it wasn't new news, but was DMatrix,right? The fact that they're really leaning into DRAM, you know, 3D DRAM packaging.

And like this is the kind of stuff that I think we as an industry need to be doing much, much more of.

Speaker 237:19

Yeah. I was going to say like your backpack stuff reminds me of what the memory people are doing. Is that an analogy that?

Memory & DRAM37:19

Host37:25

There's a little bit of that,right? Because, you know, what we're doing with our backpack and power delivery is very much, you know, a very compact 3D package,right? Where we can bring power directly to the face of our chip, of our wafer,right?

And in many ways, like, you know, the DRAM 3D stacking integration is doing something similar, but not with power, but now with memory,right? Yesterday, we also announced that, you know, we also started two years ago our DRAM stacking program for the very same reason,right?

And because I think that figuring out ways to do this kind of integration is the key to unlock the next major step forward,right? And so I don't know exactly what Groq is doing, sorry, what Etched is doing underneath the covers, but, you know, when you ask like what is the stuff that excites me the most about what we're seeing in the industry, it's really more creative, you know, integration, more creative packaging, more creative outside of the chip thinking.

Speaker 238:31

Yeah. Let's leave some hints for people. Other than DMatrix, a couple other names that stood out to you? Just.

Host38:38

Along these lines, I think, you know, Samsung has some discussions around their ZHPM,right?

Speaker 238:45

Okay. All memory guys.

Host38:47

I mean, look, in the end, there's only three main things to building a, you know, an AI chip,right? There's the compute, there's the interconnect, and then there's memory,right? And so, and more and more importantly, as we all know, for at least in the low latency space, the memory is the key,right?

The memory bandwidth is the key.

Speaker 239:09

And power and cooling not as major still.

Host39:12

No, absolutely. I mean.

Speaker 239:13

Just the redesign.

Host39:15

Those are what's powering all of this,right?

Speaker 239:17

Yeah. I'm just saying, you know, there's some physical limits.

Host39:20

But that's actually a really good point,right? I think that when we originally started with wafer scale, most people thought the biggest challenges were, okay, how do you connect all of these die together? How are you going to yield this thing?

And those were absolutely kind of fundamental challenges that we had to solve. But what turned out to be, you know, the biggest enablers was ultimately how do you power it? How do you cool it,right? How do you build it reliably at scale?

And I think that a lot of the 3D DRAM technologies are going to go through the same thing. And it's, you know, one thing to draw, you know, a PowerPoint slide with a DRAM and logic, and then it's another to actually make it work, to figure out how you're actually going to power it, how you're actually going to cool it.

And that's why a lot of what we've been doing is, you know, is trying to leverage and use all of the expertise that we've built up there and put it into our DRAM stacking approach. And so, you know, we already have solved yield at scale, for example.

We've already solved how you can actually package in a three-dimensional way. And that's, you know, problems that Samsung, that DMatrix, and everybody else are also going to have to solve over this time.

Speaker 240:40

Thank you for those comments. Last topic because we got to go. US supply chain, semiconductor supply chain stuff, you know, like China is slowly becoming completely independent of us.

Host40:51

They are.

US-China supply chain40:51

Speaker 240:53

GLM literally just like they're bragging. And then currently Huawei and what have you.

Host40:59

Hyped outside of knowing what it was on. Just model is good, doing good in, you know, open router came out, not even running on US chips.

Speaker 241:07

So the TLDR is, are people talking about it? What's what we're doing?

Host41:11

I mean, I think that absolutely. I mean, of course people are talking about it,right? Because like we're in a very difficult, you know, situationright now. The open source model market is 100% Chinese,right?

Speaker 241:28

Not 100%, but.

Host41:29

Almost. Okay. 95%,right? Most of the big models, most of the big open models that are, you know, high quality are coming from the Chinese labs. In some ways, it's great that sharing is happening and, you know, the global community is benefiting.

But obviously, that's a very strategically challenging place to be, to have such dependence on the model. Separately, as is evidenced by this, is like, you know, we have independently all of the infrastructure, the hardware infrastructure slowly being built up in the background to support all these Chinese models,right?

I think it's absolutely a problem that

globally, as well as, you know, the US and as a natural interest, that we have to continue to push the boundaries so that, you know, we can continue to compete and continue to be ahead. It's not easy. These guys know what they're doing.

But it's something that I believe is not solved by, you know, one company or one fab, but it needs government level, national interest level kind of initiative to be able to make this happen.

Speaker 242:45

So I imagine that should take place at Hot Chips. It's like a secret room of all you guys. You're leaders of our industry,right? Like you would be the guys to.

Host42:53

I mean, we have absolutely been pushing for this and we've been very supportive of, you know, any initiatives that will, you know, continue the US dominance in this space. 100%.

Speaker 243:04

Okay. You got to go, but you've been very generous with your time. Congrats on all your success. I mean, enormous. You're IPOed.

Host43:11

Yeah. Yeah. Yeah. Sometimes I still have to pinch myself to remind me that we IPOed. And it was like the largest, like, you know, semiconductor IPO in history. And it's like, yeah.

Speaker 243:21

And still early.

Host43:22

That was not bad.

Speaker 243:23

Still early.

Host43:23

That was not bad.

Speaker 243:24

No, congrats. And look forward to meeting up with you in the future year.

Host43:27

We got to ask, you know, give usultra fast, give us bigger.

Speaker 243:30

Oh yeah, give us bigger mobile.

Host43:32

No, no, I mean, give the peopleultra fast too.

Speaker 243:35

But we can probably get you in line at OpenAI.

Host43:38

Ooh. Ooh. I hope you didn't bias OpenAI too much to, you know, just keep it internal. You know, people need it. We need better models.

Speaker 243:45

It's OpenAI's decisions. We're just providing the infrastructure.

Host43:48

Hey, but give us GLM. Give us DeepSeek.

Speaker 243:50

Just give us our own rack and we'll run it. Okay. Well, thank you very much.

Host43:54

Thank you again for coming. Really, really appreciate it.