Cerebras Systems Inc. (CBRS) Earnings Call Transcript & Summary

August 18, 2026

NASDAQ US Information Technology Semiconductors and Semiconductor Equipment special

Earnings Call Speaker Segments

Unknown Executive

executive
#1

Please welcome Cerebras CEO and Co-Founder, Andrew Feldman.

Andrew Feldman

executive
#2

Welcome, everyone. Thank you so much for coming. It is with such pride and joy we see all of you here today. And to point out that while we have years of experience in building the fastest AI, managing queues out in front of our building is new to us. So forgive us for the delays there. This has been a truly extraordinary last 3 or 4 months for us. The penalty of going public. Let's all stare at that for a little while, and we can understand the cost of going public, right? This is it. Yes. Right. Right. Going public is a truly extraordinary event. It allowed me to see engineers I've been working with for 20 years. It allowed me to know they actually own a coat and tie, we didn't know this. It was a time to share with our families after a decade of work. And it was sort of humbling in the sense that we're now able to step forward in a different way and participate in the biggest of the big leagues and be part of the show. Now we were able to do this in part because we rode this enormous wave. All right. If you close your eyes and think back 3 years ago, right, AI was a novelty. All right? It was sort of a parlor trick. It was something you showed your friends and didn't use at work. And then it became useful. All right. And now it's a necessity. And in that necessity, speed, latency, shapes what we can build. And real-time AI feels like an active collaborator when you're engaged with it. In AI, speed is productivity. And as you know, what we do at Cerebras is deliver the fastest inference speed in the industry. Now when you have that speed, AI responds in real time, users do more with AI, they stay longer, they run more interesting workloads. They solve more interesting problems. Speed makes new markets and allows new ideas to flourish. But until recently, there was a trade-off. There was a trade-off between speed and throughput. Everybody wants speed, right? Nobody says give me something slower. Or said differently, if you really want to punish a naughty 12-year-old, don't take away their phone, set it to dial-up speed, right, for a week, right? This drives good behavior. All right? You had a trade-off. You had a trade-off between smart and fast, all right? If you went fast, you'd go to smaller models, more specialized models. All right. But when speed is no longer a challenge, you can have both speed and intelligence. And then you can move AI to time-sensitive work, then you can move it to new parts of the business. You can all sorts of new things are possible. The underpinning of our performance is our wafer-scale engine. And I can't share with you how many people along this journey told us it would never work. And when it worked, they said you could never make it in volume. And when we made it at volume, they said, you'd never package it, and then JP figured out how to package it. And then they said you only had 1 customer in the government. And then they said after we had a sovereign cloud, they said you don't have a hyperscaler. Now we have hyperscaler. They said, oh, you don't have a frontier lab, and now we have frontier lab. All right. That's the journey, not just of Cerebras, but for everybody who does work, that's not obvious. It is hard. It is challenging. This is at the foundation of our work. But the foundation is insufficient, right? To deliver fast AI right now, you need a chip, you need system and a rack, you need to build massive clusters and you need software to tie it together seamlessly. And a lot of what we're going to be talking about in the next hour and a bit, all right, is how we build systems, how those systems tied together into clusters to deliver industry-leading performance and set the new bar. Now speed is no longer sort of an infrastructure metric. I want you guys to really think about that. All right. For a long time, it was a benchmark. Look how fast we are. But right now, it shapes how products are built and it changes fundamentally the user experience. AI as it becomes mainstream, users expect it to feel like the best software they've ever worked with. They expect AI not to be different. It's going to be in the same category as their favorite software tools. They want to be immediate, fluid and responsive. And when the response arrives before your attention can move, you see in the problem. All right, and you keep building. And as an infrastructure builder when you allow someone to run a software that allows them to build interesting things, you're proud. That's why you build infrastructure. All right. And that's why we do this Okay. Last week, OpenAI announced a first look at GPT-5.6 Sol ultrafast mode. Ultrafast is a new service tier that runs OpenAI's most intelligent models at up to 14x faster than standard. We're testing this with select customers to start what this brings the industry for the first time. It's frontier intelligence instantly, right? 14x speed removes the speed, intelligence trade-off. You can now have both. There is only one place you can get this, frontier speed frontier Intelligence and instantaneous speed. Customers now complete more useful work per second. Now I'm going to show you a little demo because that's what I want to do, Oh, new slide. Keep the CEO on his toes. Reorder the slide deck. This is a cool slide. On the x-axis, you have speed. And on the Y axis, you have intelligence. All right. You want to be up and to the right on both. You want to be smarter and faster, and that puts you up here. And the only thing up here is GPT-5.6 Sol running on Cerebras. In the entire industry, the only thing here. Now I'm going to show you a number. What you're about to see is a speed run of a benchmark called humanity last exam. And for those of you who worked up PhD, this will be depressant because it will be fast. This is a leading industry benchmark, okay? And it's designed to show graduate-level reasoning. And we will be on the left. There are 2,500 questions. We're going to show one to start. And then we're going to go to the entire set. Let's go. Okay. We're done. First question. All right. Now they're done. Now what we're going to do is take the entire exam. We're done. It took 11 hours, 11 minutes and 26 seconds, less than half a day. The other guys took more than 3 days. As my old soccer coach could tell you, this is the only thing I've ever been fast at. So I really appreciate the applause. Okay. What I'd like to do right now is invite on stage one of the leaders in the leading frontier lab, I'd like to invite Thibault on stage. Thibault is Head of core products and platforms at OpenAI. And we'll go through a few questions and hear a little bit about what OpenAI is thinking around AI and ultrafast AI. Thibault?

Unknown Attendee

attendee
#3

It's good to see.

Andrew Feldman

executive
#4

So if you ever think you're having a good year or a good several years, OpenAI success can disabuse you of that fact. When you're growing as fast as any company in the history of capitalism has grown, right? You have a pretty good run. And what I understand right now is that you have 15 million people a week running Codex and ChatGPT agents, awesome numbers. And these products have sort of made it an enormous impact on us. I mean if you walk around our lab, and I think this is true in most companies now, everybody's got a GPT screen up. All right. Talk a little bit about your strategy. Tell us a little bit about how things have been evolving and how you think of the landscape.

Unknown Attendee

attendee
#5

Sure. I mean first of all, great to be here. It seems like we coordinated on the naming of our models in Supernova, Sol Supernova.

Andrew Feldman

executive
#6

We are. We're trying to follow with the celestial theme.

Unknown Attendee

attendee
#7

it sees more coordinated than it is.

Andrew Feldman

executive
#8

Not that coordinated. I have no fear.

Unknown Attendee

attendee
#9

For me, the strategy has been like very simple, like build the best models and then start them at scale. And then recently, we also realized, well, it's all going to be about agents. So you've got to built like the infrastructure you have agents, run at scale, run safely, have like very aligned models and then build delightful interfaces to those agents so that they can do actual work for you. And up until recently, we kind of have these trade-offs where if you wanted to get something done fast, you have to serve like a smaller model and you have to serve it at lower latency. And then you always have like this decision to make up like, hey, maybe I want something done now, so I'm going to use like a Luna-size model or like, okay, we can think a little bit more, so we're going to use Sol and Ultrafast is kind of don't have to make that decision anymore. It kind of gives a glimpse of what's to come, I think reducing the cognitive load and just building something that is delightful, just works as fast when it needs to and then can also run slower in the background, if you don't need at larger scale.

Andrew Feldman

executive
#10

Yes. I think that's super exciting, and we're really proud of our partnership. Tell us a little bit about how Ultrafast sort of fits in with your product vision and strategy.

Unknown Attendee

attendee
#11

Yes. We would love to -- we're not there yet, but we would love for Ultrafast to just be the default. I think it is a glimpse of what's to come.

Andrew Feldman

executive
#12

Write that down investment community.

Unknown Attendee

attendee
#13

As you saw in the videos, it's quite spectacular when you're in front of it. The first time I demoed to someone, they were like this must be fake. You just planted some fake data and like showed me the back end. I was like, no, it's just running on this hardware over there. This is just a little awesome company, Cerebras, they build great chips. So we're pushing very hard on reliability. Obviously, this infrastructure that powers a very large chunk of -- like in the future, like a very large chunk of the GDP must be reliable. But then also we're pushing on the frontier of like how fast we can make it. And this requires rethinking the entire approach. It's not just from the inference. I think the inference matters, but also like the model is like how do we train them? How do we think about tool calls, where do we see that new bottlenecks sort of like emerging. So we rewrote a large part of a stack for Ultrafast and it was a delightful partnership.

Andrew Feldman

executive
#14

That's really cool when sort of your compute infrastructure so fast, it pushes the software guys up the stack and then that now there's headroom and now the hardware can get faster, right, that interplay is really a pleasure.

Unknown Attendee

attendee
#15

When it started to work, the first thing that we did is like we gave ultrafast to the engineers working on making Ultrafast work and then just cranking that loop.

Andrew Feldman

executive
#16

And you told me something back there that when your engineers have a really important problem, they'll ask for Ultrafast.

Unknown Attendee

attendee
#17

Yes. Yes. Everyone wants to access the Ultrafast. We reserve it for incidents, be it like auditors or security incidents. We have those just as any other company and like Ultrafast comes in handy. Top priority projects, we always have Ultrafast like provision for those engineers and then like important research efforts as well get Ultrafast. We don't yet we reserve a lot of our capacity as well like for serving it to our customers over the API. So we don't just gobble up all of the Ultrafast capacity for ourselves.

Andrew Feldman

executive
#18

What do you think is going to change in consumer behavior when they get AI at sort of this speed and this intelligence?

Unknown Attendee

attendee
#19

Yes, it makes you sort of realize that before you were compromising and you had to switch tasks or you have to sort of like break the illusion that it wasn't really real-time and it wasn't operating sort of like at the speed of thought. And with these kinds of speeds and future speeds, this changes, right? You can stay in the loop, stay in the flow. It becomes like this real-time interaction. You can get a lot more done much more quickly. And then you just sort of like every time you switch back to like normal speeds, they're like, "Oh, wow, I forgot how slow this was"

Andrew Feldman

executive
#20

Yes. For those of you who remember, right, I mean speed transformed the entertainment industry, right? We used to go to Blockbuster and then Netflix delivered in envelopes, DVDs, right? But when the Internet got fast, right, Netflix didn't become better at delivering DVDs. They become a movie studio, right? That speed enabled an entirely new creation of new markets. They ended up buying the studio, right? And I think speed when put in the hands of entrepreneurs and developers, it allows them to create and make new things, not just be faster the things they've been building for a while.

Unknown Attendee

attendee
#21

Yes, we're excited that it's on the API because it will enable different kinds of approaches, different types of products that we haven't yet come up with. But definitely, when you're sort of like in the creation, in the flow, it feels like, okay, like now you can try something maybe draw something and then get like a few different proposals like in real time, you can sort of explore the space in different designs, so you can implement entire variations of back-end like all in real time. And it just feels like a new world is sort of like opening up.

Andrew Feldman

executive
#22

One of the things I think OpenAI has been one of the many things that you guys have been absolutely best in the world that is sort of imagining the scale that AI could be right? You guys were out building Stargate when everybody thought it was roughly big, and now everybody said, "We need a 10x bigger, it's too small." Right? Tell us a little bit about how -- I mean, educate us is how an organization thinks about scale in the way you guys do.

Unknown Attendee

attendee
#23

I think we just have deep conviction that, and we had deep conviction that the models would get better. And at some point, you reach a point where you run the model on the GPU and it's providing more value than the cost of running it on the GPU. And at that point, you want to have all the capacity in the world to just like run it. And this will continue to be the case. Like we don't see a slowdown in the level of capabilities that we're able to develop with our models. And so every time we have a new generation of models, we're like, well, like the utility that we produce on the current compute that we have increases and already vastly outpaces like the costs. And then this is a thing that we talk about all the time, right, which is like, can we get more capacity? Can We get more capacity? It's like the demand...

Andrew Feldman

executive
#24

They talk about it all the time. believe him right?

Unknown Attendee

attendee
#25

[indiscernible] fast outpaces the capacity that we have like across our entire fleet. And I don't think there is a reason to believe that, that will slow down or change.

Andrew Feldman

executive
#26

All right. And last question. We're just sort of in the first few months, for 6 or 8 months of a multiyear, your, partnership, and we announced sort of the frontier model at speed as a first look. As you look forward, sort of what are you most excited about? I mean you've got maybe the best view of our industry of anybody.

Unknown Attendee

attendee
#27

What I'm excited about is like it still feels extremely early and maybe going back to like why we were building all of this capacity and why we continue to invest and we were early. Like if you look at the adoption of ChatGPT is like above like 1 billion users right now. But really the fraction of those that use those more sophisticated agentic workflows to help them day-to-day, now those are the numbers that we have published like 15 million. That's like the number we published last week, like we'll probably publish another number end of the week, even more impressive than that. But it's growing fast, but it's still like a very small fraction of that 1 billion. And so it feels extremely early. And there's sort of like this thing right now is like you use it and if you're technical, you're kind of like used to it being clunky. But the illusion is not quite perfect yet. You don't yet have this personal AGI in your pocket that knows everything about your goals, your schedule for the week, like anything that [indiscernible].

Andrew Feldman

executive
#28

How you preferred I mean ...

Unknown Attendee

attendee
#29

Every time I try to write something, I was just like, oh, remember my tone of voice is like I have this somewhere in the file, in a skill, it's like -- that is clunky. So all of that is going to sort of reduce in something much simpler, safer as well and like readily understandable by everyone in the world, and then that's going to be something quite different to what we have today.

Andrew Feldman

executive
#30

What an exciting vision. Thibault, thank you so much for the partnership and coming on stage and sharing. Ladies and gentlemen, Thibault. Okay. We have a little video of some what customers think, both inside of OpenAI and elsewhere. Let's go. [Presentation]

Andrew Feldman

executive
#31

Okay. Let's race forward. While we love our partners in the closed source community, we also are fastest on the full range of open source models. All right, from large to small, U.S., non-U.S. on these, like just about everything we do, we're up to 15x faster than the competition. Now one of the things Thibault talked about -- one of the things that's extremely sort of top of mind for us is the acquisition of data centers. And we have data centers rolling in. Here's our data center in Santa Clara. Here's our data center in Toronto; Dallas, Texas; Minneapolis, Minnesota; Montreal, for those of you who are Canadian, I got the little thing over [indiscernible] it was pointed out to me twice in presentations previously. Oklahoma City. And for those of you who don't know what they look like when you're building them, here's our data center in Alabama that's going up. Here's our data center in Lyon, France going up. Here's our data center in a part of Norway that I can't pronounce. And here's our data center in Michelle Finland, one of several that are going up, right? We now have data centers across North America and Europe. And we've brought on 600 megawatts, right, that's online or under contract for delivery by the end of next year. It's a huge effort and it's just the start. It's not nearly enough, but we have a whole team chasing this every single day. Now what you put in data centers, all right, is much more all right, than a fast AI accelerator, right? AI is generated by a cluster, all right? It's generated by a combination of equipment, some of ours. And some of others. And what I'd like to do right now is invite on to stage someone who's sort of had an extraordinary career. At Cisco, she was, in my view, singularly responsible for the rise of Cisco in the late '90s into one of the great companies of that era. She ran the Catalysts division, which was Cisco's monster. All right. She's been CEO of Arista now for, I think, 14 or 15 years and has led that company to extraordinary success. Please join me in welcoming Jayshree Ullal to the stage.

Jayshree Ullal

attendee
#32

Good to see you. Congratulations, Andrew. What a great company. What do you guys think? Cerebras.

Andrew Feldman

executive
#33

Okay. As we -- Well, first, I began my career in networking, competing against Jayshree, and I would advise against that. For those of you thinking about building switches, do not compete with Jayshree, that is some sort of unhappy stuff.

Jayshree Ullal

attendee
#34

it's a lot more fun partnering...

Andrew Feldman

executive
#35

It is fun partnering. So look, one thing that's become clear as we sort of think about the AI landscape is the amount of equipment that goes into these clusters. And the sort of the need for it to be coordinated and work together that if your network or your fabric can't deliver the goods, it doesn't matter how fast your accelerator is you're going to get bitten. And so as we think about that, sort of how do you see Arista sort of working with Cerebras to deliver these sort of extraordinary solutions.

Jayshree Ullal

attendee
#36

Absolutely. First of all, I think you're only as good on the compute side as the network. Imagine if all your Cerebras stuff were idling and waiting for the network, right? So I'm hoping my network can pay for itself by making your compute much faster. We live in a world of you talked about 600 megawatts, you're going to go gigawatt, terawatts. We live in a world of more and more thousands of tokens. We just have to deal with trillions of parameters, but you just can't do that single-handedly. So it's really about building the best-of-breed stack together. And other companies would have you believe that everything is built by one. I think the power of all of us together is far greater than each one alone. So it starts at layer one. I'm creating my own OSI model here...

Andrew Feldman

executive
#37

No, Let carry on.

Jayshree Ullal

attendee
#38

For those of us who learned the textbook that had 7 layers, the physical power and cooling and infrastructure is powerful. I heard you talk about how you're running around trying to find this because compute power cooling is the scarcest commodity, right? And if you can get that, you grab that in any form and shape. And then, of course, is the compute accelerators. And nobody does this better than Cerebras from an inference perspective, of course, there's a few others we all know about that does some training. But here, I'm going to focus on the inference. But to make this all hum, we've had to really think about what we do differently with AI than we did in the past with cloud. In cloud, it was easy, to just say, okay, we'll just throw a bunch of capacity, do low latency, have lots of bandwidth and deal with one set of traffic, which was cloud computing. In AI, you really have different forms of traffic. The fidelity is different, the type of traffic is different, the any-to-any is different. So we had to build different types of traffic to deal with networks to deal with different classes of frontier models and application agents. And I think this whole area, especially of frontier models and agents is evolving too, because the enterprise has hardly come in yet.

Andrew Feldman

executive
#39

I mean, that's a really important point that the enterprise has been on the sideline to date.

Jayshree Ullal

attendee
#40

Yes. So we mostly talk about the neo clouds and the hyperscalers, et cetera. But I believe when you're here next time, you're going to be talking about a lot more agents, agentic AI even going into our phones, and this is going to be powerful. And behind all of that is, of course, the importance of a network.

Andrew Feldman

executive
#41

I have no questions. Keep going.

Jayshree Ullal

attendee
#42

Oh, you don't. Okay. So given our engineering background, I like to think in terms of X and Y and Z axes, right? And frankly, as a student I was terrible at the third dimension, but it's important. So when we think of this and how we work with Cerebras accelerators, we think scale up, how do we connect to as many of your, I know you don't build stands you only build dinner plate, right?

Andrew Feldman

executive
#43

That's right.

Jayshree Ullal

attendee
#44

So as much of that to get the greatest [ Radix ], whether it's 64 or 128, how much of that can we connect until we run out of compute capacity and network capacity. And then we go from scale up to scale out, how do we connect these racks together. And this is where we come in, and I think we worked very closely together. But there's another emerging market, which is how do you distribute the compute. You're never going to get enough just to be in within 1 surface area. And this is where you can have a multi-tenant capacity, multi-tenant engineering, where you can secure different connections of compute in a distributed fashion across distance. And this is what we call the scale across that goes across locations. So that's a little tutorial and networking for me.

Andrew Feldman

executive
#45

Look, I think as we sort of grown our partnership. Jayshree supported us when we weren't buying very much at all. And now we're buying a fair bit. We appreciate it. Thank you. And maybe by way of final question, what do you see in the next 2 or 3 years for the networking industry and for the sort of demands placed on it by this new type of compute?

Jayshree Ullal

attendee
#46

Well, I think at one level, we're all going to be pushing the envelope on bandwidth and latency. Just to give you guys a perspective, we were all on 10 gigabit back when you were doing networking for 10 years, right? But now we've gone from 100 gigabit to 400 gigabit to 800 gigabit...

Andrew Feldman

executive
#47

We won't tell them that you and I begin building fast Ethernet switches.

Jayshree Ullal

attendee
#48

The rate and pace of throughput and capacity is now every 18 months. We've got to keep up with your computing processes, right? So in the next 3 years, I fully see this going in completely wild directions of 3.2 terabits or 6.4 terabits it's not stopping. Now when it doesn't stop like that, you also need the physical connectivity. So the rate at which optics or cable or any kind of connections happen at good distances has to also keep up. And that's a nontrivial challenge when you go at high speeds, whether it's coherent optics or co-package or even co-package copper when you stay within Iraq. So I think the future of that is very significant. But there is one other thing I want to bring up, which comes back to, I don't want any of your processes idling. The biggest challenge going forward will be how do you keep -- retain the availability of your compute by making -- building a good network. In other words, suppose these compute cycles go away or they break down or somebody pulls out a cable, how do I restore and recover. And this is where while hardware is very important, software has to really help the recovery of this. So smart system upgrade, high availability, level of automation, analytics will be super important in the future as well.

Andrew Feldman

executive
#49

Ladies and gentlemen, Jayshree.

Jayshree Ullal

attendee
#50

Thank you. Thank you very much Andrew.

Andrew Feldman

executive
#51

Good to see you. Thank you so much. I'll say it again, do not compete against her Okay. I'm going to take 5 minutes right now and show you NVIDIA's road map for the next 5 years and ours. Okay. Let me explain. The X-axis is tokens per second per user, okay? This is how fast you experience AI. All right. The y-axis is how many tokens total is delivered through the solution, right? What that means is, it's the number of users that can simultaneously get the speed that's on the x-axis. We call that throughput. Everybody understands it's really important. Speed per user, number of users times speed per user, right. Okay. They are extraordinary in this domain. This is the graph for us. We are extraordinary in this domain. Okay. Now this should tell you exactly what the future is. This is what the road map for GPUs will be. This is what our road map will be. They will try and get faster without giving up throughput, and we will try and get more throughput without giving up speed. That is the competitive landscape. Now to this end, there is a solution that can be achieved through partnership. And that solution is called disaggregation. And by using the GPU to do part of inference called prefill, and using Cerebras to do part of inference called decode, right? You can get 10x faster than the GPU and 5x more throughput than Cerebras. And this is an extraordinarily compelling solution. Now I'd like to invite to the stage someone who I've admired a great deal in my career. He began his career at IBM, where he worked on the PowerPC and on their first blade servers, it this next part. He then went to work for Steve Jobs. He report to Steve Jobs and ran the iPhone business. He then went to AMD, where he was the technical sort of visionary behind the transformation that has produced the company that they are today. I'd like to welcome Mark Papermaster to the stage. Andrew?

Mark Papermaster

attendee
#52

Great to see you. Congratulations on this incredible event.

Andrew Feldman

executive
#53

One of the things that I forgot to say about Mark is he also is 1 of the true gentleman in the industry. Really, you should applaud that because there are not many of them.

Mark Papermaster

attendee
#54

Thank you. And to you as well.

Andrew Feldman

executive
#55

Okay. Let's talk a little bit about disaggregated solutions. Why now? Why was now a good time to build disaggregated solutions?

Mark Papermaster

attendee
#56

Well, Andrew, I think it's really a statement of this inflection point that we are at right now because you think about the massive infrastructure, and you were talking a moment ago about the whole GPU build-out, it's been so dominated by training and it had to, to get the kind of foundational model capabilities we have. And that's what's fueling this transition right now because with that kind of capability, people are finding the boundless applications that we can inference on and get real work done. And so it's a huge demand for all of us in the industry to figure out how then to accelerate inferencing tokonimics that have to be optimized, helping people get their job done. And so I think disaggregation is an innovation driven out of necessity, right?

Andrew Feldman

executive
#57

I think that's right. I think 1 of the things we forget was so you make AI smart. You make it with the training, right? But once it's made, once it's smart, we're going to use it. And we use it with the inference. And as that sort of gets more mature, we're able to attack it with specialized solutions with solutions and partnership. Tell me a little bit about how you think about the sort of disaggregated solution bringing together the best of both worlds. I've got a slide here. You have a slide here. Let's see here. How we bring together sort of the best of both worlds?

Mark Papermaster

attendee
#58

Well, that's what we love about this partnership. We know each other well. And when this problem really need to be solved it was a partnership of 2 solutions that are better together. So what have we been focused on at AMD. You see it with the Helios rack coming out at the end of this year. It's a throughput monster.

Andrew Feldman

executive
#59

Monster. Here we could show you, it's a monster.

Mark Papermaster

attendee
#60

Right. So it's 72 GPU rack, super high bandwidth, and it's everything about it in terms of how the CPU, GPU, the networking is all about just incredibly efficient throughput. And so it is a perfect solution for a broad range of computing. But when you want to also have a low latency response, and we know you guys are really serving the need of these many applications, growing applications that need latency than what better solution that disaggregated where the teams have really worked together to separate out hidden from the end users, right? It's really hide under the covers, [indiscernible] prefill where you've got to take a very broad context, and it really demands a parallel computation that the GPUs are perfect at, right? And so you can just -- you can get that economic throughput. And then when you think about decode, when you're really processing that token throughput and optimizing for ultrafast response and low latency, right, send that decode to the wafer scale engine. And this combination is a win for everybody. The users get better economics and low latency and that quick response. So it's -- again, I think this kind of disaggregated solutions, really, meeting a market need.

Andrew Feldman

executive
#61

I think that's right. I think one sort of reasonable way to think about the chart I showed you before, is that throughput drives economics and speed drives user experience. And when we can bring them together, we're in this position of sort of you can have both.

Mark Papermaster

attendee
#62

Exactly.

Andrew Feldman

executive
#63

And that's enormously powerful. Now first, Sean and I have had a lot of fun working with your team and working with Mark's teams is always a joy. So we appreciate that. But you're bringing AI across AMD, and you've got lots of parts. Tell us a little bit about how you're doing that.

Mark Papermaster

attendee
#64

Well, it's an extension of what we just said. I mean AI is being used for so many diverse needs, and that's what we focus on our portfolio. So it is a Helios rack at the top of our offering for those toughest both training and inference and this huge inference throughput that we just talked about. So that's our Instinct product line. We have our epic product line with high-performance CPU servers.

Andrew Feldman

executive
#65

We use those in our cluster.

Mark Papermaster

attendee
#66

And that's where we've been -- developed this tight partnership for years between AMD and Cerebras. And then, of course, local AI. So when you need to run, we're seeing more and more where we're getting [indiscernible] models that can run efficiently locally, often on open weight models. And then, of course, the adaptive compute and embedded. So the way we think about it at AMD is a diverse set of workloads and an open ecosystem. That's sort of essential to us. It's a rack and AI stack that's open and the ability to bring a diverse set of solutions. More and more we need innovations exactly like we're, I think, paving the way together with the disaggregated solution of AMD and Cerebras. This is a really fun partnership.

Andrew Feldman

executive
#67

Just to wrap up, and Mark came, they had a Board meeting this morning. He's racing back to dinner. So we thank him for making time. When you look out into the future as sort of the CTO of AMD, and you think about the compute needs in the future around AI and other. What do you see?

Mark Papermaster

attendee
#68

Well, I used the word this insatiable demand for more computing. And so what we see now is the fact that we're still at the early days, of AI applications, inferencing application, it's actually scary what that insatiable demand is going to be. And so what I see going forward is we need more and more innovations like we're doing together. And I think the algorithms are going to evolve. And what we're going to see is a more and more integration of diverse technologies to optimize on today's algorithms and the algorithms of tomorrow.

Andrew Feldman

executive
#69

I think that's exactly right. What I tried to think about sharing with you guys today was that we use AMD CPUs, right, to manage the software that runs on the cluster. We've partnered with AMD to build a disaggregated solutions. We use NVIDIA, excuse me, we use NVIDIA. We test on NVIDIA sometimes to see what they're doing. Now we use Arista, we use Arista to tie everything together, all right? That's the start of a solution, right? And then we're partnering with different software vendors, and we use both closed source and open source models at the top that it's taking a village. Mark, I want to thank you for coming and thank you so much.

Mark Papermaster

attendee
#70

Andrew, thank you very much.

Andrew Feldman

executive
#71

Thank you. Appreciate it. Ladies and gentlemen, Mark Papermaster. Okay. Here's our road map for the next several years. That's it. Any questions right? We are going to get 4x faster. We're going to get 20x more throughput, right? This is what we're going to be doing. Our speed will double every year. And in the second half of next year, we'll be at 20x the throughput we are today. Now to tell you a little bit about how this is going to work. You're going to hear from a collection of different people over the next little while. Sean is going to talk a little bit, Jessica is going to talk about. We've got some very interesting things for you. But I did want to sort of leave you with this, that this is where we're going as a company. All right, better user experience, better economics, all right? That's what we're focused on. Okay. And with that, I'm going to hand things over to Jessica, and we'll go to the next stage. Thank you so much, everybody.

Unknown Executive

executive
#72

Please welcome Cerebras SVP of Product, Jessica Liu.

Jessica Liu

executive
#73

Good afternoon, everyone. Now Andrew has just shared our trajectory to double raw performance every year for the next several years. So now what we're going to do is take all of that speed and accelerate the hell out of AI agents. So in last year, we've already seen Cerebras speed transform all kinds of interactive AI applications. It made AI search instant, it kept developers in the flow while coding, and it made AI voice conversations feel natural. And now with AI agents, we are capable of even more complex tasks. Because instead of just giving an agent a question, you can actually give it a goal and trust that it's going to figure out its way to succeeding in the goal you've given it. Agent can autonomously make a plan, call models, call tools, and it can revise this plan in a loop over and over again until it achieves the outcome that you want. So it is actually this thinking and this iteration that makes agents so capable. But now underneath the hood, each one of those dozens of calls incurs a latency cost. So then the smarter your agent the better your outcome. But the smarter your agent, the longer it also takes for you to get to a result. And now what if you want something faster? Well, historically, when you want something faster, you would just use a smaller model. And then you could get to your real-time interactive voice agent, and it just means that sometimes, Siri, is not going to do what you want. So on a given time budget on GPUs, you always have to choose. You can either use a faster dumber model or you can use a smarter, slower one. It's not a great trade-off to make. But importantly, it's also a false trade-off. Because at Cerebras, you can do both. You can have a smart and fast agent. And this is why inference speed is so important to agents. Because when you can run the whole system 15x faster, you don't end up with a latency debt. You actually get latency credit. You have extra time to spend on using a smarter model on more loops, even while the overall system ends up still running faster, you can do both. So now let's look at a couple of examples. Agents today can already show great outcomes on GPUs. Harvey's legal benchmark completed a whole set of common client tasks in just 22 minutes. OpenAI use Codex to completely build a design tool from scratch in 25 hours. And Cursor ran hundreds of parallel agents and was able to build an entire web browser in under a week. So those are some really impressive results. But it's still kind of a long time. Any time you're spending tens of minutes, tens of hours or multiple days for something, it is a long time before you get to your response. And so think about how much more you could do. If you could achieve the same outcomes in 2 minutes or 2 hours or a single work day. If your agents could do all these tasks 15x more quickly, you could do 15x as many tasks in the same amount of time. You can serve 15x as many clients, you could complete 15x as many projects, or you can take the speed and actually just simply use it to make working with the agent 15x were magical of an experience. You can also just use speed and speed. You can have that buttery iteration loop with the agent even if you are using a large smart model instead of waiting 7 minutes to see if the agent even did what you wanted. We can use speed to make a agentic work delightful. But the best way to see this impact is from real builders who are deploying fast agents to work in the real world today. And so we have leaders with us from Figma and cognition who are going to show us what's possible when agents can move at the speed of the people working with them. So please welcome to the stage me, [indiscernible], AI research product lead from Figma.

Unknown Attendee

attendee
#74

Cool. Thanks, Jessica. You're going to hear fast inference a lot today. I want to spend the next 8 minutes talking about why it matters for a problem space AI hasn't quite mastered yet, one that pushes on agents in ways, code and reasoning don't, design. Hi, everyone. My name is Pavy, and I lead product for Figma's AI research team. In May of this year, we shipped Figma's design agent, an agent that knows your canvas, knows Figma and works right alongside of you. Let's meet the agent. There was so much that we thought about in exploring this idea. But next, I want to touch on a particular philosophy that went into building this agent. AI should adapt to your workloads. Some of our favorite tools today are where the interface melts away, and you can focus on doing actual work. They are embedded into your workflow in seamless ways would appear when you need it the most. When we set out to build this agent, we knew designers needed purpose-built tools that serve the essentials. As teams adopted Agentic tools to build products more quickly, false choices were emerging, speed or precision. AI generation or direct manipulation. You shouldn't have to choose. We needed to create an agent fluent in Figma and native to the way teams work. We also knew that design is not a linear process. You start with an idea, build some prototype, [ hate it ], throw it out, try it again. The design process is rooted in exploration, feedback and refinement. And we have built this brilliant multiplayer canvas that supports the messy middle of a very messy process. That's our product. But many of our aI tools today have started shifting that experience. Designers are used to getting an open canvas with lots of exploration getting divergent ideas very quickly. But AI tools today focus on helping you get the highest fidelity functionality and really zero in on a singular idea. Designers are used to collaborating and [indiscernible] out in the open together. But the AI tools have created silos where the work is stuck on one person's computer and it's hard to share ideas really quickly, and be inspired by one another. And so we at Figma don't think it should be one way or another. There is room for all of it at different times, and that is exactly what we wanted to build. An AI collaborator that can help you when you need it and is embedded into your workflows. Unlike the MCP server, the agent lives directly on the multiplayer campus, no separate setup or context switching required. All right. Let's dive into the technical details. The bar for design agent is fundamentally different than most coding agents. If I ask a coding agent to write a function, there's the right answer. It compiles or it doesn't. It passes unit tests or it doesn't. If I ask a design agent to make this feel more premium, there is no compile step. There is no unit test. The answer is right when a designer looks and says, "Yes, that's what I meant. When evaluating design, it's also subjective. What looks good varies by viewer, brand and audience. Second, it's multidimensional. A design can have the perfect alignment, but terrible color contrast or brilliant typography in the wrong hierarchy. Third, it's task dependent. A social media post and a pitch deck have different quality bars. And lastly, it's contextual. The same design can be great for one audience and wrong for another Design, quality, resist reduction to a single metric. It's a bundle of competing signals and the challenge is turning that bundle of soft judgments into something we can measure in our research team. Also, as I mentioned, design not being a non-verifiable domain, it also has no ground tooth and the ground keeps changing. What's great in design today might not be great tomorrow. And that's exactly what makes this such a difficult area for our research team to work on. So we decided to help solve this problem by building our own model. Specifically, a model fine-tuned for editing Figma files, that making Figma itself legible to model in ways that aren't possible with third-party tools. With deep context of your designs, your team standards and your best practices. This model helps power our agent, which brings me to why Figma is uniquely positioned to build our own in-house custom model. While foundational models are getting better every day, they often still lack the ability to judge design quality in a reliable way. They might still have biases towards certain stylistic patterns, and we can build a specialized model just for design. Second, we have the corpus. Off-the-shelf models tend to also want to generate the average of their weight. So need to be steered. We have the data and context that allows for this to happen. For the past decade, Figma has watched billions of designs get built layer by layer as the world's designers work at their craft and our products. And lastly, we have the Canvas. One of our beta users put it perfectly. And I quote "I'm just happy that the agent is an environment I know and love versus having to go somewhere else with MCPs and figure out how to link it together." And last but not least, with inference hardware from Cerebras, speed can become our strategy. We can build an agent in your canvas that is cheaper, faster, specialized for the way designers actually work. So why exactly is speed so important for agent design? Firstly, design agents are call heavy. Earlier, I talked about how design is a messy, complicated process. The messy squittle wasn't just a metaphor for human design. It's also the architecture of the agent. Slow inference forces the agent to flatten that process into generation instead of designing fast inference lets the agent preserve the loop, and those loops are where design quality comes from. Second, you can do more in less time with fast inference. Design is full of tedious work, none really hard on your own, but together, they can [indiscernible]. When the agent makes those instant, we're not saving seconds. We are returning the day so designers can get back to the fun part of design. And last but not least, designers work in loops of seconds. Every extra second between try and see a big sum from that loop, fast inference keeps you in the flow, no more contact switching. And that's the difference between AI as a tool and AI as the way you work. Working with Cerebras hasn't just been an optimization for us. It is what helps make this product possible. Figma agent is available in beta today and will [ GA ] soon. This is a really fun way, design and collaborate and it's cheaper, faster and easier than anything ever seen before. Thank you all.

Unknown Executive

executive
#75

Please welcome Cognition SVP of Research and Founding Engineer, Silas Alberti.

Unknown Attendee

attendee
#76

I'm Silas, and I'm super excited to be here, and I'm going to talk to you a little bit about how Cognition and Cerebras have collaborated on building superfast, coding agents. And also, I'm going to highlight some of our research that powers. First of all, Devon, you remember Devon so you have the billboards in the city. This is Devon, our product suite, and we are mainly known for Devin Cloud, which in 2024 was the first -- the world's first software engineering agent. And especially in the last 6 months, people have gotten really on the cloud agent train and like running hundreds of agents in parallel, agents for hours at a time. But we have way more than that, our product suite covers the entire software engineering life cycle. We have DeepWiki for understanding and planning code basis. We have Devin Review for reviewing code and Devin Automation for maintaining code. And all this is possible due to the incredible advances in AI models. And at Cognition, use a mix of models. So we use frontier models like OpenAI and Anthropic, but we've also increasingly invested in training our own models, and I'm going to talk to you a little bit about that today, 2 of our model training projects and how we're able to run them at lighting speeds using Cerebras. The first project I want to talk about has a special place in my heart. It's [indiscernible]. And it's actually almost a year old, which is an eternity in AI era, but the cool thing about it is it was actually the first model that Cognition and Cerebras launched together and also one of the first models that our model training team builds. At the time, we had much less compute than we have today. So we have to really pick our problem as well. And we noticed that coding agents spend a long, long time, even just like finding the right files on the code base to add it again and again. And so we thought, okay, how can we make that super fast. And so we trained a small model precisely on this code-based search task. So imagine, you give the model a question, for example, here, how does VsCode efficiently implement, file watching. And then the model's task is to explore the code base and return just the list of files that are relevant to this question. And for an researcher, this is like incredible task because it's verifiable, right? Like it's like there's an objective, a list of files that are relevant, you can just grade, did it find the correct files? So we went and trained the model. And so here on this code search eval that we built for like code-based search task, our model SWE-grep and SWE-grep-mini as you see performed on par or better than the frontier models at the time. It's kind of crazy how much happened in a year, but at the time, Sonnet 4.5 was the best model in the market. And then we put this model on Cerebras and it ran at incredible speed. So here you see a SWE-grep at the time ran at 680 tokens per second. And so we've got many even at 2,800 tokens per second. So more than 10x faster than any LLM model. We optimize not just the tokents per second, but also really the end-to-end time of the agent took to complete the task, which also meant doing more things in parallel. So also at the time, a Sonnet 4.5 was the first model to do parallel tool calls. And those models would still do like one tool call at a time. And we thought SWE-grep, okay, how can we push this further and push the model to do 5, 6, 7 or even 8 tool calls and searches in parallel. And so here, you can see some of the graphs from our training run. So the -- on the right side, you see the training rewards went nicely, beautifully up over the course of the run. But the cool thing in terms of the parallel tool calling is that at the beginning of the training run, the model would maybe do like 3 tool calls in parallel max and then through our training, it goes up. And at the end, it does many, many tool calls in parallel. And this really pays off. So what we basically did is we deployed SWE-grep into our product together with Claude as the main agent. So you have to imagine like the user asked the question to Claude. And then Claude can call SWE-grep as the tool to search the code base, but then based on the results, Claude will give the final answer. And what this means is there's no compromise in terms of the quality of the answer, but you save a lot in end-to-end latency. So interestingly, as you see here, in some cases, maybe for this question, we would get like a 2x, 3x, even 4x speed up in terms of end-to-end latency for the same quality of answer. After SWE-grep, we became more ambitious and said, okay, now let's start training real frontier coding models. So here, you see one of our charts that we released a while ago. So this is like SWE 1.5 a couple of months after SWE-grep. And it was the first frontier model that we deployed on Cerebras. I think you've seen the graphs like this already earlier today, but it's just incredible the speed, the quality trade-off. You basically get quality and products equivalent to the contemporaneous frontier models, but speeds that are like more than 5x faster than anything else. And we've continued investing into this. So this is a more recent results on our in-house coding eval frontier code 1.1, and our latest model SWE 1.7 performs in terms of quality on par with GPT 5.5 and Opus 4.8. But only can run it at close to 1,000 per second, but it's also a lot more cost efficient. So what you see here is actually a cost versus quality trade-off chart. And SWE 1.7 is at the same cost and scales like a Kimi K2.7 Code, the Composer 2.5 or GLM 5.2, but achieve close to frontier in quality on the SWE. So to recap our phenomenal Collaboration with Cerebras. We're able to serve frontier models at close to 1,000 tokens per second. We're serving small models like SWE-grep at multiple thousands of tokens per second and are able to achieve a 4x speed up an end-to-end task completion. And this is not -- this is just the beginning, much more to come. And thank you so much for having me.

Unknown Executive

executive
#77

Please welcome Cerebras SVP of Product, Angela Yeung.

Angela Yeung

executive
#78

Hello, San Francisco, I'm Angela Yeung, SVP of Product here at Cerebras. The next competitive advantage in AI isn't bigger models, it's time. For many years, we've asked the question, how smart can these models get? These models are already smart enough today to impact our daily lives and to make decisions in our businesses. But no matter how smart a model is, it doesn't matter unless the answer is delivered in enough time to change an outcome. Every application, every mission-critical system has a time budget. That budget could be a few hours. It could be a few minutes or it could be milliseconds. You can think of it as a window. Within the window, an answer is useful. The answer can still influence what happens. It can shape a decision, and it can change an outcome. But if an answer arrives too late, it doesn't matter how correct that answer is. It's no longer useful and the value of that answer drops to 0. Let's take some examples from our everyday lives, Priority Uber, same-day shipping from Amazon and Disney Fast Pass. If my Uber arrives too late and I miss my flight, it's game over, even if it got me to the right destination. So we buy time back. AI is no different. Every AI application also has a time budget. And once the useful window for that application is understood, the business value of fast inference becomes very clear. Take Armis. Armis is a company that runs a code scanning service. It looks for security vulnerabilities inside code. Powered by Cerebras, Armis completed a code security scan in roughly 1/3 of time, beating benchmark frontier models and finding more vulnerabilities and at a fraction of the cost. That's not just a faster code scanner, that's a better product. And it's a product that customers are willing to pay for. It's a perfect example of when speed, quality and cost come together, you get an unbeatable product. In many cases, the clock is not negotiable. In payments, we have 50 milliseconds to accept or decline a transaction. In voice operations, 200 milliseconds before the human hangs up on an AI engine. And in cybersecurity, 27 seconds is the fastest breakout adversary attack recorded on record, and it's getting faster every year. Let's think about that for a second. 7 seconds to prevent a company-wide data breach. The smartest answer that arrives after those 27 seconds has no value because the decision window has closed. That's why our partnership with CrowdStrike is so important. Cybersecurity is a perfect example of an industry where speed is mission-critical. And frontier models only matter if the decisions show up in time. So please welcome on stage, Keith Culley, VP of Engineering at CrowdStrike.

Unknown Attendee

attendee
#79

Thank you, Angela. Great to be here. Excited to talk about the partnership. So unifying security and AI. To tell you about the partnership, I really have to go back about 17 years. So bear with me, grab a drink, it really starts with the founding of CrowdStrike, and that was actually on a plane. So 17 years ago, our CEO and Founder, George Kurtz, is on an airplane. He just took over the role of CTO of a very large computer security company. And he sees another passenger turn their laptop. And they go into something known as a mandatory boot scan. And if you're too young to remember that, it was awful. It was a 15-minute process that you had to wait as your computers scanned every single file before you were allowed to do anything. So he notices that and says, "Well, this isn't going to work for very long. People are going to come -- this is prime for disruption." And that led to the birth of CrowdStrike. So a couple of years later, George creates CrowdStrike. In our industry, time is measured in milliseconds. Your success and failure can be between milliseconds. Thinking about real-world examples. So if you have a lock on your door, you can open in a couple of seconds. That's an acceptable trade-off, right? It keeps your house secure. Little bit of friction, not a big deal. If that lock took 5 minutes, you probably aren't going to use the lock. You're going to find a reason to say, it's fine. I don't really need this lock. It's too much of a pain. So it leads to poor security hygiene. And that's the same thing when you think about security and especially AI security. With computers frustrating experience leaves the bad outcomes. With that in mind, if we start from the baseline of an acceptable window to perform inspection, what can you do when you get -- where you need more time? Well, the answer is you need faster inference. And that's where Cerebras comes. In. With fast inference, you get 5x to 10x more inspection inside the same acceptable window of time. What is that unlock? It lowers friction. It increases adoption of security. It increases the indicators, the telemetry coming in so we can correlate indicators of attack, indicators of compromise. Better security, leads to more AI adoption. And when AI is adopted properly with security, now you've realized the true value of AI at scale. So with that in mind, the metric that matters here for us is actually time to decision. It's how long it takes to decide, should I allow this action or not? Should I warn about this auction? Should I prevent this option? Less friction, more security enabled at all the layers. That's why I'm so excited about this partnership. In just a few years, we've seen AI create a sea change across the behaviors in the entire industry. But it's still just the tip of the iceberg for what has real potential. We think one of the biggest factors slowing AI adoption is the general anxiety around the security and safety of AI and agents. The key to getting past these concerns is faster inference being able to do more detection, response and decision quickly. That enables you to have low friction, adoption across the entire stack and security enabled across all your assets. And that's why we're so excited to partner with Cerebras. Thank you.

Unknown Executive

executive
#80

Please welcome Cerebras CTO and Co-Founder, Sean Lie.

Sean Lie

executive
#81

Hi, everyone. Thank you so much for being here today. When we started Cerebras, we had a vision to drastically change the landscape of compute for AI. And we did that by building the world's first and only wafer scale chip. Now this is a really big deal because we solve a fundamental problem that was limiting the entire semiconductor industry for decades. But this was only possible because we codesigned a system architecture for wafer scale, a system that could power and that could cool a chip the size of a wafer. This is our first generation wafer scale system architecture. This is the system architecture on which our current product is based. And this is the architecture that brought Ultrafast inference to the world. It was codesigned for wafer scale from day 1, and it has served our first 3 generations of the products. the CS1, the CS2 and our current generation CSI. It's a 16 RU server that sits in a standard data center rack. And today, it's deployed at scale at our customers worldwide running production workloads every single day. Now over the last few years, we, as an industry, we have learned a lot. We've all learned that scaling AI inference is no longer just a server-level problem. In fact, it's a rack scale problem. It's a cluster scale problem. It's a data center level problem. And we've all learned that to really take AI inference to the next level. We need more performance, we need faster interconnects and we need greater scale. And this is exactly what we designed our next-generation system to solve. [Presentation]

Sean Lie

executive
#82

The CS4 is our next-generation system that will push the frontier of ultrafast inference to the next level. Because with -- the fastest gets even faster by providing up to 2x faster tokens. And the CS4 was built for hyperscale, providing solutions I can provide up to 10x more tokens per watt. Now remember, today, our current generation is already running up to 15x faster than GPUs. The CS4 will be 2x even faster than that. Let me show you what that looks like. I'm going to ask a model to perform a task. I'm going to ask it to create an HTML file for the periodic table of all the elements. This might be something that you or your agent might ask a model to do when you're creating a web page. We have our next-generation CS4 on the left. We have our current generation in the middle, and we have GPUs on the right. Now before I press enter watch really carefully because you might miss it because CS4 is done, and now CS3 is done and we are waiting on the GPU. Still waiting. Now it's still going in the background. I won't make you guys wait through all of this. It is painfully slow. Now what you noticed right off the bat is that our current generation CS3 is blazingly fast at over 2,300 tokens per second. But what's crazy is that next to the next-generation CS4, it actually felt slow. Our current generation has enabled Cerebras to already be the undisputed leader in ultrafast inference. And with CS4, we will widen that gap and we will push the frontier even further. By the way, this is still going. This is possible because the CS4 was designed from ground up to run faster and larger models with up to 2x more performance per wafer and 3x more density per rack. The CS4 was designed from ground up for heterogeneous disaggregation by adding 2x higher I/O bandwidth and 2x faster latency. And lastly, the CSI was designed from ground up for hyperscale, with 50% fewer components and up to 3x faster data center deployments. Now the way we did this is with a brand-new rack scale platform architecture that we call Nexus. With the Nexus rack-scale platform, we designed it for modularity, so that it could be simpler to build, faster to deploy. And it is modular so we can innovate across power, compute and I/O all independently. In the Nexus rack scale platform, in the front of the rack is the power. And this is done with modular power supplies. And in the back of the rack, what you'll see is what we call pluggable backpacks. And there are 3 of them in the back of the Nexus platform. Each one contains one wafer scale engine. Now if we look at the backpack, what you'll see is something that looks a little bit unique. The CS4 backpack is the completely reimagined server. This is now the fastest AI server in the world. The backpack is a vertical modular enclosure that provides all of the power, the cooling and the I/O to the wafer. And we've designed it specifically to be simpler than our current generation so that it can be manufactured more efficiently, with 50% fewer components and 60% more manufacturing automation. Now what's more is that this entire backpack architecture was conceived for more rapid data center deployment because we can deploy the front of the rack, all of the power supplies upfront in the data center and then we can just drop in the pluggable backpacks on site. This allows up to 3x faster data center deployments bringing deployment times down from days to hours. Now let's zoom into the back. I love this picture. This is where the wafer scale magic happens. Because inside the CS4 backpack, we have a brand new wafer package that has a brand-new direct vertical power delivery system that provides 2x more power and 2x more cooling to the wafer. Additionally, we also have a brand-new wafer I/O module that has 2x more I/O bandwidth and 2x faster latency, and it was designed to be modular to support future upgrades as networking standards evolve. Now if we pull this all together, what you see is that the CS4 system delivers 6x higher system-level performance than our current generation. higher performance. Now this is possible because the CS4 is the first system to use our faster wafer scale engine called the WSC 3 Turbo. And the CS4 is the first system to integrate 3 wafers into a single system. With these wafers, the CS4 has 6x more memory bandwidth, 6x more compute, 6x more fabric bandwidth, 6x more I/O bandwidth at half the latency. All of it enabled with a brand-new system architecture. Now these are some truly mind boggling performance numbers. But in inference, the number that matters the most is memory bandwidth. And the WSE 3 turbo chip in the CS4 system, each wafer, each chip has 43 petabytes per second of memory bandwidth. That's 2,000x more memory bandwidth than Ruben, 2,000x more memory bandwidth in NVIDIA's next-generation GPU. Now the reason why this matters is because in inference, all of the model weights need to be read from memory over and over for every single output. And on GPUs, the weights are stored off chip. They're stored in a separate memory device called HBM. And they need to traverse this very thin and narrow memory bus to reach the compute. Now Cerebras on the other hand, because our chip is so massive, we can fit all the model weights in the on-chip memory. And we pait it with a ton of compute. And by doing so, we completely remove the memory bandwidth bottleneck. Now in the GPU world, they try really hard to work around this memory bandwidth limitation. And the way they do that is they take the model and they try to distribute it across multiple chips. They take the model experts and they distribute it across multiple GPUs in many forms of parallism, tensor parallelism, expert parallelism, all in an attempt to aggregate the memory bandwidth from multiple chips by accessing them in parallel. But that is super complicated. It has a tremendous amount of performance overhead because there's so much complicated communication between all of these chips because in the end, it's just one problem that needs to be brought back together. So the result is actually slower performance, but not just that, it's higher power, it's higher cost. On Cerebras other hand, because we have so much memory bandwidth, we can run all of the more experts on a single chip. All the experts are interleaved on the wafer memory. There's no cross-chip communication. There's no complex routing. All you get is ultrafast performance because it's simple and efficient. Now all of that complexity on the GPUs, this is what all of that complexity looks like physically. This is a Rubin MVL72 rack. And if you look under the covers, what you'll see is that you'll see thousands and thousands of cables. What a mess. Now what's really funny is that NVIDIA would have you believe that this is a really good thing, right? They talk about how they have 5,000 cables in every single rack connecting together all their GPUs, and they can provide more bandwidth in the entire Internet running through these cables. That's a good thing, really. I mean, how much does these cables cost in terms of performance overhead, in terms of power, in terms of actual dollar cost, in terms of reliability of communication. On Cerebras, all of the communication in that cable set is done on the wafer. All of the communication is done with no cables because it's all on chip. And what's more is that on the wafer -- we have 200x more communication bandwidth than all of those cables combined because it's all on chip. Now what happens when you have to off chip. To go off chip, in the CS4, we designed a brand-new next-generation wafer I/O interface. This has a new wafer I/O module that extends the fabric from the edges of the wafer. And it's designed to be modular and programmable so that we can extend it in the future as networking standards evolve. Here, we've actually taken a page out of the chiplet playbook by separating the I/O from the compute silicon, we can innovate on each independently. We have higher bandwidth, and we have lower latency because we have a brand-new direct wafer link interface. And this new wafer I/O module continues our commitment to standards based networking with RoCE RDMA over Ethernet. With this new wafer I/O module, we can run the wafer links faster which gives us 2x more bandwidth per wafer, 2.4 terabits per second compared to 1.2 terabits in our current generation. This new wafer module, wafer I/O module, also has a brand-new low-latency packet processing pipeline that gives us 1.7x faster latency through the network down to 3 microseconds from 5 microseconds in our current generation. And lastly, the new wafer I/O module has brand-new direct wafer links, which lets us connect wafers directly to one and another, bypassing the traditional network. This improves wafer-to-wafer latency by 2.5x and brings the latency down to mere 2 microseconds. Now why is all this important? To understand why the wafer performance -- the wafer I/O performance is important. We need to look at how it's used to run large models across multiple wafers. So to do that, let me first start with our cogeneration CS3. Today, our current generation, in production CS3, already runs the largest frontier models, and it runs it really fast. GPT-5.6 Sol, as an example, today already runs on our current generation CS3 at ultrafast speeds. GPT-5.6 Sol is OpenAI's leading largest, most intelligent model already running on our current generation. Now how do we do that? It's actually pretty simple. We do this by mapping the model as a pipeline on to the wafers. And this mapping is very natural because it maps to the model architecture directly, making it seamless and fast because we can keep all of the high communication bandwidth on the wafer, where we have all of that memory bandwidth where we have all that fabric bandwidth. And we're only transmitting activations between wafers. So now you're starting to see why the CS4's I/O is important because we can already run the largest frontier models on our current generation. But with CS4's higher performance and lower latency, we will be able to run even larger models of the future. Here's a graph that shows the wafer-to-wafer I/O latency versus the model size. On the x-axis, is the model size in trillions of parameters. And on the y-axis is the total I/O latency across that entire wafer plan, all added up, all combined. And this line is the CS3. And this is CS4, 2.5x faster with 2.5x lower latency because of the new I/O module. And what you can see is that even 10 trillion parameter models have a mere 0.2 milliseconds of I/O latency aggregate across the entire pipeline of wafers, 0.2 milliseconds. That's just a fraction of a millisecond. And recall that if the entire round trip latency is one millisecond, that equals 1,000 tokens per second of generation performance. So what this means is that even 10 trillion parameter models can run at 1,000 tokens per second, and the I/O is not the bottleneck. Now what's more is that all of these numbers don't even include speculative decode. So this will push even higher performance numbers. This means that with CS4's advanced I/O, we will enable sub-millisecond latency or more than 1,000 tokens per second on frontier models with 10 trillion parameters or even more in the future, all because of our optimized wafer-to-wafer I/O latency. Now I/O is important beyond just wafer-to-wafer I/O communication. In fact, we designed CS4 so that it can connect to other hardware infrastructures. We designed CS4 for disaggregated inference. At Cerebras, we believe very strongly in a disaggregated heterogeneous ecosystem where the user can choose the best hardware for the job. And this is the reason that we have hardware partnerships with AMD and with AWS to bring this aggregated solutions to the market. But to understand why disaggregation is valuable, let me show you how it works. Inference has 2 parts. The first is called prefill. This is when the model is processing the user's inputs. And the second part is called decode. This is when the model is generating the output. Now what's actually happening under the covers is that during prefill, the model is processing all of those inputs and it's trying to make sense of it by creating an internal representation of what it all means. We call this the context. Now this context is really, really important because the context is what is used during decode to generate output because during decode, the model takes that context and generate output, one token at a time. And while it's generating output one token a time, it's extending that context until the end of the output. So now if we step back a little bit and we look at what's actually happening in each of these phases, you can see that their properties are very different. First of all, during prefill, since the model knows all of the input upfront, it can process all of those tokens in parallel. This means that it can reuse the model weights over and over and over. And as a consequence, it has very low memory bandwidth requirements. Now decode on the other hand, is completely different. It's completely opposite. Because during decode, every single output requires reading all of those model weights from memory over and over and over. And so decode, because of its serial nature, requires significantly higher memory bandwidth. Now Cerebras can run both prefilled and decode. But as we all saw, Cerebras runs decode ultrafast. And similarly, the GPU can run both prefill and decode, but GPUs run decode slowly. So since prefilled doesn't need high memory bandwidth, it can be offloaded to the GPUs. And since decode needs high memory bandwidth, it can be offloaded to Cerebras and we get the best of both worlds. This is the value and the power of disaggregated inference, the best hardware for the job, efficient prefill, come on. Efficient prefill on GPUs and the fastest decode on Cerebras. But remember that context, that context that was generated by the prefill but is used by decode. Well, when everything is running on the same hardware, that context can be generated locally and used locally. But in disaggregated inference, it needs to transfer from the GPU to Cerebras. And the time to transfer that context directly impacts your TTFT because the decode can start until it has all of the context. And for very large models today, that context could be tens of gigabytes in size. That's like transferring multiple HD movies worth of data on every single request. And so now you see why the CS4 is important because with 2x higher wafer bandwidth with 2x faster latency, we directly reduce the transfer time, which directly reduces the TTFT and improve throughput. Now what I just explained is a very common form of disaggregation called prefill decode disaggregation. But it turns out there's many other forms of disaggregation. For example, attention-FFN Disaggregation. And there's many of forms that are being invented every single day. When we designed the we anticipated this. So we designed the CS4 I/O module to be a programmable universal disaggregation interface. So that is designed to support all forms of disaggregation and provide a flexible integration point with all other hardware. And we do this in 2 ways. The first is our commitment to [indiscernible] based networking for universal compatibility. And the second is by making the module programmable. We support future protocol extensions as networking standards evolve. So now let's pull together everything that we just talked about today. And let's look at the throughput interactivity landscape today. We're very familiar with this graph now, right? GPUs, can run at high throughput, but they're slow. This is our current generation CS3, up to 15x faster than GPUs. This is what invented the ultrafast segment. But with CS4, we are pushing this frontier even further by providing up to 2x faster tokens and providing solutions that can provide up to 10x more token capacity. With the CS4, we are forging a brand new frontier for ultrafast inference. And the CS4 is where the fastest gets even faster, and it's built for hyperscale. Now let me show you what this looks like in terms of -- so a little bit of IT issues up here. Let me show you what's going on here across multiple different models. This is our production performance today across many different models. Small models like Gemma, to medium-sized models like GLM and Kimi, all the way to the largest, most intelligent frontier models like GPT-5.6. Now this is the CS3 performance today with our current generation product where we are already up to 15x faster than GPUs. This is CS4, up to 30x faster than GPU solutions. This level of performance is transformative because it will completely transform user experience, up to 30x faster tokens means significantly more interactive, engaging applications. It means offline applications now can become interactive. It will completely transform our agents because of the 30x faster tokens means 30x more reasoning means 30x more agentic calls, the CS4 enables a new era of more capable and more intelligent agents. And it's built for hyperscale with higher performance and higher density, faster manufacturing, faster to deploy, so that we can bring more ultrafast tokens to the world because with less power per token, it means you can get more tokens per data center with less cost per token, you can get more profitable data centers. This is transformative and it's available now. The Cerebras next-generation CS4 is in early access right now and will be generally available later this quarter. But we're not done yet because we designed the Nexus Rack scale platform architecture from day 1 for multiple generations of products from CS4 to CS5 and CS6. This is enabled by the modular design that allows us to independently innovate on power, compute and I/O rapidly. This is what enabled us to codesign the system architecture with our next wafer scale engine, which will be in the CS5 in 2027. Our brand new Nexus rack scale platform architecture is the foundation of our road map commitment for 2x more speed every single year. And this rack scale platform architecture is the foundation for our road map commitment to provide solutions with up to 20x higher throughput by 2027. With this level of performance and this scale, we can provide ultra-fast inference to everyone to more customers, to more developers to more users to all of you. And I personally believe that we are just scratching the surface. And so I am so excited for the future of ultrafast inference. Thank you very much.

Unknown Executive

executive
#83

Please welcome Cerebras Chief Marketing Officer, Julie Shin Choi.

Julie Choi

executive
#84

How's everyone doing? All right. I'm Julie. I'm the Chief marketer here at Cerebras. And on behalf of our team, thank you so much for being the best crowd ever. And wafer and I appreciate you so much, okay? So today, we had a lot of announcements, but the main takeaway here is that Cerebras, we live to serve the bestest AI for each of you, right? We're partnering with the absolute best customers, companies, developers, partners in the world, and we want to bring the fastest AI tokens to each of you ASAP. Who here wants to build with CS4. Can I see a raise of hand. Makes them noise guys, make some noise. We want speed. We want speed Okay. So we're going to work on that. So after Thibault left, I stopped him. He's in a rush because he has to go back to work at OpenAI down the street. And Thibault and I had a conversation and Thibault like Julie, I really love the crowd there, the vibes were just insane. And we need to do more. And I said, "Thibault, I think we need to -- we need to just give these people, we need to just help them fly. So we're going to work on ways to open up more and more of this amazing frontier level speed for all of you. And folks that came to Supernova will be sending you special ways to get on early access lists and just keep us honest, all right? Okay. A little bit of logistics. After this, no more talking. This room is going to be turned into a party space. So the Midway is very famous. This is known for good sound system and the floor here is known for that thing. So we have some amazing musical talent coming this evening. So we'll come back here at 7:30 to listen to DJs, including Lucy Guo, an amazing technical founder and very talented musician. And then -- so everyone exit here after I leave this stage, go through those doors. And then we have plenty of demos from Deep Mind, cognition CrowdStrike, AMD, like OpenAI, us, all, everyone. So we have demos, go check those out, get some food and then go to CAFE Compute and make a new friend. Okay. Cerebras is here to answer all your questions about fast inference. So don't be a stranger. All right. Let's have some fun. I'll see you back in here at 7:30. [Break]

Read the full transcript via the API

You're viewing the first half of this call. Get the complete Cerebras Systems Inc. transcript — plus 253,000+ transcripts from 12,000+ companies, speaker segments, AI summaries and full-text search — through the EarningsCalls.dev API.

Get the API View API docs →

This call discussed

For developers and AI pipelines

Programmatic access to Cerebras Systems Inc. earnings transcripts and 253,000+ others is available through the EarningsCalls.dev REST API. Plans from $24.99/month — full transcripts, speaker segments, full-text search, and the recently-added /api/v1/transcripts/recent polling endpoint for ETL pipelines.