Hey, everyone. Super glad to be again on stage. It's a bit emotional because a year ago, we were all here together with Eric Schmidt, with Andrew. 12 months have passed. A lot has happened. Maybe today Andrew is the most, let's say, bankable company in the streets. You just did your IPO, and it's fantastic to see two builders working together and walking us through today, how do we go from energy infrastructure to tokens and intelligence at scale? Maybe you want to quickly run us through your collaboration and how Cerebras is working with OpenAI to power that next phase of intelligence.
You want to start?
Sure. I think maybe the way we see it is that starting in about 2025, models got smart enough to be widely used. As these models became sort of smart enough to make a meaningful impact on jobs we did, it became more and more important that they be fast. OpenAI was one of the first to recognize that, and Sam called me in the summer of 2025 and said, "We're really interested in fast inference, and we see the time is right." We began a collaboration that resulted in signing one of the largest deals in Silicon Valley history together. A deal closed in December 24th, north of $20 billion over the next several years of compute to deliver to exactly what Sachin was talking about, which was sort of a portion of the workload that is made vastly better by being faster.
Yeah, no, I think Andrew summed it up well, which is we started seeing last fall that as customers started using AI for more and more complex tasks, as the models became capable, just like we went through what happened with search in the 2010s, right? Initially, people paid for quality. Google Search was just a better answer than anything else. Then Google had to optimize for latency, right? The amount of investment Google made in optimizing for latency was what led to a lot of growth for Google itself, because lower latency is directly correlated to search revenue for Google, for example. We believe that AI is going to go through the same transformation. As it becomes part of our day-to-day work, one, first and foremost, we'll always want quality and accuracy. Immediately right after comes how interactive does it feel?
Latency is a very critical product ingredient for us going forward. That's where we started thinking about what are the options out there on the compute side to actually build such kind of capability. Cerebras has solved some really hard problems getting us there. Super thrilled with this collaboration. I think Phi-6, as we announced, is going to be available on Cerebras. It is the only frontier model that is going to run at 750 tokens per second. That's an unheard-of speed. That's probably an order of magnitude faster than anything else that's out there. I think it'll feel magical when you use it.
I mean, that's the coolest thing on earth, frontier intelligence at blisteringly fast speeds. I mean, that's what we do this for.
I think the core DNA of your collaboration is adoption. Adoption for developers, but also enterprise. Before we had a conversation on why do you see Codex as a broader knowledge work empowerment than more like a software engineering tool. Can you walk us through this, Sachin?
Yeah. I think we are seeing this internally, right? As I was showing in the data from our internal usage, Codex is the default user interface for the company now, right? People don't even interact with browsers. I know many engineers who interact with the browser through Codex because Codex has computer use. Anything we do, we basically start with Codex. It's a new user interface, and I think that's a pretty fundamental statement to make. As we looked at usage, as I said, it is going much more broadly beyond engineering. Legal, go to market, finance, you name it. Everyone's using Codex for more and more complex tasks.
Any task that has output that can be measured and you can actually figure out how good or bad a particular output is, Codex is really good at just iterating, the AI just continuing to work through it and solving more and more complex problems. As the models become more and more powerful, they can run longer and solve more and more complex tasks. The joke, HR department in OpenAI built an agent to do human reorgs. Agents are doing reorgs of human teams now inside the company. Reorgs, as many of you probably appreciate, are very complex topics. These are not easy things to do, and we are increasingly getting AI to basically tell us what's the best way to organize.
I think it says a lot that when you're owning the compute topic within OpenAI, that usage is being the one metric that you're looking at on a daily basis. Okay, every decision that I'm taking at the infrastructural level is going to have obviously huge implications for the end users who are, again, the developers and enterprise. Maybe, Andrew, can you walk us through why latency is basically the backbone of whatever decision for enterprise and developers and intelligence at scale.
I think we can begin with sort of the reverse. We could ask ourselves why there's no market for slow search, right? Why is there no market for dial-up internet? If your teenage daughter is being naughty and you shouldn't take away her phone, you should set it to dial-up rates. Right? That will change behavior. I think that what Sachin is describing is a world in which AI is integrated into the way we work. It is part of the fabric of how you do your job every day, and from big things to little things in corporate life. In that world, waiting is horrible. How long will you wait for a website to resolve before you click away? Two seconds? Three seconds? What if you used it constantly, right? That part of your job, this was the tool you used to do your work.
In enterprises, whether it's a CRM, whether it's coding, whether it's in the HR organization, we're using tools constantly. If you ask people to wait, the experience is horrible. What has been shown again and again in every study, both academic study and business study, is that if you give people fast tools, they use them more often, they enjoy using them, and they use them on harder and more interesting problems. As we take these extraordinary models that OpenAI is making, and we deliver them to a whole range of different users at real time, so they're interactive, all right? There is nothing standing between the user and the benefit from the AI, the value, the productivity. I think that's why we're seeing this sort of across the board.
Okay, we're living in an era where token maxing is basically being the very popular word we see on X. Basically, the more token you use, the more AI native you are. Let's say you're a large enterprise, how can you go above the POCs, the pilots? How can you be sure that, okay, Codex is going to be distributed at scale within those enterprise? What is the productivity metric that you are benchmarking? Sachin?
I think we are seeing it in company visible metrics. That's why I was showing the model releases stat. The pace of us releasing new models has now gone into a month. We are releasing a new model every month. I do mean it sincerely, that Codex is the primary reason that pace has quickened, because we are just able to do a lot of the work a lot more faster than we used to do before, because previously, AI research was human limited, fundamentally. We are increasingly getting to the point where recursion begins to become real, where AI is going to help, if not do, the AI research itself, right? That's the most important productivity metric for OpenAI. How quickly do we release a frontier model, right? It's showing up.
I do believe that every, obviously, part of the enterprise will have different metrics, but we are beginning to see it in the output of the company itself, on what agents are capable of accelerating.
Yeah, I don't think you should count your tokens as a measure of how AI forward you are. I think we're building AIs to do work. You should count the productivity of the work. It's an extremely sort of rough measure of productivity, the number of tokens you use. I think if you remember back when the cloud emerged, we gave every engineer, we gave them access to AWS to bypass our own IT organizations because our IT organizations were so slow, we discovered that was expensive. We also discovered that our IT organizations were in the way of productivity, both, that what we needed to do is we needed not to just sort of give it to everybody, but to think and measure productivity along the way.
I think each time Sachin talks about this, what they're seeing, and they measure it unbelievably carefully, and they show it to us, is that these are spikes first in productivity. The token use is how you get there. What you want to measure is how fast their models come out, how much smarter each generation is. If you're shortening the time and increasing the intelligence of the models, you are being vastly more productive. As you guys think about deploying these agents, that's sort of the way I would recommend. I don't know if you agree, Sachin, but pick your business metric, and I think you'll find that the agents can vastly improve your business metrics.
I think the conclusion that we got from every session over the last two day is every part of the supply chain, everything is exponential. Everything is also a bottleneck. How do you see the Agentic AI, Codex is going to boom in term of usage. What are the implication at the infrastructure level and what is the strategy on a daily basis To work with that exponential.
Sleep less.
You can tell from his face.
It's true.
both of us haven't had much sleep over the last two days.
It's true.
No, kidding aside, I think the infrastructure is becoming a lot more heterogeneous. We therefore need a lot more of every single thing, CPUs, GPUs, networking, storage, memory, you name it. I think what we are constantly doing at this point is hunting for supply wherever we can get it. The other piece is, of course, finding the data centers to put this in. Even that is a bottleneck. All of those, I don't have a silver bullet answer. There's no easy answer here. Right now, we are all trying to solve these bottlenecks. Obviously, whenever there is necessity, there's invention. We are getting creative on software optimizations to make use of these resources a lot more efficiently. That will help alleviate some of these bottlenecks, and those are important things.
We've been in this phase in AI where we are going quickly to new products and new models, it's all been about time to market. We are now getting to the point where AI is scaling, efficiency becomes important, too. Everyone's investing a lot in how to become a lot more efficient in using the hardware, whether it's memory or GPUs or CPUs.
If I do a quick transition, from enterprise usage, which is both your obsession within your companies, how can you apply that framework at the regional level with all the sovereign AI conversation that we see for every continent? How do you factor in that new geographic reality?
I think the push behind sovereignty is that this is important. That what's being built, what's being discussed here at RAISE is fundamentally important. In fact, it's so important that it could be considered a critical national resource. Once you see that, once you recognize that these are critical resources, you don't want to be dependent. That's the germination of this sovereignty push, is this is now so important, it is so powerful a technology, that nations want to be sure that they're in a position to benefit from it and not be dependent. What do you think?
Yeah, these are the factories of our age, right?
That's right.
This is the manufacturing industrial revolution of our age, effectively. These are intelligence factories. I think we absolutely need to make sure that such capability is available everywhere in the world and not constrained. Very much in line with what Andrew was talking about, how do we make sure that intelligence is not just pushed to the frontier, but also delivered at scale cheaply and made sure it's democratically accessible to everyone. I think that becomes a very important priority for everyone.
Why Europe is such a key market for your collaboration for Cerebras and OpenAI, and I think you might have an announcement you want to do, Sachin and Andrew?
Well, first, there's an enormous demand in Europe for extraordinary AI. For the intelligence that is being built by OpenAI and others, and there's a huge sucking sound for more tokens. We announced today that we're building 200 MW of data center capacity, some of it in Lyon, France, some in Norway, some in Finland, that this 200 would be finished by the end of next year. Some of it would be delivered in this year. Much of it is to meet the need of OpenAI. That we are deploying billions of dollars of capital in data center development to be sure that the fastest tokens can be delivered here in Europe, the smartest tokens can be delivered here in Europe, and this is just the start. We anticipate many more big scale deployments and big data centers here.
Based on this collaboration, when you were starting to build out your strategy on this infrastructure roadmap in Europe, was it very different? Did you have to rewire your brain on how are we going to organize, orchestrate these build-outs for serving what usage for what type of enterprise? Was it extremely different to what you saw in the U.S.?
We don't see it as different. I'd say that even whether it's in the U.S., whether it's in Europe, whether it's in Asia, enterprises are learning how to use Codex, how to use agents. The initial reaction will be, let's take the current way we organize an enterprise and figure out how to use agents in that. I think increasingly it's very clear that we will organize the enterprise differently when agents come at scale, and that's happening everywhere. What should be the size of a team? What should be the reporting structure in a team? Should I be thinking about AI workers reporting to me for specific tasks? These are serious conversations now. How should we think about which role? What do we put an agent for that role versus not?
I'd say that this is at that stage where it'll rewire how the enterprise is organized and built, and that is a universal phenomenon. I think that's going to happen everywhere. In fact, we, as a lab, are looking for innovation from the enterprises because we won't know. We don't know how to organize an enterprise in finance. We don't know how to use agents in healthcare. That's why it's imperative for us that we make sure such fast intelligence is available everywhere, that enterprises can experiment with it and come to the right answer for themselves on how to use this. That's why we put a lot of emphasis on making sure that there's no difference.
What we have available in the U.S. is also available in Europe, is also available in Asia at scale, that we can unlock innovation everywhere across the world.
I think it will be a muscle that has to be built in organizations, right? It is not clear you are going to get it right the first time. You got to go and you got to try, you got to build the muscle. It is like any other large innovation is you are going to iterate over it before you get it right. Then when you get it right, you are going to have meaningful competitive advantage.
Exactly.
Maybe some final words. What are the next 12 months going to look like? I know 12 months is a lot. I think you remember, Andrew, from last year. So yeah, what is your outtake, basically, both of you?
12 months ago, we were a private company, and we had $25 billion less in sales. My wife saw me more often. I think we couldn't have thought of how powerful the models that Sachin and OpenAI are delivering and the market has delivered right now, and that is in a year. I think their rate of change is increasing.
Yeah, I'd say that 12 months is an eternity in AI. I don't even know what will happen in three months, to be honest. I think what I can say for sure is the one constant is going to be the accelerating pace. Not just it is change, it is the rate of change is getting faster. We don't see any indication that that is going to slow down in the next 12 months. If anything, it is the capabilities of these models are going to keep improving at a very fast clip. As I said, the model releases are happening at a much faster clip. The pace at which new capabilities come are going to be astonishing. I think the bigger question will be how quickly can these capabilities be adopted for the real world, for enterprise usage, for whatever consumer usage?
I think that's going to be the next phase of our conversation. Probably 12 months from now, we are probably all going to be sitting here and thinking about these things are so powerful, how do we make sure that everyone can benefit from this, and they're actually using it in their day-to-day.
We'll probably see all our agents based on Codex, this form of agents, unlocking this productivity, this intelligence. Thank you, Andrew. Thank you, Sachin, and well done for your collaboration. Thanks for this announcement, and thanks for being here in Paris again.
Thank you for having us.
Thank you very much. Thank you.