Good morning, everyone, and welcome to Backblaze's Investor Day for 2026. Today's session is being webcast live, and a replay will be available on our investor relations website following the event. My name's Mimi Kong. I'm the Head of Investor Relations, and on behalf of the team, I want to welcome everyone for joining us today here in person and those tuning in via live webcast. Today we will have a short break around 10:45 A.M. today, just to let everyone know, and we are scheduled to wrap up by noon. Some house cleaning things, housekeeping. Today's presentation will include forward-looking statements, and actual results could differ materially from those statements. Please see the disclaimer on the screen for more details. Full materials, including today's slides, are available on our investor relations website. With that, I'd like to invite Gleb Budman, our CEO and Chairperson, to the stage.
Good morning. My name is Gleb Budman, co-founder, CEO, and Chairperson of Backblaze. Today is our first Investor Day in the five years since we went public. A few people actually asked me, "Why now?" Is there something specific about today that we wanted to share the Investor Day? The reason that we're doing the Investor Day today is because the company and the market have shifted dramatically over the last year, roughly. What I would say is over the 20 years that we've done this, right now is the most exciting time in our journey. We wanted to share that with you. Over the next 2.5 hours, this is the plan. I'm going to talk about how the market has shifted and what our role in that is.
Dan Spraggins, our CTO, is going to talk about the platform and why it's uniquely situated for AI. Anuj Kumar, our CRO, is going to talk about the signals that we're seeing with AI and how we're building that into a repeatable and scalable engine. Marc Suidan, our CFO, is going to talk about how we're changing this opportunity into a great business. We're also going to hear from Hume AI, a leading multimodal AI company, and we're going to hear from WEKA, a new partner of ours, in the high-speed storage space. Listen to how those all come together into stitching together this ecosystem of AI.
We're going to cover a lot of information today, but if there's one thing that I'd want you to walk away with, it's that not only is AI creating a ton of data and using a ton of data, but the role of data and data storage is fundamentally changing, and that's how we're going to talk through it today. AI is creating this seat at the table for an independent capacity tier of storage. We're going to explain why that is and what we mean by an independent capacity tier of storage. We believe that that seat is open, and we believe that seat is for us. Part of the way that we are ending up there is we started this company 20 years ago. We started it solving a problem that was really important at that point, which was backing up computers.
We planned to do that on Amazon S3 and store the data there. We worried the economics were not going to work. So we ended up building a full stack solution: servers, a cloud storage file system, and the operational expertise to manage all that because we had to store massive amounts of data and do that very efficiently. So in doing that, we then spent 10 years optimizing that full storage stack. We then launched it out to developers and enterprises as B2 Cloud Storage. Over the last 10 years, we've been optimizing and scaling that platform, and today it's ideal for the AI use cases that are being needed in the market. So that's the foundation from which we believe we get to have this seat at the table as an independent capacity tier.
The rules of storage are being rewritten with AI, and here are the three rules that are being rewritten. Number one, data was growing steadily for years and years and years. Basically always, right? There's now this step change that is happening where the data is just exploding, both in terms of the amount of data being created and the way it's being used. One of the interesting things with it is that it's not just the largest companies anymore. Even small new companies have massive data sets and massive data needs. That is a significant shift from the past. Number two is there used to be three hyperscalers, and companies defaulted to one of them. Now there are about 200 AI infrastructure companies out there, and that is completely shifting how companies are building technology.
Number three is that storage used to be just a thing companies used, but today it's a critical performance layer of whether companies can innovate in AI and whether they can get an ROI from AI. Those three rules play perfectly into Backblaze's strengths. As these data sets scale massively, they need to go somewhere, and they need to be efficient in where that storage is put. Backblaze's optimization of our platform for the last 20 years makes it possible to store massive amounts of data and to do that efficiently. The explosion from three hyperscalers to 200 neo clouds has two benefits where Backblaze plays in. One is that those AI infrastructure companies themselves need cloud storage as part of their workflows, and we're ideally situated to provide that to them.
The other is that all of the AI natives and AI builders using all of these companies need an independent capacity tier to store the data to be able to use them. The third strength is that as companies care about the performance per dollar of their storage system, because it is what's critical to get both the innovation out of AI and the ROI out of AI, Backblaze's optimization of that exact thing from hard drives over 20 years is exactly what they need to be able to do that. So this all supports us taking advantage of what is a large and growing data set. The data explosion over the next few years is hard to fully wrap our heads around. Right? In 2024, IDC said there was 173 zettabytes, created.
In four years, that number is going to quadruple to 700 zettabytes created in one year. Not all of that data is going to be stored. Not all of that data is going to be used. But an increasing amount of it is, because AI is also making it possible to get value out of the data that is getting created in ways that was not possible before. Not only is the amount of data getting created going exponential, but the desire to keep and use it is as well. That supports a large and fast-growing market that we get to participate in. We continue to serve the markets we have always served, the core markets where we have served disaster recovery, backup, archive, media workflows, application storage. All of those markets still exist and are growing.
That has not gone away, and that is still an important part of the trajectory of our business. AI layers on a big market on top of that and a faster-growing market on top of that. That requires a company that can scale with that. That is what we have built. We now have over 5 exabytes of data storage under management. That makes us one of the largest cloud storage companies on the face of the planet outside of the hyperscalers. We are planning to expand that capacity by about 30% this coming year. Obviously, that takes hard drives, networking, servers, data center space, power to do all of that, and having the ability to acquire that and the relationships to do that takes a certain level of expertise. On top of that, it is the actual managed storage aspect of it that is really hard.
It is the software platform that can scale like that. It is the operational expertise and rigor that has been built around that to be able to actually deploy and manage scale at this rate and this scalability. We have talked about the scale of the market. Let us talk about the second thing that we mentioned, how the rules are changing. The dispersion of these AI infrastructure companies. This is probably the most important slide that I am going to show you, and so I want to spend a little bit of time on it. For the last 20 years, technology was built inside of a hyperscaler. Generally one hyperscaler. Even when there were three hyperscalers, companies would pick one and use it and build all of their technology stack inside of that one hyperscaler.
Maybe they were inside of Microsoft, maybe they were inside of Amazon, maybe they were inside of Google, but they picked one, and they built everything inside of it. Today, there are 200 AI infrastructure companies, neo clouds, sovereign clouds, inference clouds. That is not even starting to talk about the fact that all of these clouds have different regions, including the hyperscalers. A common customer has their base workloads inside of one of the hyperscalers still. They are using databases, maybe they are using networking, maybe they are using one of the 200 different services, or 10 of the 200 different services that the hyperscalers provide. They are still using a hyperscaler for a bunch of stuff.
They go and they say, "Now I need to do model training." The hyperscalers don't have access necessarily to the latest chips or have availability of them, or they don't have the price point of them. So they go to one of the Neoclouds and they say, "Okay, I'm going to use you for the model training piece of it." That one Neocloud doesn't have everything they need, so they go to another one for some of their other model training aspects. So now they've got base workloads in a hyperscaler, data they're collecting somewhere for the model training, and then using two Neoclouds for model training. Once they've trained their model, then they want to do inference somewhere where it's optimized for that use case. So they may be using two or three different inferencing providers around the world.
Beyond that, then they may also be using different regions of these. What that means is that as that AI builder, you now have four, five, six, eight places where your data needs to go. If your data is sitting inside of one of the hyperscalers, you're paying massive egress fees every single time you send your data to any of these places. So you really need this independent capacity tier, a place to keep your data so that you can innovate with AI. It's not just about saving money, it's about the fact that you want to use the AI infrastructure that's out there in order to innovate.
One of the customers that I was talking to recently said because they switched to Backblaze, they are now enabling their AI researchers to do their model building when they feel they need to, when they have some interesting part of data and some algorithm that they want to iterate. Before this, what they did was they told their researchers, "You are allowed to run one model build per quarter." So the pace of innovation that they're able to achieve has dramatically improved by switching to Backblaze because they're able to now go and actually move their data when they need to. So this enables Backblaze to be the backbone of the AI ecosystem by providing this independent capacity tier where data can flow to wherever it needs to go. So that's for all the AI native builders.
The other part of this chart is the actual companies on this page are AI infrastructure companies. The hyperscalers knew how to build storage, but now you have 200 companies that are racing as fast as they can to become large AI infrastructure companies. Most of them know land, power, and shell. Some of them know compute. Almost none of them know storage. Yet all of them, if they're going to be serious players long-term in the AI infrastructure space, are going to need to be able to support the workflow their customers do, which means they're going to need storage. So almost all 200 companies in that space are going to themselves need a storage platform for which we are ideally situated to provide them.
Like I said, this is probably the most important chart to understand because this is all a massive replatforming that is happening in the technology industry and has only really started in the last 2 years. We went 20 years of the cloud evolution, and we are now at the very beginning of this AI infrastructure revolution. Now that we've talked about the market shift, let's talk about the third part of it, which is the role of storage itself. In an AI workflow, there are different parts. You collect a bunch of data first, then you process it to prepare it for model training. Then you build your model, then you run inference on it, then you monitor and log it and output it. Every single part of that AI process creates and uses data.
Every single part of that AI workflow needs a capacity tier. That makes Backblaze a critical strategic partner to the AI builders who are building AI workflows. Dan's going to talk more about this in his section. But the platform that we built is what supports this. The platform that we've built, we've optimized for scale and scalability, performance, and economics. It's taken 20 years of honing to get that platform to where it is today. That is a significant moat that's hard to replicate because even if you were able to write all that software today, you wouldn't have the 20 years of operational expertise that it took to hone that platform and all of the systems that support that. Again, Dan's going to talk more about all of that. I've talked about the market and how the market's evolving.
Now, we could say, well, some of this is thesis and is this going to happen? But it's not the is it going to happen. It's happening. It's happened, right? Backblaze is already winning in this AI ecosystem. We have signed now five of these AI infrastructure companies. We've signed a $335 million deal with CoreWeave to provide them storage for their capacity tier. But we have also signed four other of these AI infrastructure companies. We also have five of our top 10 customers on B2 now are leading AI companies. Beyond that, we have the smallest and fastest-growing innovators through our Flamethrower program. This is a program designed for AI startups, and we have hundreds of the leading AI innovators in that program now.
We've also signed, on the other side of it, on the larger side of it, a leading frontier model developer to switch to us as well. We're winning across the AI ecosystem, and we're doing that because these companies see the strategic value we provide them by offering them an independent capacity tier that allows them to innovate faster in AI and get ROI from AI. To summarize, AI is not only creating an explosion of data, but it's changing the role of storage. It's creating an open space, a seat at the table for this independent capacity tier, and Backblaze is well-situated for that seat. With that, I'm going to let Dan Spraggins, our CTO, share more about the platform and why it's uniquely built for this.
All right. Thank you, Gleb. I'm Dan Spraggins. I'm the CTO at Backblaze. I lead our R&D group and oversee our product roadmap as well. All right. Today, I'll be talking about a number of things. I'll dig into our architecture and how we're differentiating with our platform. I'll talk about the ecosystem. NVIDIA recently made some changes to the reference architecture, so I'll speak to that and where Backblaze fits in. I'll also talk about the storage spectrum, so how SSDs, HDDs fit in the ecosystem and where Backblaze plays. Recently, I met with the Chief Strategist of WEKA, which Gleb mentioned previously, so I have a conversation with him. You'll see a short video there. I'll then get into what we're doing today with AI, and how we're delivering features there, and then our product roadmap.
I wanted to give you some insight into what's coming. Some of that isn't public yet, so that'll sort of be fresh and new today. Okay. First off, we have our architecture. This is something that's been in process for roughly 20 years, and we continue to extend it. One of the questions sometimes I get is, "Okay, you're combining commodity hardware. Why is this any different? You're just setting up racks of hard drives." That's true, but there's a proprietary software layer on top of it that's very differentiating and hard to mimic. This is essentially a high-level overview of what we're doing, which is we break this into building blocks.
These building blocks are a combination of commodity hard drives, but then we layer on different pieces, and eventually, we call this a vault, and then we have a cluster. These clusters can get up to roughly 1.5 exabytes. When we sign a deal like CoreWeave, this is what we're doing behind the scenes in order to ensure that we can hit that type of scale. Outside of the hyperscalers, there's essentially no other company that can do this. This is very unique and again, proprietary to Backblaze. Maybe one other piece to this is the way we're doing this is we're not relying on the individual drives for performance. Instead, we're breaking it up, and in this case, into 20 separate drives and getting the performance we need by aggregating across those drives.
It's a unique architecture, it's a very valuable architecture, and it's something that we're extending even today, and I'll speak to that later. Okay. This is probably the most important slide that I'm going to present today. It's complicated, but it's really important. To the left side here, this is a screenshot of an NVIDIA reference architecture, their data center reference architecture. They recently updated this and added a category for object storage. This is what Backblaze does. Specifically, they added a capacity tier. This is how we've been talking about this for years. It's a really major and important validation of the category that Backblaze is in and what we do. I mentioned WEKA. WEKA's in this middle category, the high-speed storage, which is great. They're a great partner. I'll speak to that a bit.
But this object storage area is something to really focus on, and there's no reason we shouldn't own this category. There's no one that comes close to the scale that we can perform at. I'd keep that in mind, and I'll reference it again probably a few times. Something else I wanted to talk about is, I think sometimes there's a little confusion about what is the storage spectrum? Where does Backblaze play? Who are our competitors? It essentially breaks down to two major media types. You've got SSDs, and you've got HDDs. SSDs are extremely fast. Very fast. HDDs are very affordable, but still fast. The major difference is if you want to get close to the GPUs, the SSDs are what you need. If you want scale, which you need, you're going to use HDDs.
SSDs cost about 6x what an HDD costs. This is where you get into that performance and cost, the economics. You can imagine a company that has a $10 million storage budget. It would cost them $60 million to store that data on SSDs versus $10 million. This is how companies are thinking about us and how they scale. These two solutions are complementary. They're not competing. Gleb had showed this diagram, and this is just a click-in, which is essentially where SSDs lay and where HDDs are as well. You think about the data life cycle, and you go all the way through the process, and you want the SSDs as close to the GPUs as you can get them, because you need that speed for input and output. You don't want to bottleneck the GPUs. They're extremely expensive.
You use your SSDs, but you get that data off into the HDD LAYERS as quickly as you can, and that's where you store it long-term. This is where that 85% of the data goes. The way this works is it's not like traditional compute or storage where you store it, and you archive it, and you don't really read it again. The way this works is you actually feed the data back in. When you get into inference and monitoring, you're generating data that then goes to the HDDs. When you train the next model, it comes back around, and you feed it back into the SSDs, into the model, and then you go through that cycle again. This is how this works. Again, these two are not competing, they're complementary. Which gets into WEKA.
We announced this partnership. I think it's a great partnership and something I'm quite excited about. In this conversation, I'm speaking with their Chief Strategy Officer, and we'll sort of break down where each company fits and how we'll work together moving forward. Hi, Nilesh. Thank you for joining us today. I'm looking forward to the conversation. Doing some really interesting work. I appreciate the time. I guess to start, if maybe you could tell us about what you're doing, what's your role at WEKA, and what the company is doing, and what problems you're solving.
Great. Thanks for having me, Dan. Really appreciate the opportunity and looking forward to working with you. This is Nilesh Patel. I am the Chief Strategy Officer at Weka. I lead alliances and partnership strategy globally. I have been with Weka for almost 5 years now. Weka is building a software-defined storage that delivers memory-like latencies on various workloads and addressing the needs of data layer from GPU clusters to even some of the edge deployments as well. We, as a company, we serve AI native companies, enterprises, neo-cloud, sovereign clouds, and providers who are really driving the GPU infrastructure business today. Anybody who needs to move data at the speed of the compute is what we are able to do, and that is actually the problem we solve.
The traditional storage cannot keep up with the AI workloads and their needs, and they cannot keep the GPUs fed efficiently, and Weka is addressing that.
Excellent. What are some of the changes that you have seen? I know we are seeing a lot of change on the Backblaze side with new AI workloads. I am curious what Weka is seeing.
What we are seeing from a data layer perspective, there are a couple of interesting dimensions that the AI workloads are pushing storage. There are roughly around two axes we see, both in the performance and capacity. Both are being stretched harder than anything we have seen before. GPUs need microsecond access to sifting through the data. Otherwise, they are sitting idle. At the same time, the large amount of data sets from the actual enterprise data to vectorized data, we hear there is a 3 to almost 10x expansion of data in the enterprises. That, plus checkpointing, and then all the generated data is growing in an exabyte range, which is the hard part is not taking one axis. It is like a customer needs both simultaneously. They need to scale fast. Physical storage has lead time challenges and so on.
AI demand doesn't wait for hardware to show up. Some of those challenges are really stretching the need today, and it's putting a lot of pressure in the folks that building the either enterprise AI factories or building the neo-cloud infrastructure that scale. That's where we see the storage requirements are growing on both dimensions.
Yeah. Makes a lot of sense. We're seeing the same thing, throughput, API requests per second, just overall storage needs greater than they've ever been. You hit on a term that's especially important for us at Backblaze on capacity. Can you speak to how you and the team at WEKA are thinking about performance versus capacity with storage in AI workloads?
Yeah, that's a great question. In fact, what we see, like two walls are hitting customers at once. Inferencing and agentic workloads are hitting the memory wall. Long context, multi-term, multi-step reasoning all needing to sit right next to the GPU, faster than memory itself in many cases. WEKA solves that. We deliver better than memory latency at scale, and then there is the capacity wall. Customers need more storage now. Hardware lead times are, they don't move fast enough, and they don't move at AI speed. I believe Backblaze solves that. Instant capacity, no supply chain constraint, no location lock-in. Put together, WEKA extends its global namespace across that capacity, so you can get massive distributed scale without ever slowing down the workload sitting next to the GPU.
In a way, speed where it matters and scale wherever you need it is what I think together we are able to deliver.
Yeah, I couldn't agree more, and I think you're hitting on where the partnership makes a lot of sense between the two companies. I think they're very complementary. Maybe can you elaborate a bit on how you see our two companies working together with these performance, capacity, et cetera?
Yeah. I guess, the partnership, in my mind, lets customers scale capacity instantly and cost efficiently without supply chain delays. I think that's a big part of it. Without all that, without sacrificing the performance and the workloads. It's not really about finding a place to park your data. I think the object storage becomes almost like a, in many cases, a secondary distributed data repository. Because it's part of WEKA namespace, the way we integrate the solution that any capacity you add on Backblaze infrastructure could then very well be part of WEKA managed namespace.
Customers can now not only expand the capacity wherever they want, but they can also spin up compute and WEKA clusters that mount that secondary repository, and now they have an instant scale, not only on the storage, but also on the compute, and be able to, on the fly, get access to the data and so on. This is where, combining the two companies together, the solution can not only address some of the capacity scale needs, but also be able to deliver data to the compute where it exists. The fact that Backblaze has a distributed data access at a very cost-efficient and scale-efficient way, I think that really benefits the customers who are scaling their environment and their AI workloads.
Yeah, I definitely agree. I think it's just such complementary solutions. I have one more question, and I know you're over strategy at WEKA. I'm curious about where you think things are going over the next several years. We've seen a lot of change. Where do you see this heading as you look over things at WEKA?
I think, over the last two to three years, we saw training and model training and so on really driving the demand for data. I think that really pushed a burden on the performance vector, but mostly on the throughput and some of the capacity challenges. As we are seeing the inferencing use cases playing out, inferencing can happen anywhere. Inferencing also pushes, as I mentioned earlier, memory limits of what can be stored in the cluster itself on the DRAM. Having the ultra-low latency, petabyte scale storage, which behaves like a memory cache, is becoming a critical need.
We are able to address that, and that particular requirement is going to continue to go harder and stronger because the whole context windows are growing, the reasoning models are getting richer and richer, and the complexity and the amount of context scale that everybody has to operate at is going to grow higher as well. Along with that, the needs on the capacity and the bounds of where the data resides is also going to be challenged. With AI workload needs to be inferred upon wherever data is. The data locality is not going to be something that AI workload will respect, and they need to be served across the board.
What we see from our perspective is a need for further expanding the inferencing at scale, which is driving the demand on memory like latency and performance, need for data distributed everywhere, so a global namespace, and being able to leverage infrastructure like Backblaze, where there is a tremendous amount of capacity available to operate off of. Then finally, all the data that is being stored across the enterprise, old and new, are going to be all current and required and processed. Being able to access data wherever they decide is going to be another major trend that we'll see for the data LAYERS to come up more and more.
That makes a lot of sense. Well, thank you so much for your time. You all are doing such good work at WEKA. It's a great company. Really excited about the partnership. Again, I think the two technologies really complement each other. I have nothing but respect for all the great work you're doing. Thank you again for the time.
Thanks for the opportunity, and I'm really looking forward to working with you and the team and see what we can do together.
I think that's an excellent partnership, and again, quite excited about it. Gleb touched on this earlier, which is we're working with a lot of great companies today, and we have an existing solution with some of the fastest-growing AI companies in the industry. Probably one of the themes that you've heard is just this complexity demands around performance, scale, economics. We have a solution today that's working well for these companies, and one of the key points is throughput. So 1 Tb per second, this is one of the fastest in the industry. Retrieval time is about as fast as it gets for the HDD layer. We continue to have 11 nines of durability, and we scale at a level that no other company in the industry is scaling at. Which gets into why did CoreWeave choose us? They had a choice.
They could have built, and they seriously consider this, or they could buy. They looked at the complexity and decided they wanted to buy. They looked at the competition and they told us there were essentially two things that stuck out to them. One is our ability to scale. We've been doing this for 20 years, and I've mentioned scale a lot. Another is operational transparency. Quarterly, we release something called Drive Stats. They explicitly said this was one of the criteria that they looked at because it showed what we're doing. Essentially what we have is we report on reliability statistics every quarter, and we've been doing this since 2013 for reliability stats for all of the drives. This is very unusual. No other cloud does this. We're showing everything we have, failure rates and how we're adjusting.
They explicitly said this was a major factor in their decision because it gave them insight into we knew what we were doing. Another important piece of this is a managed storage offering, and Marc will speak to this later. Anuj will speak to it as well. This is really important for Backblaze and a major strategic change for us and a new product offering. There's a lot of talk around CapEx for obvious reasons. This plays to our strength, which is we don't have to have CapEx here. In this case, CoreWeave is handling the CapEx, and our software is running inside of their data centers. Our software is portable. We're not making custom changes for CoreWeave, and we're using the same software for the long tail all the way up to the exabyte level customers.
There's a lot of work going into this, but I'm quite excited, and I think it's really going to be a good business for us. Which sort of gets into what are we seeing now and what's the future. One of the things that we're seeing is AI agents. Everyone's talking about AI agents. If you drive through San Francisco, every billboard you see is around AI agents, right? We're seeing the same thing. There's this large uptick in the traffic we're getting. The most recent data I saw is that non-human traffic is actually growing faster than any other traffic we have on our site. We're adjusting for AI agents, and we have a whole team, and their sole job is to focus on this.
They've been focusing on things like discoverability, because in the past, things like marketing, brand, et cetera, still super important. That's not something that AI agents think about. It needs to be discoverable. So it needs to be machine-readable docs. We need to have integrations with the best open source tools. We need to integrate with AI platforms. We need to have what's called an MCP server, which is essentially a protocol for how these agents talk. We're doing all of this work, and we're seeing a large uptick in the traffic we're getting. I think we're making some good decisions here, and we have an engineering team whose focus is just on this to ensure that we're getting this right. Last up, I wanted to speak briefly to what we're working on in addition to the things that I've already spoken about.
This isn't public yet, and so now it is. These are internal targets that we have. Again, performance, scale, economics, and this is what we're seeing from our customers. We're committing to internally and now externally, 2x increase on throughput, 5x increase on API requests, and a 20% increase on drive capacity. We're actively working on this. High confidence for new deployments moving forward. These are the metrics that we're aiming for in addition to a lot of other interesting work we're doing around managed storage, AI agents, enterprise features, a number of things. But this is hard, and so maybe that's the thing that I would leave you with. This isn't something that other companies can come in and just copy. We're extending our existing architecture.
We're building on top of something we've been doing for 20 years with tens of thousands of optimizations, reducing bottlenecks. I get the question, what's the moat? This is the moat. We're working on really hard things, and extending our lead in an area that we're already leading in. Next up, I have my friend, our CRO, Anuj, and he'll cover where we're at with revenue.
Thanks, Dan. Morning, everybody. Tough act to follow, huh? The CEO set the stage for the market opportunity. The product guy is telling you we have awesome things on the truck already and even better things to come. I have been here about 4 months. What do I know? I will share that I saw this opportunity as a really market inflection point, and what I truly see is Backblaze is one of the very, very few companies that have really the scale, the performance, and the capability to take advantage of this inflection that we have. Before I go into it, I think my main section is about, as Gleb said, from signal to revenue. I think it is important I define what the signal is. For a sales guy at heart, I will say I want to break it down into 4 simple things.
What is the signal that we are going to talk about? What we need to do is where is the demand really coming from? Why is the signal really working and why you should believe that the signal is really, really strong. From there, I will take you to who are these people that are coming for the signal and what we can do to convert them, finally, into how this engine comes together to convert all this demand into something that we can repeat and scale. You are in a room full of, hopefully, math guys. I was a math major as I was growing up, and I think you will appreciate this. One of my longtime mentors said, "Anuj, life is really simple. It is just a math equation," right, if you break it down.
Let us just start with the numbers, since I am sure you will appreciate this more than anything else. What we look at in a go-to-market engine, in a signal, is what is it doing over a period of time. What I am sharing with you here is our own data over the past 12 months. This is data from our AI customers, and when you see the conversion rates, we are winning at least 10 points higher than conversion rates of a non-AI customer. The deals are landing. This is really important for us because we want to make sure we are going after the most addressable opportunity. The deals are landing 7 x. This is not an error in the slide. It is 7 x the deal size of the rest of the customer base.
What makes this signal even stronger is once they land, within a short period of time, they actually compound. This is the beautiful formula that I would love to have with every piece of the business. This compounding happens within a very, very few months. This is the total data in aggregate over the 12 months. Why do not I give you an example behind the scenes? I am truly fortunate Olya is going to join us from Hume AI as well. She is going to share her story. This is another story because this is not just a unit of one. This is from hundreds and thousands of prospects that we see in our base. What you see here is what we are coming back to the table of what we shared with you in the first quarter.
If you remember, we had said in our Q1 earnings that we had a customer that came in with AI, and that deal converted in about 11 days. What you see this is this actual customer came to us from another customer reference, and we were able to do from a trial to the first land within 11 days. Great. It's almost $1 million. Life is good. You fast-forward, less than 12 weeks later, the same customer compounded by doubling their capacity. Now we are in the third quarter. This customer's already gone to about 2.5, 2.7, was counting, almost 3 x. If you aggregate the total, we're looking at about a 4x compounding from this place that we started, from the provider that we had. Now, why is this customer putting so much data on it?
I think Dan alluded to, Gleb alluded to, AI really has a lot of ingestion, a lot of data collection, and this is a continuous stream as they are running inference models and want to have a capacity tier to take care of it, as long as they can have this available with high throughput for the base that we have. Now, we've talked about Backblaze's history in terms of why it really, really works. This is the second big piece of what I'd like to share with you. The why that I would say has worked for Backblaze has always resonated. You know from the beginning of time, we have really what we call disruptive economics. Our value proposition has been tried and tested. It's true against the hyperscalers. We also offered a very, very predictive model.
You get no egress, and that gives you a way to make sure that you can put your data where it is, but also you want to make sure that you have something affordable that you can run the models from. In the AI space, it's slightly different. What it's given us is it's not something new, but it's really given us an opportunity to expand the use case and the value that we have for our customer base. In AI, there's three specific things these customers are looking for. They're not just looking for low cost, but what they're looking for is sustained high throughput. That's what we call performance. At the end of the day, you're trying to do price economics, but you want to make sure that the performance doesn't suffer when your models are calling it for inference.
High performance, high throughput is really, really important. The number two thing is rapid scale. I think Dan talked about it in the Weka conversation. It's not just a tiering of the data, but it's also instant availability of that capacity because GPUs are very, very expensive when they're sitting idle. I think there's a whole term of token economics that's come around. At the end of the day, we don't want to keep these GPUs idle. Yes, the flash has its role to play, but hard drives and our managed storage service has a role to play to make sure that the capacity is constantly available at high throughput for the GPUs to be continuously working. At the same time, what you do want is you don't want your data to be locked into any one particular location.
I think Dan shared the one slide that Dan you shared in terms of if you take away one thing, which is all these neo clouds that really didn't exist till maybe 5, 6, 10 years ago, you have so many choices. At the end of the day, these AI builders really want to keep the data where they choose to, they shouldn't be locked in in any one place. So making sure that you have architectural freedom to keep the data where you have, make it available at scale, and make sure that it's available at high throughput for this AI customer base is really, really important. So what it's helped us do is now we see an expanded use case in terms of our value proposition, that's probably the biggest why. Now I'll probably get into the key pieces.
When I think about a go-to-market team and a go-to-market engine, what do we do, wake up in the morning and try to do from what it is. Because you clearly know there is a value. You clearly know that there is a strong signal. How do you make sure that you start converting that signal into something that you can build an engine from. The first step is really to make sure you're segmenting it right. A lot of sales organizations will show you triangles where there's strategic, enterprise, commercial at the bottom. The way we look at the world is really, really simple. If you see on the left, what really it is is the core workloads. Backup, disaster recovery, security. They have always existed, they will always exist, and there is a constant need for backup and data recovery and disaster recovery.
That stays true. We continue to grow in that market. We continue to grow above market rates. Fantastic. The other two adjacent as what is giving us the addressability of the use cases I just talked about, the high throughput, the performance, the scale, the instant capability, the ability to make sure that we can move the data, keep the data where it is, it gives us two clear segments in the market. One is what we call AI builders. These are GenAI media companies. One of them, Olya's going to be here from Hume AI. We talked about a few others in the GenAI media space or the physical AI space that are really training insane and large amounts of data to get to where they are.
I would say the other big part is the neo clouds, where we have about 200 of these neo clouds that didn't really exist about 4 or 5 years ago. But what they need is instant capacity, but also the ability to offer managed storage as a service. That is the service that we can now deliver for them, either coming to our data center, which you can today, but also as part of a managed service, which is how CoreWeave leverages us. Both of these are now giving us new addressable markets to go after in this AI space. That's the clear segmentation of the who we really address when we go to market.
From the who, the next logical question you will say is, well, how do you convert this who that you are trying to get to the how that we want to make sure that all the activities we are doing from the signals that we are seeing to convert into a repeatable, I would say, process, as well as a deliberate process of connecting with the buyers in all of those three segments. To think about it, I will keep it really, really simple. There are three main parts. We have to make sure that we are educating the community. We want to make sure that we are educating the community, we are engaging them where they are going, and we are also making sure that they experience the product as they choose to. In education, you have seen us, we have talked about it.
For the last, I would say, 13, 14 years, we have been producing quarterly Drive Stats. This is incredible information from a cloud storage company that no other cloud provider actually provides. This helps us build a community of engagement where they actually see us being really, really transparent, not only in our pricing model, but also in our development model and scale model in terms of what we do. The engagement then really gets to we want to meet where the buyer is. Today's buyer is not sitting necessarily in a physical, let us say, events. Of course, we do those. They are really engaging more in social, specifically new social channels like Reddit, Stack Overflow, and we want to meet them where they are.
That is the Flamethrower program working with hundreds and hundreds of startups on a daily basis to make sure that we can engage these startups and get them to meet them where they are and help them understand who we are and what we stand for. Finally, I am the last guy that somebody like Dan Spraggins wants to talk to when I call and say, "Dan, I am Anuj from Backblaze. I would love to talk to you about this cloud storage thing." CTOs, CPOs, CIOs, they rarely talk to the sales guy the first call. We want to make sure that they can get their hands dirty in the platform on their own with no sales assist. That is really important because we want to make sure that we are available in all these channels, whether it is Hugging Face, whether it is open source SDK tools.
We want to meet them wherever they go and however they want to engage, making sure that they can touch and feel the product and the service before we ever call out and say, "Hey, we see that there is a signal. You might want to engage with us." This is a concerted, continuous process. Think about this education, the enablement, the experience. It is something that we do on a constant basis.
We want to make sure the signal is not random, but it is something that is repeatable, something that is measurable, something that is instrumented, and we can see the results, and we can keep tweaking in terms of what it is. We are not trying to convert humans to robots, but we are trying to make sure that we can actually see what is happening and make sure that we are available in all of these channels.
Finally, as part of this, how do you make sure that all this wonderful signal we have and all the engagements that we are doing converts to something that, as we often say in sales, there is something on the truck and I got to go and then convert it. The big thing here is I have got to convert from all this demand that I see into something that is a repeatable revenue engine. This is the part where I would just say in sales life, it is just three, two, and one. I showed the three segments that we have. The demand comes through any of the three segments, and we have two motions to address this demand. The two motions are self-service and then call it sales assisted or direct sales. The self-service, you choose as you come. You can just pay as you go.
You get onto the platform and you can scale as much as you want. We have offered the service and we continue to offer it as a viable option. On the direct sales, specifically in the AI builder and Neo cloud, this is where we see the difference, where the customer does engage a little bit on the pay-go, but very quickly they want to have really high-value workloads that they want to do. That is what we want to make sure that when they test, they get the experience to convert in days and weeks and not months and years. The direct sales team is a dedicated coverage model to make sure that when we see the demand signal for high value, we attach the sales team. Both of these motions are powered by an ecosystem that is in two pillars, sell with and sell through.
I will come to it in a quick second, but the goal is that these two motions are working in tandem to make sure that we are addressing the demand that is coming in. To make sure that this actually converts from these two motions, what we have is a very strict, I would call it, sales process, and this is LAYERS. The LAYERS is just an acronym. It sounds simple, but the goal is we want to make sure when you come in, the trial converts to something of an opportunity that is landing.
We want to make sure that what you committed for and what you want to consume as part of the data is what the adoption team takes care of. Each of these are dedicated teams to make sure that we can then take you through the journey of expansion once you have got a good workload on.
I am sure there are additional use cases, all about cross-sell and upsell, and then making sure that the revenue line, the retention line stays strong. The glue that kind of pulls this all together is our support team. Because at the end of the day, you came in for a managed service. You did not come in just for software that you were running on your own. So this storage support glue really pulls it all together to make sure that we have this process, which is quite rigid, something that we measure, something we want to make sure that we are seeing the conversion rates and making sure that we can address them. But this process kind of keeps all those segments, the motions true in terms of where we want to go.
I'm sure the next logical question you'll have for me is, great, Anuj, Dan said he's going to grow 20% on the capacity, 5 x the throughput, et cetera. Similarly, as we are going to grow in the go-to-market team, I would say it doesn't really come from just infinite headcount. I think that's a bad day when I start like Marc. I just want to go and hire 50 more people because I need to double the people that we have. How do you think we're going to get there? This is, I would say, the final piece of that engine, which is to make sure that we are really building a deliberate process and a deliberate motion with a partner ecosystem. It's really, really simple. There's two pillars. There's a sell with. You saw an example with WEKA.
The goal is we are doing validated reference designs, validated architecture, so it's actually tested and integrated at source. The goal is we don't want the customer to be spending the time to try to integrate two parties when we know SSDs and HDDs work in tandem, and the cloud capacity is something that adds on to help the capacity in Flash to make sure that you can run your models. We want to make sure that this is tested at source so customers can have a validated design to go with. WEKA is just a prime example of that. Once we have those designs, what we do is we convert them into a downstream ecosystem of partnerships. These are resellers, distributors, cloud marketplaces, basically allowing the customer to have a choice in terms of whichever partner they choose, they can engage with us on.
The goal is these two work together. The more validated designs we have, the more options we have for the customers and partners to engage. This is the deliberate motion that we are putting together. This is a relatively new muscle, I would say, for Backblaze. However, we've got a good part of this engine flowing already, and this to us is really the big motion for us to scale beyond, let's say, infinite headcount that we would have. To wrap it up, as I said, I'll try to talk to you about these signals that we believe are really, really strong, and I hope I've given you some proof that the signals have some depth in them in terms of we really measure and instrument every part in terms of what's coming in and how fast it's converting.
The second piece of that is really in terms of making sure that these signals don't sit in random and we have a process to convert the signal to a demand that we can then finally convert with a process that we can then scale with the ecosystem. Just want to thank you for your time. Welcome Mimi on stage.
We're going to take about a 10-minute break right now, and then we're going to welcome Hume to the stage to do a fireside chat. Please take your time and get some refreshments. Excited to bring on stage our guests for a fireside chat, and I'm going to let Anuj take it away, and you'll get to learn more about Hume.
Okay. Well, welcome back. Hope you guys had a good break. Don't worry, this is not a sales call. I'm not going to ask you, Olya, which I am still learning how to pronounce her name properly, but really, really first, thank you for coming in here. Great to have you. Thanks for obviously being a customer and a partner. I'm sure we'll talk a little bit about Hume, but I just wanted to formally introduce Olya. I met her for the first time yesterday in person. We talked a few times before. But she's got an incredible journey in basically being, and come up to the Chief Product Officer at Hume. She's responsible for all of the strategy, all of the product. I think your portfolio has everything from strategy, product, cybersecurity, the current and the future. It's a lot.
Thank you so much for spending a little bit of time and being with us here today and talk with our analyst. Why don't we begin a little bit about, tell me, I show this audience, I talked about GenAI media, but that's really, really, really high level. If you can talk a little bit about you, Hume, little bit of the journey, I think it'd be great.
Yeah, absolutely, and thank you for having me. We were practicing names back and forth, so it's Anuj. We're getting there. I have a background in machine learning, initially for healthcare, and then transitioned through highly regulated sectors, and then made my way more to machine learning product, and then now more voice AI. Hume, as a company, is a really interesting organization because we have half of our team is AI research. A lot of what we do is we train models to improve other models. We also have data solutions, which means we need to process a lot of data. Every time that we process data, we create new data. Backblaze has been quite helpful in our needs there.
Additionally, what we have been doing in our goal to improve voice AI, because we initially started out with 10 years in semantic space theory research, really focused on affect, which is prosody. It is understanding of how you communicate. That would be emotion and how you can derive that from acoustic signals. We built models around that. We built the infrastructure around that, everything to process audio, as well as the systems to evaluate and improve other voice AI models.
Yeah. Yesterday, when we were talking, I think one of the things that absolutely fascinated me with a couple of key things that you mentioned. One, just the layers of complexity that a simple thing like voice has. I think, when you think about it is not just the amount of data, collecting transcripts, and trying to put something in. Just to understand from the tonality or how we are actually feeling when we are talking. We can hear it on the phone, but you guys are actually trying to pull all of this together. Can you tell us a little bit about that complexity and what it takes to actually put that model together?
On the surface, voice AI seems like one dimension.
Yeah.
Right? Same as text. It is predominantly one dimension. When you actually peel back the layers, it is so multidimensional. Think about us speaking right here. I have certain contexts that I came into this conversation with, as do you.
I respond to your inflection changes. I account for the background noise, of which we have none in this room. There are all of these other components within that. There is the warmth of the conversation. All of these things are actually very challenging for AI to understand. There are two core things that we solve for, which is, number one, does AI truly understand? Number two, does the user feel understood?
Right now what we have been seeing is voice AI is at an inflection point where previously you had predominantly text interactions. We type on our computers. Now the new modality that we are seeing for people to interact is voice. You have to solve for all of that complexity, and you need to ground it in real-world use cases.
Which is quite challenging to do.
True. True. That itself is one dimension, but the other dimension you also mentioned was just normal models, like why wouldn't just OpenAI, Anthropic, all of these guys, they also have a model. But you said something really interesting that caught me yesterday, which was those companies are only interested in the research, I think is what you said. They are really just doing it for raw research. As you guys are building it in layers and layers of modality, that really gets down to.
I think there's two parts to that. I think OpenAI and Anthropic and all of the large labs are fantastic.
Right. Sure.
They're also large labs are our customers, so we love them.
Of course.
But what is interesting there is they have a bit of a different focus, right? They want to push towards AGI. They want to make sure that they have the smartest, most capable models. Our goal is to improve all of the voice models, the whole ecosystem. You need models that can evaluate those models, and to do that, you need to be really, really good. I think there is the second half of it, which is as enterprises are coming up and sovereign environments are coming up, I think what we are noticing is that a lot of people recognize that they have a lot of data.
They want to build their own models as well. We actually get a lot of outreach to us about, "Hey, we want to do XYZ as this enterprise to build an AI model-
that optimizes for speech, and we want to make sure that it is performing well, and we want to make sure that it has all the emotional understanding within it." Which is really interesting, and perhaps relevant for you guys as well. Because as we are thinking about this model landscape, we are not only thinking about OpenAI or Anthropic, we are thinking about sovereign clouds, enterprises who are now trying to leverage all the data that they have to build their own AI models.
It is like models, platforms, like you are doing both things at the same time, and it is constant growth. I think I called it compounding in my section.
Yep.
I think this basically just defines it in a way that's really contextual. Fantastic. No, appreciate that. Well, let's bring it a little into how you see what you're building and some of the reasons you came to us. Did any of the things we talked about, which were the things that really resonated for you in helping you make the decision to Backblaze? Now how do you see this relationship in the context of everything that you're building?
Yeah. I think Gleb was actually walking through the journey, and I was like, "Oh, that's our journey. That's right." Essentially, we go through and we collect a lot of data. Data is obviously very important for AI. It's a hungry hippo. We collect that data, we process it, and enrich it. We actually had a lot of disparate environments where we would store our data. Not only did we want to consolidate it, but it was really important for us to not get taxed for using our own data.
Which ties to Backblaze egress, I guess, lack of costs.
Every time that we go through and process our data, and we have petabytes of data, and every time that we go and process audio data, what we create is transcript data. We have 600 + emotional tags, expression tags that we put on top of that. We segment it out. Every time we process data, we create new data.
Right.
We have to do that not just on a continuous basis, but we have client deliveries as well. Our clients come in, and they want to process a lot of data too.
A lot of data. Yeah.
When we have to do that, we need to be able to transfer all of that very quickly. What was really important to us as well was not just our ability to transfer it or our ability to store it all in one place, but the third thing was how do you make sure that we don't have to have our own continuous compute that's up and running?
We have serverless compute. What we do is we have this data store, and then when we need it, we spin up the GPU clusters, and then we transfer it over to process the data accordingly.
Right.
Which gets used for our clients or for our models as well.
Yeah. So the high throughput we talk about, the instant availability to capacity, those are all good things for you to have that readily available.
Yeah. And honestly, the Backblaze team has been quite helpful. When we were processing a lot of data in one go, they really made sure that we had really great throughput there. So it was quite helpful.
Excellent. Well, that's great. So how do you see this partnership also, I don't want to leak the stuff that you just shared with me, but if you're comfortable-
Yeah
That would be great. I did not realize that there was some commonality here, too, but how do you see this extended partnership and this relationship continue to foster?
Yeah. I was just telling Anuj that we actually also use WEKA. Typically what we do is we have our core data store with Backblaze, and then when we need to use it as part of our models and processing, we have it in our GPU cluster, and we have the WEKA storage there as well. I think as you can imagine, as the tailwinds of folks switching over to voice as a modality pick up, and they already are, we obviously anticipate processing a lot more data. It is really helpful to us that both Backblaze and WEKA work seamlessly together, and we are able to continue growing and supporting this hungry hippo.
Well, Anna, we got to make sure that the integration really works. She is our Head of Partnerships and Channels, so we just want to make sure we get all of that validated design really cooking so you do not have to do all the hard work.
Great.
Oh, excellent. I think you gave us a lot of really fantastic, valuable insights. Really appreciate you being here. Anything else you would like to share before we wrap?
No. It's been fantastic. Thank you so much to the Backblaze team, and we are excited for our continued collaboration together.
Awesome. Well, thank you very much.
Thank you.
All right. Hello, everybody. Good morning, and thank you for being here. I'm Marc Suidan, the CFO, Chief Financial Officer. Gleb told us you got to show up looking your best, so I figured, listen, I either got to do some biohacking, get more hair on the head, or pay $40 and get the AI to do it for me. I think it went a bit too far, so I'll have to do a bit of refinement there in that picture. Okay, let's get rolling. I think it's a great day. We really wanted to get out adding new faces to the discussion. Generally, Gleb and I, and Mimi have been in extensive discussions with a lot of you, so we really wanted to get a lot more of the extended team so you could all meet them.
From a financial standpoint, what this all translates to is the two things we have always focused on, more growth and operating leverage. I am going to walk through how we have delivered on what we said we are going to deliver and how we are going to do more of it. When you step back and you look at everything we promised, 2 years ago, we set out a few goalposts, and we have delivered on all those goalposts. The first one we said we are going to do is strengthen the balance sheet. In Q4 of 2024, we did a secondary offering that was oversubscribed, and we did a restructuring, driven by a zero-based budgeting exercise. In that exercise, we reduced the OpEx, and we reallocated some of those savings to invest to accelerate growth. We said we are going to re-accelerate B2 growth.
In Q1 of 2025, we started to re-accelerate B2 revenue growth. Then we said we are going to get B2 revenue growth to over 30%, which we did in Q2 of 2026. It was 34%, and we said for the rest of this year and next year, it will be 40% or higher. We are delivering on all the things we said we are going to deliver, and we said we are going to do it in a profitable way. Under the capital lease model, where our CapEx gets financed by capital leases, we turn free cash flow positive exactly when we said we would. We are well in that motion there. We always said let us use B2 revenue growth and our free cash flow margin to judge our Rule of 40 scoring. In Q2, that was 34 plus 8, so 42.
We hit 42 in Q2 of 2026, up from 18 a year before that. Delivering on the Rule of 40 score that we promised we would deliver on. What is driving that? The accelerating revenue growth, it is translating to a lot of operating leverage. Get a healthy gross margin of 63%, which is really good for an infrastructure as a service company, combined with really disciplined OpEx management, has translated to tremendous EBITDA margin improvement. The operating leverage is well in motion and in action. Underlying that is the mix shift. A few years ago, computer backup made up the majority of the business. Right now, B2 forms 62% of the business and will probably be around 75% somewhere around middle of next year. The mix shift is well in action, and the underlying fundamentals of the B2 platform is really attractive.
ARR growing 39% year-over-year in Q2. Really strong net revenue retention at 113%, and the gross customer retention at 89%, you are talking customers that stay with us with an average of 9 years. Phenomenal fundamental metrics. The business for B2 has historically been consumptive, very pay-as-you-go. As we are accelerating this growth, we felt it was prudent to start getting into longer-term contracts. Our revenue performance obligations, which are committed contracts, are gone up more than 5x year-over-year. In the marketplace, just given the supply constraints, a lot of customers actually prefer to be on a committed contract because they know that we will give them and guarantee them the capacity. That is really helping in giving us a lot more visibility in the growth and where to put our investment to fuel that growth. Unit economics.
Over the lifetime value of a customer, we deliver 60%-70%. It is currently 70%, but just to be conservative with the changes in hardware prices, because it takes time to adjust the pricing models, we are saying it could be 60. That is the value we get out of all incremental dollars from our customers. They generally stay with us for 9 years. Every cohort since B2 launch in 2016 has grown their data consistently year-over-year. You have to obviously build the CapEx upfront. It takes us in 2 years to pay back the CapEx, and then you have got the direct cost. Those are roughly, they are step function variable costs, but it is effectively data center rent and power, telecom Data center technician, customer support, sales incentive comp. That is pretty much our variable cost.
That carries throughout, but the CapEx is upfront, and then you recover heavily over the following 9 years. You look at our EBIT performance. We will continue to improve our EBITDA margin, and also operating margin, GAAP gross margin. All of those will continue to improve as operating leverage continues to feed the bottom line. Looking at the debt, we raised our convertible debt a few weeks ago. We raised $201 million. Also oversubscribed, 0% coupon rate. That strengthens our balance sheet for the next 5 years and helps us to invest. We are putting it all to CapEx. We have to put the CapEx as we have the committed contracts and we have the revenue coming in, so we got to invest in the CapEx. The difference is, instead of paying 12.9% interest expense on leases, we would be paying 0%.
If you think about the cost avoidance over 5 years, that basically pays off half the principal of the debt right there. We felt that that was a prudent move to have a stronger balance sheet for the next 5 years to build up this cash position and deploy it on CapEx, and then we would spring load and come out of that with a much stronger free cash flow margin. The metrics to watch for us over the next 2 years will continue to be EBITDA and operating margin. That is where you are going to see the operating leverage continue to improve. For free cash flows, we did turn positive in Q2, but now we are going to deploy this capital. Operating cash flows will improve, but when you pay for the CapEx in cash, you got to deduct it from those operating cash flows.
For the next 4 to 5 quarters, the free cash flows will turn negative, and then we are spring-loading the free cash flow margin to come out much stronger coming after that. When you think about all the P&L key levers you got to look at and what the outlook is for each one of those, I will start off with revenue. We said revenue will grow 40% or more for the rest of this year in 2027. By the way, that 40% assumes the same existing guidance philosophy, which is no large deals, and we are defining large deals as greater than $0.5 million. We did three in Q1. We four in Q2. Once again, this number assumes none of these large deals and assumes no customers are committing over their minimum. We are just putting in the minimum committed contracts there.
It is a pretty prudent 40%, is what I am saying, for the next six quarters. Gross margin, as we are deploying the CapEx, and of course it is more expensive, there is going to be a bit of a lag between when you deploy the CapEx start depreciating and when the revenue comes right afterwards. That lag could reduce gross margin by 300 - 600 basis points through the middle of next year, and then it starts recovering from there. Then as you get into 2028, Dan and Gleb spoke about the managed storage. The managed storage resembles a lot more a SaaS kind of gross margin. So they will start to be accretive into 2028 on the gross margin as that comes to feed in. On the OpEx side, we will continue to manage OpEx in a very disciplined way.
We will make targeted investments in R&D and sales and marketing, but OpEx as a percent of revenue on the current trajectory we are on should continue to be where it is or improve as a percent of revenue. So a lot of operating leverage there. As it relates to adjusted free cash flows, like I said, for the next four to five quarters, you will see the adjusted free cash flow turn negative as we deploy the CapEx, and then we spring-load it so that it meaningfully improves coming after that in a very healthy way under the current construct if we move back to the capital lease model. So key takeaways. We have better growth visibility supported by committed contracts and proven long-term earnings power, and we are going to keep powering that up. So those are the key takeaways for the financial section.
With that, Mimi, we are going to turn it over to the Q&A session.
Yes. So let me invite the execs up onto the stage.
We're not mic'd. We're going to need the mic.
You first, and then you can do Jason if he doesn't share.
Good afternoon, guys. Still, I guess, good morning. Ittai Kidron from Oppenheimer. I appreciate the presentation today, and great to see the acceleration in the business. Maybe a couple for me. First, on the CoreWeave transaction, clearly a landmark deal. Can you talk about the milestones that you have to go through in order to ramp the managed service in fiscal 2028? Clearly it has huge potential from a margin standpoint, as you mentioned, Marc, but clearly there's also a lot of work that needs to be done behind the scenes to get that going. So we greatly appreciate if you could provide some color on that. Then on the go-to-market side, thanks for clarifying the three different cohorts. It was very helpful to understand where the efforts are focused. Maybe a little bit more of a philosophical question here and now.
Where in those three segments do you feel your best aligned and where you have more work to do, number one. And number two, if you had an incremental dollar to put, since Marc now has a lot of money in his pocket, if you had another dollar to put in the go-to-market organization, where is the next dollar going to?
Thanks, Ittai. Let me touch on the managed storage and then actually I'll have Dan also maybe expand on it a little bit. The managed storage has two minimums, right? The first minimum is at the end of next year. We expect it to ramp basically starting in 2028. The very simple concept on it is CoreWeave comes to us and says, "We would like you to deploy X amount of storage in this location." Then we hand them a bill of materials and say, "Please buy all of the following equipment, per our specs, and hand it to us." Our people manage the equipment, the racking, stacking, and managing it inside of that facility, and we deploy and manage all of our software remotely. That's just the general concept, and then maybe Dan, if you wanted to expand on it.
Sure. We are focusing a fair amount of energy on this managed storage concept. That includes things like deployments, releases, ensuring no downtime, telemetry, dashboards, things of that nature. We're adding a number of features there. Yeah, I think that's high level what we're doing though.
Yeah. Maybe just one thing, just terminology-wise. We're calling this managed storage, not managed service because it's not so much. There are people obviously involved in racking, stacking the boxes, and there's also people involved in the sitting on our side of it, eyes on glass, managing the storage. But it really is a managed storage offering. I think it's one of the key unique differences, right? There are software platforms out there, including some open source and closed source platforms where you can use software. But the part of the reason why CoreWeave chose us for this and some of the other conversations that we're having with other neo clouds and other AI infrastructure companies is that it's one thing to just either buy a piece of software or take an open source piece of software.
It's a whole different thing to actually manage and run storage at scale. That's why we're calling it a managed storage offering. Okay. Anuj.
I am mic'd up. I think you had two questions, just so I make sure I play the questions back so I got it. I think you said in the three segments that I showed, which of them probably is growing fastest, or like where we-
Where are you best-
Where are you best fit. And then the second is if I had infinite money, where would I spend it?
Yes.
Let us just go with the first part. I think the reason I shared those three is because the good news here is there is no one segment that is carrying the weight off the other two in any way. They just have slightly different velocities, slightly different motions, obviously different buyers and different needs that we service from the same core platform. Like I also shared, the core workload, the backup, disaster recovery, it has been our most stable business. It has been there for a long time. It will continue to be of service. I think that is probably the one that we have had the fit for the longest time. And I think something Dan mentioned that I want to have recalled here is when we acquired these customers of AI builders or AI infrastructure, we actually did not build anything new.
It is important because what we realized is the value that the platform already has is really good throughput and a really good performance per dollar. I am inverting the equation. It is not price performance, it is performance per dollar. Because even when Olya was sharing, it is all about consistent throughput that they would like to have so that they can feed the GPUs when they need it. That combination is literally, as I would say, is we have these things on the truck to be able to deliver right now. Would we like to have more performance? Sure. Would we like to have more economics stage? Sure. But there is nothing impeding us from addressing this market right now. Does that help? That is kind of the first one.
We obviously have dedicated teams to make sure we take this in and we convert the demand into the things that we would like to do. If we had, let us say, more investments, and this is probably a tricky one, but I would say at the end of the day, what I want to do is the investment I shared with you was a little bit more on the ecosystem side, is, we cannot grow infinitely by just adding more headcount, having more people call, making available on more social. I need the whole ecosystem to talk about the full value proposition.
Because customers have a journey that they go through. We are not the only game in town. I am pretty cognizant of that. As they go through the journey, I want to make sure that from every angle, they actually hear of us in that combined value proposition.
That is really key. So what we stand for stays true, but they hear end to end. Yeah.
Hey, guys. Jason Ader with William Blair. Gleb, I guess for you, question is there any drawback to having your storage tier separate from where the compute is for a customer, in terms of network latency, I don't know, things like that? That sort of begs the second part of my question, which is, do you envision CoreWeave-- How you answer the first question is going to matter to the second question, which is, do you think CoreWeave long term will be more on managed storage structure, or, do you think there will be sort of still reasons to, assuming they can get as much capacity as they need on their own, do you envision that would shift more towards managed storage?
Yeah. Good questions on both of those. This is the capacity tier, latency is not as much of a question, right? You heard Nilesh from WEKA talking about the microseconds needed to feed the GPUs, right? You don't want that being far away from the GPUs because you're talking about microseconds, right? As a capacity tier, the time in general to get data off of the hard drive and get it out versus the time it takes it to get from there to the storage, the latency doesn't play that big of an impact for the most part, right? You don't necessarily want your capacity tier sitting on the West Coast in the U.S. and having GPUs in Japan, right?
But in general, the whole idea of the ecosystem is that you're going to want to use multiple providers for your model building, for your inferencing, for all these different pieces. And just by the nature of that, the data has to flow between them. It almost doesn't matter where you put the data. It's got to go from one place to the other. The first part of it is that in general, as a capacity tier, having your data somewhere, as long as it's, call it like in the same country or in a nearby country, right? It's like that kind of thing, then it's fine. As far as the CoreWeave part of it, someone asked a version of the question and let me tweak it a tiny bit and then I'll come back to it.
Someone asked, for the deal that we did, it's 2/3/4-ish on our infrastructure and about a 1/3-ish quarter on their infrastructure. And they said, "Why did they pick strategically that split?" And my answer was, they didn't, right? It's not that they said the right answer for us is 2/3, 1/3, or 3/4, 1 quarter. It was more that they looked at how much availability they thought they needed in what timeframe. We had the ability to deploy that on our infrastructure. Then they looked at how much they're going to want in their infrastructure and when, and that was that part of it for that. So it's not about the mix shift, it's more about the timing of ramp. CoreWeave has dozens of data centers.
I don't think that the likely scenario is that we're going to deploy a capacity tier scale in every single one of their dozens of data centers, right? I think that what they're going to do, just like what probably most of the 200 AI infrastructure companies are going to do at the capacity tier, is pick a handful of major regions to deploy large-scale capacity storage to service the broader set of regions. And then they may have 20 data centers across the U.S. or across Europe, and they'll pick one or two locations where they're going to stand up a capacity tier that'll service all of those locations.
Thank you. Erik Suppiger with B. Riley. On the AI infrastructure versus the
AI builders?
Builders, yeah. Seems like as you have described it, an AI company would be inclined to buy storage or capacity from you rather than going through the Neo Cloud. What is the bigger opportunity? Is it the AI infrastructure or is it the AI builders? What kind of penetration do you expect to get within the Neo Clouds?
Yeah, it's a good question, Eric. There are about 200 of these AI infrastructure companies. I do not expect that that number is going to go to tens of thousands, right? There are thousands of AI builders, and I expect that number to continue to grow, right? I will tell you that some of the AI builders we see doing both. They use us directly and they use data on the Neo Clouds through us. There are different reasons for that, right? They may use the AI through the Neo Cloud directly because they are part of the orchestration of their workflow inside of that Neo Cloud, and that is easier through their dashboard for that. They also use us directly because that Neo Cloud is not the only one that they are using. They are also using the hyperscaler to inferencing clouds, et cetera.
They want data that is sitting independent of the Neo Cloud so that they have easy access to the broad swath. But they also want to use the data inside of that Neo Cloud because they can orchestrate what happens in that Neo Cloud through a single pane of glass. We see companies doing both. Hard to fully say which of the two sides is going to be a bigger one. The AI infrastructure one is bigger chunks, right? Because they are a smaller set of companies that bite off in larger chunks. The AI builders are ones that we see having a more steady state, larger dispersion, path to growth. We talked about how we had signed. We announced one in Q1. We announced obviously CoreWeave in Q2. We have now just signed another one, right? So we are up to five.
We had talked about two before. My general belief is that for almost every AI infrastructure company, if they are going to be successful long term, they are going to need storage. It is hard to just say you are going to be a pure-play compute provider and not play in any part of the workflow that your customers use. If you think about it, the customers, like you heard from Olya at Hume AI, right?
The customers have data. It has to be. If they do not use data, you do not have AI. So they are going to have data, they are going to keep that data somewhere. If they are not keeping it with the NeoCloud because they do not provide them that opportunity, they are going to keep it with a hyperscaler or with us, and if they are keeping it with a hyperscaler, they are literally using the direct competitor.
I believe that the 200 of them are going to almost all either offer it or go out of business long term. I think we are best positioned to be that provider for them. Would I love to say that we are going to get 100% of them? Sure. But I think that it will take some time as they go up the maturity curve of what they do. There was this interesting thing which was counterintuitive. When we first went down this path, we said the ones that do not offer any kind of storage today are the ones that are going to be our best targets. The ones that do offer some kind of storage are not, because they already offer something that they are not going to want to shift. What we found was it was actually inverse.
The ones that offered storage, customers were coming to them saying, "You can solve more of my actual need." But then they started feeling the pain of not having a solution that fully solved what they actually needed to solve. So they were more open to working with us. The ones that did not offer storage yet were still just racing as fast as they could to try to figure out, "How do I get enough power? How do I stand up the GPUs? How do I make it so that I can orchestrate this for my customers? How do I get the workflow, just the basics set up?" So they are not yet at the maturity level of trying to service the broader issue the customers need.
I think as they move up the maturity curve, we will have more and more opportunity to penetrate more and more of the 200. That is a hard thing to know at this point. But I think we are best positioned to do that.
Eric Martinuzzi with Lake Street Capital Markets. I wanted to follow up on your comment there. It is really a question around back upstream on CoreWeave, and maybe it is more a question for Dan. But you talked about the scale as one of the key reasons that CoreWeave went with Backblaze, but you also talked about the operational transparency. I was just curious to know, to me, they have a pretty substantial scale. What was it about the Backblaze scale that was different? Then maybe it is more, was it brand as opposed to transparency? I would like to get a layer deeper on that.
Sure. CoreWeave, I think on the maturity curve is maybe the furthest along from what I have seen. I don't think brand was a big factor to them. They are very by-the-numbers, they are very technical. I am quite impressed with working with them. They are a great partner. The scale that they are working at is extremely large. We have spent 20 years to build 5 exabytes, going on 6. Those are the types of numbers they are talking about. They were very serious about needing that type of scale.
Even as they started to run their own pilots, they were finding that it is just very hard to run at this scale. So you have open-source solutions like Ceph. They cannot get anywhere near what is needed. On the transparency piece, that was also really important because companies say they can do this, but it is not a given.
With Backblaze showing what we have done with Drive Stats, et cetera, with our reputation, with 20 years, with very large customers, it was a real factor to them.
Was the final decision that they had already made the decision to outsource, and you were the leading candidate, or was it they were still on the fence of build versus buy?
They had already made the decision, and so they were looking for who to work with. I think another piece of this is, and Gleb sort of alluded to this, companies may initially think they are going to build it, but then they start to run into the pain of running at scale. This is where I think this operational experience has been very valuable. I think I had mentioned something, tens of thousands of optimizations. That is real. Each week, we are running into all sorts of complexities, bottlenecks, et cetera. They saw that way ahead and thought, "Okay, this is going to be a distraction." Their business is around compute, and they wanted to work with a partner that could do this at the scale they need.
If I can add one other dimension to it, I think it is also you want to factor in that you are looking at data and storage, and you want it to be in proximity to the GPU compute that they have. If I am CoreWeave, I want to put the data off my customers close to me, not just to monetize the whole thing, but to give them a whole experience. Just like Gleb said, they had storage, so they were ahead, and as a service, but they wanted to give the full spectrum of experience from inferencing to model and everything else. Otherwise, the same customer is going to put the data in a hyperscaler and then
Part of the data with CoreWeave, and that is split. Time to market to offer that full service is something they have to consider, too. All of those play into the decision-making of making sure the customers that they are getting are asking for more. They want to be a full service provider. I think at GTC last year, CoreWeave was, Jensen's words were, they are the fourth hyperscaler. If you think of a hyperscaler, that is full service stack. It is not just GPU as a service or bare metal as a service. As they matured up, they are like, "We need to have the full service stack." That service stack, if I am CoreWeave, I want to put it with me, not have the data split with three other hyperscalers.
The transparency that Dan is talking about also is, what they told us directly was they said, "The whole name of our game for CoreWeave is to get to scale, and get to scale fast, and you're the only ones we could trust to be able to do that." It is one of those where you can imagine you go down a path, you start building something, and then you are a year in, and all of a sudden their customers are demanding to get to half an exabyte, 2 exabytes, whatever the scale is, and then you have a system that does not work. That is a pretty challenging place to end up. Right? So they needed to make a bet on someone they could trust.
A lot of the team behind the CoreWeave on their storage side was a lot of the team that built Amazon's S3 service. So they really, in terms of to the maturity curve piece of it, they know what this takes, and so they were able to say they could evaluate what the choices were and pick Backblaze.
Yeah. This is Ben Brostoff . I am a private investor. I think one for Dan and one for Anuj. Dan, you mentioned NVIDIA changing its reference architecture, and I think they also had recently had a blog post on the benefits of HDDs, and it seems pretty significant, them doing that. I would love to hear your thoughts on what prompted them to change the reference architecture and talk more about that. Anuj, it also seems like that will probably open up some opportunities in the sell with movement on GTM, and I know NVIDIA has done some things with the NeoClouds, either investing in them or having rev share. I am wondering if there are opportunities like that on the horizon.
Sure.
Yeah, I think it's very significant, and a very good thing for Backblaze. I would say that why did they do this? It's similar to what CoreWeave is running into right now, which is, they sell a lot of compute. They have SSDs. It's very, very expensive. Their customers don't want to store their data on something that's that expensive, but then they need to come back to it. There was cost pressure just from customers on getting to a more affordable tiering option. I think that's a piece of the puzzle. That leads to maybe a different trend, which is, okay, if you can't give me a more affordable price, then maybe I go with one of the hyperscalers, or maybe I go somewhere else, which then puts pressure on NVIDIA as they're creating chips that are competitors.
That's part of sort of the larger strategy. Yeah, I think that's the main piece, and NVIDIA's getting ahead of this. It's a category that we've been working in for some time. No company comes close to the scale. Yeah.
I think you're spot on. It widens the spectrum in terms of the sell width and the integrations because if you think about it, there was some confusion in the market that we might be competing with a WEKA. As you can see, even like Olya has shared, we are also a customer of WEKA, which is a great thing, and I think it complements our story. It complements the reference architectures that we want to build. Frankly, now that NVIDIA has put that capacity tier recognizable in the space, you can imagine that we have some good discussions ongoing. At the end of the day, we want to make sure that we are working with these partners, actively integrating as many as we can because the customers need it.
It's a good point. The blog post that Ben is referencing just very, very recently, just in the last month, I think it was, they published one where they said, "Hey, everybody thinks about SSDs as the answer, but by the way, HDDs are actually the right answer for some of the use cases, not just SSDs." It was interesting that NVIDIA chose to publish this. I think it speaks to last year when we were at GTC, we were meeting with some of the NeoClouds, and one of them said, "Look, fundamentally, we understand that we're going to need something in this area. But in order for us to be successful, we have to follow the NVIDIA reference architecture.
When NVIDIA says this is the right way to do it, this is how we're going to build it. Because that allows for us to know that it's going to work, that allows for us to work well with NVIDIA, that allows for us to work well with the ecosystem, that allows for us to go to our customers and say, "We support the NVIDIA reference architecture, so we are going to build the way NVIDIA suggests per the reference architecture." They said, "Look, we know there's a large data set out there, and we need to figure something out. But until NVIDIA kind of puts their stamp on how to use it, we have to be cautious." The fact that NVIDIA is, I think just fundamentally looking in and going, flash memory SSDs are incredibly expensive. The prices are through the roof.
The volumes are sold out, et cetera. For a lot of the use cases, that's not the best place to do it. NVIDIA is trying to foresee where the next bottleneck for the NeoClouds overall and all these AI infrastructure use cases overall is and how to expand the reference architecture to enable that. So it is, I think, a pretty major step, and I think it's exciting that they're leaning into that. Russ.
A couple of questions. Could you segment the AWS marketplace in terms of how much is Glacier, how many other sort of SKUs they have, and how much of the market, $40 billion + that AWS alone must be selling in it, is in these other SKUs that you may not directly address? Could you talk about your selling proposition has always been ease of use, certainly transparency of pricing, and radically lower pricing. Can you talk about where WEKA fits into that, and if they offer a very different sales proposition to the customer, which is performance at a big premium? Because I thought they were like 4x, or at least that's the portfolio I'm in told me they were 4x.
Let me try and touch on a piece of that, and I'll let others add to the extent they'd like. Part of our story to customers in the past has been like you said, it's ease, it's transparency, it's economics. Anuj talked about none of those things have changed. AWS has a bunch of different tiers of storage and the customers that we heard from said, "Oh my God, it's so complicated." I remember one customer said, "If the only thing you ever do, Backblaze, is allow me to not have to navigate through those 14 things, I will be forever grateful." So it's knowing how and which one to use when and feeling like, okay, wait, okay, so I'm using this one, but now I need to access it, so now I'm stuck here. We had a customer that switched to us.
They were in the media space. They had a massive archive of media footage that they said, "Okay, I'm going to go ahead and use Glacier because this is all done. This is not edited. I'm not going to need it anymore." They stuck it in Glacier. Then what happened one day was SaaS came out. There was a media asset management system that was SaaS-based. They wanted to switch to a different media asset management system. Nothing changed about their data, but because they switched to a different media asset management system, that system needed to touch all the data to index it, which meant they had to pull all of that data out of Glacier. They said it took so long and was so expensive. They're like, "We are never going through that pain and suffering again." AI is obviously putting that on steroids, right?
Because you're often needing to retouch the data and re-look at it. One of our comments was that we are basically providing the S3 level of performance and availability and everything at closer to Glacier-type pricing. It was just like why would you have to deal with all the complexity of that before? Amazon does not break out, they don't even break out technically what they make off of S3. There are some third-party estimates around it, but they certainly don't break out what they make for each of the different tiers. But I think we're able to support a large amount of the use cases that they add complexity around between the different tiers.
We don't have an offering for the coldest thing that you just store on tape, but it feels increasingly less relevant today because it's dangerous to put something assuming you're never going to access it again, right? That's that piece of it. I think the WEKA piece of it, and I would say, obviously, you should ask WEKA more about their own piece of it. But I think they would talk about the manageability, right? It's not so much about tiers of storage, they're creating the namespace for their customers and allowing for the high-speed storage. You saw on the boxes on the reference architecture that Dan showed, there was a lot of stuff going on there. But there was a big box over here that said high-speed storage, and then there was the capacity tier, which is this other box.
This other box with capacities tier had HDDs and SSDs even in it. WEKA was sitting in this other box of high-speed storage. It's a little bit of two different parts of the data flow.
Absolutely not. Unlike NVIDIA's storage architecture, S3 is the best here. Isn't it?
Yeah. WEKA has different interfaces. Obviously, part of it is S3 compatible, which is how we work together. Hopefully, that answered your question, but yeah.
Hey, guys. Matt Calitri sitting in for Mike Cikos over at Needham. Thanks for having us and taking the questions here. On the 2Q earnings call, you guys spoke about a handful of other conversations you are having around managed storage offerings. I am curious, how do you attribute the momentum there to CoreWeave sort of being this great example of how strong your capabilities are versus this just being a new offering in general because you have never done the managed storage before? Do you have a sense of these potential customers that you are talking to, are they definitely going to use managed storage, whether it is on Backblaze or a different vendor, or is it more so they are definitely going to use Backblaze and they are trying to figure out if they want to use Core or managed?
Maybe I will touch on a little bit of the history, and maybe Anuj, you can talk a little bit about some of the conversations we are having. On the history side of this, companies have come to us and asked us to do managed storage for a number of years in the past. Not quite as far back as when we launched B2 10 years ago, but certainly over the last 5 years, it has come up as a question. In the past, our answer was no. The reason our answer was no in the past is that it requires a certain amount of scale to make it interesting to do it. Right?
Because you are talking about the operational complexity of we are going to be in your data center, we need to put people in your data center, we need to manage the rights and restrictions and SLAs and all that, and the handoffs and stuff like that. It requires a certain amount of scale to make it make sense. At our scale, we have the scale for it to make sense, but if someone comes to us and says, "Hey, we would like you to be in our data center, do managed storage for 1 petabyte," the answer is, that is just not enough scale to make it interesting. CoreWeave was the first time when someone came to us and said, "Look, we have enough scale for this to be interesting.
They are paying us about $100 million as part of the contract just for our piece of it, not including any of the CapEx, just for the managed storage. So that was kind of enough scale to make it make sense. It is basically a launchpad for us to have an offering that now makes it where at the right amount of scale, we can go and do this for others. AI is also creating the scale sizes of data sets that it becomes interesting at more places. Whereas it used to be most companies did not have the amount of data for it to make sense. Now that has become a more common thing. So that is kind of getting us up to today, and then Anuj, maybe you can talk about some of the conversations that we have had.
Sure. Just to add, there is probably three things in it. One is scale in AI is just dramatically different than the scale in any other workload. For us, the scale, but over a compressed period of time. So when they need so much capacity in a compressed period of time, in the past, you could build anything in infrastructure, no different decision like on-prem in the cloud. You could build it yourself. But if you need a compressed period of time and you had to offer an SLA to your own internal AI teams, how do you get that as fast as you can at the scale that you want? Just the scale and the complexity and the time are the two things.
Then running it as a service, because now you are starting to think about, I have got an inferencing piece, I have got a model and training piece, I have got a data ingestion. There is so much going on. Would you have the benefit of somebody that actually deals with it daily and understands what a service is about versus what the software is about? So it is just those three things. If you just combine them, you are magically in the right spot where maybe 2 years ago, it was a petabyte or 2. Now 50, 100 petabytes is a normal discussion. They are actually getting to that scale very quickly, and they can see that on the horizon. They are also trying to get that compressed in time and trying to offer it as a service. You combine those three, we are a managed storage service.
We do that at scale, and so we become a viable option for them to consider.
I think your other question was, have they decided that they're going managed storage and then they're choosing who to use, or are they choosing Backblaze and then deciding whether to do managed storage or using our infrastructure? I haven't been involved in all of the conversations. The ones that I've been involved in, they basically have said, look, we understand that we need storage as an option. We're trying to decide whether we want to just leverage you for the infrastructure and you provide it to us, or whether we want to do it in our data centers, but help us work through which choice makes sense. At least the ones that I've been involved in is more-
It's mostly that. Yes.
Hey, guys. Vijay Homan from Craig-Hallum Capital Group. I'm here on Jeff's team. I wanted to kind of separate my question into two parts. Just first for NeoClouds, why do you believe that can be an enduring solution looking 3-5 years down the line? Then for non-NeoClouds, but AI customers that are coming to you directly, are you seeing any sort of shift in terms of the types of customers, the applications? It seems like anecdotally, earlier on, maybe a year ago, you were talking about really storage-hungry applications like AI video generation. I'm wondering with Flamethrower and these kind of new things that you've put out, if there's maybe some more breadth in terms of the types of applications. Thanks.
Yeah. The two questions. First of all, why would the AI infrastructure and NeoCloud side be enduring? I think that the basics on that are if you believe the various pieces. Do you believe AI is going to be a thing? I think that seems obvious that, yes, AI is going to be a thing. Do you believe that there are going to be AI infrastructure companies outside of the hyperscalers that are going to continue to operate and operate at scale? I think that it is clear that that is the case. Whether there is going to be 200 or 150 or 300, people can argue, but clearly there are going to be significant scale of these other AI infrastructure companies around model building, training, inferencing, et cetera. Then the question is, do you believe that they are going to need storage as part of their workflow?
They have told us yes. You can see CoreWeave and others leaning into it, and I think that it is clear that the future is that, and you can see that the fact that NVIDIA has added to their reference architecture is NVIDIA believes that that is also going to be a thing, right? If that is the case, then the only last question is, in the future, are they going to build or buy? Then, like you heard from Dan, one choice they have is there are open source solutions out there. So they could try and cobble together and figure out how to scale open source solutions which have never scaled to this size before. That is probably not the best path for them if they want to succeed, right? Or they can go with someone who is proven it and done it at scale.
Their choices for that are Backblaze, and then the list gets very short. That is the path of why we believe that this is a long-term, durable opportunity for us. On the AI builders, your question was-
Oh, the types of. Yeah, the expansion. We talked about video, but more broadly it's multimodal. AI generates and uses large volumes of data. The bigger the data set, the bigger it is. The bigger the data format. Text takes up a fair amount of data, but audio, video, and images are just much larger data sets. You heard from Hume, from earlier talking about how just the audio portion is dramatically larger than text. Video is dramatically larger than audio. As we're looking at it, we see data providers. There are companies whose entire business is providing data to these model builders, providing data to Hume, providing data to HeyGen, and others that were up on the slide. Those are companies whose business it is to collect data. Some of them are scraping the internet for it.
Some of them are generating it themselves. Some of it is synthetic data. Some of it you've seen maybe like some of these robotics movies where they have cameras on people's heads to watch what they're doing. All of this is about collecting data for model building, and that entire set of data providers is a great set of customers for us because they're collecting large volumes of data. Then they need to get that data. Once they've collected it's not useful if they have it. They need to get it to the companies that are using it. Because we have high throughput and free egress, it allows them to move the data to their customers, and you heard that even from Olya for their part. That's a whole category.
The physical AI companies are another large and significant group for us because the way the physical AI companies are teaching a lot of their systems, their robots, et cetera, is through video. You've heard about world models and other things. It's all about understanding how the world operates. Again, that's just a much larger data set than text. This whole category of multimodal companies, whether it's gen AI media, physical, the data providers for it, all those categories are great targets for us because they're just large data sets. Yeah. Neha's getting her steps in.
Thank you. Matt Smith with Halter Ferguson Financial. I'm wondering a little if you had any negative feedback with the price increase earlier this year, and if not, how you're thinking about balancing either passing on cost increases in the future or maybe even taking some more margin.
Yeah. I could take the first part, Budman, and you can answer the second part. We introduced a price increase May 1, and we expected churn, both from a customer count as well as how much data per customer. None of that materialized. In fact, in Q2, we actually had more sign-ups, and the average data per new sign-up was higher. I would say the demand in the market is so strong that, yeah, we haven't seen any negative impact there. Right? In terms of future plans, Budman could expand.
Yeah. Just in general, we want to provide a great service to the customers, right? We want them to feel like they're getting great value from us. That's always been the case, and that continues to be the case. Not only did we raise prices in May, but we also have higher priced offerings, right? Like our B2 Overdrive is a higher priced offering. It's just that we're providing a lot more value to the customers, and we're charging for that. I certainly don't rule out the possibility that we'll raise prices in the future. It's not the core of how we intend to grow. Right? The way we intend to grow is a combination of getting more customers, having more of their data with us, and providing more value for them. That's the core.
But depending on where the cost of infrastructure goes, depending on the value we provide, it's certainly something we can consider again.
Oh, one more. Selfie back there, Russ.
Got it. Russ Kanga, Citizens. Regarding the upcoming platform enhancements, improving your throughput, API requests, and drive economics, I appreciate these are challenging architectural optimizations to make. Should we just interpret these developments as continued efforts to improve the platform? Or Anuj, do you feel that once these product enhancements are on your truck, that these will be an area where you will feel a further unlock from customers in terms of being able to differentiate further from the competition? Thanks.
Go ahead, Dan.
Yeah, I think from a technical strategy piece, this is a decision we made to double down on the platform. We have signals that there are customers that are very interested in this. So, we are seeing we need higher throughput, we need higher API requests, and we need more storage. So we are making a strategic decision to double down on our platform. So that is what we are doing there. Maybe Anuj can speak to the business side.
Sure. I would say it is a symbiotic relationship. Dan builds and I sell is the simplest way you can think about it. But in the way we look at the business, I want to make sure that we have enough demand, not just the signal on the customers that we can then offer whatever high throughput and when we have it. Because in a way, Dan is planning the quarters in terms of when it is coming out, and we start working with customers in advance in terms of this is the future throughput that you will get. These are the kind of API things that you could do. So we have some good close relationships with obviously some key customers that we use that to make sure that it gives us the right demand for the services that we are building. So, it is just normal working day.
At the end of the day, we want to make sure what we have on the truck converts, and we are trying to build the demand for the functionality that we want to continue to build on and scale with. Yeah. That's exciting. We have new things to go and scale with.
Thank you everyone. Thank you for joining. I hope it was a really informative day. We have lunch provided, so please help yourselves, and we also have some gifts to thank everyone for coming to Backblaze Investor Day 2026. Thank you.
Thanks everybody.
Thanks, everyone.