Backblaze, Inc. (BLZE) Earnings Call Transcript & Summary
July 15, 2026
Earnings Call Speaker Segments
Unknown Executive
executive[Audio Gap] The first time you're joining us today. Backblaze has been collecting our raw device metrics for over 13 years at this point. And the -- we published the failure rates of hard drives. If you go to our website, you'll see a maintained data set that has daily logs from each of the drives in our pool. And then we also filter those out and look at them on a tabular level, so we can talk about analysis failure rates. What happened this quarter, very exciting. So we've got over 340,000 drives now. We had about 1,000 dry failures and a bunch drive days, which is how many drives existed per day in a single quarter. So you can see the drive population by manufacturers. It's about 1/3, 1/3, 1/3, which is interesting, and we'll talk about some of the quarterly data here. . This is the big eye crunching table that you can -- if you go to our blog, you'll see that we actually have a much more friendly version of that this quarter with responsive tables. We're very excited about this. But top level here, you've got 1.24% annualized failure rate for the quarterly stats this time around. And if you want to just look at what's been happening here quarter-on-quarter for about the last year, you'd see that it's up from last quarter from Q4 2025, but down year-over-year. We've got new drives. So the 26 terabyte WDC is now in production. And you can see some pretty interesting stuff. This last quarter, we deployed a little over 10,000 drives. And of those, 9,400 or so were more than 20 terabytes. We've been keeping an eye on that population because it's -- of course, as you deploy larger drives, it's got implications all the way around. And so this quarter, the AFR for that population specifically was 85%. Don't overindex on that. These are pretty young drives. And we know for sure that younger drives fail at fewer rates or fewer -- fail fewer times. So there's that -- and then I got a couple of drives here with 0 failures, which is pretty exciting when you look at the age of those 4-terabyte drives. Those guys are, I don't know, they've been around for a real long time. If you look at the average age in months, the 4 terabyte drives are $106.1 million age in months, we do measure our drives like toddlers and months. But a little under 10 years. which is pretty cool. There's only 186 of those remaining, which says to us that they're on their way out, and you'll see that they've dropped off the lifetime table this time around as well. So lifetime data, if you -- these have slightly different exclusions here. So it will be -- they need to be hanging out for more drive days and have more drives in the population, so you'll see it slightly different. What's interesting this quarter is that our annualized failure rate is 1.39%. We've been holding pretty steady actually for the last several quarters at right around 1.29 or so. So this is a bit of a jump. But the jump is actually because some of those much older drives came out of the pool. So the 4 terabytes have dropped off a lifetime table. And like I said, those had 10 years of data behind them. So as those exited, you'll see a little bit of a fluctuation because things have changed. And then we've got 3 new drives entering this as well. So -- you'll see the 22 terabytes in the 26 have made the exclusion criteria to be tracked in the lifetime drives. So what's interesting this quarter is actually the data is always interesting, but it's a little bit something that we were able to flag when a drive was not failing that crazy. So this is the disturbance in the drive status, what I'm like affectionately calling this. But I think what's really important is that -- and something we've talked about before is that all systems that we talk about in the reporting that we're doing here, they represent a managed system, right? We're monitoring our drives at all times. So we can actually affect whether things have higher and lower failure rates if we see something we can take like proactively or reactively change it. So this is a really interesting instance of this because what happened is that we had a pretty acceptable drive model. We start seeing failures climb. And we actually saw -- it was a little bit of a confusing investigation because once the investigation was complete, the end, they ended up being 2 separate issues that were happening that were very there different, distinct from each other. And the drives themselves could experience 1 or the other or both of these failures, right? But we identified the power cycle issue early on. So we've noticed that if we'd shut them completely off, they would have trouble coming back online or maybe not come back online at all. So what we did is we took a mitigation step that we just reduced the power cycle frequency to the vaults of those effective drives. And once we did, the failure rate came down to a reasonable level. So this is obviously not real numbers here that you're seeing. But if you can imagine, what we did is reduce the risk of one of these types of failures and that left us totally fine. And then once the investigation completed, we saw that there were multiple issues happening. So from our perspective, when I went to drive stats, I was like, hey, would you look at that, that failure rate is not very high, but I know something happened. So I had to do a bunch of investigation on my end to make sure that we were properly capturing failed drives because I didn't want there to be a situation where we were under reporting where we should. And what that led us to is a really interesting lens on how we define a failure. We've talked about how we define a failure in the past and the quick and dirty is essentially that there's a C++ job to the custom program that collects the smart stats for drives each day. And then there's some things about the exclusion tables or inclusion tables that are maintained by humans that affect these things. But at its most basic, it's conditional logic. That's the driver yesterday? Is it gone today? If it is gone, and it was here yesterday, then we log it as a failure. And if we look at -- if it's a capital F failure versus a lower case f failure, let's say it that way. There's -- that's where those exclusionary inclusion tables come in. So you can imagine if you swap a drive out for normal routine maintenance. We're not going to call that a failure because the drive did not fail. We just swapped it out. So -- and then the other thing is that if a comes the serial number comes back online for the end of the quarter, then it's no longer a failure like we took the drive offline for some reason or another and brought it back. So -- there's a little bit of squishiness there. The quarter end cutoff means that like if the drive comes back after quarter end, potentially, you've got 1 or 2 false failures in there, but it's really not -- you got to cut it off somewhere, right? But what this does mean and what this -- what came into play for this specific incident that we were tracking is that if you have day 1 failures, you actually don't love those, right, because it has to exist before it cannot exist for the next day. And so -- from our perspective, I think it's an interesting caveat to make on the data set. Lots of folks look at the data set as a source of proof. So if we're under counting day 1 failures, I want to make sure that's well understood. On the other hand, how often do you get a day 1 failure. In this case, with this specific drive, there actually weren't that many, but it was something that I wanted to double check on before we came out with failure rates. But on the other hand, it's pretty rare for drive to have day 1 failures like there's qualification drive manufacturers do, we personally do qualification on the drive. So it's not something that happens all that often. But again, because this is used so widely. There you go, we've got a new caveat to the data set if you're using it, make sure you're checking longline data and see what's happening. All right. I think that brings us to resources here. at the report, of course, follow the series. These are many ways you can interact with us. We've got the drive stats home based on the website, enjoying the insiders newsletter or reach out to us directly. I drive stats at backac.com is a monitored inbox. I'm there. So I'm always happy to respond to real humans and then we've got socials in the common section. So yes, let's take this time and talk questions. Laquie, anything on your end that sparked while we were going through the data.
Laquie Campbell
executiveYes, absolutely. From my POV, I told you many times, I always get stopped at shows and when I bring up back blades they're like, "Oh, yes, drivetock. I look at that every time. And it's because in media. And when I say media, I don't mean just time in television, every company these days of a media company, whether we're talking about corporate marketing, bio I mean it doesn't matter. Everyone's trying to tell a story visually and those files typically are larger in a traditional media library a video file might carry a few hundred metadata attributes. And so now that everyone is working with AI trying to figure out how they can efficiently use it in their workflows, that same asset can generate 3,000, and that's just the Medidata. So every stage of the workflow is producing more data, more iterations, more outputs, another kind of tangible example would be of AI upscaling. So AI has made the ability to make an older show, remaster and make it look great quicker is -- so if you have a show that was shot natively in 2K and a decision is made let's make it 4K for OTT or streaming or even 8K, let's say, there's a premium tier or you're trying to future proof it. that 2K original isn't going to disappear. They're not going to track it. It's the source of truth. It's the legal master. So what we're seeing, obviously, is now you have that 1 asset that was 3 retained versions, the 2K, 4K, 8K. The 2K, let's say, it's a 1-hour show, it's -- and it's pro res, let's go with that. That might run year like 100 gigs. So that 4,000 upscale typically is going to need its own render passes going to need on mezzanines only QC deliverables. So that 1 new version is realistically going to bring 3 to 4x to your footprint per title and that's before 8-K.
Stephanie Doyle
executiveSo all of that's happening, and that's great. I mean the rent that you could be able to do that faster. However, if your infrastructure wasn't great before, you're going to start to see the tracks -- and so I think drive that is becoming even more important to this audience because they're having to bolt together really smart systems and infrastructures. The story is not whether we do cloud or on-prem anymore, that's that. It's more how do we put all of this together to work efficiently, to be performant for capacity's sake because obviously, budget is the factors of how much can we really put on this drive and it not make all of our originals disappear. Those are always going to be very front of mind thoughts. So this is why drives is really important in my domain.
Laquie Campbell
executiveI love it. Yes, it's always good to hear, and it's interesting you mentioned AI because we've got a question from Raga here in the chat. How is AI impacting back lays how do we think it will affect small and medium businesses when it comes to storage based on what we're seeing. And just based on what you are saying here sort of the conversation we had before the call, like, totally agree, right? Like people in M&E space have to consider individual drives, but also all these other moving parts of their architecture. So as far as AI goes, I think it's a very interesting conversation. And I can talk about it from a data center perspective, but I'm interested to hear what you have to say in the M&E industry or even in just like content generation at all like thinking of content broadly as multimedia files, I should say.
Stephanie Doyle
executiveYes. I mean it's just a continuous derivatives that are being created that people are thinking, "Oh, I can be more creative, our team can deliver more." But again, if you don't have that IT team, if you don't have that technical person that's thinking about that's awesome, like I love everyone being creative and wanting to grow and develop, but how do we store all of that efficiently? How do we get to it when we need it? How do we monetize it and be able to find it quickly. So that's -- all those things start to come into play when you're doing all this iteration. Another example would kind of be different grading. I don't know if people think about when I say grading color grading, sometimes, and we've all seen it the warm grade for a certain type of movie versus a blue color grade. There's a lot of testing. Directors, DPs, the direct photography, they do a lot of testing. And so now we're maybe doing AI color grading upscale and upscale. So now we have the upscale and we have the color grade. And maybe the 4 different ones. And so -- and then let's add on HDR to SDR grades or AB test audience variance. So it's amazing AI is great because you have the ability to treat these things possibly faster. But again, if you don't have the infrastructure built out and the processes to like manage it and figure out where everything is and make sure everyone is getting all of the files that they need to create, that's when it can become an issue.
Laquie Campbell
executiveYes. It's just like an explosion of data on the most basic level, and managing data, particularly in active archive has always been a conversation in this space. I think when you're talking about how it's affecting if you're saying how is AI impacting back place? I think that's a pretty complicated question because it kind of depends on what lens you view it through. So we certainly see this huge demand for data. But on the flip side, we also see access patterns for these things moving in really different ways. If you look at the network stat series that we work on, too, you'll see just a high volume of data moving all at once and then being processed, a lot of times through -- sometimes even through traditional CDM providers who are converting to or are our ready neo clouds. And that, in conversation with like traditional media tooling and where a lot of that data lives, you're seeing a lot of these integrations become super, super important to your point, Laquie, where it's like you have a lot of moving parts, and they all have to work together
Ittai Kidron
analystAbsolutely. And that's why I love the ecosystem that we've built back ways. We have amazing partners that are doing just mind-blowing things with AI and technology. And so I think that's another way that backplace is being impacted and also being part of the story is that the ones of back ways. And so yes, it it's an amazing time, but it's definitely 1 of those is like firehose -- how do we manage all this. But I think the infrastructure was built so well here that I think that we're just excited about what we're seeing and able to rise to ride with people and our partners. So yes. Yes, agreed. So next question, what's 2 kind of related ones. About how to use the stats to buy a drive or what sort of brand we'd recommend and I'll give our standard language here. We really do not play favorites when it comes to recommending drives because -- and I think this is an important part of the conversation. In some ways, what we do is sort of the ultimate litmus test for a hard drive, right? They are running at max capacity for their entire life until they die. That is what we do in a data center, right? But on the other hand, your personal use case might look quite different, right? And we also have redundancy both on the software layer and with secondary drives to be able to manage potential data loss. So the way that I would recommend using the stats is to identify what size of drive you need first and then see with what you're comfortable with, look at reviews and then build out for backup. As always, I always recommend having at least a bare minimum of 321 backup strategy, right, with parts of it in the cloud, part of it are off-site. And then if you want to have your on-site for traditional air gap that makes sense too. All of that gets complicated. Data management on a personal level is something that I have a lot of respect for. So really, you can get as durable as you want to on a personal level and your drivers are certainly part of that conversation. Yes. And I think these drive help trying to help people figure out what reliability at scale looks like for them. When they're feeding data into whether it's an automated pipeline or we're not just talking about the human editor aspect of it. People are doing a lot of modeling and training against their footage and their archives. And so undetected drive failures, mid data set, it's not just you lost a file. It's like silently corrupting the whole training model, right? So -- so yes, this is, I think, a great report for people to look at and just try to assess accordingly what's going to work for them in their current workflows.
Stephanie Doyle
executiveYes. I totally agree. All right. It looks like we've got another question from Hans -- or I'm sorry, in gotten -- here we go, guys. Excuse me for mispronouncing on there in -- so member of a group of 200 photographers, awesome, love it. You are the techy and you showed them how to back up on site, but you want to have a presentation on why they should use back ways to back up off-site. Hank with no personal motivation, I direct you to our blog, and I say that because I've written on the blog historically for quite a long time these days. But we've got quite a few articles on the benefits of off-site backup. I think when you're talking about off-site, it's really important when you think about disaster recovery and that sort of thing. But the other piece that I think is important, specifically for this audience, is a little bit what Laquie is talking about where you have lots of things everywhere, and you might want to be able to access them remotely or with different tools or in different ways. So you might want to, I think, just simultaneous and distributed access is an important part of this conversation, too. Beyond backup, those active archives that folks need to work with anything else you're thinking about it like anything top of mind for you?
Laquie Campbell
executiveI don't think so.
Stephanie Doyle
executiveYes. Yes. I think there's a lot to lot to look at here and many different lenses to look at the impact of hard drives and AI. I think what's particularly interesting to me is sort of people have started to look at their tech stack and understand it's how much they needed to move with them, right? So those active archives are always part of a conversation in my in my neck of the woods.
Laquie Campbell
executiveNo 1 wants those single points of failure. Like we're always trying to figure out how to make a workflow smarter -- and your archives, I mean there's literal gold each frame has such high value. And so it's really important to make smart decisions at this point. As you're growing. And growing and growing as you iterate.
Stephanie Doyle
executiveYes. I totally agree. All right. Well, if we don't have anything else from the common section, I feel like we've had a great time here today. I appreciate everybody for showing up. And as always, feel free to reach out if we've got additional questions, we're happy to answer them. Thanks, everyone.
Laquie Campbell
executiveThank you.
Read the full transcript via the API
You're viewing the first half of this call. Get the complete Backblaze, Inc. transcript — plus 251,000+ transcripts from 12,000+ companies, speaker segments, AI summaries and full-text search — through the EarningsCalls.dev API.
Get the API View API docs →This call discussed
For developers and AI pipelines
Programmatic access to Backblaze, Inc. earnings transcripts and 251,000+ others is available through the
EarningsCalls.dev REST API. Plans from $24.99/month — full transcripts, speaker segments,
full-text search, and the recently-added /api/v1/transcripts/recent polling endpoint for ETL pipelines.