Snowflake Inc. (SNOW) Earnings Call Transcript & Summary

October 18, 2023

New York Stock Exchange US Information Technology IT Services conference_presentation 33 min

Earnings Call Speaker Segments

Sandra Herchen

executive
#1

Hi, everyone. Thanks for joining us today. We're really excited to have you here as we're talking about the story of financing at Snowflake. Just a quick introduction to get started. I'm Sandra, and I lead the analytics engineer team for finance data at Snowflake.

Jack Peele

executive
#2

Hi. I'm Jack. I am a BI analyst, supporting the product finance team at Snowflake. So today, we're going to talk about the evolution of dbt at Snowflake, specifically how it's helped drive quick and strategic decision-making within our finance team. We're going to start with an overview of our finance support model, and then we're going to talk about how ELT within our team has evolved over the last 10 years ago or so. We're going to talk about how dbt really helps us develop quickly and cost effectively, how we hope to continue to use it into the future. And then at the end, we'll have some time for Q&A as well.

Sandra Herchen

executive
#3

So to start off, we'll talk through how our BI team support model for finance as a whole, and then we'll dive specifically into cloud finance support. So the finance data team is made up of 3 distinct sections, the analytics engineers, the BI analysts and the data scientists. The analytics engineers are in charge of building and maintaining our core pipelines, and they're also our team experts on data architecture and data model design. Each BI analyst supports a specific financial function. For example, Sales Finance, Investor Relations or SG&A. And then we also provide ad hoc support for any function not covered by an analyst. Finally, our data scientists are in charge of building and maintaining our financial forecast, including revenue, bookings, cloud spending credits. And these help support financial decision-making with data science and machine learning solutions with things like scenario analysis, impact analysis, anomaly detection and much more. Today, we'll focus specifically on our score of cloud finance as they largely function as the cost governing arm across our company and also maintain some of our largest data sets. This is also the area that Jack's aligned with as an analyst on our data team.

Jack Peele

executive
#4

So outside of myself, the product finance team is composed of 4 financial analysts. And the main responsibility of this team is to manage and forecast all expenses that go into delivering Snowflake to customers. And we also work really closely with the engineering teams at Snowflake to optimize these costs. And at its core, Snowflake is a company that is rooted in data-driven decision-making. So my role as the BI analyst in support of this team is to build and maintain all the data products that go into making the most informed decision possible, whether that be insights or recommendations based on customer usage, pricing, new product rollouts, really anything under the umbrella of improving gross margin and operational efficiency, these actions derived in a deep understanding of the Snowflake product and are backed by timely, accurate and reliable data models. So now we're going to go into a little bit about how our ELT has evolved over the last 10 years or so. So in the early days of Snowflake, we really tried to leverage the product internally as much as possible. But a lot of the business functions were pretty nascent. The finance team didn't have any sort of data support. And so all of our finance data was housed in Excel. And this worked for a while, but as Snowflake continued to grow and scale, we received more and more cost data from AWS, Azure, GCP, that Excel really couldn't handle this volume. And so that's when the finance team started hiring data people, and we made the switch over to Looker. And within Looker, we were able to build out basic revenue and cost pipelines, but we soon outgrew the capabilities it afforded us as a data team. It wasn't meeting our needs in terms of volume, complexity or code testing. And so that's when we started looking for the next best thing. That's how we ended up on dbt. And dbt is incredibly performant inside of Snowflake. We're able to use creative tables in places and then just say, look at our dashboards, for example. And ever since we made the switch back in 2019, it enabled us to do so much more than we initially even expected.

Sandra Herchen

executive
#5

So our technical architecture is landed here with Snowflake and dbt Core at the center of it all. What we do is we take our raw data from a bunch of host systems, everything from Adaptive to Workday, and we ingest that with a variety of methods. We do data sharing, we do Snowpipe and other types of connectors. And once that data lands in Snowhouse, which is our internal instance of Snowflake, a bulk of the transformation is done by Airflow and dbt Core. We also have a few machine learning models that run in Snowpark and some transformations not fit for SQL running in custom environment. Once all of that is processed, we share data to downstream teams. We write back into host systems, and we do visualization through a variety of tools, including Streamlit and Snowflake. Ever since we migrated to dbt, we've experienced development velocity like never before. So any of guys who've seen a lot of this at this point, but we have this on our lineage graph in here too. For some perspective, this is just part of the lineage graph of one of our revenue pipeline. It was so big I couldn't even find a screenshot that would fit in here, so I picked part of it. But I think this really illustrates how we've developed out support for financing dbt over the past 5 years or so. We now have over 20 pipelines across 14 dbt projects and thousands of tables, and the tables are used every single day by our finance stakeholders and reports, dashboards and investigating any sort of trends they're interested in at the moment.

Jack Peele

executive
#6

So to summarize, dbt has been incredibly impactful for our team in 3 main ways, and the first is development velocity. By bringing software engineer best practices to data, we've gained incredible oversight over our BI development, and this really helps our team iterate quickly. And the second is that dbt enables real-time or data-driven decisions. And so we have real-time anomaly detection, cloud spend forecast, insight recommendations based on customer usage trends and much more. And the last is cost governance, and we'll dive into more detail here later. But for us, and I'm sure for many of you, 2023 is really the year of cost optimization. And so fortunately, dbt offers a number of tools that really help us be the most cost-efficient data team possible.

Sandra Herchen

executive
#7

So there are endless dbt features that we love and have aided in our development velocity, but we wanted to highlight 4 of our favorites here. To start with, we use dbt Core in conjunction with GitHub to allow for version control. This allows our team of over 25 to work collaboratively on large pipelines together without stepping on each other's toes. We also have thousands of tests validating our assumptions and giving us immediate feedback on any data quality issues so we can remedy them immediately. And we also use dbt doc and projects as implicit documentation of our platform. Each project has a schema doc email file and it's filled with table and column descriptions that we programmatically read into our data cataloging platform elation. And then my personal favorite is dbt target. This allows me to easily switch between debt and product environment or different warehouse sizes. And one big thing is we're always trying to size our works appropriately in different warehouses, and this feature has made it easier than ever before for us to do so. So one big example of cost governance and data-driven decision making is our COGS anomaly detection modeling. This project was done by our data science team last year and they're build to help support cloud finance. What they did was they took our COGS data and they identified outliers in our cloud spend. Then they use a macro and dbt to create flag alerts for our finance and engineering teams. These run daily and alert on anomalies cloud spend, and allows the engineering team to investigate and remedy any overrunning warehouses that we have. Over the past year, we saved over $5 million just from these alerts.

Jack Peele

executive
#8

So if all this wasn't enough, we're doing a lot more to keep our finance team happy. And Sandra and I, we find ourselves in an interesting position where we are the analytics engineers, the BI team, incurring costs on Snowflake while at the same time, supporting the product finance team that manages our internal cloud spend. And so the reason why we think this is a unique position is because it really makes us the forefront data towards at Snowflake. It's one thing to bring quality data and information to our stakeholders, but it's another to do that in the most cost-efficient way possible.

Sandra Herchen

executive
#9

So in the last year, like many of you probably, our team was tasked with finding ways to reduce our costs. Specifically, could we reduce our internal cloud spend? While evaluating our pipelines, we realize that our largest tables were actually the ones coming from our cloud service providers. Ironically, cloud finance, the team responsible for managing our internal cloud spend, actually maintained our largest pipeline. So for some context, our COGS pipelines look like this. First, we bring in our raw data from our 3 cloud service providers and stage that data. And these stage tables only have very minor transformations from the original cost and usage reports that we're getting. So what we do is we merged some payer accounts together, we renamed some columns and add in some meta data. And after this data stage, we run through 2 reporting pipelines. So one brings in the business logic and allows us to report on COGS cost and revenue, and another is an allocation of our total cloud cost to each customer to try and get that margin level data. So back to the slide here. Just that staging portion at the beginning that I was talking about already has about 445 terabytes of data and that number continues to grow significantly every single day. So you can imagine the massive cost of the entire COGS pipeline as a whole when our starting data size is already this big.

Jack Peele

executive
#10

So when we go deeper into the cloud inventory tables that ingests this raw data, we realize that this pipeline was both far more expensive and far less performance than the rest of our finance pipelines. And so these models, to recap, in just raw cost and usage data from AWS, Azure and GCP and some of them individually take about an hour to build each day. And collectively, this pipeline takes a little bit more than 6 hours to run each day. So if we take a standard Snowflake credit, which is priced at $2 per credit, the compute costs associated with building this pipeline each month is about $6,000. And so clearly, pipelines of this volume are both a cost issue and a performance bottleneck. And from a performance standpoint, the low credit run times mean that our stakeholders can access data quickly. And then it also means it's very difficult to develop on this pipeline, because any changes or updates to these models take so long to productionalize. And then from a cost perspective, we really want to highlight 2 main areas. And the first is job credits. And job credit is a metric we use internally at Snowflake to assess workloads. And so we can think of a job as any query run inside of Snowflake. And so we take all the credits consumed by a warehouse, and we attribute them proportionately to job run on that warehouse each hour. So the longer it takes the query to run or in this case, the longer it takes for our pipeline to build, the more job we're consuming and the more compute costs we're incurring. And so even though there's a really high job credit costs associated with holding this pipeline, this cloud spend data is crucial to our product finance team, and so we need to rebuild and rerun this pipeline each day. And the second area that we wanted to highlight is data transfer cost. And so in the original version of this pipeline, all of our models were materialized as tables, which means they are rebuilt from scratch each day. And something that we do as a finance data team is we replicate a copy of all of our data models to a backup deployment in case of emergency. And so since our largest pipeline by volume was being rebuilt each day, we are also transferring our largest pipeline by volume each day, and incurring a very large amount of data transfer cost. And so for these reasons, our problem statement was pretty clear, how can we make this pipeline faster and less expensive at the same time.

Sandra Herchen

executive
#11

Originally, we approached this problem through the lens of query optimization. Our first instinct to as a consult business like Query Profile to try and identify any bad joins or slow running CTE. And we ended up not finding much here. Our joins are pretty optimized. Our CTEs actually ran pretty fast. So then we looked into replacing our computationally expensive macros with mapping tables. This helps a little bit, but also wasn't incredibly significant either. So the last thing we tried was auditing haul-ons and removing anything not being used downstream. And this had also very little impact on speed. And so it's time to attack this pipeline from a different perspective. At its core, it was clear that this was largely a data volume problem. And so our conversation changed to how we could reduce that data volume. We ask ourselves questions like, can we limit the amount of data ingested, maybe only keep 3 years of data instead of all time, or can we reduce our granularity, maybe aggregate at a daily level instead of hourly? We consulted with our stakeholders and ultimately decided that we didn't want to lose any of that data, because we wanted to track very granular changes and trends over time. And so we chose to convert our tables to the remaining solution. Incremental models. So to start with, let's make sure we're all on the same page about what an incremental model is. In comparison to a regular table that's fully reprocessed every single day, an incremental model only transforms new data and updates the relevant existing data in the table already. Once we designate a model as incremental in dbt, the first run will still rebuild the entire table from scratch, but subsequent runs will be much faster. Based on the incremental filter that you add in the table, dbt only processes that new data that you've designated and integrate that back into the table that you've originally built. And this is great with a high-volume data pipeline, because you're only updating a fraction of the entire data rather than the whole thing. The output table remains the same and you don't lose any information, but you incur less cost and spend less time constantly rebuilding your models.

Jack Peele

executive
#12

Yes. So dbt incremental models led to significant improvements from both a cost and performance perspective. By using an incremental filter and a unique key, we're able to limit the amount of data that is transformed on a daily basis. So rather than rebuild from scratch, we're updating the smallest window of data possible. As we mentioned, we replicated back up copy of all our tables to a separate deployment. So now rather than transferring the entire pipeline to a separate deployment, we're really only transferring the updated portion of this model. And so that saves us a lot of money from a data transfer perspective. And then we're also saving a lot of money from a job credit perspective, because there's less compute required to run this pipeline each day. And then this improved run time also increase performance, because we're able to get tables or these models to our stakeholders much more quickly, and we're able to develop on this pipeline much more quickly as well. And so another thing that we really like about this incremental configuration is the online schema change parameter. And so you know as we column them or added to your tables or the schema itself changes over time, this parameter ensures that our models continue to run successfully. And so in summary, our big data problem came much more manageable, but also very importantly is the fact that the scope of our data remains the same this entire time. And so despite all these improvements, we do want to discuss some challenges we ran into along the way. And the first was choosing the best incremental filter. One that optimizes performance and aligns with our business logic. And so our goal is to update the smallest window of data possible, while prioritizing data recency and data accuracy. And so we'll dive into an example with our AWS data to provide a little more context. So each day, AWS updates their Cost and Usage report, and they update data for the entire billing month to date. So that means any data that already exists for that billing month is supplanted by the new data. Now at the conclusion of every month, we expect a bill invoice from AWS with finalized data for the month that just completed. And we typically receive this bill invoice in the first couple of days of the next month. And so in these instances, we not only need to update the current months’ worth of data, but the previous month's worth of data as well. So the problem became, how do we define an incremental filter that updates just the current month when applicable, and then the current month and previous months whenever we receive a bill invoice? And now the second challenge that we ran into was defining a unique key that aligns with this business logic. And so within our AWS pipeline, we have multiple independent payer accounts. And each of these payer accounts consult or merge into one consolidated Cost and Usage report. And these payer accounts refresh at different times, and we receive their invoices at different times for each of these payer accounts. And so we needed to define some sort of unique key that would apply the incremental logic that we just discussed at the account level. And so we were able to achieve this by using a compound key, billing month and account ID. And this ensures that we have the most recent, the most accurate data for each month, all in the same model. And so this dbt migration or this incremental migration led to a lot of improvements in our cloud inventory staging pipeline. The original daily run time was a little bit more than 6 hours, that has now been reduced to just 1 hour. We've seen an 88% reduction in our monthly data transfer costs and a 67% reduction in the amount of Snowflake credits it takes to build this pipeline each month. And so perhaps what's quantifiable also is the fact that our overall sentiment toward this pipeline is much improved as well. Not only are we saving a lot of money, but we're saving a lot of time and headache from a developer perspective, too.

Sandra Herchen

executive
#13

Finally, we'll discuss where we intend to go next. We look forward to leveraging dbt features more across the finance pipeline to continue to reduce our cloud spend. Our next goal is to explore the relationship between cluster keys and incremental filters. Our hypothesis here is that in the line of the 2 could lead to enhanced query performance for certain models. We'd also like to experiment with workload-specific warehouse sizing. For example, for those known periods of diminished performance in the cloud finance pipeline that Jack was speaking about earlier, where you have to update both the current month and the previous months, if we increase our computing power and reduce our overall run time, could that actually reduce our spend? We also have a few pipelines that are easily handled by SQL, that we originally wrote in Python. And now that we have dbt Python models in partnership with Snowpark, we're hoping to migrate those models over. We're also excited to leverage Snowpark to facilitate new machine learning workloads and continue to champion data-driven decision-making alongside our stakeholders. If you have any questions about anything we talked about today, feel free to reach out to us via e-mail or connect with us on LinkedIn and we're looking forward to continuing the conversation. Also, I think we have some time for Q&A, if anyone have any questions.

Unknown Attendee

attendee
#14

Do you have a strategy for implementing for refreshes at your incremental models?

Sandra Herchen

executive
#15

Basically, only one we need to. We have, like I said, thousands of tests, every time we make something incremental, we make sure we're including a test that will alert us if the data come fresh for some reason or like some historical thing has changed somewhere, and then that will tell us that we need to do a full refresh.

Unknown Attendee

attendee
#16

Can you shed some light on your slack alerts from macro? How were you able to achieve that?

Sandra Herchen

executive
#17

Do you want like the whole...

Unknown Attendee

attendee
#18

What's the API like, you have macro, I think you have a macro who send who Slack Alert or something that you're trying to do with it?

Sandra Herchen

executive
#19

Yes. Yes, so I mean you've done some Slack Alert stuff, would you like to talk about it?

Jack Peele

executive
#20

Yes. I mean I didn't necessarily build the back end of it, but what we do is we basically have this query set up or this system setup, where we define what we want to detect as an anomaly. So maybe we see, we take a trailing 28-day average, and we see something that exceeds that after a certain barrier. And then any time that happens, we'll ping a Slack channel that we define as a team, maybe it's called COGS anomaly alerts, and that will pretty much give us the date that we see it, the reason why it's been detected as an anomaly and give us a message that we can then go into that model or wherever that data is housed and see, is it an issue with our data models, or is it perhaps an issue with something that we need to go address with a separate team.

Unknown Attendee

attendee
#21

[indiscernible]

Jack Peele

executive
#22

Right.

Sandra Herchen

executive
#23

Yes, exactly.

Unknown Attendee

attendee
#24

When you do the incremental table, when did you decided to [indiscernible]?

Jack Peele

executive
#25

So we typically tend to use an incremental table when there is a significant amount of volume. So on tables that have a little bit less volume and are pretty performant, when they materializes tables, we don't see a need to materialize them incrementally. We actually have seen kind of a reduction in performance if we try to incrementalize small tables. And so if our pipeline is pretty performant, it completes in minutes rather than hours, we don't really see a need to use incremental models. But otherwise, in this case, where we have this really expensive pipeline, this really slow pipeline, we see this as a really good candidate to try incremental models. And if we see a significant performance improvement, we'll stick with it. And there are times where we've experimented with incremental models and not seen the kind of performance improvement or cost improvement that we wanted, and we've gone back to table materialization.

Unknown Attendee

attendee
#26

Do you have any kind of guideline on [indiscernible]?

Sandra Herchen

executive
#27

We probably won't look at it if it's under like 1,000, 1,500 seconds for run time for the table, just based on how long our pipelines usually take. But if it's pushing above that, we'll consider it.

Unknown Attendee

attendee
#28

What are the core path for [indiscernible] you guys focused on?

Jack Peele

executive
#29

I think a lot of the tests that we do are ensuring -- I mean they're pretty uniform, I guess, across all types of our models. So, a, we don't want to have any sort of unnecessary type of duplicates. The second, if we really want to confirm that our unique keys are working the way we want them to, another thing that we're checking for and one of the reasons why defining the best incremental filter is one of our challenges is ensuring that if there is an update, say, the data a year ago outside the scope of our sort of incremental build, are we recapturing that data? And so we have test that attempt to try to capture that sort of look back, make sure that we have an AWS data say the correct assembly in each of our billing months and stuff like that. So some of it is pretty unique to the business logic that we're using in that model. And then some of it is pretty uniform in terms of like duplicate observations, stuff like that.

Unknown Attendee

attendee
#30

[indiscernible]

Sandra Herchen

executive
#31

So our team actually isn't in charge of that. Our IT team handles it, but I believe it's put in to S3, or I guess it's in S3 and we bring that in. I would get either Snowpipe, something like that. And I know that's incremental to, for sure, it's huge.

Unknown Attendee

attendee
#32

I guess that's kind of related to my questions, like what is the organizational structure of the team against like analytics engineers settings in east domain and something that you have data engineers actually populating?

Sandra Herchen

executive
#33

Yes. So Snowflake has data teams all across the company. So we're on finance data, there's IT, sales, marketing, et cetera. A lot of the ingestion stuff or like financial systems, tends to be owned by other teams. There's a census team that handles anything from like Workday, stuff like that. We'll sometimes own Snowpipes to bring that in and sometimes they do it. It just depends. But anything that already is in a raw table, we'll handle everything after that. Anything that's a core pipeline, revenue, contract information, stuff like that is the analytics analyst. And then each analyst owns all of their miles for their specific area.

Unknown Attendee

attendee
#34

[indiscernible]

Sandra Herchen

executive
#35

Sorry, I don't know if I heard that completely.

Unknown Attendee

attendee
#36

[indiscernible]

Sandra Herchen

executive
#37

Yes. I mean, there's definitely a difference between like reports internally versus what we report to shareholders, right? Like our earnings report and stuff. So anything that is going to investors like -- or for us to release on our earnings call, we have a whole data model structured for that as well where we snap at everything. So we make sure we're only reporting what we reported last quarter, stuff like that. But then in terms of internal reporting, oftentimes like you want the most up-to-date data, regardless if it's GAAP revenue or not, right? Stuff like that.

Unknown Attendee

attendee
#38

[indiscernible]

Sandra Herchen

executive
#39

Yes. So we use dbt tests, but we write them all ourselves. I don't think we're using any external packages. But we also have built a lot of dashboards on top of tests. And we have -- we're making our data quality observability dashboard right now, which is like a central spot to go and see like anything failing in our pipelines, any data that haven't run. But we have hundreds of eyes on these numbers every single day.

Unknown Attendee

attendee
#40

[indiscernible]

Sandra Herchen

executive
#41

Yes. We do all of that, and there's always like human component too. If that number looks strong, we should go investigate it, yes.

Unknown Attendee

attendee
#42

[indiscernible]

Sandra Herchen

executive
#43

Yes, that's a great question. Our team was definitely not the only one tasked with reducing our costs. Like if you think about a lot of the data science teams probably have even bigger data that they're running. And so this is like a company-wide thing. We definitely have a central team that was tracking all the different projects for this, and it was definitely like everyone was on board to try and work on it. This is just the biggest example from the finance side that we did.

Unknown Attendee

attendee
#44

[indiscernible]

Jack Peele

executive
#45

I think as -- by supporting the product finance team, it's more that we'll see areas -- we'll see like large volume areas of data or large cost areas of data, like what we're incurring to store data. And so we'll more attack those that kind of like show up on our radar and attack them specifically. Our scope just as the finance data team is really obvious, like which one was our biggest bottleneck just because this cost pipeline is so much slower and more expensive than the rest of them. But in terms of like having like, I guess, a uniform metric for costing out things, we typically go by what we're -- the cost we are incurred from like the cloud providers or where we store our data and things like that. So we have like a cost number, but the standardized, we were using this Snowflake from a compute perspective.

Unknown Attendee

attendee
#46

I wanted to ask if you guys experimented Snowflake [indiscernible].

Sandra Herchen

executive
#47

Yes, that's a great question. I get asked this one fairly frequently. So we looked into it a little bit, and we ran into some issues with having dbt refreshing our tables in full every day and causing the stream to fail. So hopefully, that we can figure out how to deal with that. I do think it would probably work better with something incremental since that's not being fully rebuilt, but we haven't done anything yet. Definitely something we're trying to investigate.

Unknown Attendee

attendee
#48

[indiscernible]

Sandra Herchen

executive
#49

To be honest, I don't think I've investigated enough too.

Unknown Attendee

attendee
#50

[indiscernible]

Sandra Herchen

executive
#51

Yes, let's try it out later. Anyone else?

Unknown Attendee

attendee
#52

[indiscernible]

Jack Peele

executive
#53

That was the anomaly detection that we built out.

Sandra Herchen

executive
#54

Yes. So -- that number didn't come from us. That's from the data science team. They work with us, but they're the ones that did the actual project. So I think -- I would assume they look into all the overrunning warehouses and the longer term of those link and came up with some in that way, yes.

Unknown Attendee

attendee
#55

What's the framework you planning to bring in, you mentioned at the end of the year [indiscernible] dbt as well?

Sandra Herchen

executive
#56

I mean the goal is definitely to integrate into dbt. But I can't speak to that very well. Again, I'm on the analytics engineering team, that's more like what our data science team is working on. I'm sorry. I wish I had a better answer for you. Definitely, if you follow up, I'm definitely to ask them, and share the answer with you though. Yes.

Unknown Attendee

attendee
#57

[indiscernible]

Sandra Herchen

executive
#58

Yes. My understanding is like because we were on it so early, like 2019, I think that's when cloud came out, to my understanding, that's about -- or I would assume that's why we picked it and we kind of stuck with it and ran. Everyone across Snowflake uses dbt Core, runs on the dbt Cloud, that I'm aware of. So, yes, but I wasn't here then also, again, happy to follow up if you want more explanation. Yes. Anyone else? Think we're way over time. So. thank you guys. Thank you.

Read the full transcript via the API

You're viewing the first half of this call. Get the complete Snowflake Inc. transcript — plus 251,000+ transcripts from 12,000+ companies, speaker segments, AI summaries and full-text search — through the EarningsCalls.dev API.

Get the API View API docs →

This call discussed

For developers and AI pipelines

Programmatic access to Snowflake Inc. earnings transcripts and 251,000+ others is available through the EarningsCalls.dev REST API. Plans from $24.99/month — full transcripts, speaker segments, full-text search, and the recently-added /api/v1/transcripts/recent polling endpoint for ETL pipelines.