NVIDIA Corporation (NVDA) Earnings Call Transcript & Summary

October 26, 2023

NASDAQ US Information Technology Semiconductors and Semiconductor Equipment special 57 min

Earnings Call Speaker Segments

Unknown Executive

executive
#1

Good morning, and welcome to our webinar. Before we begin, we wanted to cover a few housekeeping items. At the bottom of your screen are multiple application widgets you can use. All of the widgets are resizable and movable. If you have any questions during the webcast, you can submit them through the Q&A widget. We will try to answer these throughout and at the end of the event. A copy of today's slide deck and additional help materials are available in the resource list. We encourage you to download any resources or bookmark any links that you may find useful. You can find additional answers to some common technical issues located in help widget at the bottom of your screen. And finally, an on-demand version of the webcast will be available approximately 1 hour after the webcast is over and can be accessed using the same audience link that was sent to you earlier. Now I'll pass it over to you Amanda.

Amanda Saunders

executive
#2

Thanks so much, Isabelle. My name is Amanda, and I'm going to hosting the webinar today. We had thousands of people registered to attend this session. And I think it's a really strong indication of where we are in the journey of learning about generative AI, to really thinking about what it means for enterprise businesses and for enterprise applications. So this session, as Isabelle said, is going to look at the enterprise demands that we see around generative AI and what considerations you and your team need to be making, as they're building out these applications. I'm joined by Don who's going to cover a lot of the technical details so that you understand what you and your team need to think about to make sure you're building your applications with the enterprise requirements in mind. I think it goes without saying that generative AI is going to be a powerful tool for the enterprise. We know this. We understand that it's going to help augment our work. It's going to help assist us with our day-to-day tasks that's going to make us more efficient and more productive. The latest research estimates anywhere from $2.5 trillion to over $4 trillion is going to be added to GDP annually from this new technology. So it's important to understand how and why we should be adopting this across our business. In the future, workers, nurses, creatives, all of us will be surrounded by generative AI applications, by LLMs that are going to help us in our business in the work that we do. It's going to augment, it's going to make decisions, and it's going to make us more efficient and more productive in the long run. I mentioned earlier in the session that the focus on generative AI has really changed, and I think this quote from Fortune Magazine certainly shows that. The race to adapt AI has evolved. We're moving on from simply being first and building these incredibly powerful tools to understanding where the gaps are, where some of the inefficiencies are and where we need to improve. And so for enterprises, who are adopting generative AI, it's about more than just adopting what exists, but really tailoring these AI recipes to our own domain-specific understanding, understanding our data, connecting it to those data sources, providing it with that powerful information that it needs to differentiate the business applications we need to build versus those that are being built by the rest of the world. All right. So with that in mind, let's look at the steps to get started with generative AI because that's going to frame the rest of what we're going to talk about today. So the first thing is all about identifying that business opportunity. This is about understanding where you have the biggest impact to make, whether it's for your customers, for your employees. Choosing the right business opportunity is actually going to drive the rest of this effort and make it a lot easier to roll out generative AI in your business. After we choose the opportunity, we've got to look at what are the teams that we have in place. Of course, we have domain-specific expertise generally within our own organization, but it's also important to look at those AI teams. Whether or not those are internal today, maybe you work with a partner of upfront, but building out that expertise down the line will help you as you continue to roll out generative AI. The next thing we need to do is analyze the data that we have for training and customization. And Don is going to talk a lot about where that fits in the pipeline, but this will typically be our most valuable data, the data that differentiates our business from every other business out there on the planet. That's what's going to make these use cases compelling. We need to invest in the infrastructure. The POCs that a lot of folks are doing on generative AI are great, but they very quickly scale. As you think of new use cases, as you look to deploy this and roll this out faster within your organization. So investing in the accelerated infrastructure you need today as well as what's going to drive this in the future is important, and looking at more than just the cost of an individual system. When we're rolling out generative AI, we're thinking about data center scale, we're thinking about cloud scale, and so you need to be looking at the overall cost efficiency, energy consumption of the infrastructure you're investing in. The final piece and probably the most important piece is also developing that plan for responsible AI. There are best practices. There are tools out there. There are lots of folks, who are investing in this area. Having a plan for your company when you get started will help you just grow in a way that means you're sticking to your company's values and your company's brands, and it can continue to evolve, continue to develop, but having it in place from day 1 is just so important. All right. With that, let's go over to Don and learn more about building these enterprise applications with generative AI.

Donald Chouinard

executive
#3

Thank you, Amanda. So at this point, you've seen the types of exciting things that can be done with generative AI and you started thinking about the people aspects and the resources that would need to be brought together to create such a thing for your organization. So the next thing you want to do here is start learning more about the tech stack. And this slide pretty much serves as a table of contents for the rest of the presentation and go through each of the technologies that I want you to be familiar with. You're going to go through a stage of data preparation and then training your model and then deploying it and using it. Now when it comes to data preparation, the saying is garbage in, garbage out. So if you feed your AI model, a diet of garbage, then it will not produce great answers. So you've got to get really good sources of data publicly, privately. You've got to have enough data for what you're looking to do. These models are notorious, the more data you can give them the better. And then you need to go through, not only getting the data from documents, databases and other applications, but you end up cutting it up into little pieces and then those data fragments, those have to be curated. You need to be sure that all the cleanliness is there. And we're going to talk about all the different aspects of data curation. Now your data is all set. Now what you need is you need an AI model. Now there are many of them out there. It's very expensive to train a foundation model. And so the chances are very good that you're going to be able to find a foundation model that suits you just fine. Are you in a data center? Are you mobile? Does cost matter? Does cost not matter? Are there certain performance characteristics that you need to hit? All of these things are going to go into the selection of the actual AI model -- foundation model that you want to use. And you probably don't have to build it yourself. But if you do have to build it yourself, you can definitely do that with tools that NVIDIA provides. Let's say, you selected a model that someone else spent $20 million developing, and it's now available as open source. And so you get that model, now it only knows about generic things. It only knows about language in general in the world in general. You need to take -- it's like a little kid. What you need to do is you need to take that AI model, you've got to send it to school. You've got to teach it, your domain specific vocabulary. You've got to teach it more about the types of skills that you want to see it exhibit for you in your application. And that's all called model customization plot to it. It's like 4 different types, and we're going to go into the different types of model customization. After you customize model, you're going to try it out. You're going to hook it up to your application. It's going to be making calls on the model. You're going to be looking at the results of that and determining whether that's where you want to be and just keep iterating until you have it right. Once you've got it where you want it, then you're going to take your foundation model -- your trained foundation models. Now it's really becoming yours. It's really learning about your business and you and your users. And now you need to deploy it. You need to deploy it potentially at scale where it's going to be used in an operation called inference. There's 2 things you do with models. You train them and use them for inference. Inference is when you're sending them prompts and getting results out of them. So you deploy at scale, and you need to use parallel hardware. There's a lot to that, and we'll talk about the deployment pipeline. And then the last thing is -- think about these models is they're creative. They're generative. They use the data you gave them and they produce very interesting, but sometimes concerning results, sometimes potentially harmful results or wrong results. These are all called hallucinations. And it's your job as the application developer to contain the hallucinations, right? You've got to constrain the genius and that's where guardrails come in. Guardrails is essentially a model that you put in front of your model to keep it in line. It's like a monitor for the model to make sure it doesn't say things it shouldn't say. So -- we'll talk about guardrails and how those are implemented as well. So let's get into some of these subjects in a little bit more deeper detail. Okay. So let's take a look at the data acquisition piece. You are going to need this data for training an initial model or for fine-tuning a model or for doing what's called retrieval augmented generation. We're going to get into all of these details. So let's just first take a look at where the data is coming from. It could be public sources. There's a lot of good hubs out there that have excellent data that you can be taking advantage of. A lot of it is free, but more likely, you've got data that's inside your organization, inside the firewall that you want to be taking advantage of. So that's going to exist in things like documents such as word documents or PDF, excel spreadsheets, things like that. Now you can't use the documents whole. What you end up doing is you end up pulling them in and then cutting them into a little fragments. And then those little fragments are used to give generative answers to people later on when the applications up and running. So the other place that you can get data from is internal databases. You could do queries to them, get those answers set, cut them up, save them away and use those for training. And the last thing is that you might want to talk to certain internal applications that you have that are mission-critical applications, that are running your business and get information from them -- from the APIs that they provide to you. So all of these are good sources for data. And when you're taking the data in, these are the things you want to be looking for. Probably, the most important is the second one, which is personal identifiable information because that can come with some severe economic penalties if something like that happens. So do not release PII. That's like job one. The other thing is you want to look at the intellectual property rights of that data. Is this data that is owned by you? Is it owned by someone else and you need to pay the money to use it or come to some sort of agreement with them? So make sure you've got all that squared away because you don't want to do all this work and then run into a legal problem about the data that you used. The other thing is you need to think about like any application -- any server side application, you need to think about access control, who was given the rights to do, what to this data, read only, read/write, delete? What are they able to do? And make sure that your application always uses the least privilege to go into the data, to get what it needs to do its thing. Don't be having a high privilege process laying around accessing all kinds of data and hoping for the best. That is a horrible strategy. You need proper access control that is founded in a proper IAM, where all of the user identities are kept and make sure everything stays secured. The other thing you want to look for is in the data that can sometimes be biases, that can be creating inaccuracies. They could be creating hard feelings. They could actually be harmful to people, some of the things that are in your data. So you want to make sure you scrub those out and look through the condition of your data generally, make sure it's in good condition, that it's in good shape to go ahead and be doing your training with. And not only before you do your training and while you're doing your fine-tuning, but also months from now, you need to think about your data sources. And who is responsible for maintaining them, who kept them current, who make sure the old stuff came down, the new stuff went up and so forth. So all of these are important considerations. So Don, tell me more about that data curation, tell me more about what we have to do to do the hygiene on our data to make sure it's as good as it could be. Well, let's take a look at the picture on the left. You've got internet scale data sets out there. Those have been downloaded, and you were able to detect the language that they're in. That's actually a big thing right there. A lot of documents are in the English language. If your applications being deployed in an area where English is not the predominant language, take that into account, make sure that your model knows the local language well enough. And -- anyway, then you're going to have to go through and you're going to find some text within the document that needs to be reformatted. So we have a tool set that lets you go in and do all kinds of text reformat and cleanup type operations easily. The other thing we're going to do is we're going to go through and we're going to look for like a blatant redundant information. And also, just kind of like fuzzy. It's like very, very close enough that it's duplicate information and we can knock it out just based on fuzzy duplication rather than exact duplication. You're going to want to go in and filter the quality of information that's in the document, potentially taking some information out. You might want to combine some of the information that's within your data sources, do data blending or -- in addition to deduplication there. When you're all done, you can use it to train your model. And what you're going to find for the graph on the right is that accuracy goes up. Why did the model's accuracy go up after it was trained? The reason is because the model understood your vocabulary, therefore it could understand what the user wanted to know, and therefore, it delivered a better result. You've got to teach the model your vocabulary and you've got to teach at the skills that you want it to exhibit for you. Okay. So once your data is clean and all scrubbed up the way you need it to be, then you're going to use it to do 1 of 2 things. You're either going to train your own foundation model, which is highly unlikely or you're going to take a model that someone else pretrained and you're going to do customization of it. That's about the 95% case. Most people are just doing customization of a pretrained foundation model. Now if you can't find a pretrained foundation model, then you're going to have to make your own. And that's going to be very, very expensive. It can run into the tens of millions of dollars. You need ridiculous amounts of data to get the foundation model trained up, and that's going to require a lot of hardware, a lot of expertise, a lot of time, a lot of money. And in the end, you're going to have a nice foundation model out of it. The path that most companies are going to take, though, is just going to jump in and find an appropriate foundation model for them. Now what makes it appropriate? Well, what are your constraints, what sort of throughput do you need to have, what sort of latency do you need to have? Are you -- is this AI application going to run in a data center? Is it going to run in a cloud? Is it going to run mobile and some sort of robot that's walking around? Is it going to make a big difference? Does cost matter? Does cost not matter? These are all things that will go into helping you to determine the foundation model that you want to choose. And once you choose that, then you need to train it with your domain-specific data. This is absolutely vital, to get it to understand your business, to satisfy your users and to give you the return that you want on the entire effort. So when you take a foundation model and you train it with your domain specific data, what you end up with is the exact same foundation model architecturally, but it's smarter because you've taught it all of your particular information. Now you have a custom model that is ready to be deployed and to satisfy your employees and your customers. As it turns out, when you're training artificial intelligence models, you end up having to do a lot of matrix multiply operations, and the nice thing about matrix multiplies is that those operations can be done in parallel if you have parallel hardware. So NVIDIA has spent a lot of time on this in terms of creating the GPU chips and creating tensor cores to accelerate these types of operations. So the first thing we're going to do is we're going to go in, when you're doing your training, either of the foundation model or you're doing customization of a foundation model either way, we're going to go in and we're going to find opportunities to do operations in parallel so that we can use multiple GPUs that might be in one system and -- so that we can use multiple systems that each have multiple GPUs in them. So these are the different parallelization techniques that we are using. The first column is basically exhibiting that we need to cut the problem up into pieces. The second column is indicating that we then need to take those pieces of computation and to send them out across the infrastructure so that they can all be done in parallel and to manage all of that. And then the third column is saying that sometimes when you're doing training, you don't need to recompute everything. Sometimes you can just recompute some of the values and that way, you get the best of both worlds. So we've pressed forward in the world on all of these fronts. We're the world's leaders because we need you to be able to take the most advantage of the underlying GPU hardware so that you can get your AI model trained cost effectively. Okay. The organization is in place, the data is all scrubbed, the infrastructure is appropriate for what we're going to do, and we're not going to create a foundation model. So let's look at the slide on the left. We're going to select a foundation model that's appropriate for our usage scenario, and then we're going to do what's called model customization. There's 3 types and we are going to go into those. Once the model is customized, then you have instead of a generic foundation model, you have your enterprise model. This is a model that you're going to use to move your business forward. Maybe you want to do some supply chain forecasting, maybe you want to do some financial modeling, maybe you have a sales pipeline analysis exercise that you want to go through or some sort of legal contract discovery. Summarization that needs to happen. Customer question-and-answer is a really big application area. Whatever you're looking to do, those AI applications are going to call into your trained enterprise model. Now the way that model gets trained, there are more techniques than this, but let's just look at these 3. The first is don't mess with the model, don't change its weights, don't change its biases, only change the prompt that you're giving it because if you give a model, a really good prompt, you can get a much better result out of it. You can give it some context. You can give it some vocabulary. You can give it some document fragments to work with. And the model will do so, and it will produce a better result for the user. And you haven't changed the model at all. You're only messing with the prompt that goes into the model, which is then going to result in the answer coming out. The second thing you can do is you can start messing with the model. You can go in and start tuning with these custom data sets, and that is going to change the weights and biases of the nodes within the model. And it's going to get smarter at the information that you're training it with. You need to be careful that it doesn't get dumber at other things than it used to know when you first met it. So you need to be careful with supervised fine training. But one of the most important things about supervised fine training is that the model is going to learn your vocabulary. That's crucial. And also, it's going to start to become a little bit more proficient in the skills that you want it to be exhibiting in your application. So that's fine-tuning. You can do it with label data, you can even do it with unlabeled data. There's a lot of aspects to it. We have white papers on this. We have blogs on this. You can become an expert on it, just come and visit our appropriate resources. And the last thing I want to tell you about is reinforcement learning from human feedback. This is a big deal. This is how ChatGPT became so popular, so fast. It was just very, very friendly to humans. It was giving answers and an interactive back-and-forth format to humans that the humans are resonating with. And so this is something that's going to be very important. Once your model knows it's vocabulary and is getting good at skills, then you want to give it some RLHF where, basically, you've got humans and you're saying, this was the prompt, these are the 2 answers, which one do you like better? I like A. Okay, noted. Okay. Here's another prompt. Here's 3 answers, which one do you like better? You like E, beautiful, duly noted. And you just keep going on that way. Now it doesn't scale because it's humans. But it's a very good source of information in terms of making sure that your model is going to be optimized for the problem domain in which you're going to deploy it. Okay. Now I just wanted to double click on these customization techniques, if you will, and go into this in a little bit more detail. Let's start with the graph and look at the y-axis. What we're looking at here is the amount of data that will be required for this particular customization technique, the amount of compute that will be required and therefore, gives you general feeling for the level of investment that you're going to need to be making. And along the x-axis are just 4 different customization techniques. Now in terms of prompt engineering, this is where you're not going to change the model at all. This is going to -- what you can do is what's called few-shot learning, where you give the AI system some examples. It's like monkey see, monkey do. When I do this, you do that; when I do this, you do that, okay? Now I do this? What are you going to do? You do that, right? Exactly. Monkey see, monkey do training. So that's few-shot learning. You can also have 0 shot learning. And sometimes you might want to take the answers that come back from the AI model and you may want to check with another AI model. Or you might want to ask the AI model that gave you the answer to check its own answer and to give you an assertation about its accuracy or its confidence in that answer. So that's a chain of thought, you can have tree of thought. All of these types of things can go on in the area of prompt engineering. The advantages of prompt engineering are that you're going to get good results. You're going to dramatically improve accuracy and satisfy your end users at the lowest investment possible. And it takes at least amount of expertise because you're only messing with the prompts. The cons are that you can't go as deep as some of these other techniques that we have been talking about. Now with prompt learning, what happens is we actually put a little model in front of the model that can help do things a little bit more than just formulating strength, which is what we're doing when we do prompt engineering. So we've got this thing called prompt tuning and another technique called P-tuning. And the result here is that you can go in and you actually are not -- again, you're not changing the parameters of the model as much as you're very carefully selecting certain parameters that you want to be changing. Okay? So it's not a wholesale changing of everything. And so this is great because you can learn new things without forgetting old things. So -- and the next way to go, let's go all the way to Column 4, is with supervised fine tuning, sometimes just called fine-tuning or something like RLHF. This is where you're going to be changing the parameters of the model. It's going to cost more money, it's going to take more hardware, it's going to take more expertise, but the idea is that you want to go through this to get the new vocabulary pounded into the model and to have it really develop those skill sets, really hauling them, the ones that you want to see it exhibit. And then the third column is just kind of like the in between. This is like supervised fine-tuning, but we have found ways to reduce the time required so that we're introducing new layers into an existing model and changing only the values of those nodes. This means that you're not changing the whole model. You're only teaching it new skills that you can get really good at, but it's never going to forget the old skills that it came in with when it was just a low refoundation model before you started teaching it and making it be your enterprise customized model. That's it. You've got 4 levels that you can deal with everything from, don't mess with the model, just mess with the prompt, to mess with the prompt more effectively with prompt learning, to mess with some of the parameters of the model to make it smarter, to mess with the most parameters of the model. It will be the most expensive and get you the best results, but you need to have the most expertise to actually pull all of that off. So there's a whole spectrum that you can operate on when it comes to customizing your model. There's a big difference between the generative AI application that you are intending to field for your organization and the ones that are out there that are just generic, right? When you go to a generic foundation model, be it ChatGPT or Llama 2 or Bing Chat, whatever, and you ask it a question like, "When did I last send a payment to my credit card company," it has no idea what you're talking about, right? And so the foundation model typically responds with something like, "I was trained 2 months ago, I was trained in September of 2021. I don't know anything past that date" or whatever it's going to tell you to let you know that it doesn't have the information that you're looking for. Instead, what enterprises want they want to take their custom data and to make that available to the large language model so that when a customer comes in and asks the exact same question, they can get a proper answer. Your last payment was sent on May 27, 2023. Perfect. So what we're doing here is we're taking a generic foundation model, and we're making it our own by giving it the custom data that makes all difference in the world. Domain-specific data is used for more than just training your model on new vocabulary and providing it with incremental knowledge and teaching at new skills. Your domain-specific data can also be ingested and put into a vector database. And then those little fragments, those little document fragments can all be considered by the large language model in formulating its response to the end user. So let me go through the diagram on the slide here, so you can make sure we're all following this. On the left, you've got a domain-specific question that comes in, in the form of a prompt. Now prompts don't go directly to large language models, right? ChatGPT is not a large language model. It's an application. It's an application that runs as a service. And if you can get to it, you can send it prompts. It talks to a large language model to get the answers. So your prompt comes into a program, which I am going to recall -- which I am going to call here, a retriever program. The job of the retriever program is to understand the prompt sufficiently, right, so you'd be able to parse it, you'd be able to get to the embeddings, get from words to vectors. And once the retriever understands the prompt, it can then go to a vector database where all kinds of document fragments have been stored with all of their appropriate indices and can do what's called a semantic search to find other things within that database that are similar to a particular concept that's coming through on the prompt. Now all of these similar little document fragments that could play in the answer, play a part in formulating a really great answer. They all get sent to the custom large language model. How do they get in there? They come in through the prompt. You've got 4K, 8K, 16K, you've got a certain context size that you can feed into the large language model in order for it to give you a great response. And one of the things you can put into that context is not only your prompt, but other background information, examples and document fragments that can be used to create an answer. And those actually get used to generate the answer and off it goes to the human. So this is called retrieval augmented generation. And the nice thing about it is that it can use the same data that you used for supervised fine-tuning. That would be fine. But typically, for fine-tuning, you use a little bit of data just to teach vocabulary and skills. And then the bulk of your data gets sucked into a vector database where it's going to be used to give better answers in this technique called retrieval augmented generation. In other words, the large language model is able to augment its answer by using information that got handed into it through the prompt by the retriever exactly. So that's how all these concepts come together to -- RAG is a very powerful pattern that's being used out there for AI-enabled applications today. So you found a custom model, you customize it with your own domain-specific information, you took even more of your internal data and put that into a vector database so you could do retrieval augmented generation and you start doing some evaluation of the whole thing and you realize that it's awesome. And so now it's time to deploy it at scale into production. And so the first thing you want to think about when it comes to the deployment is you want to take a look at the actual model itself. Is it little? Is it big? Does it have a few dozen neurons? Does it have hundreds of neurons? How many layers? How many parameters? Is it dealing in? And therefore, based on those things, it will have a certain memory footprint that will be required to hold it. And that's very important. You don't want to run out of GPU memory. And so you're spilling over into system memory. It's like missing system memory and having to go to disk. It takes forever to get something off of a disk, and it takes forever to get something out of system memory compared if it were already in the GPU's memory. So can the model fit in GPU memory? We can break a model up and make it fit across multiple GPUs. That's part of the NVIDIA magic. Anyway, you want to look at the model attributes. The next thing is you want to understand some of the aspects of the use case. Like is this a type of thing where someone is going to give me a video and then I can batch those up and then just give them transcripts back or is this going to be more of like a real time? Someone is going to be coming in and asking interactive questions. So your architecture is going to vary depending on the real time or the batch nature, right? The other thing is, typically, you would -- an application will come into the foundation model, the large language model -- your enterprise large language model through HTTPS. So it's just standard Internet protocol, secure and you make the call, you get the answer. You make another call, you get another answer. But if the latency on that is too long, if you need to get down to 30 milliseconds or less, when it comes to voice, people do not like even the slightest amount of delay. So if you start getting into those situations where the latency is going to be super important, then don't use rest APIs. Instead, you can come in through a gRPC mechanism. That's going to be a lot faster, lower latency for you. Now you've always got a trade-off, though. I could build a system with very low latency that had terrible throughput or I could build a system that has tremendous throughput with a horrible latency. Everyone waits a long, long time and then all those people get satisfied right away. So that's another way you can do it. The other thing about the use case is you want to understand like with ChatGPT. When you give it a prompt, it starts giving you the answer right away. It gives you something to read. While it's formulating the rest of the answer, right, and predicting those next tokens, you could be reading the ones that I already gave you. So that's called token streaming. And if that's something you want, then make sure that you've got that all in there. Now based on the attributes of your model, based on the attributes of your use case, you can now go ahead and put together your infrastructure. Is this something that you're going to run in your own private data center? Is this something that you're going to run at a cloud provider? Is this something where you can just like get some virtual machines together that have some GPUs in them and this thing will be off to the races or do you need to reserve some serious DGX capacity? Yes, we're going to always have the spectrum out there as to what people need for compute, what they need for memory. And then you always want to do the typical things that an engineer would do, which is you start thinking about the fault tolerance of the system because people are going to fall in love with it. They don't want it to go away, start thinking about the load balancing as the things scales and more and more people come on, how are you going to bring more and more capacity online to meet the demand. And then when demand goes to last late at night, can you constrain that capacity down and save yourself some money. So elastic load balancing. Typically, you could use things like Kubernetes for this since it does fault tolerance and load balancing and takes care of nonvolatile storage capability and a bunch of other things that Kubernetes just going to do for you. Also, security is going to be an important concern, as it always is monitoring the system so that you understand its health at any given time. And it might not just be an AI system sitting there by itself using data that's in the vector database. It might be constantly reaching out to other data sources within your organization, other applications within your organization. And so what are those, what are the integrations that are going to be required. Once you've got all that figured out, then you've got a production-ready deployment at scale that will deliver the results on your enterprise AI application that you hoped for and that your organization are hoping for. So your infrastructure is in place and you're ready to go. Well, you don't have to piece together dozens of open source applications and snippets and libraries and things in order to get all of this working. What we've done is we've created some optimized kernels that are all ready to go, to run your AI application and to make sure that it's accelerated across all of the GPUs that you want to give to the cause. So we can use multi-GPUs that might exist across the multiple nodes to run this model at scale for what we call inference. So what's happening is there's going to be a communication happening where the AI application is going to call on the large language model to do inference. And then that model in order to give a good result, it's going to need to do a lot of communication among all of the GPUs to figure out and do the generation that it needs to actually deliver the result that you're looking for. So that -- that's how it all works together. We've got premade containers. You can just come to NGC and download those or even from your cloud provider, a number of different ways to get a hold of our containers, and they've all been tested. They all have all of the appropriate software within them so that you know that you've got a production-ready bundle that you can go forward with there. So now you've got your application deployed. It's scalable, it's reliable, it's secure. Everything is good. You just need to be careful about the answers that it's getting out. So this is where we come to the technology known as guardrails that I mentioned earlier. So as you can see on the left in the -- on the slide on the left, the types of things that guardrails are going to save you from when people try to get off topic and ask the LLM things that are inappropriate for the way it's been trained and the intention that you have for it. It can also -- guardrails are important to look for things that could be safety concerns and catch those answers and make sure those don't get back to users. And lastly, anything involving -- you can say to a large language model, you can jailbreak and tell it to get out of its confines and break loose and do things for you. So you want to have another application that's kind of like the comp that's making sure that there are no security breaches happening. So on the left is prompt coming in. And that's coming from either a user or an API, and it's going to the thing in the middle, which is a large language model. And you need to be sure that you've got guardrails in the loop there to make sure that everything is always the best that it can be. So that's your last component to a fully functioning enterprise-ready generative AI application.

Amanda Saunders

executive
#4

That's great, Don. I think Don did an amazing job walking through each step of that end-to-end pipeline when it comes to building, customizing and then deploying these generative AI applications in an enterprise setting. And I think it gives us a lot of food for thought on how we build production-ready generative AI. Because as much as we want to take advantage of the tools that are out there, we want to make sure that when we're deploying these applications, they have that rigor, that security and of course, any support needed to run these in production and run them at scale. NVIDIA, we've been in the AI space and have been leading the way on generative AI for many years and investing heavily across this pipeline. So I wanted to talk a little bit about what makes working with NVIDIA unique, and when it comes to specifically these customized and enterprise-grade generative AI solutions. And the first one is exactly what I just talked about and exactly what Don just covered really looking at that end-to-end pipeline, ensuring that we've got a complete solution that helps you go from customization, taking a pre-trade model, customizing it with your own data, connecting it to your own data sources and then deploying it out in the market. We have acceleration techniques that we use across this entire pipeline and that we've worked to help build into the ecosystem of solutions that are out there to help you build that comprehensive generative AI solution. Accelerated performance. This is what we're known for. The accelerated infrastructure tools that we have, allowing you to do multi-node, multi-GPU training as well as inference. Because unlike other workloads, this isn't just a heavy workload at training time. We also have to make sure that we're deploying it in the right way. So I think we've done a ton of work with our teams, with our software and our hardware to make sure that we're providing that total upper end of performance and throughput that actually ends up delivering the best TCO on the market. Production ready is an important value prop for us and for our -- the organizations that are running at scale, particularly when it comes to deploying inference applications. You want to make sure that these are secure, that they're optimized, that they've been tested along the full stack so that you have that support security and API stability. And we have a solution called NVIDIA AI Enterprise that makes this possible. Now a lot of the organizations we work with, they're adopting generative AI wherever they're building today so whether that's on-prem, in the cloud. Some of them are developers who are working on workstations, but they want to be able to make sure that they can run the generative models they're building wherever they need them. And again, that may be where they live today, but it could change in the future. They may want to deploy it in multiple locations. And so that ability to run anywhere and still get that accelerated experience is critical. And all of this combines together to that increased ROI. When we're looking at the total cost to build these generative AI solutions, it's not just about the investment in the infrastructure, it's not just about the software that we run on top of it, it's about the teams, it's about the expertise. And by pulling together a full end-to-end pipeline for our customers, we're able to improve that return on investment and deliver the best possible results. So when we're looking at this end-to-end pipeline, these are the components that are required to build generative AI. And I'm going to talk through a little bit about each one of them and some of the tools that NVIDIA has to help. Of course, there is your training data, and Don mentioned how important this is. We have data curation tools that allow you to streamline your training data, get the most prepped and clean data so that it's ready for that AI training so that you can get the best possible results for your models. Accelerated computing. Like I mentioned on the previous slide, we allow organizations to run wherever they want, whether that's in the cloud potentially using our DGX or DGX cloud solutions, but also in any of the major providers where we've got accelerated infrastructure already in place today. We also work all of the with the major system builders out there in the market to ensure that we're delivering the high-end accelerated infrastructure required to run wherever, whether that's in your data center, maybe that's even running at the edge, running on your desktop. We have solutions that allow you to run across the board. We also have invested a lot in the training and inferencing tools. We know that this isn't just about putting a GPU into your system and immediately speeding things up. We need ways to accelerate not only the training and customization but also that inference so that we can run at scale. Our NeMo framework is a end-to-end cloud data platform that allows you to do everything from that data curation, training and customization to accelerated inference. We've taken all the libraries and SDKs that NVIDIA has to offer, and we've actually packaged them into that solution. We also have additional solutions for the life sciences space with BioNeMo, same thing, LLMs, but designed for the biopharma industry. And of course, we're known for our graphics. We have our Picasso service, which is an AI foundry for driving -- generative AI for visual content, whether that's images, 3D, video, you name it. And then, of course, the final piece is AI expertise. NVIDIA as well as many of our partners have built this AI expertise in-house to help you get started. So whether you're looking for a partner to help you build your first application or you mean help scoping out what kind of infrastructure is going to be required, we have people who can help you get going and really bring these enterprise solutions to market. All right. So we've made it to the last slide before we get to the live Q&A. I just wanted to leave you with a couple of next steps for your generative AI journey. First and foremost, we have our LLM Developer Day coming up on November 17. This is going to be a half-day virtual event. So you can log in from wherever you are. We've got hands-on practical sessions all around advancing LLM application development. So whether it's you or your team, whoever is working on developing these applications, this is a session not to be missed. We're going to have all of our experts who work with organizations every day, sharing best practices, giving tutorials, providing demos and helping you really accelerate that development within your organization. Next, we have our video solutions for generative AI. I touched briefly on a couple of our solutions. But if you really want to deep dive into the hardware, software and services that we offer, you can go here and you can learn all about them. And then, I finally wanted to touch on the generative AI chatbot workflow. So Don spoke about retrievable augmented generation. This is a great new technique that's out there in the market and it's enabling us to connect that enterprise data to these LLMs. The workflow is a resource reference design for building enterprise co-pilot assistance, chatbots with retrievable augmented generation. So a step-by-step guide that will help you get the setup for your environment. So we will be releasing that shortly and cover it a lot at the LLM Developer Day. But if you're interested in getting notified as soon as that's available, please click here on the link or the link in the resources and get started.

Amanda Saunders

executive
#5

Hi, and thank you guys for tuning in and continuing to ask amazing questions in the chat. I know the team behind the scenes is actively answering as many of those as possible. But there were a couple that we wanted to address to the group here. So I'll just -- I'll jump in and Don, I'd love you to give some of your thoughts. I think you covered this briefly, but just really succinctly, like when do you do fine-tuning versus when do you do training from scratch? What's your best recommendation there?

Donald Chouinard

executive
#6

Right. So think about doing training from scratch is that it's going to require huge amounts of data. This is in the news all of the time about how much data is required to train a model from scratch because you've got to teach it language. You've got to teach it logic. So you've got to be up for it. If you have the resources to think about creating your own foundation model, you're in the minority. 95% of enterprises are going to pick a foundation model that's appropriate for their usage and then they're going to begin to just find train it. And it's like starting a ball game when you're 95% the way to the goal line, okay? You want to start with the foundation model and then do your supervised fine training so that it understands your vocabulary and your tasks and then do some RLHF so that it's very human friendly. And when you deploy it, please deploy it with guardrails.

Amanda Saunders

executive
#7

Yes, absolutely. Yes, I think a great example of a customer who did find the value in the training from scratch is Bloomberg. So there's a lot of information about what they've done BloombergGPT online. But essentially, they didn't need a super large model that knew everything. They needed a focused model that just understood their data that came from their platform and their sources. And so they were able to get better results without having to have an extra-large model, which I think is great. The next one...

Donald Chouinard

executive
#8

It's really fun working at NVIDIA because we know who's building the biggest, baddest model you could ever conceive of right now. And it's not just one company either, right? When you need to build a foundation model, that's going to have billions of parameters and you're going to feed it terabytes of data -- terabytes of data, you're probably doing business with NVIDIA and you probably use NeMo because that's exactly what it's all about.

Amanda Saunders

executive
#9

Yes. Absolutely. So the next one -- and this one, I'm going to -- I'll give an answer, and then Don, you can chime in. I wanted to address this, what's the typical effort in time a company needs to set up a fully working solution from data preparation to deployment either on-prem or in the cloud? So I hate to say it, this answer is it depends, and it really all depends on what you're trying to accomplish, what your end goal is. I will say, again, if you've got the expertise in-house, you've had applications that we use in our company, they can be set up in as little as a week. These typically are pretty standard chatbots. We're using foundation models that are out there in the market, and then we're training them on data that we have, and those things can get spun up quickly. Then comes the rest of the process, which is you built the chatbot now how do you fit that into your processes, how do you fit that into your organization. So I'd say for most of the organizations we work with, there's probably at least 3 to 6 months of development time that really goes into understanding what's the use case, how we're going to build models for it, what's the phased approach. And so to get out that first MVP, is probably in that 3 to 6 months' time frame. And then as the use cases get more complicated, the time can go up from there. So I hope that's valuable even though the answer is it depends.

Donald Chouinard

executive
#10

I've got two cents to put on top of that. My readings have been showing that it can typically be 80% of the effort just on getting and cleaning up all that data. And as we all know, when you want to field a nice new application for your organization, dealing with your IT department is also a significant amount of time and communication that will be [indiscernible] your own internal IT and getting that data clean. Those are the huge part of the project actually.

Amanda Saunders

executive
#11

Yes. That's a good point. All right. Another question I thought was great for the team to hear. Can you give a rule of thumb for when to use RAG -- I'm actually going to break the question a little bit differently. When to use RAG, when to use fine-tuning and when to just use a pretrained LLM with no adjustments? So I don't know, Don, if you want to dive into that?

Donald Chouinard

executive
#12

I will. Yes, sure. Yes. It's an excellent question. It tees up all the points I wanted to make. Thank you so much for the question. So if you ask a generic foundation model, anything that's particular about your business, it's not going to return the answer. If you ask it about something new, it won't know, something external that's out on the Internet, it won't know. So you need to open up your language model to the world. Now the data that your language model becomes exposed to and learns, part of that is so it can learn vocabulary and skills, but you can also take that data, cut it up and put it into a vector database so that when you deliver answers to people not only do you know what they're asking because your model knows the vocabulary because you trained it, but you can deliver really good answers because you're using document fragments that are your own. They're like the choice, little snippets that should all be brought together into that final answer. So hopefully that -- I would not use too much RAG on top of a generic foundation model because like if you don't understand the question, you won't get the best document fragments to put into that potential answer.

Amanda Saunders

executive
#13

Yes. And I'll give an example because again, we build a lot of these in-house. We use LLMs a lot at NVIDIA. We have an application that's built on our announcements and our marketing [indiscernible] data, right? All the information that we put out and people want to be able to search that using natural language. That's something that you can use a pre-trained LLM. The LLM understands language. It can read through our content using RAG when we retrieve the right documents for it. And so that one doesn't require any retraining. Now other applications where there is proprietary data, maybe on our solutions, maybe on our bugs reports, things like that, yes, absolutely. You definitely need to train the model so that not only does it know the information, but it's delivering more relevant answers. So there's less sort of work that the human has to do on the other side. So I guess those are maybe a couple of examples. Again, another fantastic, it depends question...

Donald Chouinard

executive
#14

And that's why when people ask this question, should I do supervised fine-tuning or RLHF? And the answer is, you want to do both. Because the supervised fine-tuning teaches your model, your vocabulary and then the RLHF, now when you're giving your humans a choice of 2 answers and asking them which one they like better. You're giving them a good choice to make. If you took a generic model, you'd be giving your human like 2 iffy answers, and they pick the one that's not as bad as the other. If your model is trained, then the human is looking at 2 pretty good answers and pick them the better of the 2. So they work together.

Amanda Saunders

executive
#15

Absolutely. And the final question I'm going to answer with our LLM Day. So there was a question about best practices alongside the NVIDIA tools, our developer tools and things like that, where can we learn more? So obviously, a lot of information available. And I think we put a lot of the links into the chat as well as in the Resources section. But if you are interested in best practices and how to use and accelerating your own development process, please join us at LLM Developer Day on November 17. It's a great interactive set of sessions that are really going to provide a lot of that detailed technology and how to's so I highly recommend that. And with that, I know we're at time. So I want to thank the hundreds of people who joined us today, and we'll see you on the next one.

Read the full transcript via the API

You're viewing the first half of this call. Get the complete NVIDIA Corporation transcript — plus 248,000+ transcripts from 12,000+ companies, speaker segments, AI summaries and full-text search — through the EarningsCalls.dev API.

Get the API View API docs →

This call discussed

For developers and AI pipelines

Programmatic access to NVIDIA Corporation earnings transcripts and 248,000+ others is available through the EarningsCalls.dev REST API. Plans from $24.99/month — full transcripts, speaker segments, full-text search, and the recently-added /api/v1/transcripts/recent polling endpoint for ETL pipelines.