NVIDIA Corporation (NVDA) Earnings Call Transcript & Summary
February 8, 2023
Earnings Call Speaker Segments
Unknown Executive
executiveHi, everyone. Thanks for joining us today for our live webinar, Drive Innovation and Speed Up Scientific Workloads. Before we begin, we wanted to cover a few housekeeping items. At the bottom of your screen, you can find various widgets for use in your event. Once open on the screen, they are resizable and moveable. If you have any questions during the webcast, you can submit them through the Q&A window. We'll try to answer these at the end of the event and in between each speaker. A copy of today's slide deck and additional help materials are available in the resource list. We encourage you to download any resource or bookmark any links that you may find useful. Here's some tips that can help make this event as best as it can be. To maximize the quality of this audio stream, please close any open applications aside from your browser window. Also a good old-fashioned browser refresh can cure many ills. And so if you hear audio sputters or the slides seem to be lagging, give that a try. You can also try opening this event in a different browser. If you encounter any other technical issues today, please let us know in the Q&A box, and we'll help you troubleshoot. Now with further ado, we'll turn the event over to our speakers to begin the presentation.
Earl Joseph
attendeeThank you very much. This is Earl Joseph with Hyperion Research. I want to welcome everyone to joining us. These are exciting times with lots of new technologies, new interesting ways to construct computers, make them more performance and productive. And today, we're going to be talking quite a bit about DPUs in the industry. For those of you who don't know us, Hyperion Research does a lot of market and technology studies. We also operate the HPC user forum, and we welcome everyone to join us on that. It's where we, 4 times a year, get vendors, buyers, users altogether to talk about best practices and new technologies. And with that, I'd like to give you just a quick overview of what we see going on in the market today and kind of our forecast for the next 5 years. So for the 2 charts you see here, the first one is the on-premises purchases of servers, storage, middleware applications and services, and we see the HPC market growing very robustly from just over $32 billion in 2022 to over $40 billion by 2026. So a very healthy growth on the on-prem part of the business. The lower chart is showing us the growth in cloud computing. And what I mean by cloud computing, this is the money that people spend to run their HPC, AI, quantum computing problems in the cloud itself. And we're seeing extremely robust growth here of over 17% a year over the next 5-year period. When you add these 2 together, the market is actually going to be greater than $50 billion in 2026, which is just amazing. But there are many other segments in the market that are even growing at higher rates, in particular AI, with machine learning, deep learning, those workloads are growing at over 22% right now. So some really strong growth. So here are our top 10 predictions that we have for the coming year. We just finished these. As I just mentioned, the strong growth. There's a lot of issues around the supply chain sector and government policies and so on. But today, we're going to talk about prediction 7 and 8. And those 2 predictions are #7 is that we expect systems architectures to actually split into multiple types of systems. And we're going to have some systems that are highly optimized for traditional computational work. And we're going to see ones that are going to be much more optimized for memory and data-intensive work, and then those that for AI machine learning and deep learning. So we think the world of the future data center will not be to have one singular system, but they have multiple systems that can address these different architectures. And we see these divergent requirements really driving the architectural focal points and the requirement for interconnects, improved storage systems and getting in and out of the storage system, moving data around, especially with the growth in model sizes and the growth with machine learning and deep learning are just going to explode in the marketplace. And so let me start with prediction #7 here. As far as the architectural changes, it's a continued prediction from last year, but we think there, users and the systems designed for HPC users are going to factor in new requirements. There is a lot of new workloads. I mentioned the AI ones, big data. But in addition to the growth there, the amount of workloads, application areas are just phenomenal. We're seeing some areas where within 60 days, a new concept or idea hits the market, and it has, last I heard, 15 million users on one case. It's a push for a faster time to solution. If you're doing things more in real-time or near real-time applications, we're seeing those applications coming on board, too. And a lot of interesting new areas of research. And for the system decisions, we only go to purchase a machine or a vendor goes to build them, there's going to be a much larger and diverse set of building blocks. We're seeing all types of new processors, GPUs. And then you see new types of applications of processors and DPUs and different parts of a system. And as I mentioned earlier, we think we'll see multiple smaller systems. But when I say smaller, these are still large systems. It's just not one singular one. And then we're also seeing the use of the cloud fitting into almost all sites, overall architecture and operations in the future. And then in addition to this, these heterogeneous systems will incorporate data-focused areas versus processing intensive designs, a lot of node different configurations. For example, the most extreme I've heard recently is with Amazon. Amazon right now, if you want to run the HPC job in the Amazon cloud, you have over 450 different processing node configuration to choose from. So you almost need an AI system to figure out which is the most optimal way to do that. And then infrastructure accelerators like DPUs to offload some of the processes, some processing accelerators into the storage and as far as moving the traffic round are going to have just a much more impact and you're going to see them much more frequently. And then we're also going to see some smaller systems targeted to very specific application. I mentioned here, AI and big data, I'm also mentioning within that realm, a very specific area, as far as the application area. So we think a lot of data centers will have the split of the large system maybe into 2 processor memory one or a third one and then a number of small ones that are very application-specific. And with that, I think I'll hand it to Mark.
Mark Nossokoff
attendeeThanks, Earl. I'll touch on kind of related to the last prediction, but the -- as far as different architectural perspective, we are definitely seeing a much increased focus on storage and interconnect, in particular, as a new architectural focal point for any number of reasons: performance, both the bandwidth and latency aspects to -- with the ultimate goal of improving time to results of applications and simulations. And with the emergence of the newer modern workloads, data-intensive workloads, complementing and even maybe even more so dominating systems and the traditional compute-intensive workloads, the interconnects and their ability to be able to handle a much more diverse I/O workload profile is becoming even more critical and will be much more increased focal point for architects in defining their systems. I wanted to touch on just a couple of points from our annual site surveys that are related in this area. We've been looking at part of the architecture, interconnect, networking architecture, differences in the -- in whether there are separate independent networks, node-to-node networks, node-to-storage networks, and converged networks. I wanted to provide a snapshot of what's happening of the most recent site survey on converged networks. And it's showing that the -- very strong leadership of InfiniBand interconnect with still a large adoption of Ethernet as well as the interconnect technology. What is interesting is a fairly strong beachhead adoption of the 400 gigabit line rate and aggregate rate for the interconnect is being adopted and emerging, which is quite interesting how it appears to be being quickly adopted. Apologies for that. We're also looking -- and we've seen GPUs become emerging, and we're also interested in how people's current adoption rate is and future adoption rate is of other offload type of accelerator architectures, including DPUs. We see today that there is somewhat of an adoption of about half of the sites surveyed are adopting some other form of acceleration, whether DPU, APU, other types of processing units. But we see moving forward into the next currently planned procurement that not only is just the alternative acceleration types of architecture is going to be adopted, but the adoption of DPUs themselves, our interest and adoption is growing quite rapidly as well. That closes out our intro and market overview section. I do want to invite people to submit questions in the Q&A within the platform, and we'll be answering them as we go. We will also be -- as we go, as they come in as well as at the end of each presenter, we'll field questions and then we will close the session with a more roundtable Q&A as well. And with that, we'll move on and invite Gilad Shainer from NVIDIA to join in with his presentation on NVIDIA's HPC networking platform. Gilad?
Gilad Shainer
executiveYes. Thank you. Thank you very much. So what makes NVIDIA unique? What makes NVIDIA unique? And I would say that NVIDIA is essentially a computing company. A computing company in the way that we build the hardware elements, the silicon devices for CPUs and GPUs and switch networks and NICs and DPUs. The system software that runs on top of that, the software development kits that enable our customers to write their applications on top of our platform and be assured, by the way, that future generations of the devices that we build will carry what they already written and a further improved performance and capabilities. We created platforms on top, and then we create applications that runs on the entire system. So when you send everything in a full stack approach, we built computing platforms, computing systems. We're using those computing systems in order to understand the next generation of applications, the next generations of performance bottlenecks. And with that, we're building the devices that will take us to the next generations and so forth. On the networking side, the NVIDIA HPC networking platform includes the Quantum-2 InfiniBand Switch, a 400-gigabit per second switch platform that introduces in-network computing. And in network computing is one of the key elements in the network to actually analyze data within the network closer to where the data is and that can accelerate applications in a very dramatic way. The ConnectX SmartNIC with 400 gigabit per second interface, PCI Express gen 5 in different kind of accelerations around data movement. BlueField-3 DPU, which I'll talk a little bit in more details about it, which runs at 400 gigabit per second, PCI Express gen 5 and including in-network computing engines for storage, for infrastructure, for security, for simulation such as NPI for AI workloads and so forth. Our Spectrum-4 Internet switch, we're creating Ethernet platform for high-performance computing Internet platform for AI. Our customers that utilize or use Internet for their infrastructure, for AI infrastructure, will be able to leverage from Spectrum-4 to build their communication network or communication fabric. And the ability to manage to orchestrate, to monitor, to look on telemetry information and so forth from the fabric from the network from the system itself. So building the full stack, looking on the full stack, looking on the applications of learning today, the ones that we want to run. We identified several performance bottlenecks that impact workloads, impact our workloads, which we aimed or plan to solve with, first, in-network computing; and second, with the DPUs. The first one is be able to achieve overlapping, overlapping between compute and communication. Running simulations, multiprocess simulations includes compute cycles and includes communication cycles and be able to hide communication cycles behind computation cycles, be able to achieve asynchronous progress or overlap in between compute and communication is something that was desired for many, many years. It's kind of the holy grail of HPC. And finally, with DPUs, we are able to achieve something like that, and I'll show you examples later on. Load imbalance, one of the key performance bottlenecks. When we're doing multiprocess simulations, it's not that every GPU, every CPU finish the work at the same time, and then you can do multiprocess communications or multi-process synchronization. It's not a fact. And therefore, a slow CPU, for example, can slow down the entire system. By utilizing the DPUs for managing the communications, by managing NPI operations, for example, we can finish those operations even though CPUs might be still busy. And we'll then be able to reduce the effect of loading imbalance and improve application performance by that. Jitter is another area. The infrastructure workloads running on a CPU create jitter, and the jitter affect performance of the system. It can be jitter because of storage processes, infrastructure management processes and so forth. And that's another area where the DPU was designed to. The DPU was designed to run the infrastructure processing to separate the infrastructure demand from the application domain, to run storage, to manage storage access. And by moving the infrastructure processing outside of the host to write on the DPU that enabled to reduce the jitter or almost eliminate the jitter effect and enable to have CPUs and GPU fully dedicated for running HPC simulations to AI workloads. Last and not least, multi-job performance. More users running HPC, we are hosting more jobs or running more jobs on a supercomputer. We are hosting more users on HPC cloud. We need to make sure that one job, one user cannot impact another user, another job performance. And this is a problem that exists today. And with the combination of e-network computing telemetry that run on our platform as well as DPUs, we can use reinforcement learning capabilities in order to isolate jobs performance from another job performance and provide the best performance to a job or process regardless of what other jobs are running in the same system. So the DPU is one of the key elements that enable us to overcome those bottlenecks and dramatically improve application performance or the full HPC or AI system performance. The DPU combines the entire ConnectX inside. So it does include the entire NIC functionality inside. BlueField-3, our latest DPU includes the ConnectX 7 inside. So it's a 400 gigabit per second device with PCI Express gen 5. And it includes 2 areas of computing programmable elements, 16 anchors within the device and a data path accelerator. Data path accelerate can do quick things in real time on the data, and then the anchors can run or host multiple processes or elements of services that can enable to migrate the infrastructure processing to the DPU to deal with storage, to accelerate simulations, and accelerate AI applications. And that's what we do with the DPU. The DPU runs fully RDMA, so you can touch the host memory, you can get to the host memory completely bypassing the CPU or GPU or accessing GPU memory, bypassing the GPU in that sense of or bypassing the CPU. And it can access remote host memory as well, bypassing the remote CPU. So the DPU can actually move data locally and remotely and bypassing the CPU in that sense and do it in a very effective way. And by doing that, to completely manage NPI operations, for example, or help with NCCL operations. And with that, as I mentioned, to overcome the bottleneck that I described before. So in-network computing is the key. And BlueField is the key for overcoming the performance bottlenecks that I mentioned. In-network competing on the InfiniBand side exists also on the switch. The Quantum-2 switch includes data reductions, small message reductions and large message reductions. And those large message reduction can support virtually unlimited concurrent large message reduction risk that runs on the network that improved reduction operational performance and in turn, of course, increased performance of AI training, deep learning and NPI applications. BlueField, I mentioned, includes the programmable units around the security around storage, around NPI, around NCCL. And BlueField-3 compares to previous generation, improve the compute capacity of it by around 3 to 4x, improve the memory speed that it has, and it's become an important factor for building the new generations of HPC or AI platforms. So here are a couple of examples on what we do on BlueField related to NPI, related to simulations. Here I took in, for example, a multiprocess operation, one of the most complicated one, ones that actually it's more sensitive to anything that happens on the system and impact performance dramatically of different workloads. Here is an example of Alltoallv. So Alltoall with vector. On the left side, you can see the performance benefit of moving that operation to the DPU. And on the right side, you can see a nonblocking in Alltoallv performance and the impact of moving that operation to the DPU. You can see that with the blocking operations, we're almost reducing that time or improving that performance by 2x and for the nonblock operations, it's even more than 2x. Not less important, we're able to achieve full overlapping, full overlapping between compute and communications, which means that we can completely hide communications behind computations and of bringing the Holy Grail of HPC to HPC finally. What does it do to applications? It's increased data center or HPC center performance in a very nice way. So at the beginning, we started with FFTs when we start the work on the DPUs and demonstrate more than 1.2x on FFTs, and we'll continue to improve that performance, continue to optimize FFTs. We did an example on a wire code, demonstrating more than 1.2x performance and recently working with Max Planck Institute on Octopus, which is a physics and chemistry code, demonstrating also more than 1.2x performance improvement. Those improvements are due to the ability to move NPI to the DPU and then the DPU will be able to manage and execute those NPI operations. By having that, we're going to accelerate almost every NPI application out there. Almost every scientific computing applications will be accelerated by around 20% because of using the DPU. So those are a few examples at the beginning, and there will be more examples coming out shortly from us and from our partners. Now what does it mean more than 1.2x performance improvement? It's not just the performance itself. It's bringing BlueField into the server and you're using BlueField instead of the NIC for the connectivity and for the infrastructure processing and NPI accelerations. We're also increasing the performance to TCO, performance to dollar by almost 1.2x. And also improving performance per watt. And we know that power has become one of the limiting factors, one of the elements that determining what we can build. And of course, we want to reduce power consumption in many data centers of power limited. So bringing the DPU into the AI and HP systems also improve the performance per power, improve the performance per watt by almost 1.2x. Those are great elements here. It saves millions of dollars for building -- of building the HPC in data centers. And for the BlueField is something that we'll see become the de facto element in every server, CPU servers or GPUs. So with that, if there is any quick questions, happy to take. And if not, I know that there's other presenters that -- who are going to present cool things today.
Mark Nossokoff
attendeeThanks, Gilad. There is a question that has come in, and it may have been answered after it came in, but I'll go ahead and share the question that came in. How do you envision using in-network computing to improve simulations in HPC?
Gilad Shainer
executiveYes. So in-network computing means that we are moving data algorithm that can be better executed on the network instead of doing that on a CPU and a GPU. Reduction -- data reduction is one example, and data reductions improves simulations and deep learning applications in 20%, 30% and 40%, which is amazing. So this is one area. The in-network computing capabilities on the DPU, the programmable in-network computing on the DPUs enable us not just to improve security and improve storage access, but also NPI. And you saw the results of doing -- running NPI on the DPU instead of doing that on the host and achieving more than 1.2x performance improvement. So that will help to improve performance for data centers that obviously limited in expense. And of course, power care data centers, DPUs is a must for them.
Mark Nossokoff
attendeeAnd quickly, there was a secondary question to that. Is in-network computing an alternative for synchronization? Or is it better using a DPU? And if so, can SHARP be used for that?
Gilad Shainer
executiveIt's not one versus the other. It's actually -- you're using all of those. So SHARP is running on the switch and does the data reduction there. The DPU uses the switch for doing the data reduction. So data reduction elements are managed by the DPU. The DPU is using the hardware elements, the in-network computing elements on the switch in order to run data reduction operations such as reduced and unreduced and [ bigger ] and broadcast. Other operations such as AllGather and all-to-all, those are running completely on the DPU in that sense. So once you combine the network competing on a switch and in-network computing on the DPU, you get amazing results.
Mark Nossokoff
attendeeOkay. Great. Thanks, Gilad. If any more questions come in for Gilad, we can submit them and hold them for the Q&A at the end. At this point, let's move on to our next speaker, Dr. Jithin Jose from Microsoft, who's going to talk about InfiniBand network computing technologies at Azure.
Jithin Jose
attendeeThank you, Mark. Thank you. So in this talk, I want to mainly focus on the NDR and some of the performance systems with NDR on Azure. So let me please start with the different breakthroughs, I mean, over the last 2 or 3 years. So I mean, in 2019, like we started the very modest 20k core MPI job, it was for the cloud. Later -- I mean we offer like HDR 200 gigabits per second IP, and we try to scale out more like [ ADK, MK ] jobs again, first, for cloud. And later like we built multiple different SKUs with HDR, like focusing on the traditional HPC applications and also AI plus of applications. And then eventually like -- I mean, we, according to the top 10 supercomputing rank and again like top first for cloud also and many more in the top 4 key ones, right, for the different types of workload. And I'm happy to share that these are actually numbers on a public cloud. It's not like any private or any dedicated cluster that we have. So any more like should be able to like go try them and then reproduce these numbers, right? And again, like end of last year, see, like we announced in the RFQ, I mean that is the HBv4 AMD Genoa SKU. So I mean in this talk, I want to focus a little bit more on that, like how we design this SKU, what are the motivations behind the different design choices and also some of the preliminary performance systems, right? So yes, so this is -- I mean, the general architecture of our scale, like HBv4 or HXv1 Genoa-based AMD SKU with NDR-400, of course, right? So as you can see, right, so we are -- we picked the 96-core based solution here. So I mean 96 cores per socket and we set the number of new marker socket as 2. So 48 cores for NUMA and then 4 of them, right? So -- and we selectively -- we selected a few cores from the different NUMA nodes and dedicated those on the host, right? So the Hyper-V, the host stages, which coordinate all the rollouts, the performance monitoring or monitoring different telemetry, et cetera, those are completely done on that, right? And so there is a complete isolation between hypervisor and between VM, right? And the VMs actually get the full -- a full control of those cores, right? And of course, NDR is actually exposed over SR-IOV. So all the features that -- I mean, that NDR is offering. So different features that Gilad mentioned, like we can call it the performance results, right, but all of them are visible to the end users, right? And until we can come up with like different designs based on the different workloads, like Earl mentioned in the very first slide, right, many of the architectures are now being like focused based on workload characteristics, right? So just to support from that, right? So we have 2 different configurations for this SKU, right? So one like 700 GPU memory configuration and another one, the higher recomputation because certain plus of workloads, as you know, right, like they need a higher memory per core requirement. So that's the main situation. And again, another dimension here is like we can selectively pick core that are from the different CCDs. So you can have a full 176-core VM, and you can also have a like a 24-core VM, right? So you're selecting the different cores. So moving ahead, so this is the -- I mean, let me particularly introduce a different IP fleet at Azure, right? And all of these VMs have fully on VM per house. So it's not like you're sharing in the house between different VMs, right? So you have a full control, right? And the main reason is that for many of the IT workloads or multi workloads you want to scale out, right? And if you want to do that, like why not start from the node itself, right? So that's the main motivation. And I mean active overview of the different VM family, like we have HPC or traditional HPC-based work cores. Those are just by the H series. And then we have the M Series, which is for the AI or training/utilization skills, right? So those are the NCBs. And here, I mean in this program, particularly focusing on the HBv4 or HXv1 SKU, which is the NDR based SKU. And on the right, I mean, have quickly noted on a few of the IV features, right? I mean, I have mentioned, I mean, a few of them in the previous section. But I mean, I just want to reiterate that since I mean the VMs actually exposed through SR-IOV, I mean there is no feature. I mean we are able to like expose all the features to the VM also, right? And let me quickly jump in through some of the performance numbers in the interest of time. So basically, this is particularly the NDR performance results. These are at the MPI level. So you can see right for -- I mean, for the MPI core latencies, we are able to keep -- I mean, in the 1.5, less than 2 microsecond range for different transports. And if you look at the bandwidth single node. So directional bandwidth or bidirectional bandwidth, right? We are able to hit close to the peak theoretical bandwidth, right? And like these numbers are pretty much comparable to the results also. And this is another view of the performance, right? So in this case, so we -- typical NDR cluster like we did like an all fair performance, right? So just to simulate and there will be multiple users and all of them will be like doing operations, right? So I mean, in this case, all -- I mean, like, say, for example, 2 different nodes and say that will be like the N square different pairs for N minus 1 by 2. So about 10k different pairs. And then we -- like N by 2 different pairs communicating at the same time, right? So you can see there are 2 different bands for the latency. So I mean, these 2 different bands are because of -- so we have a topology of a 2-layer factory, right? So basically, the [ night ] systems can -- between any 2 nodes can be like 2 nodes or 4 nodes. So this matches exactly with that, right? So the lower volumes of 2 cores and then the other one is 4 cores, right? So the maximum latency that -- I mean, any -- between any 2 nodes is in this range and then this is the best case, right? And then on the right-hand side, we have the MPI bandwidth or bi-directional bandwidth distribution, right? So here also, you see the [ topology ], right? So these are strictly for bandwidth intensive workloads right? So with respect to you, like where -- how you picked the nodes, right, you were able to hit the peak bandwidth, right? And we quickly move on to some of the chart performance. So this is again NDR performance on a public cluster. So like Gilad mentioned about the SHARP V3 and which has the 2 different protocols, the LLT and the SAT protocols, right? So you can see for the different scale, 16 nodes or 32 nodes or all the way to like 128 nodes. Yes, you can see clear benefit with the SHARP report in fidelity on small messages and SAT on the large messages, right? And then as you scale, like you can see that things are actually getting much, much bigger, right? So really nice results with the SHARP NDR. And then -- yes, so this is another, I mean, that's by price. I mean there can be perceptions that the cloud, like the performance can be variable at different times, right? So in this situation like this is the 128 node [indiscernible] in 2 different input considerations, one with a relatively small size, like 8k [indiscernible] node. And then here is like a 20x input model node. So yes, in both the cases, I mean like these are different runs on the same cluster. So we are seeing important -- we are seeing a steady performance like plus probably a 2 percentage. Yes. So -- and yes, so with that, let me conclude, it's time also. So I mean I just want to follow up a few points. We are actually taking a generic HPC approach rather than having and building different types of SKUs, right? And it doesn't really matter like how many different bond predictions you have, like we want to really focus on the HPC workloads, right? And then we build SKUs targeting the different HPC workloads, right? So yes, for example, the different types of SKUs, I mean, based on the AI, based on HPC. So the different [ actions ] there, right? And I mean we also focus a lot on the time to market initiatives. For example, like even in the case of the HBv4 or Genoa-based SKU. So when AMD announced the Genoa-based, like we had a cluster ready that can take customer workflows, right? So like the other previous point on like time to solution, right? So that's the motivation we have with those kinds of initiatives. And again, like partnering with some of the customers, for example, OpenAI, right? So like [indiscernible] or like their training models, right? So they have like key customers on that. Yes. So yes, that's all. And we had a lot more different scaling results, et cetera, like showing really nice scaling patterns. In the interest of time, I'm skipping those. But with that, let me conclude. And if there are questions, maybe we can take it now or maybe towards the end.
Mark Nossokoff
attendeeThank you, Jithin. I appreciate that. Again, I invite attendees to post and submit questions in the Q&A window. I'll start with one. Jithin, how does Azure InfiniBand performance compare to bare metal systems?
Jithin Jose
attendeeSo it is very close. Say for example on the bandwidth, there is no difference at all, right? So we are able to get the peak performance. So on the latency, I mean -- so we had recently on India, we had a comparison with bare metal on our VM. So I think like there was like about 20 -- I mean, 20 nanoseconds are different. But I mean, we are able to match -- we are close with SR-IOV, right? There is not much gap between bare metal and VM, right? So yes, I mean, and these kind of differences, right, I mean you won't be able to even see at the application level, right? And I also want to follow that, I would say, if you compare it with another cloud provider, right? So these latencies are in range of like 10x higher, et cetera, right? So, yes. Very close to it.
Mark Nossokoff
attendeeGreat. Thank you, Jithin. Okay. Let's move on to our next presenter, DK Panda from Ohio State University to share his experiences with InfiniBand and BlueField from the MVAPICH perspective. DK?
Dhabaleswar Panda
attendeeMark, thanks for the introduction. Again, all of you, it is very pleasure to bring some of the latest things what we are doing with the DPU and InfiniBand. I think as most of you know, I mean this MVAPICH project has been there for many years. We started almost in the year 2000 when InfiniBand just came into the market. But over that, over the last 22 years, we are continuing it. I'm very proud of my team. And now we have support for all the different networks. And currently, this stack has been used by more than 3,300 organizations in 90 countries. And just from our website, we crossed like 1.65 million downloads. So it is not just the download. We are actually working with a lot of people. You must have heard of this NASA's DART mission. In fact, the nuclear fission research. All these things, we have been very happy to work with our collaborators at the Lawrence Livermore National Lab to really push the frontier of science and engineering to the next level. Recently, what we have done, this is our just release if you have been following, we are on a new 3.0 release series. And this is where like we have actually brought also OFI and UCX support. So with this, actually, we have it on our page and then MIP plus, which is basically we'll be trying to replace our previous GDR and MVAPICH2-X. And through this, actually, we support all the interconnects. Now InfiniBand, OmniPath, Rocky, Slingshot, any of these networks we are able to support. But today, what I'll be talking briefly about 3 things, especially with the BlueField DPU. I'll start with offloading non-blockchain collectives. I think Gilad indicated some of the results. I'll try to share some more. And also not just communication offload, I'll also try to show how you can offload some computation also, okay? And especially there, I'll try to show the ideal training. What kind of numbers are we getting. And within the BlueField architecture, especially with the software approach, I'll try to bring this new GVMI interface we have just started working on, and I'll show you some preliminary numbers with that design. So this is the kind of a very high-level architecture of the BlueField DPU, some of you might have seen it. So I call it like in a very layman's form, executive and secretary. So you can think of, let's say, your host processors are executives and these arm cores are secretaries or assistance. So we have to now rethink in our programming paradigm how we distribute the task, who does the task? And can we offload communication? And can we also offload computation? And can we also do both? So these are the kind of the directions we are exploring. So we have the basic MVAPICH2 library. We also have an MVAPICH2-DPU library. It is being jointly done with another commercial spin-off called the X-ScaleSolutions. So these where I'll try to show some numbers there, we have offloaded like a nonblocking MPI all-to-all, allgather, Bcast. Gilad indicated about some of these non-blocking things and non-blocking operations actually allow the competitions or the, let's say, the collective you introduced the Ialltoall. While the Ialltoall is being progressed by the arm cores like the systems are progressing, the CPUs or the executives are free so that they can actually perform with the competition. So that brings actually overlap of competition and communication, and that's what you see here. This is like on 32 nodes, 512 process here and 1,000 process here. The basic MVAPICH2, is like the red bar here, you can see like you get very little overlap, like eager, runover. There are different protocols because you are trying to do pooling, you are sending TCTS, you don't get the overlap, whereas MVAPICH2-DPU, suddenly the overlap totally goes up because the arm cores are like your assistance, and they are taking care of all these overlap, okay? So the main processors are totally free, and that leads to very good performance benefits. And Gilad showed some Alltoallv, that is the vector variant. Here, I'm trying to show Ialltoall, which is the personalized kind of things. So now you can see what we are trying to show is the compute plus communication, okay? So that you see in the nonblocking, the idea is that you should be able to -- while the communication is happening, competition can proceed, and this is from the issue standard micro benchmarks from my group. You can actually run these experiments yourself also. And as you can see for different messages, we are able to -- when you have the both competition and communication together, we're able to give you like almost 22% benefit here. And that is for the like Ialltoall nonblocking. And we took it to the P3DFFT. The P3DFFT has been modified from the blocking to nonblocking. And it also has Ialltoallv. So here, we are just trying to take advantage of the Ialltoall. And here, you can say for like the different grid sizes, 512 process. And right-hand side is like 1,000 process. We are able to give -- accelerate it by almost 21%. So this comes very close to what Gilad was showing, like you had 23% core benefits. So that is just like think of like a communication we are offloading. So the question is, can we also offload competition, okay? And this is what we have explored on the DL training. So we have a packaged called X-ScaleAI-DPU. So here, you can think of like the standard DNN training. Now think of like in addition to your executives to host cores, I have some DPU cores. So while this training, can I divide the workload between what CPU cores do and what the DPU cores do? And then can we get some boost in the performance? And that's what this package does, and it actually supports the VAPICH and let me show you some numbers. So here, this is the with CFR data set. This is only 32 nodes. You can see we are trying to train ResNet-20v1 model, and we can give to almost like a 17% benefit in the DNN training on 32 notes. And similar kind of things, we also see here, this is the training of the ShuffleNet model on the TinyImageNet, and we are also trying to see 13% improvement. So you can think of like these cards are already there. You're trying to utilize their infinite functionality, but you can also utilize the ARM functionality to really accelerate your training workload. And finally, I'd like to show how do you do this? So there are some differences. So in the -- up to last year, this from the NVIDIA with the hardware, there was only interface or APIs were provided only to do staging transfer. So that means, let's say, I want to move some data from node 1 to node 2, in the basic host, RDMA and all I can move it. Now when the DP was there, I need to do an RDMA read and then I need to do an RDMA write. So I'll call it like a stage design. And that design has been used for the last 2 years. But now there is a new interface from NVIDIA called GVMI. And this GVMI, as you can see, you can basically instruct the DPU, both the source address and the destination address of the RDMA. So you don't need now any more copy. And so that means directly we should be able to, in fact, improve the performance of this offloaded communication. And here, I'm just trying to show initial results. These are very fresh, just even taken last week. As you can see, this is a 4 node 8 PPN. Left-hand side, like this is the pure host. And then once you offload that stage design, if you see like a 512k, 1 megabyte, 2 megabyte improve, but look at this -- the golden lines, that is the new GVMI, it is trying to really improve that communication latency. And in fact, for very large messages, the stage design will not be good, as you can imagine. I mean, we have to do one more copy. That is not good. That's what is being shown here, but the new -- the GVMI design is able to really reap the benefits here. So with this on the right-hand side, you see the overlap. Both the design stage design as well as the GVMI design, they are able to give you the overlap much more higher than what the host can provide. So that means using this new GVMI design, there is much more potential to overlap and also to have better capabilities of offload. And that's what we are working on. These are very early results and might be in the coming weeks, we'll be very happy to share some of the actual application level numbers and all are all different collectives and see how this new GVMI interface can push the benefits of the offload to the next level. So with this, I'll close here. If there are any questions, I'll be very happy to answer.
Mark Nossokoff
attendeeGreat. Thank you, DK. Appreciate it. Again, invite folks to commit questions through the Q&A window. DK, have you found that the DPU capabilities are inherently any better with the specific networking technology? And if so, why?
Dhabaleswar Panda
attendeeFor the timing, I mean, is the DPU only is the NVIDIA DPU. That's what we are working on. Of course, in the market, people are talking about IPU, there are also these option. There are a lot of other -- vendors are coming up with these technologies. But as of yet, we have not got time to actually work on some of those things. But if the principles will remain the same, if you remember my executive assistant paradigm, I think all the vendors are following the same kind of paradigm. And if we can show in one network, we should be able to show it in the other networks also.
Mark Nossokoff
attendeeOkay. Great. Thank you. Okay. Let's move on to our last presenter, Sebastian Ohlmann from the Max Planck Computing and Data Facility, who will share his experiences on overlapping communication and computation with the institute's Octopus code. Sebastian?
Sebastian Ohlmann
attendeeYes. Thank you very much. So going from the more general to the more specific, I'd like to highlight how, in a bit more detail, those DPUs can have an overlapping communication, communication for a specific code. So let me first introduce where I'm working. I'm actually part of the Max Planck Computing and Data Facility. And that's basically across institutional competence center of Max Planck society. So the central computing and data center of Max Planck society, which is Germany's biggest fundamental research organization. We offer a lot of services, but we mainly have 2 big HPC systems, Cobra, which is a 12-petaflop CPU system; and Raven, which has 5-petaflop CPU partition and 15-petaflop GPU partition. And so I, myself, I'm part of the application support. So we collaborate with scientists and try to develop and optimize their codes that take a big share of our computing cycles. And that's why I'd like to show you the Octopus code, which is an electrostructure code. It's a real-space TD-DFT code. So that's the functional theory, which is solving the Schrödinger equation. And it's among the biggest users on our systems. And that's why we have an interest in optimizing it and also to see different ways of optimizing it. And for the derivatives, it uses finite differences, and that's important also for the parallelization and [ dusters ] using stencil, 25 points. And it's parallelized using a classic domain decomposition. So you have the other dimensions as well. It's not that relevant here, but I'd like to highlight how we do this in order to show you really how overlapping competition and communication helps in this case. All right. So what do we do? We actually split the application of the stencil. And that's, I think, just a classic example of how you do this because when you split the domain, you basically need to have the points at the boundary that are present on a different rank or a different processor and to have those goals available in other 2 stencils. And then you can split the application to -- in a part that doesn't touch these ghost points and then to the other part that does touch it. And the exchange of the ghost point, this can be overlapped. And I think that's a classic pattern that's used in other codes as well. And that's why I think the results we will get from this are also something that can be generalized to other codes. It's implemented using a nonblocking MPI function. So Irecv, Isend, Waitall. Could be done also within Ialltoallv. It's basically the same, it's just coded by hand. And so first, as we start the communication pattern, then we apply the stencil to the inner points. We do the Waitall and then we do the rest. And now the point is how can we exploit this overlap. Well, first, I'd like to state that just using nonblocking MPI communication doesn't mean that there's any progress in the background. And this is quite important because that's kind of a trap you can easily fall into. So it means that the communication always needs to be driven. So the progress needs to be driven somehow. And there are several possibilities. And the first would be to do this inside the application, and there are several techniques that are known to work. So you could, for example, call MPI testing inside the loop that actually does apply the stencil, but that's quite tedious to program because you have to mix in the application, different responsibilities, applying the stencil, doing the communication. And so that's not really desirable. It would work, but it's not that desirable. The second possibility would be to do this inside the MPI library using progress. So do this on the CPU. And the advantage is you don't need to change the application, but the disadvantage is you need dedicated CPU resources. And that potentially decreases the application performance. So the problem is that you need to sacrifice some DPU cores that drive this progress, threaten -- drive the communication. But of course, those cores on the CPU cannot do the application work. And we did some tests with the implementation of Intel MPI, and it turns out that it is actually quite difficult to achieve a benefit because in the end, the overlap you achieve needs to be so big that you still gain something, although you sacrifice some of the superior sources. And so that's why the third point here is quite interesting mainly to offload the communication to SmartNic, such as the BlueField DPU. It's very promising because today, you don't need to change the code. So the idea is that all of this is done below the MPI level so that the user doesn't see anything. And that's an external resource. And this has -- also said before already, it's basically hosting its own cores, so it can drive the communication without the CPU needing to sacrifice some resources. And that's what really can overlap the communication and the computation. And that's what makes it also very interesting for us. Currently, we are still exploring this together with NVIDIA, and Gilad showed some numbers from the study. And I think there's still some potential to get even more there because it benefits the other collectors we can use in the court and some issues with [indiscernible]. And so I think in the end, for such specific applications, as I have highlighted here, the DPU can be very promising to really drive the overlap from communication and competition. And with that, I'm at the end. I'd like to thank you. And if there are any questions, I'd love to take them.
Mark Nossokoff
attendeeGreat. Thank you, Sebastian. I think we have time for one question. And I'll share you -- you indicated you don't expect that there's no co-changes in moving to offloading to BlueField. I mean how big of a lift and effort has it been? Or do you expect it to be? And are there other any work considerations besides the specific MPI code changes?
Sebastian Ohlmann
attendeeNo, I think the work from the application side is not really there if you can already take advantage of overlapping computation and communication. So not all applications can do that. But if you have that pattern in the code with nonblocking communication calls and overlapping that with the computation, then you should be fine to benefit from it. It's more that I think it's still in development in the MPI leverage to, at the lower level, really benefit from this. And so that's why I think it's unattractive from duplication point. It's just that I think it needs a bit more time to mature from the implementation of the MPI libraries.
Mark Nossokoff
attendeeOkay. Great. Well, thank you, Sebastian. And with that, we'll move on to the closing remarks. I want to thank on behalf of Hyperion Research and NVIDIA, thank all the attendees, and we'll move on to the closing now.
Unknown Executive
executiveThank you, Mark. Once again, I'd like to thank you all for joining this event. I'd like to thank Hyperion, NVIDIA, Microsoft, OSU and Max Planck for -- and their speakers for presenting. An on-demand version of this webcast will be available approximately 1 hour after this event ends and can be accessed using the same link. Thank you again for joining us, and have a great day.
Read the full transcript via the API
You're viewing the first half of this call. Get the complete NVIDIA Corporation transcript — plus 248,000+ transcripts from 12,000+ companies, speaker segments, AI summaries and full-text search — through the EarningsCalls.dev API.
Get the API View API docs →This call discussed
For developers and AI pipelines
Programmatic access to NVIDIA Corporation earnings transcripts and 248,000+ others is available through the
EarningsCalls.dev REST API. Plans from $24.99/month — full transcripts, speaker segments,
full-text search, and the recently-added /api/v1/transcripts/recent polling endpoint for ETL pipelines.