Intel Corporation (INTC) Earnings Call Transcript & Summary

February 9, 2023

NASDAQ US Information Technology Semiconductors and Semiconductor Equipment special 146 min

Earnings Call Speaker Segments

Unknown Executive

executive
#1

Good morning, good afternoon and good evening, and welcome to the oneAPI hands-on workshop series. Thanks for tuning into today's episode, Accelerated Machine Learning for AI Solutions. I'd like to introduce today's speakers, Praveen Kundurthy and Bob Chesebrough. Bob is a developer [indiscernible] Intel with over 20 years of experience in software development, optimization on Intel platforms and AI. Praveen is a developer of [indiscernible] AI and oneAPI, who has over 15 years of experience in software development, optimization on Intel platforms. My name is Austin Webb, and I will be your host. A note about our platform we're using here. You can access the biographies of our speakers in the top left-hand column. The following section contains downloadable resources, including a copy of today's presentation. Across the bottom of the screen, you see a series of icons. This will give you access to the slides, the Q&A and the survey. You can also resize and move around all of these windows to best fit your view. You can hover over the video box to turn on close captioning. If you have any issues with the presentation sound, slides or other, please note it in the Q&A and a producer will help you out. You can post questions at any time in the chat window, Praveen and Bob will be responding in real time, and they will go over these questions at the very end, time permitting, of course. Due to the pandemic, we are hosting these workshops from the confines of our own homes. So please bear with us should we face any technical issues, dogs barking or children crying. And with that, Bob, I will hand it off to you.

Bob Chesebrough

executive
#2

Great. Thanks a lot, Austin, I really appreciate it. Let's move right into the discussion here. The agenda for today, just so you know whether you're in the right spot. We're going to be going over the -- an overview of oneAPI, including the oneAPI analytics toolkit. We're going to be talking more specifically in my chunk of the presentation about the Intel -- for Scikit-learn to do traditional machine learning things very quickly. We'll be introducing a mechanism called patching to you. We'll also be introducing the data parallel control, which allows us to use those same Scikit-learn modules or libraries on Intel GPUs. And we're going to be explaining the compute follows data approach, how to target the Intel GPU for those algorithms and then I'm going to turn it over to Praveen, and Praveen will be talking about the NumPy, data extensions, and he'll be covering some specific algorithms pairwise and K means showing how to do that using NumPy. And so this will be a hands-on lab. We intend to walk you through some of the Intel Exchange for Scikit Learn and some of the pairwise distance stuff using NumPy as a hands-on thing. The thing I do want to cloud is that I don't know actually how many people are online right now and able to get on the DevCloud if you are already registered for the DevCloud, go ahead and log in. I'm going to be walking you guys through how to register, how to log in, how to launch the Jupiter notebooks, all off that. DevCloud is a sandbox, how to do the GitHub cloning and that you'll need to get the codes. And then we'll start playing But if you're -- if you have a DevCloud account for oneAPI, go ahead and log in. So the learning objectives for today, you will be able to describe the value of the AI analytics kit good for, described the value of one of the subcomponents Intel Exchange for Scikit-learn. You learn where to get this toolkit and you'll understand how to apply this Intel extension for Scikit-learn for those really do see the common suspects in Scikit-learn the PCAs and the SVCs and the random forests and so forth to do them either on Intel Xeons or core processors, but especially on Xeons, or Intel GPUs. And so with that, I'm going to go next -- we're heading into I'm going to describe the overview of the toolkits and just the what's oneAPI after all. And then we'll start going into more of the hands-on thing to kind of get you set up, and then we're going to go into more theories. So we hopefully, we can hide some of the latency of you guys logging in and doing the get clones and stuff like that while I'm doing the rest of the lecture. But to set the context, again, just to make sure you guys know that you're in the right spot what we're talking about today is oneAPI. Now oneAPI is a one programming model for multiple architectures and vendors. It is a standards approach. It's an industry initiative. It's not just an Intel thing. It's an industry initiative based on open standards and specifications. And it really includes a unified library language and libraries that over full native code performance. And since it's based on these standards, for example, sickle by the Cronos Group underneath the hood, all these libraries are essentially built with sickle. And so even though we're doing the Python versions of these things, the wrappers in Python and the libraries that are at a higher level than that, that it's inheriting all that sickle goodness under the hood. And what it allows you to do is to deliver a native code performance across a wide range of hardware, CPUs, GPUs, FPGAs, AI accelerators and so forth. So this is kind of the idea. What Intel has then done is that we've created a reference implementation of that oneAPI initiative. And our products are currently cast as toolkits. And so we'll see this as we go forward. So the one toolkit we're going to be talking about today is the AI analytics toolkit. However, there are other toolkits that we have available. We have Internet of Things, IoT toolkit, high-performance computing toolkit, rendering and so forth. And so we've got the open VNO. But today's talk, we're going to be focusing on the AI analytics toolkit. So with that kind of overview, we're going to keep kind of drilling in so you'll understand a little bit more of the context. So Intel's oneAPI ecosystem is built on this rich heritage of CPU tools, but now extended, expanded to cover XPUs as well. And so -- the idea here is that for a long time, we've had tools covered mouth Cronos libraries and data analytics, libraries and all those kinds of things, and they were primarily aimed at Xeons and so forth. But that's been extended to cover Intel GPUs and other accelerators. So this is the idea. It's -- we're leveraging our taking advantage, we're utilizing the underlying sickle approach, the open industry initiative approach and then building these reference kits and libraries on top of that. So the Intel oneAPI analytics toolkit specifically, we're going to be focusing on my part of the talk, anyway on the Intel Exchange for Scikit-learn, but I want to give you an overview of what's in the toolkit. So think of the AI analytics toolkit as a compendium of smaller libraries. And so -- and they're grouped or arranged it for -- at least for purposes of conceptual thinking in these different domains. So on the deep learning domain, we have the Intel optimization for TensorFlow, the Intel optimization for PyTorch, Intel neurocompressor, and I won't talk that much about Model Zoo, but these other things -- these Intel extension for TensorFlow and PyTorch, for example, are -- you get extra speed up for both training and particularly for in France with -- by utilizing this toolkit. Now let me just say that we work very hard to up stream all of these optimizations that we do into the stock versions of these tools. So TensorFlow has a lot of our optimizations already. PyTorch has a lot of our optimizations already. So when you run those on Xeon, you get good performance. But the cadence which those communities adopt our changes and our optimizations is typically slower than what Intel creates. And so the cadence at which we're creating these things allows us to have extra value add by installing the AI analytics toolkit just by virtue of the fact that we've got all the recent changes there. So if you want the latest and greatest fastest you use analytics tool kit, but rest assured that those optimizations do get pushed into the stock versions of those tools. The Intel mirror compressor is a really useful tool for taking deep learning models and to optimize them so that it does the pruning and potentially quantization and different optimizations to influence very quickly on -- particularly on Xeon, on CPUs, but it's just good for inference. And so the other thing that we optimize here in the machine learning domain is XGBoost, which is a very popular both classifier and regression algorithm. And so we have optimized version of that, and we also work with that community to give our optimizations there into their stock versions. But then the other tool that I'm going to talk about today is the Intel extension for Scikit-learn, and so I've circled that one here. So we do have accelerations for data analytics in terms of Intel distribution for Modin, which is a very similar tool to Pandas, but you can use a distributed computing and really getting some amazing speed ups using Modin over Pandas. And so that's -- we won't cover that today, but that is another toolkit. Another part of the toolkit are the optimizations that we provide essentially through the math library, the MKL for NumPy, SciPy, Pandas and so forth. So you can get really wonderful speed ups by using those universal functions or aggregations or reductions and so forth, function calls to replace your Python loops that do the similar stuff. And then NumPy, Praveen is going to be taking us through the data Python and NumPy, -- so we'll be covering that as well. So now just motivation to achieve acceleration on Intel current and future GPUs and CPUs. So you're ready for future innovations from Intel, and this can be achieved typically with just a few lines of code. And so this is our segue into kind of getting the lab set up. This is my placeholder for the Intel DevCloud. So what I want everyone to do is go to the devcloud.intel.com, oneAPI gets started. And so what this allows us to do on the DevCloud is that you can learn data, C++ if you wanted to. You can use these oneAPI toolkits that's what we're going to be doing today. You can evaluate workloads, you can prototype your project, you can build cross architecture applications, test them out or with different hardware. Oh, gosh, let's try the Intel GPU over here, let's try an FPGA over there and so forth. So the DevCloud would be a good sandbox for you. So I'm going to show you -- I'm going to switch to my screen share here.

Praveen Kundurthy

executive
#3

Yes, sir. let's a little bit take time and some questions here, okay? So first one question is -- where's my chat. Yes. Sorry, if I mainly use TensorFlow with OpenVINO, will oneAPI help me further? Or does oneAPI already use oneAPI's optimization.?

Bob Chesebrough

executive
#4

Well, the -- you're going to get a lot of good news just from doing the OpenVINO for sure. I think there may be some ancillary benefit that you would get from using the oneAPI toolkit as well. Because underneath the hood, since we have these optimizations that are built in to the TensorFlow, it would be good to be using probably the Intel version of TensorFlow through the toolkit, the Intel extension for TensorFlow, I'm sorry. So that would be my best answer to you. So -- now that said, in order to do that, we don't have any kernels on or instances on the DevCloud right now that cover both tools simultaneously. So if you've done OpenVINO in the past and be careful of this when you're doing the lab, you probably have an instance of DevCloud that is where you log in and you get all the OpenVINO stuff ready to go. The one I'm going to be walking you through here is the Intel DevCloud with oneAPI where it's going to be focused on the oneAPI analytics toolkit. You can't create a custom kernel that does both. But for those people that are wondering about this OpenVINO thing. The one thing I just want to warn you is that please follow this oneAPI version of the DevCloud for right now for this particular lab. You can do what you've asked with both things, but you probably have to curate your own custom kernel that has both projects in it. And that's outside the scope of today's lab.

Praveen Kundurthy

executive
#5

There's one more question on the bare metal installation. So you can go ahead and -- so DevCloud gives you all the toolkits and everything in place, so you don't need to install. But -- so there's a question from Miguel saying that -- I would like to get some guidance to those under Lennox who went VS code. Yes, you can use a VS Code on an open tool, right, Lennox open tool, but all the workshop material is all using Jupiter labs. So you'll not be able to run all the Jupiter labs on your VS Code, but there are some instructions online how you can actually integrate DevCloud with the VS Code, maybe you can try that out.

Bob Chesebrough

executive
#6

Yes. I'm not really going to be covering that in this lab context using the VS Code. We're just going to go directly to a Jupiter notebook natively, kind of on the DevCloud. So yes, -- but it is possible to do, and we have colleagues that do that. So follow up in the user forum and ask questions, you'll probably get a quick answer for how to do that. So -- I'm going to jump in and share my screen and have -- we're going to play follow the leader now, and we're going to go to the Intel DevCloud. So Praveen, can you see my screen?

Praveen Kundurthy

executive
#7

Yes. I can see your screen.

Bob Chesebrough

executive
#8

Awesome. So what we're doing here is we're going to go to this devcloud.intel.com, oneAPI gets started. And so for you, this is where I would recommend. And if you haven't already go ahead and click and roll over here. Now I'm not going to actually enroll because I'm already there. But if you click and roll, you just fill in your first name, last name, you do all these things and you submit it, takes a minute or 2 to register a new DevCloud user for you for specifically the oneAPI. So if you've already done this with OpenVINO, but you haven't done it with oneAPI, you need to do the oneAPI version, okay? So you may have a different account for the oneAPI. But just go ahead and do that, don't use the OpenVINO one because that would be too complicated for today. So now what you're going to do if you've already enrolled, you would click sign in. And in that case, what you do is you just fill in your e-mail and password and it will log you in and it will take you back to this page. Now if you scroll down to the bottom of this page, all the way to the bottom, there's an icon that looks like this Jupiter notebook thing over here. Now if you signed in, instead of saying sign in right here, this blue box, this blue box will say launch Jupiter lab or something like that. And so if you signed in, -- so you have to register first, then you sign in, when you sign in, if you click this, it's going to take you to something that looks something like this. And let me just show you what this would look like. It will typically look something like this. So there'll be a navigation bar over here on the left, and then it will have this launcher right here. And this launcher you may not have as many of these are what custom kernels look like right here, these little blue and yellow little plus signs. But the standard that you will get, if you're brand new to this, is that you will have a Python 3 kernel and we'll be playing in those kind of things a lot today. You'll have a PyTorch kernel and you have a TensorFlow kernel. You can also launch consoles, you can launch a terminal and we'll be doing this in these labs to day too, we'll be launching a terminal, where we can just start interacting right away with the system, okay, and do a PWD and so forth. And we're going to be doing that today as well. Now if you lose your launcher, you can always get back by just hitting this blue plus sign right here, and that will give you a new launcher. And it's kind of like your home base so that you can see -- oh, yes, let me get to this. Now you can turn this navigation pain. Let me just share this a little bit bigger here. This navigation bar is controlled by this folder over here. So if I click it or just toggle it back and forth, you'll see that my folder structure appears or disappears. And so I may be doing that today to give us more real state to look at. But when I'm doing DevCloud work, I like to see my folder structure. It's just like a security blanket for me. I'd like to know where I'm at. So you'll typically see me do that. So the first thing that we're going to do today is that I want you to launch a terminal, so we're going to go here. Yes.

Praveen Kundurthy

executive
#9

Sorry to interrupt. Can you just show how to launch the Jupiter lab again, there are some questions on that.

Bob Chesebrough

executive
#10

Yes, yes. So if you come back, so you're going to be at the devcloudintel.com, oneAPI gets started. And if you've signed in, so you click to sign in and you fill in your e-mail and password and all that stuff and you sign in. Then when you -- it will take you back to the same page. But this page when you scroll all the way to the bottom, it's going to look different, slightly different to this. It will still have this icon right down here on the lower left. But instead of saying sign in to connect, it's going to say launch Jupiter lab. And so you'll just click that button when it says launch Jupiter lab, and that takes you instantly to kind of this, right? So that's how you get there. So I want to leave with you and Praveen, if you can just share this link for the hub repository that I'm using, I'm going to be showing people how to begin the process of doing the git clone. So yes. So this is the idea. We're going to do a git clone. Now I'm going to show you where this lives. If we go to GitHub, you go to Intel software, machine learning using oneAPI. If you go to this...

Praveen Kundurthy

executive
#11

I put in the link.

Bob Chesebrough

executive
#12

Great. So Praveen's got the link. So you'll see this page. And if you just scroll down the read me, you'll see the instructions right here of how to prepare to run on the Intel DevCloud. So you'll just follow these instructions. We're going to make a directory called ML1 API. I've already created a directory called ML1 API. So we'll just bear with me, CD to ML One API like this. And you won't do this because you're doing this brand new. But I'm going to remove the machine learning directory here. I'm just going to start from scratch just to show you guys how to do this. So I'm cleaning out some work that I had done before. So here we go. So it's as if I've just done this one line, I made a directory. I ceded into the directory, the director is called ML oneAPI. Now what we do is source a version of the oneAPI analytics toolkit. And so what we'll do here is we're going to source if that toolkit, we have several different versions or copies of it from different chunks in time. We're just going to launch this one from 222.3. So this is one of the latest ones from 2022. We're just going to source this. And what happens is it's behind the scenes setting up all the pointers and all the -- setting up your local system to understand in this log-in node of this cluster on DevCloud that we want to use this particular version of the DevCloud. So the next thing we're going to do is we're going to do is we Condo activate base, okay? So this is all kind of standard stuff if you're used to using Anaconda whether it's Condo activate base. And what you'll see is that your prompt changes. So that now you have IntelPython 3.9, for example, right here. And that means we're using all these latest things. We're using the VPL. All these things are available to us, the MKL, Modins, neuro compressors, all these things have been sourced for us, right? And so -- now what we want to do, I want to make sure we're going to -- this is going to make sure that you have a kernel available to you that knows about this instance of the oneAPI toolkit. So I just want this to appear in your Jupiter notebook as something you can download and click on. And so we're just going to do this setup of the kernel in order to do this little command right here. And so this little command right here is going to say, "Hey, we just source the 2022.3 version of the oneAPI toolkit. Let's go ahead and make sure that, that's going to be available in my drop down box.

Praveen Kundurthy

executive
#13

Bob, can you show the new terminal again, the launcher and sourcing that's...

Bob Chesebrough

executive
#14

Yes. Here's what the launcher looks like. So you're right here. And then if you want to launch a terminal, you come down, just scroll down and there's this terminal button right here.

Praveen Kundurthy

executive
#15

So can you show file new launcher?

Bob Chesebrough

executive
#16

Yes. Like this. Yes, you can also do it this way, instead of hitting a blue plus sign to do a new launcher, you could also say, file new and do a kernel. For example, that's another way to get there.

Praveen Kundurthy

executive
#17

And there's a question asking about sourcing. They want to see the sourcing command again. How you source 22.3, right?

Bob Chesebrough

executive
#18

Go to the GitHub Intel software machine learning using oneAPI, just keep this in a separate tab like I'm doing. So you can go back and forth. But the source command is right here. You're going to source from glob development tools, versions oneAPI 2022.3.1 into oneAPI and then the set bars. And so once you do this, -- that's the thing that we're going to be using. We're setting that as our base engine. And so this then is going to be saying, well, make sure that, that base engine is accessible by a kernel selector inside the Jupiter notebook. So we'll just do that -- and so I've done those steps, myself just now. And now I'm going to do the git clone. This very repository. That's my next step. So after I've done all these other things, I mean ML One API, I'm going to do this git clone, and this will take several minutes. And so that's why I kind of wanted to get you guys started on doing this now. And then we'll just set into the machine learning. We'll do a PIP installs from requirements and so forth. But -- but essentially, we'll be ready to go. So I just wanted to get you guys started on this journey, while we go through the theory, okay? So here we go. So I'm going to stop sharing my screen on that part. And you can see my slides Praveen now?

Praveen Kundurthy

executive
#19

Yes, Bob, you're back to slides.

Bob Chesebrough

executive
#20

Great. So here again is where that GitHub link is. So it's at GitHub Intel software machine learning oneAPI, Praveen, put it in the notes there. So you can a look at it. So this is where you're going to go to see how to set up your lab. Now that getting out of the way because you guys are all working on that. So what we're going to do here, the procedure for ignition for this acceleration. We're going to be talking about it both from the terms of the Intel CPU and also for the Intel GPU. So the basic idea for Scikit-learn is that we have to import what we call a patch, and I'll be teaching you guys how we do that and it's just one line or two lines of code to do this. It's easy. Then we're going to -- in order to use the GPU, we have to have both things. We have to do the patching. So the patching that we've got to tell it, we're using Intel exchange for Scikit-learn, we have to patch. But then we are going to do what is this idea called compute follows data. And I'm going to be showing you guys how to do that. But we have to do both things. We have to do both patching and applying the data apparel control library to do the compute follows data. So all of that will be done as we're going through these workbooks here today. We're not going to get through all the workbooks that you're get cloning. We're going to just cherrypick some that would be the most salient for you guys to understand how to move forward. And then if you ever want to come back and do some extra playing, all that code is there, but we'll just be doing a handful of notebooks today. So okay. So what kind of algorithms are accelerated here with Intel exchange for Scikit-learn. Scikit-learn has a lot of traditional machine learning algorithms built into it. And we don't accelerate all of them. We accelerate a subset. But they're the kind of like the common suspects, right, the usual suspects that you would think of the things that you tend to go to the watering hole time and again. The one that's missing here is decision tree. We haven't accelerated that one. But we do some other ones. And I'm going to tell you about some other ones and why that might not be actually a bad idea and why it's actually probably smarter than we're doing what we're doing. But here's the list. So we have DBScan. This is for clustering, okay, very useful for looking at data points that swirl around in a couple of different manifolds or whatever hard to separate out using something like K-mean, DBScan can be very, very, very handy. We accelerate that one. We accelerate K-mean. It's more of the go-to standard clustering algorithm that people tend to think of. And so we've got a very fast versions of these. You're going to see how quick these things are today. And I'll show you a link that shows just a benchmark against the stock versions of -- between using the Scikit-learn version versus just using the stock version, how fast algorithms are for either training or fitting or for predicting. We accelerate Nearest Neighbor. Now Nearest Neighbor is actually -- this instance of it is an unsupervised learning one. So it's kind of a -- have predict function. It's just a fit. And it -- you can find nearest neighbors very quickly using this tool. Principal component analysis. This one, I use just all the time. I beat that one up a lot. It's very, very, very handy. You'll see that in some of the live exercises today. [ Disney ] is a -- so well, Principal component is used to reduce the dimensionality of your data to find the most combinations of columns that really let you just select a tiny subset of columns to actually work with. Disney would be similar, but is based on nonlinear kernels and -- it's not as what do I want to say, deterministic. It will give you a different answer when you run different instances of it. But -- these are the ideas. If you're familiar with we have an optimized version. -- we're familiar with PCA, we have optimized linear regression, linear aggression, elastic net regression which basically encompasses both L1 and L2 regulators. And then if you just want [indiscernible] or whatever, we've got accelerated versions of those, plus other ones that I'm not even talking about. These are just like the most high-level hitting that people have heard about algorithms. Logistic regression is another now, odd-late logistic regression is really more of a classification algorithm. Oddly, not really regression. It uses logistic regression under the hood as a classifier. And so K Nearest Neighbor, we've got both a classifier and a regressor, so you can -- the difference is classifiers, you're going to have discrete data and you're going to be predicting a discrete class. Well, it could be even continuous data. K nears stay on the regression side, you're predicting a continuous outcome, not discrete. So same thing with the random forest classifier and regressor, support vector classifier and regressor. And then we have these other assert all finite. So sometimes you're dealing with data sets and you want to do these low-level things that say, we don't have any right, so you can do these things sort of all finite and all these kind of things. And they're much more efficient, it's much faster when you're dealing with large chunks of data.

Praveen Kundurthy

executive
#21

Do you have any transformers in one of the questions?

Bob Chesebrough

executive
#22

Not in this. So this is not really transformer based. This is -- but none of Scikit-learn really is. That's more of a deep learning concept. And I'm not really uncovering that here, but certainly, since most of those transformers are built on top of something like all the goodness of the oneAPI is applicable to that, just not in our talk today. And then the ROC AUC score, that's the receiver operator characteristics -- curve score -- that's a very commonly used one. We accelerate it. We accelerate Train Test split. I'm sure many of you used Train Test split. We accelerate pairwise distance. Now we have some caveats on it. We're going to do a lab here today on this. But maybe you've not pairwise distance. So I'm going to show you, I think, kind of a compelling example of how you might be able to use this for comparing time series type data kind of like stock data and compare 500 pair of stocks and see which ones are most similar. So we'll be going through all of this today. So these are kind of the usual suspects here. And so it's all predicated on this concept called patching. And so what is patching? Well, patching is a way of setting up, think of it as a listener, okay? You say, okay, import, the ability to pass to import from Scikit-learn X, import patch-sklearn, and then when you call that function, patch-sklearn with 2 parentheses, you call the function. That's kind of -- think of it this way, it's like listening. And what it's listening for in the rest of your code is the thread of -- execution goes through your code. It says, okay, if you ever run across an import statement to Scikit-learn, I'm listening. And if we have an optimized version of it, we'll use the optimized version of it instead of the stock version of it, okay? And that's all patching is. It's kind of like that listener and it's just saying, okay, I'm waiting. Oh, they're starting to use Scikit-learn, "Oh, I've got a fast version of PCA. I'll use that. I got a fast version of SVC, I'll use that. And so it's just a way with 2 lines of code of getting the best optimizations, okay, from Scikit-learn extensions. So you don't have to make any changes to your code. It uses the same underlying API. It's like drop-in replacements for it. You just don't -- you literally just put that import statement and patch with the parenthesis, you do that. And now you're -- all your calls to Scikit-learn functions. If they've been optimized, it'll use optimized version. If it's not up optimized, it will use the stock version. And so then you can also control it with unpatching and so forth. So we're going to be going into the details of how this works. But here's on the command line. Now if I wanted to patch an entire Python application, I could just run the command line if I've installed the Scikit-learn extensions and learn either because I've installed the AI learn X tool kit or I could explicitly only just go out and PIP install the Scikit-learn EX library itself. But if that's on my system, what I could do is I could say, Python-M, SK Learning X, my application. Think of it as a verb, and you're just saying what this means is I'm applying a patch to the entire application. Now if you're in a Jupiter notebook, which is what we're going to be playing with today, what you can do is just say, from SK Learn Ex import patch SK learn, and then you do that function call. That sets up that listener and then it's auto magic from there, okay? So you can also unpatch. So if for some reason, you find some pathological algorithm that with respect to your small data set size, typically, it happens with small data sets. You're not getting the performance benefit that you want from patching. Well, then you can unpatch and you can unpatch at high-level ways or in really fine-tuned ways, and I'll be kind of covering more here on this slide here. So you can just do the patching as I've described before, or you can explicitly say just patch SVC or just patch SVC and PCA or random forest or rock AUC score, whatever, Train Test split. So you can explicitly turn on and off lights the way you want them to, right? So you can explicitly patch things and you can explicitly unpatch things or you can just unpatch everything, you can patch everything. I mean like this gives you so much control. So really, the advantage of this is you shouldn't really lose any performance by using this. Because if you use the patching and you find in some particular kernel, you put everything up by 10x or 100x. But on this one thing, it slowed me down by 50% or something like that. If you somehow find some pathological then you can unpatch it, just that one thing, and then you get your performance back. So you should only be able to get the same performance or better almost as a guarantee by using this, right, if you're smart about it. So -- so this side, you can get the list of the optimized functions that we have available. You can just get the patch names itself. And then this is what it looks like in code. So you -- the one caveat here to be using it, and I just want to make sure of this. Remember, this patch SK learn thing. It's a listener. So you need to turn the listener on before you do your imports from Scikit-learn. And as long as you turn the lister on before the imports of Scikit-learn, not the invocation, like you need to import first, then it knows when you start using that function that you're going to use the optimized version. Here's an example for Train Test split. Now you noticed that I did not have to change my train test split function call in the slightest, whether I use the stock version or whether I use the version from Intel exchange for Scikit-learn. As long as I do this listener up here, it will get the right version, and I call it identically. So it's literally hardly any impact to your codes at all other than importing these 2 lines of code, okay? So that's kind of the advantage of it. So this is what that looks like. And then this is just an example that you're going to see today where we're using support vector classifier on a CPU. And just some code snippets shows how you use it in context. So this is kind of what you're seeing in the wild, like give us how you'd actually see it in the code. And then this is some of the performance overstock version that we got from using it patch versus unpatched. And so I want to talk about a different topic now. So it's very important that you understand the concept of patching because patching, you're going to use that, whether you're using the CPU or the GPU what have you to get that best performance out of whatever random forest or what have you. But to get access to the GPU, we're going to be talking about the data parallel control library, okay, DB control library. Now this is just like a high-level screenshot view of it. There's sort of like 4 domains here that we take advantage of it that we govern by use of the DB control. We have the DB control Tensor. We can -- we have algorithms to control the memory to control program. And then we have library -- I can't read it hardly anymore. I'm trying to remember, but it's the library sickle interface is what it is. And so DB control, what it does is it provides a lightweight Python wrapper over the subset of the Cronos sickle APIs. So if you've taken any of the courses that Praveen has done on this topic with [indiscernible] specifically in the past or any of the sickle trainings for C++ and data parallel C++ sickle, all those kinds of things, then this is like the Python wrapper around that, okay? And so it describes basically how you can select the device, how do you select a GPU. And more than that, it boils down to how do we bring sickle to Python explicitly in these other things, Scikit-learn and TensorFlow and these other things, some developer somewhere created a library and they promised you, yes, we were using Scikit-learn under the hood, and you can prove that it is done that way. But in this case, you as a developer now have the opportunity to take your own custom code and bring it and make sure that you're using sickle even for custom stuff. And so the DB control library is intended to provide that common run time to manage specific sickle resources such as devices, unified shared memory, sickle based python packages and extension modules. And we've got these 4 things I just described that we cover. But the main features presently provided by DB control are the Python wrapper classes for the main sickle run time classes mentioned in the specs, the unified shared memory manager to create Python objects that use sickle unified shared memory for data allocation. The DB control is available as part of the One API Intel distribution of Python IDP. Once One API is installed, DB Control is ready to be used by setting up IDP that is available inside the One API. So it's just -- in my -- if I were to just give one quick summary of it, which is really a gross oversimplification. I think of DB control as that library that allows me, I interface through it to the DB control Tensor very frequently. And so this is my way of telling Python how I'm going to interact with a GPU. And so it's my interface to the GPU for the intel GPUs from Python and therefore, from the Scikit-learn extension stuff that we're doing. So it's a way of targeting the GPU, okay? So the intent here in this workshop too, by the way, is I'm going to show you how to target the Intel GPU. But given the nature of how many people we have on and just various things that are in a large setting like this, I'm not trying to say that you're going to get super-duper performance for [ retarding ] this. You may, you may not. That's not my goal. My goal today is to show you how to do it, okay? And then -- and particularly as we get more and faster performance GPUs coming online as they're starting to grow and turn on, then you'll know how to do it and then you can look at performance later. But right now, that's what my intention is just to show you how, okay? So to show you how what we're going to do is I'm going to show you how to use compute follows data. So we're using a support library. We're going to cast our data to a device tensor, and then we're going to send that to the device. So I'll show you how that looks like in code. But any enabled, that means a patching, any enabled Scikit-learn methods involving that tensor will be computed on that device, typically during that call to your fit function or maybe your fit predictor from you predict. And then the results are returned to the host. And this is called compute follows data. Let's -- we'll take a look at some examples here in just a little bit. But I did want to just say as a caveat that we don't have as many algorithms optimized for the Intel GPU as we did for the CPU. So here are the pared-down suspects on the GPU. We accelerate DBScan, K-Means and PCA. Those are like -- you need that for breakfast. That's your common everyday breakfast. We just have that as a default standard. Linear regression, K Nearest Neighbor, random forest regressor, logistic regression, random forest classifier and K Nearest Neighbor classifier and support vector classifier. So we have a few regressions and we have quite a few classifications that are optimized as well as PCA, okay? So a little bit of how to use the thing and then we're going to start seeing the nitty gritty here, okay? So this is just kind of setting things up. So here's a way that you could use the DB control library, just print information about the device that you're targeting. So you can see that we have a UHD graphics card, okay, or whatever, we're running out of [indiscernible]. So it will tell you all the devices on that system. So here's the one way to do that. You can just iterate through all the devices and then print the info of them. If you wanted to use this for turning Scikit-learn on for specific thing, then the first thing you need to remember is you do have to import the library DB control. You do have to get the patching ready. So you have to do it from Scikit-learn X, import the patches Scikit-learn, or unpatch, but mainly patching is what we're going to be concerned about. And then we could -- this is just adding that piece and right now, we're just printing the info about the devices. Here is one means you can use to just iterate through all your devices on a given compute node. So it would go through and say, well, how many processors do you have, what kind of accelerators do you have? You have Intel GPU here. You've got -- it could be any number of accelerators that are on a given compute device node, a computer. And so we can iterate through these, and we can just select if we want, specifically the GPU device, and that's with this line of code right here. We could just select the GPU if we wanted to. And then I remember what it's called, I just say GPU device. And then I just set myself up a little flag that says, "Hey, we did have a GPU in the system, set of variable that says that, that's true. So this is one way to do it. There's actually some ways in code where with really one line of code, we can get rid of this block and just get me the GPU, and we can just specify and do that, get the default device basically, okay? But this is, if you want to be more explicit and have more control. This is a way to do that. And then what we would do then we first have to do the patching. So we would say apply the patch and we're going to -- since we set the listener up, and now we're going to be applying this to DBScan. We import DBScan, it's looking for a way of using the accelerated version of DBScan. Then what we do in order to actually use the -- this on the GPU, what we do is that we take our data, our ex data, okay? And we cast it as a DB control tensor. And so when we cast it as a tenser, one of the things that we specify is the type of device that it is. It could be GPU, it could be GPU 0, what have you. And then this is the unified shared memory type that it's just boiler played out, just used device as a default right now. And so what you're doing is you're binding the data with where it lives. So right now, the data will live on the GPU. And we have a variable that's a tensor called X device now that any computations that I do on X device will be done. The compute follows the data. It will be done on that device. If the device is tied to the GPU, the computations on X device will be done on the GPU. If X device was bound to a CPU, then the computation would be done on the CPU. It follows wherever that tensor is located, right? And so in this case, the tensor is located on a GPU. And so when we use the fit function and we pass the X device to that function, then that computation is then done on the GPU. Now hypothetically, if I were to take the X value here, and I were to pass that to the fit function, then that would be done on the CPU by default, right? Because this X data is not a GPU tensor, it's not bound to the GPU. It's just your NumPy array. That's on your, your -- on your host, on your log-in node or on your CPU that you're running on, you have a variable that's a NumPy variable called X. Well, if you were to pass x to the fit function, then that it would know to do well, well, that's located on the CPU, and it would do that computation on the CPU. But since we're passing it X device, an X device is bound to a GPU when we pass X device to the fit function, it knows to do that computation on the GPU. So this is why it's compute follows data. If the data is on the GPU, the compute follows there. If the data is on the CPU, the compute follows there. So that's how compute follows data is. Now in order to do this, you want to cast your data from your NumPy formats into that DB controlled Tensor, both for your X train, your Y train, potentially your X test, Y test, depending on what you want to do with it, but you would cast all those things as DB control tensors. Now some algorithms such as predict and fit-predict, return back to you in a normal circumstance on Scikit-learn, they would normally return to you a NumPy array. Well, so in terms of doing this on the GPU, it's going to return to you instead a DB control Tensor. So if you're calling predict or Fit-Predict, which normally would give you back a NumPy way, now it's going to return back to a DB control tensor. So now we need to cast that using this tensor to NumPy. We need to cast that back to a NumPy array. So in this case, I told this data the X train device and Y train device are on the GPU. It's a tensor now. So x train device is a DB control tensor. I pass it to the fit function. [indiscernible] to do that on the GPU. I get a result back, CLF. And so that CLF thing, I say, okay, it's an object, and so I'll call the method of called predict on that object. I do the predict on x device -- x test device. This is also on the GPU. So this computation is done on the GPU. I capture that result. And -- but it's a DB controlled Tensor. Most functions and most things that you deal with in Python don't know what a DB controlled tensor is all about. And so let's cast it back to NumPy so that we can manipulate it more easily. So what we'll do is we'll just take this DB control tensor that was retract to us. We'll pass it to the conversion, DB control Tensor to NumPy and now we get a NumPy array back. And so it just strips off that information about where it was -- that it was tied to the GPU, and you just get your numbers back. So with that, you can pretty much tackle anything now. I've shown you kind of the ideas here. And this is just sorry that this -- I didn't catch us on the slide somehow in the formatting, I lost those, but I've down here at the bottom. But this is how you use it. And the idea here is that the fit and you don't have to cast that object that has returned. But when you do the fit-predict, if it returns NumPy array, then you do need to cast, you do need to cast for predict, you do need to cast for the fit transform, let's say, from principal component analysis or just even the transform. Any of those things that were normally returning in Scikit-learn it's now going to return to DB control tensors, so you have to convert it. So that's it. So that's pretty much how you do this thing. And so hopefully, now let's go back and start doing some of the lab stuff now that we've got some of the theory out of the way, and I'll show you how this lives and breathes for real. So Praveen, can you see my screen now?

Praveen Kundurthy

executive
#23

Bob, there are still some questions on the playing your customer environment, like running the system requirements.text. If you can just a couple of minutes and just wait again.

Bob Chesebrough

executive
#24

Yes. Let's do that -- let's do that. So here we go. We should be in the -- we'll just CD here to -- we just did the clone. So let's see, CD to machine learning One API. If you do an LS, there's going to be a requirements.text file right here. And what we're going to do is, we're going to tell PIP to install that. So what we'll do is we'll just say PIP install from that requirements list. And it's just doing a few things, seaborne and a couple of other things that we're going to need. And so it returns back fairly quickly. So there we go. And from there, now we can start going into Chapter 1, and Chapter 1 Module 1. And so from here, and I'm just going to go ahead and run all right now, and we'll start discussing some of what we're -- what we're going to be doing here. What we're doing here is we're going to use the cover type data set, the cover type data set. And this is -- can be a huge data set. I'm just doing a small subset of it, just to show you. But what we're doing is we're predicting the forest cover type, okay? So it picks a 30-meter by 30-meter cell of ground somewhere with a whole bunch of independent variables that are derived from the U.S. Geological Survey and forest service type guys. And the PH of the soil, a whole bunch of a wide variety of ways that they characterize soil types, okay? And then from the soil types, they want to be able to predict the kinds of cover that you would grow on there, maybe [indiscernible] ponderosa pine or Juniper or what have you, right? And so that's kind of the idea of the data. And so this is just -- this first chunk of code here is just we're just going to be refreshing that data. So if you want, I went ahead at one point and put a timer in here because I think we did use train test split in here. And I got a little bit of a speed up. It was not anything to brag about necessarily. It was faster. It was great. So -- but this wasn't the intent of the lesson. This is just -- we got to get some data to play with. And in this case, I'm going to take the data and just use 1 quarter of it. So I'm just kind of skip counting the amount of data that I'm using. And then I'm just going to do a train test split. So I have X train, Y train, X test, Y test and so forth. And then I'm just going to set up some parameters for the -- I think in this case, I'm using the K Nearest Neighbor function. And so K Nearest Neighbor is your Scikit-learn, it's your friend. And so what we do is we invoke the listener first. So this is -- you'll see that you need to insert the patch. It maybe there may not be -- it must be here now. And normally, I erase this so you have some work to do. But I think I left it for you already filled in. But this is -- you have to do the patching first as the point of this, right? That sets up the listener. The listener then is listening and says, "Oh, I've got a faster version of K&N, cool. And then -- right here, we're just going to show you how you do this just on the -- basically the CPU, notice we're not doing any DB control or anything like that. And this is just showing you the results of what we get, okay? So this is the F1 scores and recalls and so forth. Now I want you to learn a little bit about patching and unpatching here. And so I don't want you to remove this code. I want you to be able to say, well, let's unpatched specifically because maybe I've done experiments in a cell up above somewhere where I turn patching on. Let's make sure patching is off to start with, so we understand the baseline. And then let's run this and we'll capture time for this, okay? And then we'll look at our accuracies and so forth. And then since we've done that, now we can do the same thing, but now we can apply the patching. And this is just going to show you, in this case, I got a 56 X speed up between the patch version. Now this is not using the GPU. This is all just doing the CPU stuff, right? And for this first module, Chapter 1, Chapter 2, we're just going to be doing CPU stuff. But -- so we get this -- that's pretty good day's work, 56x. And it turns out if I use more of this data set, not just the quarter of it. Typically, I get even better performance because it's -- you get that efficiency over a larger chunk of data. So this is just how you would use this for K Nearest Neighbor classification kind of a thing. On a fairly complicated dataset type. And so cool. So the motivation here for using Scikit-learn is the performance gains you can get here. This is, I think, for training, but you get -- there's an equivalent curve that was generated. And I think I specified in here somewhere the URL where that's dictated. Yes, right here, information about acquiring the toolkit. If you click on this, you can see the graphs, the latest graphs that they've done for these. But what this is, is -- this is doing the training. This is looking at the training phase and if you're doing DBScan. Well, it's highly dependent upon the amount of data you have and the complexity of the data, right? So you do DBScan, 500,000 rows and 10 columns. We get a 4x speed up over stock version, okay? In terms for Scikit-learn is roughly 4x faster. Take DBScan on the 500,000 roles and 3 columns, 5x faster. DBScan, I'm just telling you how to read the chart. DBScan 500,000 rows, 50 columns, we get 322x speed up, okay? This is for training. And so you get some really amazing speed up, some that are, hey, it's nothing to sneeze at. Even getting a 2 or 1.4x speed up, and this is for the data set, 1 million rows, 2 classes using random for us. We get a 1.4 speed up on training. But you don't sneeze in any of this stuff because all it costs you as being 2 lines of coding. Okay. Now let's look at the same thing for inference. So for inference, here, I'm getting a 4.1 or 4.9 322, that's training again -- inference right here. Inference. So the same thing, K-mean, 10 million rows, 50 columns, 5 clusters, I'm getting 72x speed up and so forth, right? So some things are more sensitive on inferencing side. You get more speed up on inferencing. Some things you get more speed up on the training side. Some things you get the speed up for both. And so one of the things that this does is it's going to help you to test this out yourself, right? So you can try different patching strategies and how do you patch surgically like this, how do you patch surgically with several things you want to turn on, how can you unpatch surgically, just all the stuff around patching. And so what we've done here is I've created a little test harness that you can use. Now this is -- what we're going to do is we're going to run a function called comparison. And what it will do is it's going to go for a whole bunch of algorithms that we've chosen, we'll specify an algorithm, we'll specify an estimator, so that's basically the algorithm with the parameters applied, whether it's -- excuse me, whether you're using whatever parameters for random forest or for support vector classifier, PCA, like how many clusters, you expect or how many dimensions. All those things will be in parameters. So that would be the estimator, which is based on the algorithm and then the data. So we have this idea of coupling essentially the algorithm and the data together, okay? And then we're going to time it, and this will be the comparison. Well, that's a function we're going to call is what I just showed you. But what we're going to do for the data or for the control mechanism, the test harness that we're going to use is in this case, I want to use SVC with a linear kernel. Well, I do this import right here. I'm going to be importing for SVM, SVC, I'm going to give -- it has a parameter called C that we're going to set to one, but you can play with it yourself. You can try a linear kernel or some other kernel. So this is giving you the ability to play, right? And then we're going to use Scikit-learn's ability to generate synthetic data sets. So we're going to make a classification data set. We'll give it in this case, 20,000 rows, 30 columns, 3 classes of data. 3 of the columns are going to be informative and others they're guaranteed to be independent columns and then specify a random. So this is -- this first entry here says we've got an algorithm and some data. This second entry says for SVC with real basis function kernel, certain sized data, certain parameters, which algorithm do want to use. Here, we're going to try to adjust a regression with certain sets of data, right? Million samples, whatever that is. KNN classifier, we are going to do the same thing and KNN regressor, and so forth, right? We're just going to keep trying an algorithm and we're just going to play with the data sizes and complexities of the data by using the Scikit-learn ability to create synthetic data. And then we'll just -- after we've defined all these things, we're going to run it. And what I want you to feel here is the burn. I've started running this thing when I went into that long [indiscernible], okay? It's still running. This is the unpatched version. This is the stock version of this thing. It's one thing to see that performance on a graph. It's another thing to feel it in the nerves in your legs as you sit on your chair, okay? This takes time. And it's uncomfortable. This is what the stock version does on this -- on all these data sets. It's running through all the test cases I gave you. and it's saying, okay, for linear regression, do this kind of data set. But you can tweak and play with it so that you can make that data as much like your data as you want, right? And you can modify to try other algorithms. But this is going to give you a comparison, both patched and unpatched. This is the motivation slide, okay? And it takes a little bit of time. We're still waiting here. Is it still running? Is my kernel still running? Yes. I'm up to 88% completion now. So -- and I'm sorry, in a way that it's slow. But this is what you feel and how long it takes when you're using the stock version for these kinds of tests, okay? This is the stock version. Now we're going to do the same thing, but we're going to patch it, and here's how, right? So I set up that listener ahead of time like this, and then I call the get cases. And that's all I had to do is insert these 2 lines of code and then repeat. Now I'm already up to 62% and that's just from the time that, that other one completed. We're almost done, okay? So this is going to be quite a bit faster for of these algorithms and you're going to see why and how directly as we analyze the result. So just let that complete what time do we have -- how much time do I have left, Praveen, like 15 minutes? Before I turn it over, Okay. So I'll pick up the pace, but the code runs as the code runs. So let's just -- I'll be a little patient on this one. And we'll just get to see what happens with results. But already, the nerves in my legs are saying thank you for running this thing quite a bit faster. This is the speed up that we got from that test. SVC, the stock version took 45 seconds, the accelerated version took 2 seconds. For SVC with the radio base function, 21 seconds versus 2 seconds. logistic regression, 8 seconds versus 1 second. 22 seconds versus 4 seconds. Now this is for fit. This is training, okay? Some of these were slowed up again with on training, so we didn't see that much benefit. But now let's look at the same thing with respect to the inference side. And so with respect to the inference side, we get something like 20 to 40x speed up something like this for SVC. For k-Nearest Neighbor, we're getting 1.3 compared to 3. Logistic progression happens so fast, we couldn't hardly measure it here linear regression. So you're just seeing the overall benefit. This was plotting everything together, both training and fitting together. You'll just see that this is -- or any number of these algorithms. And now you can go back and you control the right -- you can go back and put in your own algorithm.

Praveen Kundurthy

executive
#25

Can you show the kernel, the separate kernel, you picked up to run these road books oneAPI 2022.1, right?

Bob Chesebrough

executive
#26

Oh, yes, yes. Yes. So to select the kernel, you go up here. And since we did that little trick with kernel, you're going to select the oneAPI 2022.3.1, right here. Why did it not so select. Yes, there we go. Okay. So there's that. And then I could rerun it. Now there is a lab here. I'm not going to go through this right now. This is a lab that if you -- this is just for read-me instructions for you, okay? The goal here is that it's going to have you do some things on the command line, these kind of things on the terminal. You're going to go to the terminal and you're going to do these things. And what you're going to do is you're going to convert a Python. I mean, a Jupiter notebook right here, you're going to convert that to a python file using some of these scripts. And then you're going to use the command line, patch method, Python-M, Scikit Learning X on this entire notebook. And so not even changing any lines of code in it. You're just going to do this online, and you can go and get to feed ups. But then you can go back into the notebook itself here and you can turn on some patching and turn off some patching. So you get used to how to do surgically unpatching and patching and doing wherever you want. So you're guaranteed to get performance. I'm not going to do that right now because we're starting to run out of time. So I do want to just show you just a couple of other things. This is using K-means' Chapter 2, the first notebook, just showing you how to do patching specifically to this. In this case, it's the IRS data, we're doing K-means. What you're going to see is that the patching mechanism that we're talking about, it doesn't matter whether we're doing or we're doing K-means, DBScan or Classifier or whatever it is. The idea is simple. You do the listener first, then you do the import. And every one of these things is basically the same way that you'd ever use it. So it's here for you to play with. It's a great little learning exercise for learning how to patch and unpatch, but I'm not going to belabor it. There's another thing here. This is for using the support vector classifier and you're going to -- same kind of idea. And for the sake of time, I'm not going to go through those, but this is what the lab is. This one, I may share with you just a little bit -- this is the pairwise distance. Now for pairwise distance Scikit-learn -- yes?

Praveen Kundurthy

executive
#27

Again. There are, yes, still are questions on the usage of the kernel, you can a little bit on the...

Bob Chesebrough

executive
#28

Okay. So what we're going to do, if you're in a notebook and let's say that by default, whatever happened, you were -- had this kernel selected, Python, oneAPI. This lab may or may not work for you if you have this selected. So what I did is I guaranteed that it works because I know it absolutely works in 2022.3.1. So in order to select that, you just go over here, right click -- or I mean just left click on your kernel, whatever this is. And this little location right next to the open circle will give you a drop-down list that lets you select various kernels. Now I have a bunch of custom kernels. You're not going to have this many. But hopefully, you do have if you follow the steps on the GitHub, you will have oneAPI 2022.3.1. And so you select that one, and you'll see that it changes right here. Do this. And then all you have to do is go ahead and run the notebook. So I'll just do a run all cells there. And that's how you would do that. So this is just something I want to explain. This is using pairwise distance. Now pairwise distance is this idea that you have a bunch of dots, a bunch of vertices somewhere in your data set. You could be 1 dimensional, 2 dimensional, it could be N dimensional, okay? And you want to find the distance matrix. It looks like this. So how do you read this distance matrix? What this says is the distance from 0.1 to itself is 0. The distance from 0.1 to 0.2 is right here, D 1 2. Distance from 0.1 to 0.3 is right here. It's a symmetric matrix. The distance from 1 to 2 is the same as the distance from 2 to 1, there's a symmetry here. But that's just the distance from 2 to 1. So this is just all the permutations of all the different payers that you would have, and it computes a distances. So there's an algorithm called pairwise distance that will compute that distance matrix for you, and then you can use it for various things. One of the things that you can use it for is, for comparing shapes of stocks, for example, or shapes of any kind of time series data. So the idea here is that I want to come up with an example of why you would use pairwise distance for the things that we accelerated. We didn't accelerate all of pairwise distance. Pairwise distance, if you use euclidean matrix, so it's 1 of those parameters that you pass, we don't accelerate that. And we don't accelerate the Manhattan distance. What we did accelerate is the cosine distance, and we also accelerated the correlation distance. So what I wanted to do is come up with an example to show you, well, who cares about cosine distance on pairwise distance, who actually would do that? This is an answer to that. Here's 1 way you could do it. You could use it to generate a portfolio of stocks. So this is just a random walk, synthesizing 500 stocks, how ugly, too much information. But it does kind of show that starting over here around $600 per share, some of these random walk stocks increased quite a bit in value and some of them actually cut in half or 1/3 of their value. So it's just trying to say, yes, it's roughly stock-like behavior. Let's just look at 4 of them, right? Some go up, some go down, it looks kind of stock-ish, that would be believable, right? But this is just random walk stuff. And then what I do is I draw with my finger some curves or I can draw mathematically, if I want to, which I did here. But I'm going to draw some curves and say, if I had this curve right here, in blue, which of this mass of stocks, which are the 4 or 5 or 2 or 1 of these stocks, most closely match this shape of this line? Or let's try this line, which of the stocks follow that pattern. This is a cyclical pattern or this pattern, right? So I have a shapes -- I have a shapes array that has these things in them. And so this is shape 0, and this is shape 1, this is shape 2. And so what I'd like to do is to use -- and I'm going to do the patching. This thing runs so fast, it's crazy. But what I want to do is to take that shape. In this case, I inverted shape 0, and I said, find all the stocks that are as close as possible to this stock using the cosine distance. And so to do that, it was just a pretty simple method. I just reshaped the data to put all my -- this is a portfolio of 500 stocks with 12,000 rows, representing 1 year's worth of stock data taken every 9 minutes or 10 minutes sample. And so I just reshaped it to put all the times as columns and the stock symbols as rows. It's a little weird maybe, but that's what I did. That's what this is saying. And then this is just a way of basically doing arc max. It's a way of doing a 4 -- find the 4 largest instead of just single largest distance. I just call pairwise distances on the shapes and that data. And I just give it the metric cosine, I'm guaranteed that we're going to be using the fast version from Intel, and then we plot it. And so sure enough, it says, plot the original data as a shape, plot the 4 stocks that are similar. And so here is one way and one means in which pairwise distance can be used to accelerate some kind of interesting stuff. So anyway, I just offer that to you. I need to switch gears really quickly here and just show you how to do patching with the DP control library. So I'm going to move on to that subject now for the balance of our data for our time here. So let me just go ahead and start running this. So here, what we're going to learn is -- I did give you just some information. I encoded it into a CSV file that you can pull into a data frame to understand which things did we accelerate with respect to the GPUs and what are they good for, what are they -- what are the advantages, disadvantages? So you can kind of query that in play. That's not really the main intent. The main intent here is to have you realized that you do need to cast your data. That's exactly how you do compute follows data is that you specify what your data is and tie it to a specific GPU, GPU 0, and then you do your computations. And this is what I showed you on the slide about the fit, fit predict, the transforms and all those things. you do have to do some casting on the way back out sometimes, right? It depends on the data that you get back. This is how you invoke the DP control library. And in this case, it's going to be printing the version of the DP control library where you're running. So it's very simple. You just import DP control. Remember that we have to do our patching, right? And then now in this case, I was just going to show you -- I think just the standard version, I'm not sure if I did the -- on this one, I don't think I actually did the conversion. So this is just showing the baseline without doing the GPU, and now we're actually going to do the GPU. So we're going to import the library, do the patch. And now here, I've got -- I can just say, look, find the device, just select the default device. And if it has a GPU on this note on the DevCloud, what will be returned is a GPU device. And so I can say, just give me the default one, capture it as a variable. And then when I do that casting of the array, I can just say, look, just use DP control Tensor array. We're just going to cast it from NumPy right here, this is my data, cast it into a Tensor using the device we just captured here. And that's it. And so then what I can do is I can call K-means with that data. So this compute follows data. This is on the GPU, we cast it. And then now this is just a DevCloud right now, the current iteration of DevCloud. What I'm doing, the log-in node that you're playing on right now doesn't have a GPU on it. But there are other nodes on the DevCloud system that do have GPUs. And so what we're going to do is we're going to write this file. We're going to write this entire cell to a file. And we're going to bundle it up and ship it to a node that does have a GPU and tell it to execute that and give us the results back, okay? So to do that, we're writing this cell. So this cell doesn't actually execute because I have the cell magic right here that says right to everything in the cell to a file. But when I execute this line, what it will do is it will write to that file. And then this just says, "Hey, take that file and send it off to the GPU." And so we submit that job. It comes back. And sure enough, it did its calculation. We were doing something on the Intel graphics card. Okay. Great. It can return back a result. I'm getting an error here. I don't know why it doesn't matter. I got my result. So there's something in the cleanup with that's returning, and I haven't tracked that down. Don't worry about it. The main thing is, are you getting your results, did you get your values from the GPU back to your log-in node and your Jupiter session? Yes, you did. So this worked. And so this is doing the same thing, but with DBScan. And so the idea here again is that we have to set up the DP control library, you import it, you do the patching. Here's my data that I'm going to send it. And now what I'm going to do is I'm going to select the default device again from DP control library. It says, yes, okay, I got a GPU. And so I just used that when I do my casting of my NumPy array, do my Tensor. So now x device is a Tensor. If I call the DBScan fit function on a Tensor attached to a GPU and that computation for DBScan will be computed on that GPU. I get the result and then I can print components and so forth. And so we'll just look at this and -- okay. Well, I can troubleshoot now. I had some error, and I've seen this happen, frankly. This was running perfectly last night. And once this morning, occasionally, I would get this error and I do not know why. Just rerun it and you'll likely get that result back. I don't know why intermittently I'm getting that result, but this runs successfully if you are patient with it. And I don't know what's going on right now with that system. I can't explain it. But I think it has something to do with configurations on the DevCloud. And I don't quite understand yet. But this is what you do. And so this -- and to just show you now misbehaving -- yes. But here we go, K-means we're getting results back from the GPU. DBScan, here, okay, when we ran at this time, it came back with the values. Here, we have a linear regression, came back with some values and all these things look good, logistic regression. So we did all these calculations just now on the GPU, and that all works successfully, so getting a little error. But that's how you do it. And so Praveen, I need to start turning this over to you probably pretty quickly, right? So we're pretty much done. If you want to play with some of these other algorithms, some of these other notebooks, 5.2, I would use the -- not don't use the old one, use the newish one. And -- or if you want to do this gallery of functions. This gallery of functions is, again, going to just show you which things. There's -- you'll see. But this is -- you've seen how to access the GPU through Scikit-learn. So that was my objective. We've accomplished that. Praveen, I think maybe it's time to head it back over to you, so you can cover your part of the talk. Actually, I'd probably going to talk about this thing, too. So this is the AI reference kits. The AI reference kits, we just want to call your attention to these. We currently have AI reference kits and it's an effort by Intel and one of our ecosystem partners, Accenture, to provide a recipe to vertical-specific AI opportunities. So we invite any other ecosystem partner here if you're on with us today to do the same thing or you can what's there already. We currently have 22 reference kits on GitHub with 6 more being posted next week. In mid-July last year, Intel started to release a set of trained AI reference kits to the open source community. And so these reference kits give people a starting point to create basic models so that you don't have to go through the exercise of having to figure out your own. And for us, for Intel, the best way to drive adoption of developer tools by customers and our partners is using them in a real-life environment. And our reference kit is kind of like a recipe that demonstrates how a specific challenge is being solved. And the benefits when using the oneAPI AI Analytics Toolkit and many of these components that we've been talking about when you're developing these optimized version. So it kind of gives you a short kind of like the notebook that I'm giving you now, but it covers many other things. So there's at least 22 kinds of applications that have been optimized and put out there for Accenture. So go ahead and look at those, those are the reference kits. And then here's just some of the kinds of paradigms that we cover. So regardless of whether you're doing data processing, model training, hypertuning, inferencing, the reference kit will include guidance on how to use those and how to do that the best way. So the reference gets contained a solution brief, which is an overview of the value proposition describing the problem with the solution as a developer guide with recommendations for the frameworks, the algorithms, data processing techniques, hypertuning, quantization, deployment, all kind of stuff. Code repository, so that you get the GitHub good snippets. Platform architecture, a guide for best-performing compute architecture for the reference kit and then benchmarking results. And so using these reference kits can really radically reduce your time to deploy because we did all the heavy lifting. Let's say that the work is for about 70% done and all that work used to take quite a bit of time that you would have to do, but now like 70% of it is done by this thing, you can just -- it's not work that you have to go searching for things and making things work and deciding things because that also takes expertise when you have to which parameters and different things. So start with these reference kits. So new to this, in particular, use this reference kit because it shortens your learning curve. It gets you started quick. And so that's why I believe the reference kits are going to really help to reduce the starting costs significantly for you. So -- and then if you want to get the Scikit-learn, the oneAPI AI Analytics Toolkit for your compute that's not on the DevCloud, here's where you can go to get that. If you want to go to do some of their -- there's links that we provide also, I think, I'm there for how to get the -- just the Scikit-learn extensions by itself if you want to. But now I just want to save some time, Praveen, for you to cover the data parallel essentials for python. So hopefully, I've given you enough time.

Praveen Kundurthy

executive
#29

Thank you, Bob. Can you hear me?

Bob Chesebrough

executive
#30

Yes.

Praveen Kundurthy

executive
#31

Yes, there's a question on the GitHub link to the reference Bob, if you can please post a link that would be awesome. It should be on the oneAPI source.

Bob Chesebrough

executive
#32

Yes. Yes. I'll pull it up. I'll answer that.

Praveen Kundurthy

executive
#33

Hello, everyone. My name is Praveen Kundurthy and Bobby is my colleague. So we'll be doing this together. So I'll be doing the data parallelization for Python part. Bob just showed the [indiscernible] if you remember, let me try to show the AI Essentials Architecture Toolkit. -- right? And I'll show you where, yes. All right. Please -- if you remember this slide, right, Bob, he is talking about this, and he introduced the Intel extension of Scikit-learn. And if you go to the down part, if you see the Intel optimized python. So I'll be talking about the number data parallel essential for python. It's a data parallel extension for python and we'll see what this is. And we'll see the similar algorithms like pairwise distance and K-means using the data parallel essential for pythons. We use the same module, similar modules like DP control, to offload device, and we'll quickly cover that portion, right? Let me switch quickly back to the slide that we want to cover. Let me scroll. All right. So this is the data parallel essentials for python architecture. So the main point here is if you are actually familiar with the SYCL programming right, the basic SYCL. SYCL is a heterogenous -- Khronos heterogeneous programming capability probably. So like -- so you can use SYCL to write parallel programming code and offload your device. And that capability is brought the Python world. So the goal here for data parallel essential for python is how to bring the Python programmers, how to use the oneAPI programming model, right? And it is a fit of packages, implementing a common programming model for XPUs offering from the python. And the packages in this data parallel essential for python provide the necessary building blocks for developing Python packages that use SYCL, right? I talked -- we talked about DP control, right? DP control is the module that actually helps in offloading our python work to a device using the capabilities of DP control that can select a device that we can offload to a device using multiple API, right? And we also talked about the implementation of Tensor library, right, which is based on Python and API standards. There is a DPNP module, which is called data parallel NumPy, which is a drop in replacement and this actually -- all the NumPy calls are being executed on a GPU device. And the last thing is the extension from Numba, which is called Numba Data Parallel Extension. And this is exactly like how you use programming -- GPU programming model, like you've got the kernel call you use all the advanced features like indexing, synchronization, using the group barriers. If you are a familiar GPU programmer, if you are familiar with GPU programming, then the Numba Data Parallel Extension is the way to go with, right? So let's see what we bring up with the upcoming -- so Bob also already introduced about the compute FOLLOW set approach, right? So it is based on the python array, API standards. And it's a python way to specify on what device, a computational kernel actually executes, right? So if you see here, I got a DPNP array of -- x is DP dot array and I'm passing in the default device here. I'm not passing any parameters. So this particular DPNP provides are a array reconstructor, which is called dpnp.array that have a parameters like device that uses what is the type you can use it. Is it using USM type, and you can also pass in the SYCL queue, which is used for work management -- execution of the work management on the device. So these parameters actually allow you to specify where exactly is data placed and then -- so once you place the data, the execution happens on the same data. The second box, including simple code array, I'm passing in the dpnp.array and the constructor as 1 2 3 right, like array. And I'm passing in the deep device which is a GPU, right? So this will be executed on the GPU device. The only thing that doesn't -- that is not supported here, right? You see I got 2 DP -- 2 NumPy arrays, right? And these are on 2 different devices, GPU of 0 and GPU of [ 7 ]. So this will provide run time, this will provide error. This is because compute follows data approach cannot work on the data that is executed on 2 different devices. The data should be on the same device and the device, the data knows where it is and the computation happens on the same device. So explicitly, you don't need to call DP control dot select the device. If you are using this compute follows data approach, you explicitly don't need to call the compute follows -- sorry, called the DP control device. You can just specify the queue or device and then it automatically runs on that, right? So there are 3 ways you program and take the SYCL parallelism in your python world, right? The first one is called the automatic offload approach we use, if you see, I'm using the special decorators called at njit, which is numba JIT. JIT is just-in-time compiler And if you are familiar with NumPy, it's a very popular parallel programming approach using python, right? And so using the njit and if you got explicit NumPy senior array. So if you're saying njit and then you've got explicit NumPy, you anything, right, specific to NumPy. Then your code is automatically offloaded to a device that you actually offload to you. So you create njit decorator, you got your -- you got lots of NumPy arrays in your function, right? And you use the DP control and select the device and, let's say, your -- this will be automatically offloaded. So the other way is called explicitly using the loops, explicitly using p range in the loops. So you got a loop, right? If you see here, I'm still using the njit decorator. And here, I'm calculating a simple distance, which will be the distance between 2 points, right? So if you see, I am using explicit for loop and the main important difference is, I am using a p range. Instead of range, I'm using p range, which will sent for panel execution. So use a p range in njit decorator. This will also be sent for parallel execution on a device, specifically if you are using DP control to target a device. And the third way, the famous very familiar with GPU programming. The third way is called the open-style kernel programming, right? And this programming, right, as I mentioned, it is very common in GPU programming world using this type of programming model. And if you see the differences, I'm using the kernel decorator and this is similar to a numba's other GPU back ends like numba.cuda, numba.roc. These are exactly similar. And the numba DP pack kernel decorator is provided in the numba depix package. And similar to -- if you're familiar, if you attended our previous oneAPI and SYCL classes. So this is -- if you use this kernel decorator now, you got the capabilities to call all the low-level sickle kernels that can be written directly in kernel -- in python. So several advanced -- GPU few advanced features like indexing, right? You get the global ID. You see I'm getting the global ID. You have 0 global ID of 1. So I'm getting the single index in the X and Y iteration space, right? And synchronization, how you synchronize things because there's -- in the GPU programming world, right? You send a parallel execution, you need to have a proper synchronization mechanism to avoid any database conditions. You create fences and call atomics, all these GPU features are here if you are using the kernel -- open-style kernel programming -- and we'll see...

Bob Chesebrough

executive
#34

Before you complete the slide or whatever get too far down the road, Bradford has a question, a couple of them really. Is it possible to overload a device? Is there a limit to array size that you can pass.

Praveen Kundurthy

executive
#35

So the first question is, is it popular to -- is it possible to overload a device? I didn't exactly get that question, but let me understand it this way, right? So you are -- I'll see that. I'll show you the code. So once you use the, let's say, the OpenCL Kernel Decorator, right? So you call the function. And then you'll have a you'll -- then you use the DP control .dot select default device. So the select default device, what it will do is it will select the best available device on your system. If it's have -- it will have a GPU and if it sees it is the best developer device, it will offload directly to that. So you can use 2 different context. You can use device DP controlled dot device context of CPU and take the second part of my, I have got 2 workloads. I can use the second -- I can use -- I can create another DP's control context and offload to a second device. So if it the question is. So yes, I can split my work across different workloads and offload to different devices if I got multiple devices on my system.

Bob Chesebrough

executive
#36

You just have to be careful not to combine the results don't try to -- like your previous example. Don't try to be adding a result from GPU1 onto the result of GPU2 directly like you'd have to think about that a little bit more carefully. But yes.

Praveen Kundurthy

executive
#37

What is the second question, Bob?

Bob Chesebrough

executive
#38

It has to do with any limits to the array sizes. So I don't know if this is in the context of if I shared memory or with respect to explicitly passing the data. But are there any array limit size can think of?

Praveen Kundurthy

executive
#39

No, not exactly, but within the limits, you are good, but if there is anything that's like exceeding, let's say, I remember if I'm using like a dataset size of 26 or of 28, which is huge. Like, sometimes I've seen like exceptions throwing up, but you are good within the limits of a programming practice. Some, if not all, but sometimes when I'm using like a huge data set like 30s and stuff. You may see some exceptions showing up, but you are mostly covered. Okay. So with that, let's quickly look at the simple examples, right, that I'm talking about. This is using njit decorator and automatic offload approach. So it's simple. I'm importing DP control, if you see, and I'm importing numba and then this is a simple L2 distance calculator, which computes the distance between 2 points. And I'm using the numba.njit, as I mentioned. And if you see, I got multiple NumPy expressions here. So it's using subtractions, square root of, NumPy square of, substraction and then -- and then I'm summing up those squares and then I'm doing a square root of that, right? I'm using all the numba expression to calculate this L2 distance, right? So this will be automatically offloaded until I have this particular section, where I got the diverse DP control .dot select default device, right? So I'm using the DP control module and I'm trying to get the best available device on my system. Or I can pass in DP control dot select GPU device. If I already know that there is a GPU available on my system. But this is like a fallback mechanism, the default device, right? If you don't have a GPU, so you don't want to throw some run time exceptions being happening, right? So the easy -- the better way is selecting the default device and that way you're providing a fallback mechanism, and you're actually calling the DP control dot device context of that particular device and then you're passing in this decorator function, which is L2 distance of kernel, right? This way, this particular function is offloaded to device and all the parallel executions will happen on the device. So this is the first way. The second approach is using p range and exactly similar to what we did before. I got import njit and p range from the numba. I'm still importing DP control, right? And I hear -- I got a for loop here, if you see, I got a for i and p range of, right? Length of B, I'm calling explicitly a p range function so that this loop will be sent for parallel execution on the device. Exactly similar to what we did before, we are using the DP control dot device context and we are offloading this to the device to a GPU, specifically if you have a GPU. If you don't have a GPU, it selects a CPU, right? And the last approach is the GPU style programming, like the openCL style programming that we talked about, is the kernel decorator. And if you see I'm been putting the DP control here, and I'm also importing extra module called NumPy depix. And I got the kernel decorator here, right? So the first one is indexing as we talked before. So I get ix depix get global idea of 0, which will give me the single index in the complete. So it will be a single item in the complete iteration space. So I'm looking through all the elements in my array, right, using this. And I'm doing a simple vector addition here. CFI is AFI plus PFI? I'm passing in 3 data items. So this is a [indiscernible]. And -- you see the common way of kernel vacation is you have a driver and you'll pass in the global size and the local side. This is all of your familiar with GP programming. So you pass in the global size, which will be the total number of work in your GPU. And the local size will be the total number of items in a [indiscernible] group. So this is a little bit out of scope of this, but consider this as a familiar GPL paradigm, how you call in a GPU function, right? -- and then you're calling the DP control dot device context and passing in the GPU. And my driver function is being invoked from the device context.

Bob Chesebrough

executive
#40

We're starting -- yes, yes, we're starting to queue up a few. So we need to provide a mechanism. Ataliba Miguel, ask Praveen, could you point me to the right bespoke support at Intel, where I can get installation of running oneAPI smoothly? So maybe we can get a -- at the end, maybe we can share an e-mail address or something where we can connect with Ataliba and point him to the right people to get maybe pointing to the forum, I'm not sure what the right answer there is. And then there's Edmond Bett, has a question that says, can you add to devices array dynamically. For example, C-style array of 10, can you add to a size to become 20 dynamic collections like you would get in Java, where the memory is managed, where it's not controlled. Is there a way to do like a dynamic expansion of the arrays?

Praveen Kundurthy

executive
#41

So not explicitly, you can do that, but what you can do is you can -- like there is a concept called Unified Shared Memory. You can create a device memory, right? You can create a device memory, and you can pass in the device memory explicitly to the device. It can pre-device memory, and that will only be accessible on the device, but will not accessible host. So that particular device memory can be then explicitly transferred from the host to the GPU. And then you can add that. But on the device side, you can't actually on-the-fly add more items to your array.

Bob Chesebrough

executive
#42

Yes. Yes, that's great.

Praveen Kundurthy

executive
#43

So -- so it's the first question he's trying to install oneAPI on a separate bare metal machine where you got some issues? Or what is that like? DevCloud doesn't need any installation. If he's on DevCloud, he doesn't need any custom installation of oneAPI-based kit. But if you only required on the bare metal, right, what I suggest is to make a cleanup of your previous installations, and we can provide the links to the oneAPI-based toolkit. Let me provide the instructions of how to install oneAPI. Okay. so that you know you can...

Bob Chesebrough

executive
#44

Yes, I didn't understand the question. I think that you're right. I think you're on the right train of thought there, Praveen. So yes, the DevCloud, if you're trying to install it and do all these things as a custom kernel, that's probably not the approach. Just follow what I did in the GitHub link and I can repost that GitHub link right here. If you want to do it bare metal, then I'll fish around for the link that will describe how to do that, and I'll provide that here correctly. But...

Praveen Kundurthy

executive
#45

You're just providing the link to the -- I'm covering the link to the bare metal installation of oneAPI. So here is the link, right? So if you got the options to -- you installed base toolkit, you install AI Analytics Toolkits. So just you've got the option to install either of those, so install that. You can use the offload installer, online installer or you can use the command line installers using APT or whatever is required, right? So there is a lot of -- you can actually follow the instructions and they also for you to coin question like if you're seeing any issues, contact support and stuff. So that's the best way you can -- if you are installing at bare metal, that's the best way to proceed forward. Okay. I hope I answered the other questions, Bob. So let me -- so if we look at....

Bob Chesebrough

executive
#46

Just in 1 right here.

Praveen Kundurthy

executive
#47

So we have seen the 3 different types of parallel programming using numba, data parallel for python, right? So let's see how we can do a simple pairwise distance using one of the decorators. And we'll also see how we can do a K-means using the decorators, right, and offloader devices. So it's just simple examples and we'll wrap up the session doing couple of lab exercise in the DevCloud and we can end this workshop, right? So the first is you very familiar with the Euclidean distance in machine learning world, right? It's distance between 2 points, and we know the distances, hypotenuse whole square is side square plus side square. So this square will be square plus square, say, this example here, I will do 2 distances, square plus square to the square root of that, right? So it's a very basic math that we already know. And this is the common way of computing distance, which is L1 distance, where you've got -- you're actually -- it's called city block distance, where you are computing the actual distance between the 2 points, right? As you're vetting the [indiscernible] distance, but this is not become popular because you may have multiple blockages in between, and that's not the best way to calculate. But these are some cases where you can use L1 as a standard distance calculation, right? So the pairwise distance, let's say, this is our data set, right? And the data is distributed in certain way on a certain axis. I got X1, X2 axis. We can see it visually. And our final goal is to find a distance between all these payoff points. So I need to find a difference between 1 and 2, 2 and 3, 1 and 3, 1 and 4, everything, right? So if you see my distance from D1 to D2 is, right, D12, which will be PX1 minus PX2 whole square plus PY1 minus PY2 whole square to the square root. Similarly, the same formula applies to distance between 1 and 3, distance between 1 and 4, distance between 1 and 5 and this distance between 2 or 3. And finally, our goal is, we create a distance metrics, right, with all the distances from each point to another point. And if you see where the 0 is right distance between P1 to P1 is 0 -- from P2 is 0, right? So that way, we got all the diagonal elements as 0, because this is distance from 1 point itself. And you get the distances when we form a distance metrics. That is our goal, right? To form a pairwise distance matrix. And we are seeing this L2 distance right, which is as we have seen hypotenuse whole square is side square plus side square. But the main point from this slide is, Intel extension [indiscernible]. Bob actually was showing, it does not support this particular L2 distance. It only supports cosine and correlation matrix, we quickly look what that is, right? The cosine distance, you've seen Bob's workshop examples where he is trying to show a timeline series, where he's trying to find similar stocks, which is closest to a particular type of model, right? That's exactly how we are using cosine distance. We are trying to find dot product of 2 vectors, and we just see what is the nearest product of that, right? What are the nearest points, nearest angle to that. So let's see an example. So I got 2 vectors, A and B, which is 3, 4 and 5, 12. And my cosine distance will be -- if you see the length of these vectors will be square root of 3 square plus 4 square, which will be 5 and the length of P is 5 square plus 12 square which will be 13, right? And my cosine distance will be 1 minus cos theta of that. Which will be the dot product, right. 1 minus 3 into 5, which is 3 into 5 plus 4 into 12 divided by the length of the vectors, which is 5 into 13. And now I got my cos theta as 0.969, and the cosine distance is one minus of that so that you're identifying the nearest point,closest pair to that, right? So cosine distance is 1 minus 0.969, which is. So this is the logic behind cosine distance. And this is well used in the Intel extension of [indiscernible]. So we can follow this approach. So there's other way called correlation matrix, which will be similar to that and the calculation of the correlation distance will be based on the mean calculation here. We use a distance calculation and the dot products. That will be based on calculation of mean and the nearest mean values. And that's the other way to -- the other metric that you can use in a Scikit-learn pairwise distance. So let's see a simple example of using pairwise distance using the DP kernel, right? So as we've seen the examples, my distance is a computation of distances between each individual point in the complete data set to the other point, and we are creating a distance matrix. The same example here, we will be doing the same. I got an importing DP control, and I'm using the number dot depix dot kernel decorator right? And I got x1, x2, which is the set of data points, right? And D is the distance metrics. So my I is a single index in the complete titration space. So I'll have 2 loops here. The first loop is the outer loop will be calculating the distance between the dimensionality of the pairwise distance. So let's say it's 3 dimensional, right? It actually calculates the distance between them. And inner loop, we calculate this each point to the other part. So that's a very common distance calculation logic that we've seen multiple times. And finally, what I get is a D matrix, which is -- which got all the pairwise distances like we have seen in the slides, right? And finally, I'm calling this particular driver function, which is pairwise distance, calling the DP control dot.device context of a GPU. If I don't want to do that, I can call with DP control dot select default device as a fallback -- this as a fallback mechanism. There are multiple ways you can offload GPU. You can use a device context and you can create multiple contexts to call different GPUs or you can select 1 GPU device submit all the work. So there are multiple ways you can do. This is one of the ways. So you create a DP control dot context and you offer a open CPU and you're calling the actual function, right? So this is the pairwise distance. We go and run an example on the DevCloud and see how this works. But -- and quickly, we will also walk through the K-means. We already know what K-means is, but just a quick recap, it's a unsupervised algorithm and we are trying to cluster 2 different groups of points, right? And so let's say, we got case 2. We got age and income defined here. And -- so the first is, we assign cluster centers randomly, 2 cluster centers. And what we do is, we initialized these 2 clusters and we're going to identify the distance between each point to the cluster, and we are actually color coding this. And what we are doing is we are placing these points nearest to each cluster, right? So in this, I'm color coding this. So the points near us to this color are closer to this, right? And the other cluster is a different color coded. So we are trying to actually put the points nearest to this closest center, right? And now as the next iteration, we are moving each center to the clusters mean. Then we move a cluster center again to the cluster's mean again. So this is the second iteration, and then we start the same iteration same process of moving the -- again, the points to the closest to center again, right? And when we do this, we keep doing this until we see a point of convergence, right? So this is the third iterate And finally, there is no change. So finally, the points are all converged. So we got a final cluster center and the points associated to the center. And we find other -- so if this is case 2, you've got 2 cluster centers and 2 clusters right? So this is the basics of any algorithm. And what we see is, we see a K-means implement. We have already seen the K-means implementation using Scikit-learn, right? Bob already showed that. So simply, you're applying the patch and your calling the K-mean syntax. There's other way around call using the numba JIT decorator again using, this is using data parallel python. Sorry, this slide is a little bit clustered, but the main point to show in this slide is I'll show it to you in the DevCloud, the actual code. But the main point from this slide to take away is right, you import the DP control. And we got 3 different njit we are calling here, right? The first one is determining the euclidean distance from cluster center to each point, right? It is group by cluster function will group all the points to a center. The second section will point to cluster and it keeps updating the centralized position, as we have seen, right? Until the points get converged, second logic will update the and keep the iteration continues, right? And the final one is, we got this logic to actually to implement a algorithm. And then we are calling the DP control dot device context, and we are actually offloading to a device, right? And -- if you've seen here, I'm using the njit decorator and I'm using p range for each of the for loops in this, which will be sent for automatic parallel execution.

Bob Chesebrough

executive
#48

So Praveen, if you are gone too much further, admend. Bet has a question. So he's asking about the numba depix kernel decorator -- at numba Kernel instructs run time to run on the GPU. Is that the idea?

Praveen Kundurthy

executive
#49

So unless you call use compute follow set approach or you call the DP control dot device offload. So the DP kernel decorator will be ready for offloading to device. So if you don't select any device, then it actually defaults to the default device, whatever you are executing on. And what are the functions that are not available, right, they will be just using the same resource, let's say, it's available running on a CPU, you own offload device. It tries to do all the parallel programming paradigm on the CPU, but let's say, there is a synchronization of SLM, right? So CPU doesn't have any shared local memory. And it may keep some of those and do the local resource allocation and stuff. So that's why it works. So we make sure you explicitly call offload to GPU, even you call it decorator, right? The purpose is to upload the device. But if you want to take all the advantages, make sure you call the DP control or device context or you follow the compute follow set approach, where you are explicitly mentioning where your data resides.

Bob Chesebrough

executive
#50

So really, it's kind of the combination. Yes, you use the numba depix kernel decorator as well as the DP control, what is similar to what I've shown with the Scikit-learn control is the gateway to the GPU.

Praveen Kundurthy

executive
#51

Exactly. So let me quickly run through the examples on DevCloud and we can wrap up this session. So I hope you're able to see my screen, Bob?

Bob Chesebrough

executive
#52

Yes.

Praveen Kundurthy

executive
#53

Let me see where I am Yes. So this is the DevCloud that Bob has already showed, right? And if you -- just to recap, right, you created an account, click on launch, Jupiter lab, right? And you should see the Jupiter lab coming here, opening up and so you can see file new terminal, everything the same, right, file new terminal. The only difference, right, from these course to what Bob taught is, I don't need to have a separate installation. I don't need a separate custom kernel. Because everything is -- I'm using the latest kernel and it should work as it is, right? So you don't need to create a separate kernel and stuff. But the GitHub link that I'm going to use for the workshop will be different. That's the important point. Let me share the GitHub link, which will be -- let me paste in the chat window. So samples get and once you clone that right? Let me go back to my screen again. And I got this on my -- already I have it on my DevCloud. So you can see, Git clone. Don't do it on the base folder, you can create a directory and just clone the complete repository. So -- and once you clone it, so I got this here -- 1 minute. Yes. Once you clone it, you go to oneAPI samples. Navigate to AI analytics, Navigate to Jupiter, numba depix essentials turning. And all the modules that are required, right, for this understanding this is are here, go to click on welcome.ipynb. You can click this folder to maximize your screen. And there are complete details of all the modules here. So you can talk -- if you wanted to understand what numba is, click on this. Whatever the -- all the details I provided in the slides, they are clearly detailed and you can actually run the examples. So let's say you've got an explicit parallel right? You're using numba expressions here. You see I got this explicit parallel kernel. So I got p range here, right? I've got a for loop, I'm using njit decorator, I'm importing DP control, and I'm doing a simple vector addition here. I'm calling device DP control select default device. I'm printing the device information. And what I'm doing is with DP control dot device context of this device that I selected, which will be the default device and I'm printing the results of 2 areas, right? Let's run this example and see how it works, right? So I submitted the job to the DevCloud. So the way it works on a DevCloud is it's based on mechanism, right? So you create a job and you use the qsub command to submit to a GPU compute node on the DevCloud. So this will be submitted to a GPU node and results will be executed around the GPU node, and you'll get the results back to where you are. The default is a log-in node what we are in, and it will not have a GPU. You'll have to use a script called Q, which is actually in-built in these notebooks. And my simple Q script is, if I go here -- if you see my Q script. The main essense is here, which is [indiscernible] GPU node and I'm passing in the script, which compass my Python application on the DevCloud. So this is just for information for you to understand what's happening on the DevCloud side. So you don't need to actually worry too much actually on the Q standpoint, but this is information purpose. So if you see, it is executed on the internal HD Graphics device. And you see this is level 0 GPU and I'm adding 2 areas with coming as to, right? So this is a simple example. And the second example is automatic offload using numba expressions, you calculate the pairwise L2 distance using the njit decorator. And if you see the square root of x1 minus x2 plus y1 minus y2, logic is being implemented using the NumPy here. And same thing, I'm selecting a default device, blending the device information and calling the L2 distance. So I already ran this, and I can run this. And should not take much time, and you should see that actually this year, and I can show you the results here. I pre-ran to make sure we save some time during the workshop. And this is my automatic offload. Yes, I think I didn't do it, but let's make sure -- still coming up. And once we get the results, we can actually go back to an example on pairwise. If you see, this is actually called Intel GPU, and we got the pairwise distance L2 distance has 3.314, which is device. So go through that in your free time, right? And you can -- there's a of implementing matrix multiplication. So you can go through that. This is a lab exercise using kernel decorators. So please go through that and try to get your hands on with how you use the kernel decorator and perform matrix multiplication on a device. There's 1 module on completely on DP control, so which you don't want to miss. So this is DP control intro. And please go through this DP control simple examples too. So this example, if you see, I can show this here, right? So this is selecting the Q and passing into device, right? If you see, I got DP control dot SYCL device, right? Printing the device information and you're printing all the information of the device. So the first example is creating default device. So either I can use a DP control dot SYCL device or I can create a DP control object of select default device, right? So this will be using the default selector, the systems the best available device. Or if you know you got a GPU, you can say directly select DP dot SYCL control GPU or DP control GPU device, right? And you can clear explicitly a fallback mechanism, is not present, you can provide a GPU and CPU, right? So there are multiple ways. So go through this module and I run this example and let me run this again. You should see all the information on it. So similarly, while it is running, right? You got to SYCL Q to create multiple contexts. And I was talking about unified which is out of scope of this, but I was talking about that, right, how you explicitly create a device using memory USM device object. There was a question asked like how we can create a data on the device, right? So you can create either a shared device -- shared memory on your host, which is accessible both on host and device. Or you can create a USM device, which will be only accessible on the device. So there are examples how to do that, right? So -- and there's examples on memory management. So I can go through that in your free time and understand the concepts of DP control. But we saw the code that is required for this workflow, which is basics, right? Using the DP control, we just ran the code. Let me see executed. Yes. So we run the code and we got the Intel HD graphics device, right, as a default device here. So -- and we have seen examples of using njit decorator using parallel offload approach and kernel decorator, right? So let's quickly run the pairwise distance module. The pairwise distance is the third module. And you can see the pairwise distance implementation. I used -- I explained everything in this excise. But the main important things to consider are this lab exercise is in such a way that you can adjust your data points as per the command line parameters, right? And if you see the 2 power 28 initial data size, I'm repeating 1 time, -- and all these are passed as come online to me. So the d is dimensionality of pairwise. So it's a 3-dimensional, I pass in 3. If I'm using unified share memory, I'm passing unified shared memory. If I want to do steps or 3 I'll show you an example here, right? So let's say my pairwise distance of GPU, kernel, say this one, right? So this is my the actual script file that I'm coming online parameters I'm passing. I got steps of 5, size of 10 -- so I'm starting 1024. So my data items, there are 1024 items in my data set, right? So the increments by 5 steps. So 1024, 2048, 4096, 9048 step. Each reputation -- each one is done 5x. So we are actually loading your GPU. So you're actually getting trying to play with hard -- they hit hard on your GPU and you don't get maximum parallel capabilities. And then you're passing in the dimensionality as 3, right? This is 3 diamonds array. So this is an example, and let's look at the pairwise distance with the kernel decorator right? And so there are all these examples. The first one is using njit. Second one is using GPU. Targeting GPU using the numba decorator? And the third one is using GPUs using kernels using the kernel decorator, right? And I already talked about this logic, but let's go through that quickly, right? So I got 2 data points, X1 and X2 set of data points, right? So if you see 1024 is the number of points in each data set. These are distance metrics. Get global idea of 0, you get a single index of each item. And then we have 2 loops, which are performing the pairwise distance logic and finally computing the distance metrics, right? So that's our final output. And now explicitly calling DP control as device context and you're passing in the device selector, which will be a GPU as true. And you're passing this kernel, which will be executed on the device, right? And this code -- I ran this code and it's executed on the device. And if you see, I've got pairwise distance. The kernel side is 1024 and you can see the time it took to calculate that. And similarly, 2048, 4096. It went through 5 iterations and you're calculating the time, just for information purpose, how much time it took to perform the pairwise distance calculation, right? And once you do that, we just plot the results, let's say, you compute the pairwise distance, and we tried the results. If you see my results graph. So there are 2 sets of points, blue and -- purple and sand, right? So you're seeing that we are trying to match the nearest purple icon next to the sand icons, right? So these purple is nearest to the these 2 nearest to this purple. So we're trying to exactly nearest points using this approach. So just to show that we can calculate the pairwise distance and plot the points as per the logic that we created here, right? So -- and then there is also -- there are lots of tests we made. Unfortunately, we don't have enough time to cover all these today, but so I ramp, you're familiar, it's a profiling tool. And what we did in this is we ran a simple data set, which the lapse time is 0, right? And we increased the data load size. Now the GPU time is 12.5%, you increase the data size, right, from 2 power 20 to 2 power 28. You'll see the GPU time elapse 12, which is still slow, and it's asking that, hey, you can still increase your GPU load, right? And you see that you stalled at is 51%. So it's saying that your GPUs are very idle. And then we still increased our approach. And now we're trying to hit hard again with now the -- if you see the last example, the GPU time in now it's 40% here. So I tried -- we tried multiple things. And finally, we set it in such a way that I hit my GPU in such a way that now we're not seeing their message that it's stalled, right? So I see in the GPU time as 58 seconds out of 70. But the problem here is we got a L3 bandwidth problem here. So it's not using the reuse in the cash, right? So the I am hit now I'm an L3 bandwidth bound. And these are all GPU familiar concept, if you are familiar with it. So it's 95%. So the way how I can fix this is using comparable locality, which is using the shared local memory using the global memory, I load it from the shared memory, all these items from the shared memory. So that I did use the latency happening from the global memory to the actual device, but I'm reading from the local memory. That way, I'm using the -- I'm reducing my L3 bandwidth and optimizing more my application. So quickly cover this. And pairwise difference, and we can quickly cover the key means exactly the same logic, and we'll run an example on that. Same exact approach. This is K-means, right, and let's run the K-means algorithm. Let's run the directly the plot is plotting right? So the same logic for K-means, we talked about that. So we got -- we are here grouping the points by cluster. The first example, we looked at the logic up there. We are computing the are being high rated and we're existing the distance, right? And the points are being adjusted as per the placements until there is no convergence, right? So all these logic is clearly explained in this example here. And I'm updating the rights. And finally, I'm using the kernel approach. And what I'm doing is I'm actually calling the function, if you see my DP control dot device context, but I'm calling these particular kernels and calling the K-means algorithm. And if you see, the same logic of the data sets arrangement, right? If you see my GPU, I got the similar command line parameters I'm passing. Steps 5 size, 104, repeat is 5, right? And if I run my graph, if you see there are -- I've created some -- tried to create some 10 clusters. And these points are adjusted like as per the cluster. So they've got different colors. And I can see for 5, 6, 7, 8, 9, 10, I adjusted all the points as per the cluster. And so we have followed the K-means algorithm using the data panel as Python. Finally, the same and adviser reports, increasing the data size, trying to make sure your GPU is handled properly, trying to reduce the bandwidth, everything the same. So if you see the last approach, we are using 68% of the GPU times lapse is busy here. And yes, finally, we got L3 bandwidth bound, but the GPU time is much better here. So the same approach, increasing your workload size and on with this examples and get your hands on.

Bob Chesebrough

executive
#54

Praveen, we've got a question from Alhakamao, He's a little bit puzzled. He wants to see you actually run this code in the Intel cloud, you want to see an example of it running. Can you click run on and just show. I think that's what he's asking for. And what I was saying, too, that he can always grab this, run it on his own instance of the DevCloud and play with it and see it run I think he wants you to see you run it, yes. Yes. it's running right now. He's actually hitting the GPU node with this right now. So it's submitted that to the -- because this log-in node that he's on right now does not have a GPU. So he's got to submit that code as a bundle, submitted off to a node that does have a GPU. It's -- that GPU now is running that code, that node with the GPU is running it, and then it will return the result. -- maybe that's some of the confusion. I don't know.

Praveen Kundurthy

executive
#55

If you see this, the same thing. K-means, the size is 1024. You see the millions of operations per second, and you can calculate the time. And then once you run this, right, so you got the compillation done on the DevCloud on a GPU node. And once you run this, we are plotting the exact results on your here, showing the points -- showing the positions and points -- each point associated, which are to the particular position.

Bob Chesebrough

executive
#56

So what happened there was we sent that code off to the GPU. The GPU computed it. It sent back -- well, it wrote a CSV file with some of the data on that local GPU node. The file system -- the FDS file system copied that information over to us into our log-in nodes we applaud it, and he was plotting that result. So you are actually seeing the code running. It's just gets hidden because it's being submitted to GPU node.

Praveen Kundurthy

executive
#57

So if you -- if you see the comment, the actual script that we are actually sending it to the GPU, right? So this is a script, Python, kernel graph and we are passing the kernel parameters, right? So this particular script file, is being used by this Q script. And this script offloads to GPU, right? A value is GPU here and x script is this particular script, we are passing into the GPU. So we made it easy to the developers so that they don't need to worry about all the qsub mechanisms and they just on the -- you just run this cell.

Bob Chesebrough

executive
#58

if you wanted to -- if you could add a dash I to what he just showed there, what Praveen just showed there. to do an interactive node, and then you would get a login, you would actually logged on to that login -- I mean, not to login, but to the GPU node itself, the CPU that has the GP built into plugged in. And then you could run all this stuff by hand. What you can't do is run the Jupiter notebook there. You can run the Python code there. So that would be another thing if you want to with it and see it run and kind of get more of a feeling on that actual node, you could do that, too.

Praveen Kundurthy

executive
#59

If you see this, I'm trying to get hands on a general and I run this interactively. And if you see this command, as Bob mentioned, I used the minus I command. And I got an interactive GPU node and I can type in [indiscernible]. And you see this device, this node has got an Intel GPU node. It's at a CPU node and FPG platform. If we exit out and let's say, no, I'm actually on the log in node and I say [indiscernible], I don't have a GPU node now. It's just a CPU and FPA emulation. So by default, you'll be on CPU. But if you want a GPU, you need to have a qsub command or you log in interactively. All right. So we can wrap up the session. Let me stop sharing the screen.

Bob Chesebrough

executive
#60

And yes, because we're -- we've got...

Praveen Kundurthy

executive
#61

We did the hands on the rand yes, we can wrap up the session. So we have seen me and Bob covered, right? What are the challenges of programming in world. We've seen how to achieve on Intel CPU and current future GPUs using internal Scikit-learn. And we've seen how we use the Scikit-learn, apply the patches with simple few lines of code and take the extra advantage of what were the stock version of the Scikit-learn algorithms. We clearly did look at the data parallel control module and the data parallels [indiscernible]. We performed code walk through on DevCloud of Scikit-learn algorithms like KNN, K-means, PCA using the Intel extension for Scikit-learn. We looked at the pairwise and the K-means algorithm using the njit kernel decorators on a GPU and CPU, right? And we have seen most -- many of the results there. So with that, we are good to take any more questions. Let's see if we got any questions that we can address here.

Unknown Attendee

attendee
#62

I haven't seen any other since I'll look questions on the code actually seeing the code brand. I think we've kind of caught up with all the questions that I can remember. I don't think there's any extent that we have to address. But if we didn't catch one of your questions, please type it again. Please feel free to fill out the short survey. It doesn't take very much time, but the information that provides us is invaluable. We really crave feedback for this workshop. So please do that. And if there's any further questions, we'll give just a few more minutes for you to ask before we wrap this up.

Praveen Kundurthy

executive
#63

Yes. Thanks,. Thanks very much for taking your time. And the session went a little bit over, like we have 25 minutes over. But thank you very much for taking the extra time to listening to our walk through.

Unknown Attendee

attendee
#64

Thank you, Bob and Praveen for the workshop today. A recording of this presentation will be e-mailed out tomorrow. You can also view a replay right after we conclude using your attendee link. A quick reminder to please complete and submit the short survey. This is your chance to help shape the series, tell us what you want to hear about and what we can do better. Thanks to all of our attendees for your great questions and engagement. You can check out our event website. There is a link in the resource box. This will have a full list of all our upcoming and past workshops and webinars. Thanks so much for joining us today, and we will see you next time.

Read the full transcript via the API

You're viewing the first half of this call. Get the complete Intel Corporation transcript — plus 248,000+ transcripts from 12,000+ companies, speaker segments, AI summaries and full-text search — through the EarningsCalls.dev API.

Get the API View API docs →

This call discussed

For developers and AI pipelines

Programmatic access to Intel Corporation earnings transcripts and 248,000+ others is available through the EarningsCalls.dev REST API. Plans from $24.99/month — full transcripts, speaker segments, full-text search, and the recently-added /api/v1/transcripts/recent polling endpoint for ETL pipelines.