BioNTech SE (BNTX) Earnings Call Transcript & Summary
October 1, 2024
Earnings Call Speaker Segments
Karim Beguir
executiveHi, everyone. Welcome to AI Day. I'm Karim Beguir, I'm the CEO of InstaDeep, the AI company of BioNtech Group, and it's a real pleasure to welcome you for this very first edition of AI Day. So we have an exciting program to share with you today. So first, before we start, obviously, as BioNTech is a listed company, we're going to make a series of forward-looking statements, and these are the disclaimers that come with it, as you probably know. And today, really like we're going to be deep diving into the work we're doing between InstaDeep and BioNTech to really deliver on the promise of AI. So we have a very excited -- exciting program today. So we will start with a few introductory remarks and really like to discuss the work we're doing on scaling AI capabilities, but also provide you with concrete examples of the work that is being done between InstaDeep and BioNTech. Some of the work is actually presented for the very first time. So we're quite excited about that. But before anything, I'm very excited to welcome Ugur Sahin, the CEO of BioNTech, to present. Ugur?
Ugur Sahin
executiveCan you hear me? Yes. Okay. Thanks for the introduction. And Welcome, everyone. So let's get started today. This is -- what we want to do today is not to provide you an advertisement about AI but we really want to accomplish 2 things. At the beginning, we want to make clear why we are doing what we are doing. And the second is why we really need this AI capabilities to be able to do what you want to do. So let's start with the goal that we want to accomplish. BioNTech was founded in 2008 with the goal to change the way how cancer patients are treated. And the motivation for that is a scientific biological motivation. So you all know that the whole industry has spent billions and billion dollars every year into cancer treatment. And the problem is, of course, we are making progress, but actually, vivo cure is still the exception for cancer patients. And the reason for that is the core reason for that. And the root cause for that is that every patient has a different cancer. And this is depicted in a simple slide here, showing how cancer establishes. It is based on gaining mutations in normal cells. And these are random mutations. And so healthy cells acquired mutations. And there are 3 billion locations in the genome where these mutations can happen. And then sequentially, the cancer cells acquire more mutations. And we have 2 fundamental challenges in cancer. One is that every patient has a different type of cancer, and this is even more complicated because we have all our transplantation antigens and recognition of cancer by T cells means that every patient has a different type of immune system. So that is one variation. But the second variation is that every tumor cell within a patient is different. So we call this in the scientific community. We call this interindividual variability. And the second is intra-tumoral heterogeneity. This is known for 20 years, but it is not really addressed by today's treatment. And if you want to address that, it is -- it becomes clear that is an extremely complicated situation. So that means every cancer treatment for every patient is a battle, it's a new battle driven by the complexity -- driven by understanding the complexity of the disease. So that means one question will be or is how can even every cancer set is different, how can we develop treatments that address as many as possible of this tumor set? And the second question is cancer is evolving. So cancer is adaptable? Can we somehow predict how the evolution will continue so that we get an understanding where the treatment not only works, but how the tumor is going to react to that. And given that this is affecting cancer is affected by every year by affecting more than 20 million people worldwide, this becomes now a high-level computational question. So that is something that we want to accomplish to create solutions to address that. Our pharmaceutical strategy to address that is combination therapies. Combination therapy is very simple. One is immune modulators. We know that the immune system is able to recognize cancer, and we are developing next-generation immunomodulators. And we have powerful molecules that can activate and modulate the immune system. We are not going to talk about this. The second are targeted treatments, molecules like antibody drug on new case, where the drugs, the chemotherapy is delivered the tumor and not only the target positive tumor cells are dying, but also the target negative tumor are dying. But there is one additional element. And this is our mRNA vaccines. Our mRNA vaccines provide us the real opportunity to customize to tailor the treatment according to the genetic profile of the patient. And this provides us the opportunity to really ask the question. We have 20 different clones can we develop a vaccine that address this 20 different clones. So this is fast forward, the way how we see cancer treatment in the future. The cancer treatment in the future will be starting on the top, getting clinical samples from the patient, doing the clinicalonics. So that means analyzing the genetic changes in tumor cells. So the data generated here are about 4 terabytes of data for each patient. And that requires really AI and machine learning algorithms to come to the right conclusions. And of course, if we understand the situation, we need to make decisions, and we need to have treatments. And these are our drug toolbox. We have our mRNA therapeutics, including the mRNA vaccines, but also many included antibodies, cytokines and so on. We have engineered cell therapies. So that means engineering the patient cells to attack cancer. We have anti antibody conjugates against new targets. We have T-cell receptors. These are the receptors, which T cells in the body used to recognize mutations or tumor antigens. And we have small molecules immune modulators. Many of the molecules are in variant. So that means they can be applied to many different patients. And this is the way how cancer treatment works today, so that means a treatment of certain antibody is applied to 20% of patients who have this target. But some of them are absolutely personalized is our personalized vaccine. So that means we are combining off the shaft trucks with personalized treatments. And we treat our patients. And in the future, we will not only treat the patients, but we will monitor them and see how they react. So that means we have to combine a few skill sets. So deep genomics and immunology expertise to analyze the patient data. That's what we are already doing today, but AI gives us the opportunity to do that in a much deeper and faster scale. Individualized treatment platforms to address the interindividual variability. We spent 30 years in developing our mRNA pharmaceuticals. We are not going to talk about this, okay, but this really means that we need to have this cutting-edge technology to ensure that we can address different type of targets in patients. And then AI and digitally integrated drug discovery and development. So when we started BioNTech, we had background in machine learning. We did computation on the medicine. But we did that like biologists would do that. And we wanted to do it really in the way how AI researcher would do that. That's the reason why we partnered with InstaDeep, to have really not only some AI capabilities, but the most -- the cutting edge, cutting-edge AI capabilities that have been developed for the specific purpose that we need to address. And of course, in-house manufacturing is another scared. So this was my intro, I would like to now to call Ryan Richardson.
Ryan Richardson
executiveThank you very -- thank you, Ugur, and welcome, everybody. Thank you for coming. I'm going to be very brief and -- but I just want to give a little bit of context to how these 2 companies came together, BioNTech and InstaDeep, some of you may know BioNTech, some of you may know InstaDeep. But what -- how did our pads cross? So just a little bit of a historical perspective on that. So as Ugur mentioned, BioNTech was founded in 2008. And just a couple of years later in 2011, we introduced our first computationally designed mRNA cancer vaccine. And just a couple of years after that, we took our first mRNA personalized cancer vaccine into the clinic in 2014. Incidentally, the same year that InstaDeep was founded. In 2017, we transitioned our personalized cancer vaccine platform to a fully in silicon process, meaning that we use algorithms to select the neoantigens removing human intervention on a patient-by-patient basis. And so far, since that introduction, we have used AI to select thousands of neoantigens across hundreds of patients that have been treated with our vaccine. Our pads crossed in 2019 and when we started project work with InstaDeep and at that point, it really was project by project, but it quickly escalated from there. And in 2020, we formed a joint AI lab where we underwrote a sort of long multiyear commitment in the bio AI field to establish infrastructure, dedicated personnel and a joint vision, that being quickly escalated. And already in 2022, and we, alongside Google and other technology investors invested in InstaDeep Series B, and that was soon followed by a broadening of the AI work across BioNTech platforms, and that culminated in the 2023 acquisition of InstaDeep by BioNTech, and today, we operate InstaDeep as a wholly owned subsidiary based here in London. So 2 companies, 1 mission. BioNTech, over 6,000 employees headquartered in Mines Germany, with a mission, as we were mentioned to harness the immune system to fight cancer and other serious diseases. InstaDeep, over 370 employees based here in London with a mission to productize disruptive AI innovation. And I think what's important here is these 2 companies, these 2 forces have really joined together now under the rubric of one common mission, which is to build a leading AI-first personalized immunotherapy platform and to leverage the breakthroughs that we obtained in the process across the full value chain. And for that value chain, deep dive, I'm going to turn it over to Karim.
Karim Beguir
executiveThanks Ryan. So thank you so much for the introduction. And so today, we're going to be, as mentioned earlier, showcasing the capabilities we have, but also showing you concrete examples of how we are applying those to BioNTech immunotherapy pipeline going from like labeling of like medical samples, RNA, DNA sequencing, proteomics identifying targets, protein design and so on and so forth and also like lab operations. So we're going to deepdive straight into it. And the first part of the presentation is going to be about AI capabilities. As you know, to deliver world-class AI, you need lots of capabilities. You need compute world-class innovation, you also need sort of like platforms that can allow people to use these tools, these powerful tools easily. And so we've done a lot to develop those capabilities at InstaDeep and BioNTech and this starts with compute. So today, I'm very happy actually to introduce you to our new supercomputing cluster, which is coming online in Paris, France, and that's we call Kyber. [Presentation]
Karim Beguir
executiveAnd really congrats to the team who worked super hard to bring Kyber online, which is the case today. And so for this supercomputing cluster, we spent tremendous amount of time designing and optimizing the cluster. Most companies don't do that, but we have expertise in this and so to tell you more about this work and the engineering and software aspects as well, I'm happy to invite Nacef and Alex to come here.
Nacef Labid
executiveThank you, Karim. I'm Nacef Labid. I'm Head of Infrastructure at InstaDeep. So -- let's dive into the specs of this new computing infrastructure. So Kyber is composed of 14 racks, bringing altogether more than NVIDIA A100 GPUs, more than 86,000 CPUs and 1.7 petabytes of first NGM storage -- so this design that we created is repeatable and expandable. So yes, and this new computer infrastructure is near exascale. So it provides half an exaflop of computing power, which brings us to the top 100 -- in the top 100 worldwide of compute and infrastructure and specifically in the top 20 of H100 GPU cluster, specifically. So we put our expertise and it's actually our third deterioration, creating in-house infrastructure. We designed this cluster in a way that it's an in-house design. It's repeatable. It's predictable. So all the regs that you saw on the video are identical, which provides us with less advantages in terms of maintenance, in terms of predictability of the cost of the power usage, the cooling needs in the data centers. So this paves the way for a huge cluster and huge and possibilities of expansion in the future. So it's a consistent design that we created internally validated by NVIDIA also. It's optimized for large AI workloads. And as I said, it simplifies maintenance and expansion in the future. So not only we are -- we've built on our expertise in hardware but also in software, and especially, we built internally our own platform for orchestrate and AI workloads, large massively distributed workloads, and it is called AI Core. So it manages -- it allows us to manage our day-to-day business of training machine learning models from the compute infrastructure to the projects, the user, the security and the workloads themselves. It's fully tailored for our own usage and for our hardware infrastructure, and it's built on open standards. So bringing hardware expertise and software development expertise brings us lots of benefits. And this supercomputer infrastructure is we built it for our engineers, so that they find it available whenever is needed. It's -- you know that it's very hard nowadays to get hold of GPUs and even in the cloud with all the demand happening with LLM. So yes, we have this huge power in our premises, and we'll be exploiting it. So also this design, bringing hardware and software gives us lots of flexibility. As I said, it's fully tailored for our needs. Also, there is no vendor looking in-house design. And we built it in a way that it's -- we can repeat it, we can expand at any time, and we can also provide it to other -- so yes, it gives us predictable costs. predictable energy consumption, whenever we want to expand in the data center, we know exactly where we are going. And yes, it serves lots of capacity management issues. And also, it's cost efficient as we calculated that it makes us save 50% on equivalent cloud spend. So thanks, everyone, and happy to pass it to Alex.
Alexandre Laterre
executiveThank you, Nacef. So I'm Alexandre Laterre. I'm the Head of AI Research at InstaDeep. And now that we have seen how we're growing our computing power at InstaDeep, I want to dive a bit more on how we intend to use that faster for what purpose. But also more importantly, why is it such an exciting moment for InstaDeep and BioNTech. So you might have noticed that the recent advancement in artificial intelligence has mostly been driven by what we call the scaling lows. So the scaling lows, empirical load that's been observed, stating that the performance of modern AI systems such as LLM, large ones models, growth as we increase the amount of data is being trained on the amount of compute and the size of the model, so the number of parameters. In practice, what does it mean is that -- if we scale an existing system here OpenAI GPT3, we scale it at an unprecedented level with more compute data and number of parameters, well, we get more or less more intelligence for free. So we don't need to necessarily innovate from an algorithm point of view, but simply by scaling, you get better performance out of an existing system. However, scaling those do not come for free in a way. It's no small feat to try to leverage the scaling loans because the scale at which you operate is really tremendous. So you will scale very large networks that are terabytes of data, so to train them efficiently, you were going to have to split this model in small pieces and show that spread that across a pool of hardware accelerators, so a very large cluster. Because this is split across different hardwire accelerators, there are networking issues that arise. So terabyte of information will be shared across this hardwire accelerators. So you have to make sure you're not bottlenecked by this communication, you actually spend your time not waiting for information, but can actually perform a useful operation. Also, this new later, they are so large, they perform billions and billions of operations per second. So you have to make sure you squeeze the most performance out of each piece of hardware. So you have to code very low level and be close to the compiler such that you can squeeze the most performance. So this calls for a very advanced engineering solution. And at InstaDeep, we have been using this and apply that to 2 very concrete applications, which I want to dive into now. The first 1 concerns reinforcement learning. So reinforcement learning is a part of artificial intelligence that focus on learning from trials and error to solve an optimization problem. So as opposed to commercial machine learning that use a predefined data set. Here, we reinforcement learning will leverage a simulation engine to turn compute into data. And here, scaling laws supplies as well, meaning that a key ingredient to the success of reinforceful learning is how quickly you can simulate your system of interest and generate data to learn from. So at [indiscernible] developed sophisticated solution for that here. You can see an example of diagram showing how we split it the part that simulate the system of interest and learn from it in a very efficient way across hardware accelerators that also allow us to scale horizontally. So when applied to a concrete product we develop at Instadeep, the results are better, cheaper and faster. Better because as you scale the compute, you can see that the performance of the system goes up, up to 50%. Cheaper because out of each piece of hardware, we really squeezed the most performance, which translates into a linear scaling of the system. So as we scale the amount of hardware the amount of data we can generate growth linearly without being bottlenecked by the communication. And then finally, it's much faster. So here is an example of compared to a baseline legacy system. We can see that as we scale the hardware, we can increase the speed or cut the time it takes by more than 200 times, which means that an experiment that we're taking maybe 16 hours before now takes 5 minutes. So you can imagine how empowering it is for scientists and engineers to have that in-house. It really accelerates the scientific discovery process. So I hope this keeps -- sorry, I have another use case actually. This concerned generativeAI. So this is the latest -- our latest innovation on ported language models. It's actually going to be the topic of the next session. So I'm not going to dive into it. But I want to show you that scaling lows are in action, again. Using our in-house software stack, we can see that by scaling a model from 150 million parameters to 15 billion we see that the training loss decreased as we scale the size of the network. And that's going to translate in better downstream performance on some downstream task of interest. That's what we've been developing in-house and our software stack reach on par utilization of the hardware as the latest Meta Lama 3.1 model. So I hope this gives you a sense of what can be achieved if you combine let's say, an advanced software stack with a growing computing capacity at InstaDeep. We really think that's going to accurate how quickly we can do scientific discovery at BioNTech in this study. So without further ado, Karim, if you want to join?
Karim Beguir
executiveThanks, Alex. And really, I think I want to really emphasize this. Everybody is talking about large language models. And the difficulty in those is actually being able to deploy those kind of like computational workflows at very large scale. And as you have seen, we have the hardware capabilities, but we also have the software integration and the wealth of expertise, having trained those models for many years to be able to do that. But this doesn't stop here. And actually, at InstaDeep and BioNTech, we're doing also lots of fundamental AI research. And this is an exciting area. And so today, we're very happy actually to introduce you to our latest model -- GenAI model, which is a genuine innovation. This is not a GPT style auto aggressive model. This is not a diffusion style model. This is an entirely new concept by Bayesian Flow Network developed at InstaDeep and I'm very happy to introduce Alex and Bora to give us sort of a description. Thank you, guys.
Alex Graves
executiveThank you very much, Karim. So my name is Alex Graves. I'm a research scientist at InstaDeep, and I'm going to be talking about Bayesian Flow Network here. Okay. So a lot of you have probably seen this video before. This was generated by the Sora generative AI model. It's a caption to video model. So it takes a text caption like the one at the bottom of the screen here. and uses that to generate a high-resolution video. And it's really quite remarkable just how much this model has learned about the real world, which you can see from this video. You can see that it's learned more or less about the way people walk about the way light reflects of wet pavement, cities look and so forth. And it's done all of this purely by crunching data. So it's looked at millions and millions of images with associated tax descriptions and somehow learned to build a bridge between them. And so it's important to remember that this kind of thing really was unthinkable just a few years ago. Like this is really -- these are really sort of recent breakthroughs. We've all got quite used to seeing them in the news, but it's important to keep in mind just how sort of game changing these kinds of technologies are. And really the thrust of this talk is to say, well, how can we take this technology and apply it to scientific data, like the kind of data that we care about here. And I think in terms of thinking about scientific data, it's important to demystify these models a little bit. They can sort of look like magic. You type in a prompt, you get back an image, you get back a poem, it just seems like this machine is doing something kind of really incredible. But under the hood, it's basically -- they're basically very large, complicated parametric probability distribution. So in principle, they're not actually that different from the kinds of statistical model scientists have been using to analyze data for many decades now. But -- so -- the question you might ask then is well, if these are just statistical models, then why is it only recently that we've had these big breakthroughs. And the key point here is -- they are models of not just 1 or 2 variables, but millions of variables simultaneously. So basically, they're modeling a joint distribution over very many variables at once. So for example, a generative model of phases like the one that generated this phase here, that's basically trained on lots and lots of images and it's essentially learning the joint distribution of all of the 1 million-plus pixels in each of those images. And when you run this as a generative model, you're simply just -- you're sampling in statistical time. You're picking a sample from the joint distribution and the reason that's difficult is that all of those variables, all of the pixels in this case, they're all interrelated. A face is roughly symmetrical. The color of one eye generally matches the color and the shape of the other eye and so forth. And so if we imagine kind of translating this to the scientific realm, there as well. Of course, we have extremely rich complicated data sets where there's an underlying -- there's a very complex system, and all of these variables influence one another in a complex way. And that is the power that we want to kind of bring from the world of generative AI into the world of scientific data. Okay. And so we don't just -- in general, we don't just want to pick samples from a joint distribution. We want to control what the model does. So in the same way that you control chat GPT by typing in a prompt joint model of images and text. You can get this kind of control quite in a straightforward way. It basically comes down to what a statistician would call conditional sampling. Essentially, you fix one of the modalities and you generate the other. So if you have this joint model of images and text, if you fix the images and generate the text, you have an image caption model, captioning model. that will tell you in text what's in the image. If you do it the other way around, fix the text and generate the image, then you have a generative image model, a caption to image model. But underlying both of those things, those seem like 2 different tasks, but underlying them is the same shared joint distribution. And that's really like the key, let's say, paradigm that we're aiming for here, which is whatever data we've got, whatever complicated scientific system we're attempting to model, we will just try to think of it as a huge collection of interdependent variables and learn a joint distribution over all of them. And then once we've got that joint distribution, pretty much anything we want to do with it later will boil down the conditional sampling. They'll say, let's take some part of the data and fix it because we know what that is and generate another part. And that might sound a little bit abstract right now, but we'll see later in this talk how that can be applied concretely to protein modeling. Okay. So then the question becomes, which model should we choose. And there are several contenders out there. We've already heard maybe the fusion models mentioned for diffusion what was used to generate the video at the start there. There's auto aggressive model. So this is basically large language models, the GPT family. They're all -- they all follow this auto aggressive principle, which is quite simple. It really just says, predict the next token given the previous ones. And then there's more of like mask prediction type approaches, which are often known as burked style models, where you sort of hide part of the data and attempt to predict the rest. Now I won't go into the details of these models, but maybe the key point for our purposes is they all have pros and cons. They're all good for some things and not so good for others. And in particular, there isn't really a model out there that is good for discrete, continuous and discretized data. And we feel that this is an important -- this is really a problem in scientific fields because typically, what you have in science, so unlike when you're just generating images and you have quite homogeneous data sets, in science, everything is very heterogeneous. You have text, you have labels, you have charts, you have time here, you have all sorts of measurements. And if you want to follow this joint modeling paradigm, you need something that can handle all of those at once. So enter Bayesian Flow Network. This is a new sort of type of generative model. The paper -- the original paper was published last year by myself and 3 colleagues, one of whom Tim Atkinson is also a research scientist at InstaDeep and so we've got this basically -- over the past few months, we've been really pushing the development of this model for biological data, specifically for protein data. And it's maybe somewhat of a technical point, but one of the key advantages of Bayesian Flow Network sort of differentiates them from diffusion type models, is that when they generate discrete data, they do it in a continuous way. And conceptually, this is because they operate not on the data itself, but rather on a set of beliefs about the data and even for discrete data, the beliefs that the model has about the data can be kind of encoded in a continuous way. But basically, in practical terms, what that means is we can use advanced gradient based sampling techniques to do conditional sampling across all of these modalities across discrete continuous discretize. And so taken together, we think that this is really an exciting opportunity to -- it's really a good model for this task of jointly modeling a wide variety of heterogeneous scientific data. So that is a sort of very high-level overview of what Bayesian Flow Network are. And now I'm going to hand back to Alex, who's going to talk you through a little bit how we can actually specifically apply these to protein data.
Alexandre Laterre
executiveYes. Thanks, Alex. So indeed, I want to share with you how we see and how we started deploying Bayesian Flow Network to scientific data. And here, our vision is very clear. We want to propose a unifying framework across modalities and data type, meaning that at any time, we would collect all the data we can we can put our hands on and train a joint distribution across this data type and modalities. And at inference time, we would kind of prompt the model. We would query condition sample from it to solve a specific task, right. So perhaps here you have as a way of example, are showing proteomics on the slide and some modalities that are relevant in that context. Obviously, the sequence of a proteomics is well defined by the sequence of amino acid, which is simply a sequence of discrete variables, right? You might hear also about the structure, right, because that's highly coated with the function of the protein. But you might also have automated data you care about predicting or conditioning your model on, for example, the good terms annotation if you care about the function of proteins, the species or the organism protein comes from. Perhaps one step further, you have access to experimental data from the lab and you know, for example, on which antigen and antibody binds to. So you have all this information, and we want to unify them and train a single model that can then be conditional sample at test time to source specific task. And talking about this task, if we take for instance, a sequence that you want to predict the structure that's actually the protein folding problem, which actually motivated AlphaFold, which completely revitalized the field of artificial intelligence for biology. Perhaps you care, like I mentioned, about the function, and you want to predict the go terms based on the structure and the sequence of one step further, you know an antigen and you want to generate the structure and sequence of an antibody that tightly binds to it. So that's our proposal. Each of the tasks could be sold with the task-specific models, and that's generally the approach that people take nowadays. But we want to propose this unifying framework that's going to train a single model across all the modalities and then you can prompt it differently at test time based on your task. But if we just take one task in particular, just sequence generation is not even clear. As Alex mentioned, which type of approaches are the most suitable. You might want to use auto aggressive models if you want to do denovo generation, so generating 1 minerals at a time in a sequence. However, there are of limited usability for conditional generation. So perhaps you want to generate the CDR loops based on the framework region of an antibody, their auto aggressive models are not amazing. You might want to use a mass production model instead, but that's not good at denovogeneration. Finally, you have what we call the discrete decision, which are great in the continuous viable case. But for discrete variable, as Alex mentioned, they require much more engineering. It's very difficult to apply. So first proof of principle, let's say, we want to apply Bayesian Flow Network to this problem of sequence generation for proteins. And cutting straight to the results. Here, I'm going to present the results we get for our ProBFN model, which is simply been trained on a very wide variety and large database of proteins. We compare here to alternative approaches. So for GPT is an ultra-aggressive model. EvoDiff is discrete decision model. And we see that the performance are much improved, like in terms of naturalness. So the protein that are generated by ProBFN are much more likely to have been found in nature. It's also much more diverse. So the protein generated cover much nicely the let's say, the realm of known proteins. And finally, they are highly novel. So we don't memorize the data set, but we can generate novel sequences. And that's tremendously important because as a scientist, when you do antibody design a protein design, you want to make sure you focus your attention on the region of Intervest. So these AI tools should help you focus your attention on the key regions that might be relevant for you. And what relevant means is it should be novel, diverse, yet variable and natural. This is the goal here for this PlugBFN algorithm. Now we went one step further. We started folding the sequences and look at the 3D structure of the generated sequences because it's highly correlated again with the function of that protein. What we observe there is really interesting as well. We can see that some of the structure generated have 1 domain, but the vast majority has 2, 3 or even 4 domains. And that tells us that Plug BFN generic proteins that have let's say, coherent interaction between the regions that are far apart in the sequence space but that will close in the 3Dstructure pace. So PBF can generate proteins as a whole and not just local structure that makes sense. We also look at plenty of different annotation where it comes from the 2 of life. And there again, we see that we nicely cover the species, the type of structural motive and so on. We summarize all the spending in a paper. We just actually released for this event. So please have a look. We'll also look at how to do zero shot conditioning of the CDR region based on the frames for antibodies. So please have a look it's online out there. But we actually haven't stopped there. We carry out, like I mentioned, our goal is not just to model the sequence of amino acid, but we want to model all these modalities in data type, right? We want to move away from the concept of you have one task for simple protein folding, you would generate -- you would create a data set to in a model and then have scientists, let's say, access their model and front it. We want to put the model first. Gather all the data you can find and put your hands on train a single model. And then based on the task of interest, you would prompt the model differently, right? So that's really the objective. That's what we then started doing. And we took from -- we started from the ProBFN model, and we trained an antibody-specific models, that is highly multi-model, meaning that we also model a lot of the biophysical properties and so on. And that brings us really at the next step. It's really in our next-generation model. So as opposed to showing you more baseline and metric, we thought it would be useful and more interesting for you to see how we can use this model in practice. And that's why we'll be calling a Bora on the stage to show you how to use our next generation of protein language model.
Bora Guloglu
executiveFantastic. Thank you, Alex. So yes, I'm Bora, I'm a research scientist here at InstaDeep, and I will be introducing [indiscernible] is our first multimodal model for antibodies. And what we've done here with [ Abiyafenix ] is really model all of the attributes that we think are important about an antibody. So we're actually modeling 36 different data modalities or data modes, covering sequence, genetic and biophysical properties of an antibody. And our aim really is to kind of hark back to what Alex was just saying to create a model that is flexible in the tasks that it can achieve. And this really is so that we can empower scientists with tunable flexible generation depending on the scenario at hand. And the way we do this is essentially, we don't think about the antibody as just kind of one big entity. We actually look under the hood. So -- while we normally have the kind of FV region composed of the BH and the VL with the CDR loops, there is actually a lot more kind of that's going on here. And we essentially unpick this, we undo the stitching and model all of this in one big joint distribution. So we break the sequence down into 14 different attributes. We cover genetic lineages that contribute to the actual sequence of the antibody including species but also specific genes. We cover biophysical attributes that we know correlate with developability of an antibody as a therapeutic, and we also are interested in the length of the different regions. So we cover this as well. And what this then allows us to do is really get this flexible approach to modeling the antibody sequence space. So what we can do with this, for example, is say we have this antibody here. We're happy with it. But we think that the 3 loop that the green one is too long, and we want to shorten this and redesign it while keeping the rest of the antibody. We can do that. We can condition the model, so fix all the data modes that correspond to the regions that we want to keep. We can also request a specific CDR-H3 length and then ask the model to generate our samples. And we get exactly what we hope for. So everything stays the same, but we have a redesigned shorter CDR-H3 loop. That's obviously a very simple toy example. If we actually look at a scenario that is more akin to what we might see in the lab, we have here the generation of a library of anti-HIV antibodies. So these are antibodies that recapitulate the properties that we find in a common class of anti-HIV broadly neutralizing antibodies. If we were to design this library in the lab, we would have to go through a multistep kind of design process, but with AlphaNPI-X, we can essentially compress this into one. What we're interested in specifically is the HP genetic lineage, the light chain locus, the species, the L3 length and also developability. So we condition the model on all of this, and we just search through the space within the confines of this. And what we find is that AlphaNPI-X is really good at effectively searching through the space. So when we compare the rate at which we find full hits, we are actually more than 5,000x more effective than just looking through natural repertoires. But we also know that these antibodies are still diverse. So when we're searching through this space, we've essentially effectively set the H1 and H2 loop sequences and structures. And the L3 loop links. So we can see that, that all looks very similar. But we've not conditioned the model at all on the H3 loop. So here, we see remarkable diversity in sequence length, and also structure. And that's precisely what we want when we design a library. You might have a completely unrelated task, heavy and light chain pairing. So in this case, we might be starting with a heavy chain that we are really happy with, but we want to find other light chains that are able to pair with it. And essentially, this time, the task is different, but we can use the same model just by changing how we query it and how we interact with it. So this time, we would condition on the heavy chain sequence. Maybe again, we're interested in the light chain locus and the species and also the biophysical attributes. So we can condition on this as well. And in effect, what we're doing is we're taking the heavy chain and we're looking for solutions. So we're looking for light chains that are able to pair with this heavy chain. And what we find again is this diversity. We find light chains that have different lengths and sequences and shapes -- but we know that the model still respects the conditioning information that we provided. So the heavy chain sequence always looks the same. And the light chain actually shows sequence biases and length biases that are consistent with that heavy chain. And that is exactly the same model. We've not changed anything about the model. It's just the way we interact with the model and because the model has learned this rich distribution of the underlying data. And with that, I will ask Alex just back to summarize what we talked about. Thank you.
Alex Graves
executiveThank you. Thank you, Bora. So yes, just to very briefly summarize. I've introduced the sort of new class of models, Bayesian Flow Networks. And hopefully, motivated why it is that we think they're a really good choice for the kind of data that we're looking at here. We've talked about some of the early results we've already got with protein modeling and how promising they are -- and what you've just heard from Bora is -- really gives you a flavor of just how kind of tunable and steerable these types of models are. There's so many things that we can do with them. And of course, we're really excited about where we go from here with this. There's lots of data we still haven't tapped. And so with that, I will hand back to Karim.
Karim Beguir
executiveThanks so much, Alex. And really congrats on the amazing results. This is fantastic. And yes, Ugur.
Ugur Sahin
executiveWhat we are doing here is amazing also because we are -- with our models, we are more closer how nature works in the way of the independence of optimizing domains of keeping some domains constant and changing some other domains. So I believe with this type of optimization of our models, we are more likely mimicking what is happening in nature.
Karim Beguir
executiveAbsolutely, Ugur, and what's exciting is indeed with this model. We can do conditionality of any variable against any other variable. Obviously, here, we've sown it in a few examples with sequence and structure. But this is really like, we believe, a profound sort of like breakthrough, and we look forward to working more and importantly, deploying it also in the lab with our BioNTech colleagues. But I think for me also like what is refreshing is this is not all about scale and LLMs, there is still room for fundamental research. And InstaDeep and BioNTech, we have 45 AI researchers inventing new algorithms. Now algorithms compute. That's what we've seen. But importantly, we need to deploy and make those breakthroughs, those powerful models accessible to all our colleagues internally, but also externally. And so for that, I'm very happy to invite Arnaud and Julia, who are going to introduce us to the DeepChain platform.
Unknown Executive
executiveThank you much, Karim for the great introduction. So hello, everyone. I am Julia and I'm a product manager at InstaDeep and I've been leading the product development for the DeepChain team. So you've just heard from super advanced innovation. These BFN models are incredibly exciting, like Ugur just mentioned. And what I think is super unique about InstaDeep actually is that we're able to combine state-of-the-art research with the most advanced engineering to enable the delivery of AI tools that can directly integrate into the R&D pipeline that you see here. And so I'm going to run you through a couple of examples of these tools. In particular, BFN, so you've just heard about BFN. These kind of models can be used for de novo antibody design, but also can integrate into the optimization part of the pipeline. Beyond VFN, we also have tools in the world of genomics in particular, we have created models called the nucleotide transformers, which can be helpful for prediction tests such as spicing or gene expression and today are supporting BioNTech in the identification of new targets. Finally, one super exciting application has been the development of assistance. We go to the next step. I don't know if this is working. So the development of assistance, which can be used as fenalone AI tools that can support scientists with natural language so you can directly interact with them, and they can process tools in the background, but they can also be integrated directly into the labs, like Karim was mentioning, and we'll see later today a demo about this. So today, we're super excited to be upgrading DeepChain and to be launching the next generation of models in the platform. DeepChain, in particular, is a single platform that combines our AI expertise with our life sciences in order to empower scientists to deliver the next generation of novel therapeutics and other biotechnologies. As a first instance, we are releasing our flagship model. So the ones you've just heard about the state-of-the-art generative models are going to be on DeepChain. These are super exciting models because they enable the generation of sequences that are very natural lag. They are very structurally coherent. And as you saw, they can generate sequences upon certain conditions and certain parameters that the user can decide. So that opens up a lot of opportunities. On the other hand, we are also releasing our nucleotide transformers and segment NT. These are our DNA foundation models and specifically segment NT can be used for predictions at the single nucleotide resolution. And that can be extremely valuable, especially for applications such as spicing whereby, for example, if you take a DNA sequence, a single point mutation on the sequence can actually have a lot of impact in the downstream process such as transcription. And so it is vital to get that kind of resolution in our predictions. Beyond that, we're also -- we've also been training our models with contact lens that go up 50 in length and also with our performance drop. And finally, even though some of these models have been trained with human genomic sequences only, they can also generalize across different species in a zero-shot manner. Diving a little bit deeper into these NT models. We've consistently shown both through our studies that have been peer reviewed and published in major journals, but also through independent studies that have taken our models and benchmark them we have seen how we consistently outperform other models in the space. But we're not only state-of-the-art. We also have an enormous amount of traction in hogging pace. In particular, we're one of the most downloaded genomics models in the world. And as of this morning, there's been more than 700,000 downloads across model sizes. And so that opens up a wide range of opportunities and specifically the opportunity to go out there and speak to researchers and ask them about which applications they're using our models on, but also understand more about the challenges and the pain points that they're facing. And this is why today, this has inspired us to not only release our new AI models but also release capabilities. Capabilities for our users to take these models and build on top of them and then scale them for their own applications. In particular, our first capability is around an optimized setup, so we've heard from Alex earlier today that our models are becoming and models in the LLM space, they're becoming bigger number of parameters, but also in the amount of compute that you not only need to access, but also orchestrate. And so on DeepChain, we're doing that for you. In particular, you can now access our very optimized workflows with a few simple lines of code. We have already shown how this kind of setup is already delivering value in a specific application for a silicon design of regulatory sequences, whereby we have been able to increase the inference speed up to 7x and reduce the cost by half. And this has been super valuable for the team we've been working with because previously, they had been trying to integrate these kind of models in production and had been struggling due to these kind of running time requirements. And today, they'll be able to do that with DeepChain. A second capability we're releasing is the ability for our users to customize our models. And so now they'll be able to take one of our models and customize them with their own data for their own application. We also have an example of this. So we've been testing around the fine-tuning of our model in spicing use case. And so in this case, we have seen how compared to an external implementation, we were able to improve performance by at least 1.5x. Finally, one last capability that actually Arnu will be introducing in a second through a demo has been the release and development of assistance. And so these assistants are tools that you can interact with through natural language, and they can support you in connecting multiple tools at that we're developing here at InstaDeep. So now I think we'll be jumping on to the demo in a second. I don't know if you're able to switch the screen amazing. Okay. So allow me to introduce you to the outgraded teaching platform. We have here the model page. We can get access to different models. For example, we've got the DNA models here, the segment NT that we've been discussing. Here, we have the specific information about its data sets and different training parameters. We've also got our protein models that have been introduced like ProBFN and BFN. And finally, we can jump into these fine-tuning capabilities we've been discussing. In particular, you can start and you run by clicking on this button. You can now select which model you'd like to use for fine tuning. We're going to be adding a name here, AI Day for this run. We'll be selecting the downstream task for gene expression. And now we'll be uploading our fine-tuning data. So in particular, for fine-tuning, you need file the contingent your sequences and then you have your training labels. In addition, you can also upload your validation data, so both your sequences and your labels. And then down here, you've got some parameters that you can fine tune, and this is actually the part where scientists can experiment and has different parameters, different type of data. So biologists are really the experts in the data. While we can handle the compute and the orchestration on our side. And so here is where they can actually start experimenting. So now we're going to start the fine-tuning run. The data has been uploaded. And now we can go on to the runs and actually keep track of this processing around here that we just started on AI Day. So what we see here is the different parameters, and the plot is being updated live with the results that are being processed. But because usually fine-tuning takes quite a few hours to be able to achieve good results. We're going to use a run that I started actually earlier this morning that succeeded and also model for AI Day. And so here, we see how -- so the pubs have been completed, and we're able to use this model now for inference to be able to show how it works in practice. So now we just added the model to our list, so it's the one here. We're going to be jumping on to the CLI for you to see how we can run these models. So here, I'm going to start by typing DeepChain models to see the list. So we just got the list of DeepChain models. In particular, this is the 1 here we'll be using today. And then to run our models, we're going to be typing DeepChain run, then select the model that we just moved to our list, then referencing our sequences that we want to use and our output file that we will be sending the predictions to and we're going to ask it to wait, so we can see what's happening. So now what we did is just send some sequences to our systems and what's happening in the background is because we're using a fine-tuned model for gene expression we're going to be seeing for each sequence that we've sent through, we're going to be seeing a few values predicted for gene expression and the different tissues. So we just got back the results. Here there's a snippet of those results where we can see the GID associated with these different values of gene expression. And now we will be using -- so this is where we can see the full results. So we'll be using this file to be able to evaluate the results. We have a function here on DeepChain where you can evaluate -- you can have DeepChain evaluates, yes. And then we can reference our results that were just computed and the list of labels. Sequences under score labels is where our labels are. And so now our results pop up here. We see the performance improvement of our fine-tuned model that we've just showed case compared to a baseline. And so you can see -- as you can see, if you really pay around with those parameters and your data sets, you're able to achieve improved results with our models and customize results to your application. So now I'm going to be passing it on to Arnu on dark mode. Go ahead.
Arnu Pretorius
executiveThanks, Julia. So we're going to move on to our AI system. And my name is Arnu Pretorius, I'm a research scientist at InstaDeep, and I'm delighted to introduced to you today our AI agent called Layla. Layla is deeply integrated into all aspects of the DeepChain platform, including running models, performing analyses, calling internal as well as external tools and much, much more. So to showcase some of Layla's capabilities, we're going to run through a hypothetical scenario of a user who's new to the DeepChain platform and interested in analyzing DNA. So I'm just going to clear here at the slide, it's clear this and we can go to Layla. So being new to the platform, we might begin by simply asking Layla which models can I use for DNA. Layla then knows to call the internal API associated with the DeepChain platform and give us a list of available models that a user could be using to analyze DNA, segment NT, a multi-species version of segment NT and a whole host of other fine-tuned models based on the nucleotide transformer series built by InstaDeep. Now being new, this user might not be familiar with these models and might be interested in knowing what they can actually do. So we can ask Layla what is SegmentNT used for Again, Layla called the internal API, but now fetching information specific to segment NT, giving us the reply saying segment NT used for detailed genomic analysis offering single nucleotide resolution predictions for various genomic elements. So this might sound interesting to us, so we can go one step further and try to do some analysis. So if I just pull up a DNA sequence file here, I can upload this to the platform, and you'll see this under the attach files being shown over here. Now I can tell Layla that I have uploaded a sequence and target using the @ symbol. Once this is done, I can then ask Layla, can you please segment it for me? Layla now knows that it can call the custom segment NT model and give us input the DNA sequence that we've just uploaded. It provides us a segmentation text output saying that at specific indices within the sequence, we can find exons, introns slice doners, splice acceptors as well as UTR elements. But not only this, this gives a scientist a quick sort of overview and a feel for what's going on, but they might want to actually build a much more custom pipeline in Python, some sort of larger scale batch to workflow. And as Julia showed, we have the CLI for this that can really be useful. And to get started, Layla provides us with a command, we can actually directly copy and paste to run this analysis through the CLI. So I'm just going to pull this across and here, if we just quickly look, we have this DNA file here that we uploaded, and we can simply pace this, knowing that this is exactly the DNA file. We want to analyze with the output being specified. When we run this, the CLI goes and fetches the job does the analysis by calling the model and provides us with the output. So this takes a few -- just a few seconds. And here, we can actually take a quick look at this output file. And what we can see is it has a whole host of probabilities of certain regulatory elements being in certain places. And even though this is not nice to look at through the terminal, but in a pit in workflow, you can very much use this easier for further analysis. But we can now go back to Layla and simply ask if we just want to again have a sort of overview of what's going on, to visualize the results for us. Can you please lot it for me. Layla knows it has an associated plotting script with segment NT and can run this to give us an interactive plot of what's going on. So on the Y-axis here, we see certain regulatory elements and on the X-axis is the nucleotide positions in the sequence. And what we're seeing here is probabilities of certain elements being present at certain positions, and we can actually zoom into specific regions and see that here, for example, nucleotide A in position 1490 has a probability of 92% of being associated with a tissue and variant promoter. And this can be a nice tool to really improve the productivity of biologists to quickly get going with the DeepChain platform. So I hope that shows you some of the capabilities that Layla provides, but we'll see much, much more later on. So just to conclude. We have a whole series of Layla models built on top of Meta Llama 3.1 and they come in different sizes, including the 70 billion and the 405 billion models, all internally fine-tuned by InstaDeep. And we want to stress that Layla is more than a chat bot. It has expert knowledge of biology, integrated with powerful tools and the capability to really reason and make decisions as well as learning through constant feedback. Thank you very much, and I hand back over to you, Karim.
Karim Beguir
executiveThank you guys. Thank you so much for the live demo. And as you can see, those tools are super powerful because if you're going to waste time as a scientist or biologists to code those models, this is maybe not the most optimal use of your time. Everything is available on the DeepChain platform. And so today, we are actually releasing the upgraded version of DeepChain with our most powerful models, including BFN including the functionalities that you've seen with Layla. These are all available on the DeepChain platform, which is made accessible both internally within BioNTech Group, but also externally. So we are really happy to partner with you. So this concludes the first part of the presentation. We've seen what we've done into building our supercomputing capabilities, the AI innovation that is coming from our research teams and also DeepChain as a platform to make all these innovations available and sort of quickly accessible. So we're going to jump into the second part of the presentation, which is really about looking at concrete use cases showing you how we work within BioNTech Group -- InstaDeep and BioNTech colleagues working together to make this progress, which is actually important because this is really where rubber meets the ground. This is where we drive innovation that ultimately can translate into saving lives. This is the spirit. And so with that, we're going to start with histology, and I'm very happy to introduce Youssef to tell us about the work he's doing.
Youssef Ben Dhieb
executiveThanks, Karim. Hi, everyone. I'm Youssef Ben Dhieb. I work as a senior machine learning engineer at InstaDeep. And today, I'm going to walk you through some of the cool stuff we are developing with the histology department at BioNTech. So the histology is a core component of the immunotherapy pipeline. And one of the critical tasks that pathologists need to do at this stage is labeling digital slides of tissue. And as we scale up, pathologists are facing a heavy and growing actually workload. And to understand more of this challenge, let's take a look at a typical histology image. So notice how we can zoom in from the broad details -- broad overview actually to the cellular details. Where we can see the details of each individual cell. And these images are actually very big and have a lot of details, but you need like a full tennis court filled with 4K screens to be able to visualize a single image with all its details. And labeling these images manual lead is very time consuming and requires a lot of attention at every like magnification scale to label the image. So how to solve this challenge? The idea here is to harness the power of AI and develop AI tools that will allow pathologists to become more efficient and very fast in labeling this. So allow me to introduce the first tool that we developed, which is the AI-assisted tissue annotation tool. And this tool actually is a collaboration between the AI and the pathologist where the precision and the speed actually of the pathologist is enhanced. And we can take a look at how this process actually work or how this tool work. So first, this is a mutation when a pathologist does it for just simple to red cells when you do it manually. It's quite precise, but it's a bit slow. Now when you use AI for that and use our tool, you just draw a box around these cells, and it will automatically segmented in no time. We can take, for example, another region, like the white region here, the background. And with just drawing the box, it was segmented. Now if you take a more complex region or area and the model doesn't recognize it, you can quickly like click on the area you want to exclude or include and the AI will understand that, and it will correct itself automatically. So by deploying this tool to our pathologists, actually, we're able to achieve a fivefold increase in the speed of the pathologist -- and at the same time, we didn't lose like the quality, but it's quite the opposite. We actually got a better quality because pathologists were able to even refine the annotations at different levels of magnification. Thank you. So based on the success of this tool, we actually developed a second tool, which does the segmentation of the whole slide image and actually, it does it in just one click. It segments the whole slide with all its details. And how we achieve that, we actually use a state-of-the-art vision foundation model. And then we decompose the hold slight image into small patches, and we transformed the problem from segmentation into classification of patches. So like it's -- each image is decomposed into like millions of purchase, I think, like 5 million patches per slide. And we are processing them in parallel like hundreds of thousands in parallel. And with that way, we can very quickly like undertake the full slide. And then when you group them back, you will see it as a segmentation instead of classification. So you can see here when we zoom on our region and we try to classify like the patches -- patch by patch. Then when you zoom out you actually see how the tool is progressing and segmenting all of the slides in no time. And actually, thank you -- actually, with this tool, we are able to achieve 100x speed up compared to the first -- compared to the manual allocation actually of the of the annotation of full -- like whole slide by pathologists. And yes. And here, my colleagues will show you more good stuff on the other stages of it.
Karim Beguir
executiveThank you so much, and I hope that this shows you like the power of AI vision tools. to help. And this is increasing the speed of data accumulation, which is also useful in many ways within BioNTech and beyond. And so after actually handling like medical tissues, classifying the different areas you do have naturally DNA and RNA sequencing. And as hinted before, we've developed like very advanced models in genomics, and I'm very happy to have Thomas, Marie and Maren present the work we're doing in comment on that.
Thomas Pierrot
executiveGreat. Thank you so much, Karim. My name is Thomas Pierrot. I'm a research scientist at InstaDeep. Very excited to be here. And I'm meeting the team behind this segment entity and nucleotide transfer model that already presented. And what we'd like to do now is to have deep dive into what's happening on the road and also explain how we actually average these models in the immunotherapy pipeline. And as we mentioned, we are pretty excited as this model gets a lot of traction. They are now pretty popular in that sphere. And probably the best way of speaking about this more is to think about ChatGPT, Emma, Gemini, any of these big language models, but instead of joining them on English or on any other language, we actually turned in on DNA. And here, we average the exact same technique called [indiscernible] supervised learning, which is an amazing technique because it's as you to run from any data type of data without needing label. And in this case, what we do is the tractor correct a lot of genomes can correct genomes from lots of different individuals, genome a lot of different species and train these models almost in a never-ended fashion out of these genomes. And as Alex mentioned, scale is key. So we first scale the data by going through lots of individual, lots of species as much genomes as we can, but we also scale the models to billion parameters. And probably the most impressive results is that even though these models have been trained with resourcing any knowledge reducing any able about DNA, they actually acquired some of the genomic knowledge during the training. And what I'm showing here on this slide is we're looking deep-diving into the activation to the barriers of one of our biggest mode, the 2.5 billion parameter, [indiscernible]. And we see that actually the model acquired some basic genomics knowledge training. In the first year, you can see it can already make the difference between coding and non-coding regions, but what's even more impressive is that as we go through the areas, this kind of representation becomes more granular, and you end up in the final area with a very granular presentation with lots of different elements where you can capture, for instance, UTRs. And that was to explain why then this model can be fine-tuned as you explained to solve lots of different tasks in genomics with a very high precision. But we didn't appear with the team and say, okay, so far, we took inspiration from NAP to be this first generation of model. And so why not to be the second generation take inspiration into computer vision. If you look into computer vision, people know are training very impressive foundation models through segmentation. I'm thinking about anything models from Facebook. And the way they were they simply take all the images they can find, they actually train the model to a segment on the image everything they can find like to understand the different elements in that case can be the cars, the people, the road and so forth. We do the exam section on DNA, but instead of working in 2D, we work in 1D, and we train our model to segment a DNA sequence and to find in the sequence of the genomic segment it can do. To do that, we worked out with the team together a very high-quality data set of millions of annotations over the human genomes, but also lots of other species. we looked at many elements. I have a few example on this side, routines had [indiscernible] a seat promoter and an answer that's going to regret the expression of genes, but also cathodes of genic elements. And refining this nucleotide transform model that we presented on this annotation to bring them to a new level. And my colleague, Marie is going to give you more information about the performance of these models.
Marie Lopez
executiveThank you, Thomas. So I'm Marie Lopez. I'm the genetics in charge of the AI applied team here at InstaDeep. And actually, I think that Layla been doing a great job explaining the performance of the segment NT model. However, I'm going to give a fuller picture here. And what we see is that for each nucleotide model is able to predict the probability of belonging to each of these genomic element classes here. And of course, we have some annotation that are related to function or critical functional elements such as protein cottages, but the model is also able to actually predict elements that are usually more difficult to map such as regulatory elements like enhancers and promoters. So to give you kind of a visual explanation about what this model is doing, you can see here on the top of the DNA sequence of 50,000 base pairs. And you would see like 3 different genes that are encoded here. And what the model is doing, as you can see in the line -- in the first line of the truck is that it's accurately predicting the protein cutting genes under those. And if you look in the other trucks, you can see that the model is able to differentiate transform Exxon, so getting a better accuracy at detecting gene architecture, for example, detecting spices in white as well and also detecting more progressive in answer, as you would expect in this DNA sequence. And what is even more impressive is that all of those together, prediction represents 700,000 different probabilities and the model is about to a put them in actually less than a second, which is an incredible speed and precision delivery that is given by this model. And this is actually an incredibly powerful tool for genomic annotation and research. And Maren is going to expand a bit more how this is actually used inside of the pipeline.
Maren Lang
executiveThank you. I'm Maren Lang, Senior Director of informatics Research and Development, BioNTech Mines. And I would like to show you one example application of the segment NT. And this is about alternative splicing. Splicing is the event where the parts that are coding called Exxons splice out of the pre-mRNA the entrance, the entrance, which are the non-coding parts are spliced out and the exons are joined together, Exxon's are the part of the coding parts. And this is pricing and -- but not all exons always are part of the final MRN and therefore, the protein. So that is the process that is called alternative splicing. And this is a normal process taking part in each healthy cell. And here we showed unhealthy data that we can detect those events with the segment NT. And here, we see the splice donor accept simply not that it's simply the task from Exxon to intron or intron to Exxon, which parts of the splicing events are addressed here. And we show here that the segment and he performs much better than the splice area, which is a state-of-the-art tool. And related event. That is simply the extent interim detection is also addressed by the segment and, as you just learned, it can predict many different tasks. And the splice AI is performing worse here because it's an indirect task and -- but it can be addressed as well. So we are much better for splice event detection, but also for excellent intron detection. So this has been done on heavy cells, but we know that alternative splicing is an event that is very complex and can be easily disrupted. And this is what -- therefore, it is associated with cancer and many other diseases. And we wanted to see also on cancer data, whether we can detect these events. And so we checked -- we fine tune the segment NT to detect to detect tumor antigen candidates, which present possible targets for the -- for immunotherapy and yes, this is what we did here. And you can see here, I have no point, sorry, that the segment NT also for this task performed much better than any of the other tools. In any event here. And this is great news because we can use the segment NT also for this task. And yes, this shows that we can use it for the alternative splicing prediction. And you can also combine it with other methods that will be shown now by Nicolas, Daniel and Mike, who will show us AI-enhanced protein.
Karim Beguir
executiveThanks, Maren. So really exciting application of segment entity. And like Maren said, really like we are deploying AI end-to-end on the pipeline. So we've seen the visual AI modality, learning from pixels. -- here, we are learning from nucleotide sequences, but we're also working on protein and proteomics and bringing in multiple modalities together here. So we're going to see an example of mass spectrometry and how we can use AI to identify potential targets with Daniel, Mike and Nicolas.
Daniel Rothenberg
executiveWonderful. Thank you, Karim. My name is Daniel Rothenberg, and I had the pleasure of representing the proteomics team from BioNTech. And today, myself, along with my colleagues, Mike and Nico will be talking to you about how we can use AI to supercharge our target discovery efforts. So we'll start with some basic immunology first. intracellular proteins are processed and presented on to MHC complexes. And at BioNTech, this is important for a couple of different applications. First are the T cell targeting RNA vaccines, where the RNA enters the cell, is translated into a protein. And then that protein can be processed to presented into epitopes onto the MHC complex. Second, and in the context importantly of oncology is looking for tumor-associated antigens or TAAs, and these are proteins that are expressed specifically in cancer cells but not healthy cells. And just like all proteins, these proteins are also processed presented on to epitopes such as to MHC. Now the field has kind of coalesced around the same set of TAAs, and you can think about PRAM, your Majes [indiscernible] and HPV-derived proteins. These are all important, but the problem is that these really -- if you focus on just these targets, it limits the breadth of the therapeutic population in terms of the population as well as the number of disease indications. And so discovering new targets is of the most importance to broaden the range of the population that we can treat. And epitope presentation is so important because MHC presented epitopes are the immune systems window into the intracellular Protium. And so in order to get a therapeutic response, these presented epitopes must be recognized by T cells and then these T cells have been activated and that gives you your therapeutic benefit. So how can we validate what epitopes are presented at the cell service. Well, being from the proteomics team, I love mass spectrometry. I'm biased and so I think it is mass spec. And indeed, mass spec is the current state of the art for detecting, identifying and quantifying MHD predicted presented epitopes. However, the challenge is that mass specs don't just spit up peptide sequences rather they give you a mass spectrum, which represents a biophysical fingerprint associated with that peptide, but it's not the peptide sequence. And at BioNTech, we have a massive database of mass spec validated MHC-band epitope peptides for studies that we've performed internally as well as publicly available data sets that are external. And so this database has over 200 million spectra in there, but these spectrum needs to be decoded into peptides. And so using commonly available heuristics we're able to run these through search algorithms, and that leads us to about 1.8 million peptides in our database. Maybe these peptides can further be maps onto specific genes, and that gives us whether they're useful or not. And so a lot of these hits are not particularly interesting because they're not tumor specific. But also in our database, we do have tumor-specific TAAs, such as frame, such as made, some of those other ones I talked about before. However, you can see that we still have a lot of spectra that have gone unmatched using these basic heuristics. And so where we can turn to AI here is to supercharge our search and find new peptides using novel search algorithms, and that would find novel targets that can expand the population that we can possibly treat. And so for more details on that, I'm going to turn it over to my colleague, Mike.
Michael Rooney
executiveI'm Mike Rooney. I lead the computational biology team at the Cambridge, Massachusetts site of BioNTech, and I've been working closely with Daniel for the past few years to get the most at this data set and use it as a tool for target discovery. And one thing we realized early on is that we need to bring in AI-based methods. And the key application here is to use the AI and to help us validate whether our peptide identifications are correct or not. So there are 2 key examples of how we do this. The first is in the upper left, where we're looking at the retention time of our peptides. So we can get very high accuracy predictions of the retention times. And so if we see any peptide that deviates from our expectation that we know pretty likely that's a false positive identification. In a similar spirit, we can look at the way the peptides fragment in the mass spectrometer. Each peptide it goes in and sit with high energy, it breaks into pieces. We can predict the intensities of these fragments. And if we can match that predicted the serve fingerprint matches the predictive fingerprint strongly, we know we likely have a good idea. Otherwise, we know it's probably incorrect ID. So we can do this one by one, but really practically what we do is we run this across the entire data set. And we use it as a way of getting deeper and getting more confident and vacations. So running this on our data set, we can see up to a 200% increase in the number of peptides that we can recover per sample. So how do we use this data? Well, what we're really interested in is what genes are producing more peptides and tumor samples than normal samples. So we've evaluated the systematically across all 20,000 genes in the genome. And on the plot on the right, you can see the count of how many times each genome seen in normal samples versus in tumor samples. And we have this really interesting population of genes in the upper left, which are seen many, many times in tumor samples and never seen a normal sample. So we, of course, we know what those are. We do see genes like mad in that circle, but there's others that are not widely known as tumors shed antigens. -- we have done that. We are doing follow-up experiments currently to try to 0 in on these. We use a different workflow that's lower throughput, but much higher sensitivity, so we can be sure about -- we don't want to be a lower-level expression in normal tissues. And so we can test that directly, and we're getting some hits here that are looking really promising. And so these can go into vaccines or they can go into TCR-based therapies. And there, we also have some really interesting computational developments ways that we can discover TCRs de novo using computational approaches as well as use rational AI-guided optimizations of TCRs to make them more sensitive to antigen. So that's ongoing. We'll hopefully present on that soon. But today, I'm going to come back to this target question. And one thing I over is that even with these great AI-based methods, there's still a huge fraction of the spectrum that cannot be confidently identified. So there's various numbers floating in the field. But kind of best case, we're identifying we still have 55% that we cannot identify. What could these be? Are they like non-coding RNAs, -- are they circular RNAs, and gonds retroviruses, post-translational modifications, no unusual splice junctions like Marin was talking about, there's no consensus in the field. But what there's is a need for tools to identify these because these could be good targets. These could be cancer specific. So Nicolas is going to take us to the next section where InstaDeep has created a new tool in [indiscernible] that really zeros in on these Spectra in effort to figure out what they are.
Nicolas Lopez Carranza
executiveThanks, Mike. My name is Nicolas Lopez Carranza, and I lead the BioAI team at InstaDeep. It's a pleasure to be here. As Mike says, between 55% to 75% of the peptides available in mass spectrometry database cannot be identified. The main issue here is how traditional mass spectrometry target deco search works. It relies on a target database here in green as well as a TCO database, which is derived from the target database by scrambling the target peptide or reversing them. then the algorithm scores all of the peptides and keeps the best matches and calls for those peptides using the Deco database as a way to control for false positives. But what happened if we develop an algorithm that does not rely on a database to call for those peptides. Here, we are talking about the de novo peptide sequencing, and you see how a sequence-to-sequence AI model is translating between an MS2 Spectra into a peptide, as you see here in the picture. The great advantage here is that we do not need to rely on a database. It's a simple translation model as we would be translating from English to German. So that's why we partnered with DTU, Danish Technical University to develop in Stanovo, the novo peptide sequencing with deep learning. We trained this model on 33 million peptides from the Proteum tools database, and we actually developed 2 models. The one on the top right is the autoaggression of the peptide where we given the input spectra are the coding 1 taken at the time as ChatGPT works. The issue with an outer aggressive model here is that once we made an error at the beginning of the sequence, we cannot recover from it. That's why we developed also in InstanoPlus, where we use a deficient decoder to avoid this issue and we improve the performance. Regarding the results, we did manage to increase an immunopeptidomic data set by 40%, as you see here in the picture. We also made this model available for the community to build on top of it, and you can find the publication and the code available on Giant. Without further ado, I'll let back with Karim to continue with the innovations.
Karim Beguir
executiveThank you, Nico. And congratulations on the great work in Senova that we have open source in partnership with DTU. So this is an example, obviously, again, of enhancing the quality of the data we have by AI-assisted labeling, which I think is very exciting. But things don't stop here. Obviously, we also work a lot on protein design, and we're going to speak a bit about that. And I want actually to share the first joint like press release we did with BioNTech Back in November 2020, when we announced our strategic collaboration like Ryan Richardson said, and importantly here, we mentioned that InstaDeep DeepChain platform would be deployed on multiple tasks, including Ribomab,like working on BioNTech Ribomab platform. Well, I'm delighted that after productive collaboration between BioNTec and InstaDeep, we have results to share, and I'm pleased to welcome [indiscernible], who's going to tell us more about this.
Unknown Executive
executiveHi, everyone. I am [indiscernible], I'm a research engineer and teammate at InstaDeep and it is certainly satisfying to see that what we announced 4 years ago actually has become a reality since. And maybe let me remind that Ribomab is BioNTech platform for mRNA and coded therapeutic antibodies for cancer and infectious diseases. But what problem do you try to solve here actually. Among therapeutic antibodies, co-expressed and bispecific antibodies hold special interest. However, this requires a precise paring of the constituting heavy and light chains. Let me illustrate that with the example of co-expressed antibodies. Let's say, we would like to provide a patient with antibody A and B here simultaneously. Each of their constituting heavy and line chains would get translated from their respective mRNA into protein separately and only later would assemble to form the antibodies. Now if this assembly process goes with antibodies that only differ by either variable domain. In most cases, this will result in [indiscernible] construct. And in the end, only 12.5% of the correct antibodies will be obtained. This is why being able to control the pairing process of heavy to heavy and heavy to light change is critical. This is what we set out to address, focusing first on heavy to light chain paring. Our approach to this program has been to engineer the interface between the constant heavy one, CH1 and constant light CL-domains antibodies. We set out to introduce mutations that would yield so-called NEO CH1 and Neo CL domains that one of the antibodies could be equipped with. Such mutations are called orthogonal mutations because they both seek to enhance the affinity between the new CH1 and the OCL domains, all while abrogating the binding between NEO CH1 to [indiscernible] and wild-type Neo CL 12. Now this protein engineering problem is actually a multi-objective combinatorial optimization problem where we search the gigantic solution space to find the optimal set of orthogonal mutations. This is the type of whole that InstaDeep has lots of experience with. Yet here, it came with its own set of challenges. We had to properly estimate the binding energies of the correctly paired and mispaired domain assemblies. We had to estimate the impact of mutation on the stability of the heavy and light chains. We had to structure model the mutations and to gain a deep understanding of the key interface interaction to help steer our models through the gigantic solution space. All of which we could achieve, thanks to our DeepChain platform and to an efficient in silico in vitro collaboration working hand in hand with our BioNTech colleague. We did an amazing job on the in vitro part here. And not for the results. Well, we were able to achieve more than 90% correct pairing matching the best patented designs in the market. We did validate that antibodies equipped with our domains remain the full functional activity, and this is how InstaDeep helped by BioNTech acquire the technology required to develop next-generation bispecific and coexplaced antibodies. Thank you very much.
Karim Beguir
executiveThank you so much, polite. And congratulations again on this amazing results obtained with the DeepChain platform and in collaboration between BioNTech and InstaDeep. So we are now at the last presentation of this AI Day. But last but not least, I'm pleased to have Sven come also and present with me.
Unknown Executive
executiveYes. Thank you, Karim. First to myself, I'm Sven [indiscernible], Director for Global R&D automation at BioNTech located in Mines. And I would like to give a very short introduction about the connection between AI and automation. So with automation and AI. So we have the great potential really to revolutionize the way we do or we work in R&D. So we have in lab runs that, together with in silico runs can really in a closed optimization circle accelerate the scientific discovery. But on this way, we also have different challenges. So we have R&D that is constantly changing. So we have to react also with automation on these changes. So we have to solve the contradiction between automation and flexibility. We have a high complexity in the combination of automation and science. This needs to be solved. And we also need to keep the transparency for our scientists who need to have control about the experiment and also on their results to give the transparent to them back. And so with the assistance of artificial intelligence, we see opportunities to overcome these challenges to create transparency and also to be fast changer also with automation system. And this then really helps us to unlock the full potential of laboratory automation. And with AI, we see capabilities. We can have information discovery from different resources, not only from the machine fast protocol development and change of the automation and the protocol needs to be set up on the machine, machine error diagnosis supporting also the troubleshooting while implementing all the cross-team interaction between engineered scientists and all the AI experts. And in total, this also supports the whole change management to create transparency and full control what we are doing in our labs Handing back to Karim.
Karim Beguir
executiveThank you, Sven. And so obviously, like AI being highly capable offers an opportunity to improve automation in the lab. And so to give you a flavor of that, Actually, we're going to go live into mines and the tech lab of BioNTech mines. So we have David with us here, and we are very excited to show you a first, which is the DeepChain platform with a Layla agent, but in the lab. The idea here is we want to inject all the intelligence that AI is capable of into practically useful workflows, like Sven described and this is a challenge. And we are actually tackling this challenge. So over to you, David. Please introduce yourself and just go ahead for a demo, you have 5 minutes.
Unknown Executive
executiveHi, everybody. Can you hear me okay?
Karim Beguir
executiveVery good.
Unknown Executive
executiveHi, everybody. As Karim said, I'm David and I lead the RNA optimization group here in InstaDeep. Right now, I'm in the tech lab, which is fintech center for laboratory research and innovation. I'm very excited to show you how we fully integrated our Layla AI agent with BioNTech laboratory machines to provide any help that our lab scientists need, which we call Layla lab. This page shows you an overview of all the machines in the tech cloud. Layla in the lab allows scientists to see exactly what's going on around the organization and find out what they need to know immediately without interrupting their colleagues. Here, you can see our telling us a live real-time update of what's happening on other machines and easy to understand and relevant way. Let me start by saying how that works. Our AI system is connected directly to live feeds from each of the lab machines. Let me show you a couple of these. Here is the live feed for one of our machines, the open tons machine. And if we look here, we have a live feed for another one of our machines, the [ Tecan. ] As you can see, the data is extremely technical. It's also generated constantly building up to thousands of lines. It's impossible for a non-expert to quickly find the information they're looking for in these kind of lots. So I understand everything that's going on, you would need to and understand technical manuals like things and many, many more Layla is able to use contextual information such as these manuals, information from within the company and also data from informational laboratory management systems to convert the roll on streams that you've just seen into concise, relevant, understandable information that you can see on screen now. But equally importantly, it can do is a high throughput and in real time. Let's have a look now at how this can help Lifestyle during their day-to-day work. What we're seeing on the screen at the moment is a summary page for tech and machines. At the bottom is a summary of the job as a whole. It represents the most relevant information. For example, we can see which protocol is currently running who requires this job and other information about the job. In the middle, we've got the state of summary, which is a real-time feed of information in an easy to understand for about the job as it's actually executing. And then at the top, we have the next action information. This shows the next manual action, the scientists will need to form on machine allowing scientists to quickly and easily sign the day and be prepared to carry out each action efficiently. This summary page presents what we think is the most commonly needed information, but the power of AI is that he can answer any question a scientist has. Layla has a track in space designed for lab use with speech recognition. Let's have a look at that right now. Layla, what is the current machine and what is doing? The current machine is Tech Labs can 1080. It is currently experiencing an error while running the protocol I-AL073. The error occurred during the step pickup plate with stained cells from incubation. The machine is unable to grip the Labware 96 well around CoStar 1 at the specified source location suggesting a mechanical failure or misalignment of the gripper. Well, with AI, you don't always know exactly what you're going to get back. And in this case, I think Layla gave us slightly more information than we needed. So I'm going to jump on and look at how we can actually solve this error that they been telling us about. To help us, you can see that the summary screen has now changed to a view that specialized for the error analysis. On the top of the page, we can see a concise summary of what the error actually is. The key piece of information is that the Gripper didn't define what it was looking for in the right place. Underneath this, we've got the impact assessment. In this case, it's telling us who needs the job and what the potential issues are if it's not so quickly.
Karim Beguir
executiveCan you actually hear me? Yes, David?
Unknown Executive
executiveYes, absolutely.
Karim Beguir
executiveAnd here, you're showcasing like Layla as an agent that can read -- but in the sake of time, perhaps you can show us the writing capabilities of Layla, basically the agent acting in the lab, if you can demonstrate that like you should be thankful to Layla.
Unknown Executive
executiveOkay. Absolutely, Karim. So if I go back to the home page, I can show you a view of all of the different machines that are currently operating in the lab. I'll just zoom out, so we have a little bit of a better view of them. And I'd just like to say later, congratulations on an excellent job done today. I think it's time to disco.
Karim Beguir
executiveAnd as you can see, actually, Layla or AI agent in the lab directly from Deepen into the lab actually is controlling the machine. So we've been working very, very hard to integrate those capabilities not only to read from the machines, but to write to the machines and you can see we had the robotic arm waving. We had the lights flashing. This is showing you like the potential that intelligent AI agents can bring into the lab. So this is a very exciting time. And importantly, this is leading to actually time savings and efficiencies operationally, and we're going to have Michael tell us more about this.
Unknown Executive
executiveHi, all. My name is Michael [ Downs. ] I'm Director for Digitalization of Scientific Labs at BioNTech and since we progressed a lot in terms of time, I keep it short and sweet. So just to give you one metric where it can save actually time in the lab. If you have common errors where the lab technicians know what they're doing because they happen each and every day. We cannot talk too much about efficiency gain. But if you have uncommon errors where you need to consult the technical manuals like David showed them in his presentation, if you can easily use up an hour to analyze the error and just technology can definitely speed up this error analysis a lot. Talking about next steps and where we should go to with this application. What makes this technology demo unique is the level of abstraction, you have above your normal lab automation procedure in terms of semantic information. So what is the system actually doing? What is it for? Who is it doing? It for and it contains a lot of interesting use cases for different kind of users like the lab operator sell. There's part of it for the scientists, for the lab manager maybe, and we should carve out the different technologies we have here to have the best-fit use case for the user where it needs actually to. And next steps are, as I said, carve out the different user group requirements. Secondly, our digitalization backbone, we are already running at BioNTech, which is celebrity information system, electronic lab notebook and applications like this, we should connect this application to enhance the level of details, it can provide for the actual samples, for example, it's processing. And last but not least, we should scale up to different lab devices. Currently, it's pretty much focused on liquid handling devices but they're also mass spectrometry devices. There's a whole bunch of analytical devices and other devices which make total sense to be included here. Thank you. Back to Karim.
Karim Beguir
executiveThank you, Michael. And as you've seen, we're quite excited by the potential here, and there is still a lot to do in the future. So we're going to be working very hard with our BioNTech colleagues in the lab. And hopefully, in subsequent additions of AI Day, we'll be able to show you no progress, but this is a very exciting time. And as you can see, Layla is already in the lab, which is super exciting. And so this takes us to the end of the presentation. We're very happy to have with us Ryan Richardson, and we're going to move into a Q&A session, fireside chat. So we're happy to take your questions. Maybe, Ryan, you want to come on board?
Karim Beguir
executiveThanks. Awesome. So I hope you enjoyed this session. we've covered a lot, as you can see. But this is actually a small snapshot of all the work that's going on between InstaDeep and BioNTech productizing AI, not only innovating, but deploying with the colleagues in the lab, in the different R&D projects in the future also on the sort of industrialization and production. So we're very happy to take questions, and thank you, Ryan, for taking the time.
Ryan Richardson
executiveAbsolutely. Thank you. So questions.
Unknown Analyst
analystThank you for sharing. Really interesting, actually. So my question is going to be around Ribomab. And I just wonder how close are we now to kind of getting therapeutic levels of antibodies delivered as an mRNA. So is the AI actually or machine learning providing features of an antibody or imparting greater stability that would actually make it a much more realistic possibility because -- and also from the perspective of cost of goods because obviously, that's also one of the other drivers about bringing our -- not RNA, but maps to the marketplace for therapeutics.
Karim Beguir
executiveThank you for the question. Perhaps I can say a few words on AI and you can discuss the different programs. So -- if you look at where we are from an AI standpoint, I think the best comparison is to look at where we were in natural language processing, basically understanding language in roughly 2020. When you had GPT-3 coming. So those models are starting to become seriously powerful. They are not perfect yet, but that gap is being bridged. So I believe we're going to see tremendous progress in the coming years, and this is a very exciting time to be working at the intersection of AI and biology. Now on top of that, we actually shown -- your question is on antibodies. We've actually shown very exciting results for the first time today. First, on our BFX antibody generation system, which is currently state-of-the-art and in particular, with an ability to condition to different chemical, biochemical properties, that's exciting. So that's happening on sort of like the variable regions. But on the structure of the antibody itself and have it M&A encoded, this is exactly what we've shared with our results on the Ribomab platform using DeepChain. So I would say the progress is happening now. And the slope at which things are accelerating is definitely going to increase in coming years. So yes, it's like GPT in 2020 and hopefully, many exciting things to share in the future. So that's in genAI and perhaps Ryan, you want to share on specific programs.
Ryan Richardson
executiveYes, it's exciting that we can -- that we already can see these kind of results with the application of AI. But in terms of where the therapeutic platforms are, we actually are already in human testing in the clinic, with both our RNA and coated antibodies and also RNA encoded cytokine. So to your question about do we get therapeutic level of translation? The answer is yes. And what we're doing is we're using a liver targeting LNP to deliver the RNA and effectively turn the liver into the manufacturing engine of the therapeutic, so yes, I think we're seeing very encouraging results. I think more work to do. But this could open up a whole new space of therapies across a wide range of potential targets. Again, RNA-encoded cytokines and multispecific antibodies and T cell engagers, we've already taken into human testing.
Sam Fazeli
analystIt's Sam Fazeli from Bloomberg Intelligence. I'm not quite sure where to begin, but I'll limit to 2 questions. Half exaflop is a pretty significant product, right, in terms of computing power. You clearly have very big ambitions for this. This can't be just for BioNTech to use in terms of its scale. So what is the ambition? What is the business strategy here going forward? Should analysts be thinking about modeling InstaDeep as a revenue line but significantly for BioNTech. And then just to that last point. With regards to your pipeline, clearly, the pipeline currently is populated with a lot of products that have -- that -- some of which you've in-licensed, some of which are internal, clearly, the personalized cancer vaccine, whichever way you want to call it, individualized new antigen therapies, they benefit from AI today, I'm pretty sure. When would we see data to start getting excited by from work that's emanated from this collaboration.
Ryan Richardson
executiveYes. Let me start, and Karim, you can add. So when we did the Inside acquisition, we identified a couple of key value drivers. And one of them was, of course, cost efficiencies associated with effectively internalizing what was, at the time, even our largest AI technology and service provider, more than a service really technology solution provider. We also saw advantages to -- or synergies in terms of integration and capability building. But I think the fundamental value driver was really in our core business of developing novel vaccines and therapeutics, right. That's where -- that's our business model at BioNTech. And we saw the potential application of InstaDeep's technology solutions across the whole -- across our platforms across the whole, let's say, value chain, especially in drug discovery as the primary reason for this acquisition and what we think is so powerful about the combination and also unique to the industry, right? Because fundamentally, BioNTech is about discovering and developing new therapeutics and vaccines. And so -- and we do think, as you've seen, maybe you got a glimpse today that we think that there's multiple applications in that discovery arena across platforms of InstaDeep technologies and capability. So I think, obviously, that's a long-term value creation engine -- we don't currently split out InstaDeep in terms of financial reporting, but I think it is worth noting that while we haven't really talked about it today much, InstaDeep also has a third-party business outside of the BioNTech relationship. And maybe, Karim, do you want to say a few words about that.
Karim Beguir
executiveAbsolutely, Ryan. And the idea is that we want InstaDeep to be a leader in AI. And the way to do that is actually to develop core innovation in AI that obviously applies to the biological pipeline and strategic objectives we have with BioNTech, but also that can have applications outside biology -- so we do both. And we found that this is actually a variable way to operate because the same technology can apply to multiple use cases, and we've seen multiple, multiple proofs of that. For example, if you look at like designing a new protein, this is a combinatorially explosive kind of problem. We are leaders in industrial optimization within biology and outside biology, and these add up together. So really, the objective is to continue to be a leading power in the world of AI, continue to invest, continue to derive new innovation. And from that point of view, our new supercomputing cluster is a must and it allows us for more flexibility for the type of workflows. So we will use it for biologically like compute-intensive applications, of which there are actually many, but also in other types of AI research work we do, which ultimately benefits the progress we do as BioNTech Group company. So that's a bit like the spirit of what we do. And I think also like in the idea of like sustaining a dream team of AI talent, we found that this AI identity of InstaDeep is the right approach. And as you can see, we're constantly pushing the frontiers of innovation. We published many papers as well. Last year, we had more than 25 plus research papers in AI published at major conferences, Nature journals also something where we publish -- so it's an exciting time, and we are continuing to invest in those capabilities, but importantly, making sure all the work we do does actually benefit our BioNTech colleagues in the lab, in the industrial processes. And in a sense, this is kind of like having the full vertical from pure AI compute algorithmic innovation to in the lab testing, in vivo clinical trials and the others. That's the goal.
Yaron Werber
analystKarim and Ryan, a question from the webcast here from Yaron Werber. Cowen. So 2-part question. How does DeepChain take external input from academics and other groups to fine-tune the model? And then secondly, do you foresee other BioNTech companies using this open source model and how can we at BionTech keep some things proprietary to maintain a competitive edge?
Karim Beguir
executiveSure, absolutely. So we're very excited to have DeepChain available both for internal stakeholders within BioNTech Group, but also outside BioNTech for external parties. So the culture of collaboration and scientific sort of like work together with investees is very strongly established at InstaDeep and BioNTech, so as you've seen, for example, our instanovo protocol that was presented today was developed with DTU, technical invest of Denmark. So this is the kind of spirit -- and there is room for joint innovation there using public data sets and things become proprietary when you're using specific proprietary data, for example, within BioNTech. And so that's the spirit with which we operate. We develop core technologies on open source potentially in partnerships. And if you want to go the level beyond, which is use your specific data, make sure with all the sort of like privacy guarantees that come with that. This is something that we offer, and we have extensive experience doing that. So that's a little bit the spirit behind the DeepChain platform.
Ryan Richardson
executiveYes. And to address the second part of the question in terms of other biotech companies and how we balance that. It's interesting, when we started working with InstaDeep in 2019, we were really one of the first major BioNTech clients, right? Karim and team had built an extensive client list in other domains in the tech industry and industrial sectors. But we were the first BioNTech or one of the first. And I think it's interesting, initially, we were the main BioNTech customer. And I think that the question there hits on actually a very interesting point. But obviously, you can see that inside has amassed a very abundant wealth of expertise very quickly. And we've seen with our own eyes at BioNTech, the pace of learning of the models over the last couple of years. It's just been extraordinary. In fact, it's one of the reasons that we decided to make the acquisition. Because we reason that as these models, effectively the brain of new drug discovery and design as they progress, they got smarter and smarter, but that needed to be a core competency inside BioNTech, right? Because we're not unlike many big pharma companies were not just acquisition entities and commercialization engines alone. Our fundamental business is in drug discovery and innovation. And so we felt that, that needed to be in-house. So I think that leaves the door open to also expanding a X BioNTech business in the future for InstaDeep. And I think that's something that we certainly could do. And I think if we do decide to pursue that, I think it won't be a problem to balance our internal domains with external because of the breadth of application that some of the technologies you've seen on display here today, Karim?
Karim Beguir
executiveAnd from a technological standpoint, this is absolutely feasible. Think about it as like you have large-scale language models or innovations like BFN, those are trained on public data at scale. But then you can fine tune them and Julia earlier showed an example of fine-tuning with the specific data of a particular company. And so this creates a model that is more advanced for a very specific use case. So as a platform, we can definitely engage with multiple stakeholders while making sure everybody's data stays completely protected and useful only for the person who brought in the data. And so this is a model that works very well. It's exactly the same as like if you go to one of the large LLM providers as an enterprise customer, and you're like, hey, can I fine tune or refine your model on my private data, but I do not want you to learn from my private data and give it to somebody else. -- and the large like providers, whether it's Google with Gemini or Microsoft, OpenAI with GPT, 4.5 offer this service. So it's exactly the same, but applied to biology, where we are a leading force and working together with BioNTech, we're actually capable of providing experience on use cases that can be used for others in different sort of like task in biology. And biology is quite vast. So it's totally not a problem.
Unknown Analyst
analystMy name is Harry, reporting for Time Magazine. You've described Layla as a genetic system with expert level, biological knowledge and toll use ability. My question is what measures do you have in place right now to ensure that this tool isn't used by nonexpert actors for malicious use cases?
Karim Beguir
executiveSo it's a very good question. So until now, Layla is actually available internally and in testing. But importantly, like the main use case for Layla is improved interaction with the system and improved tool use. So for example, like you can call Layla to like do a routine on segment NT like has been shared. Those models are open source, and we are restricting users of Layla to tool use and sort of like database collection and others. So we are having like a very stringent process, making sure that the system cannot be used for other things before we release that. But as a sort of like conversational interface to very powerful tools such as BFN, segment NT like we have seen, there is a great sort of use case for Layla, even in the lab where we've seen even like with having like a text-to-speech capability for people working in the lab, which we believe democratizes AI because if you look at where we are in terms of like biologists on one side and AI machine learning experts, those rarely intersect. And so Layla is sort of like democratizing access to powerful tools, but which are tested and open source, and we are constantly exchanging with the online community to make them better. But -- so we're very careful on that. But to answer your question, specific tool uses and increased sort of conversational capabilities to lower the bar of -- like the entry bar for using these systems in a safe environment.
Unknown Analyst
analystJust to clarify, [indiscernible]
Karim Beguir
executiveYes. We -- the tool used here is not like we do not allow Layla to go, for example, on the Internet and use multiple things on multiple models. -- it's only to the models, actually, in this case, developed by Instead or that have been sort of like open source for a while and validated by multiple parties. But we do indeed restrict the list of tools to only validated ones.
Unknown Analyst
analystAnother question from the webcast here this time from [ Elliott Bosco ] of UBS. This is about prioritization. So the question relates to how BioNTech prioritizes the incorporation of InstaDeep technologies throughout our development pipeline.
Ryan Richardson
executiveYes. So I'll try to -- I'll take that. So it's -- we don't actually differentiate directly between InstaDeep-derived molecular structures and non-InstaDeep or AI-derived structures. I think the goal that we're striving for is to embed AI where it makes sense to do so. And as you've seen today, at least you've got glimpses -- we actually think that the applications across our platforms are quite broad. So we talked about personalized cancer vaccines being an obvious use case where we're using AI to design each and every vaccine in terms of the neoantigens or mutations that we target per patient, but you've also seen use cases for off-the-shelf drugs. Ugur talked about our model, focused on both off-the-shelf drugs, more traditional in that sense, still using novel technology and individualized therapies. We see applications in both. So it really comes down to performance, right? And we oftentimes will take AI-derived molecules into the wet lab into lab testing versus non AI-derived molecules, and it's a battle of for performance.
Karim Beguir
executiveActually, there's been many cases where we have AI generated, say, constructs coming from InstaDeep and expert design constructs from the BioNTech colleagues, these are anonymized, and then we take them to the lab and we see what works. And the feedback from that is that the best is actually mixing the domain expertise of the biologists at BioNTech with the AI expert. And if those collaborate, that's when you get the best results -- but obviously, AI capabilities are increasing constantly, and this is something to keep in mind. And I think we had a follow-on question from a gentleman there, yes, before, yes.
Ian Johnston
analystIan Johnston, Financial Times. Just a question on Layla. How do you tackle potential hallucinations within the model and what impact that could have on its use in the lab? And how broadly within the lab, is it likely to be used across all indications.
Karim Beguir
executiveIt's a very good question. I would say, hallucinations is sort of like a well-known problem. But if you look at the evolution of the latest models, this is increasingly less of a problem and how we tackle it. is by having like very solid guardrails in terms of like pump engineering and the other. To give you an idea, Layla can accommodate context window of 128,000 tokens, like roughly 100,000 words. So you can use this to give a very powerful context. We use it, for example, to give the context of every machine. If you saw in the live demo from the tech lab in mines, we showed you actually like an entire manual. This has been entirely uploaded into Layla, which then can provide very accurate sort of context. And so where we are today is really like we see that this technology is very productive. And as we've seen with Sven and Michael, who are experts in Lab Automation at BioNTech, this is already useful today as a transparency into what the systems can do. the question about like what's coming next is really about, we believe continuing to integrate the technology into the lab, but really having that loop with the lab experts to see where this is most useful. The first feedback we got is that actually, this is useful to create radical transparency of the state of every machine and sort of like troubleshooting. But obviously, as time passes, we're going to see more and more use cases. But I think what's exciting is to show that this is actually possible today. Very often, we think about large language models are only like sort of like smart Q&A partners, but as you've seen today, including like the action capabilities of Layla, there is a lot more that is going to come soon, and we believe that the people who are going to be the best at this are the people who are iterating with real lab technicians, experts to make sure this is deployed the right way.
Manos Mastorakis
analystManos Mastorakis from Deutsche Bank. So we had a lot of fantastic stuff most of it, most of which, I believe, is in preclinical. In terms of the preclinical kind of processes and how you discover drugs. Some of the things that we didn't hear is how BioNTech is using AI across the operations of the company. So when it comes to manufacturing or how the companies run in general or how clinical trial data is analyzed. Could you give a bit of color on how you use AI with or without InstaDeep' help in those domains. That would be helpful.
Ryan Richardson
executiveYes, it's a great question, Manos. I think -- so the primary use case is and has been to embed AI in drug discovery, right? That's where we see again, an ability to combine our therapeutic platforms on 1 hand, which are largely very novel with the AI capabilities that Density brings to bear. And that's where we see truly profound disruptive potential in terms of developing or discovering new drugs. But you're absolutely right that we see broader potential across different domains of the business. And -- there are examples there are projects that we didn't highlight today that are happening in other areas. For example, we're looking very closely at areas of how we can make clinical development more efficient. Right, how we can select patients more efficiently, how we can write protocols more efficiently. Just to name a few, of course, manufacturing, which is also a core competency for us, both personalized RNA, also bulk RNA and even cell therapy, we do in-house. We think there's multiple applications there along with supply chain. I think in those cases, though, we at BioNTech, we also are aware of the fact that there's other external service providers or technology providers that might have certain domain expertise or a certain focus. And so I think where we think we can have -- we can competitively differentiate through an internal solution, -- that's something where I think InstaDeep is very well placed or where there's a particular problem where InstaDeep has deep expertise. I mean, there's actually quite a few of those areas, especially across the industrial automation arena. But we're also going to use external providers, too. So I think we're in the fortunate position to be able to kind of choose what we -- where we sort of invest in InstaDeep to build capability where we might rely on a third party that has an existing business and it's a mix across the company.
Karim Beguir
executiveExactly like the goal is not to have InstaDeep deployed on every potential AI use case within BioNTech to give a concrete example, if it is a system to better manage like HR, for example, it's probably better to have an off-the-shelf solution rather than have the sort of limited capabilities that InstaDeep has in terms of personnel, number of projects deployed on that. So we will aim to really move the needle, look at what is strategic for BioNTech, where we can bring that extra edge in terms of innovation, compute at scale deployment that makes a difference. And so we constantly look at sort of like the value of having the InstaDeep deeper on specific projects versus taking off-the-shelf solution, and we're very pragmatic about that.
Daina Graybosch
analystMaybe we'll take 1 final question from the webcast. This is from Daina Graybosch of Leerink. On cancer vaccines, can you help us understand how the city platform could assist and not just identifying neoantigen mRNA sequences, but also to predict the neoepitope immunogenicity.
Ryan Richardson
executiveWell, so how we can use it to predict neoepitope [indiscernible], I mean that's a fundamental application of AI that we've had ever since we went to in silico process. It's trying to predict MHC Class 1 and 2 binding affinities and immunogenicity for a variety of antigens and to be able to apply that on a per patient basis, right? So that's a very fundamental use case that we're already using AI for. I mean in terms of how InstaDeep can help us do that better. I think maybe Karim, you can talk about some of the latest advances in models that might be applied?
Karim Beguir
executiveYes, absolutely. I think this is a very rich topic. And there are multiple ways where you can have sort of like AI-led improvements into the current pipeline. So while I can't disclose specifics here, definitely, like you have to look at AI as a capability that you can apply across the pipeline. Today, we've shown you several examples. There are other also use cases we're working on, and this is -- if you look at the capabilities that AI offers, they are increasing extraordinarily fast. For example, like the lab demo that we've seen is something that would have been strictly impossible a couple of years ago. Now it's possible. So we constantly reevaluate. But yes, could we do something about personalized cancer vaccine, absolutely and from a technical standpoint.
Ryan Richardson
executiveAnd maybe just to add one point to that. So I think if we look at sort of what's the holy grail in the personalized cancer vaccine space in terms -- as it relates to again, one of the unique aspects of a personalized cancer vaccine that's powered by AI is, again, the ability to harness and create data assets. right? So, so far, the industry has largely built these models, train these models on sort of hundreds, maybe a couple of thousand patients of data. There's a lot of overlap in the data sets that different companies have used A lot of the data sets are actually shared across university, public-private consortiums. I think if we fast forward, I think if we, as an industry, are able to harness the power of data as we treat more patients. I think that could truly unlock a pace of learning that could be exponential to an extent. Again, we've never had in the industry, a model where for a given drug modality, where the more patients you treat, the better the modality gets right? That's not been the paradigm in the past. And I think with personalized cancer vaccines that could be the model in the future. And of course, AI is going to be an important ingredient to help us get there.
Karim Beguir
executiveAnd absolutely. And like Ryan said, the limiting factor in biology today, and this includes personal cancer vaccines is really data. So the more data you have, the more you deploy AI to extract more data or make a better use of the data you have with a few examples we've seen the better things are. So it's all about like going through this virtuous loop faster and faster, which is what we are doing between InstaDeep and BioNTech. Thank you, guys.
Read the full transcript via the API
You're viewing the first half of this call. Get the complete BioNTech SE transcript — plus 251,000+ transcripts from 12,000+ companies, speaker segments, AI summaries and full-text search — through the EarningsCalls.dev API.
Get the API View API docs →This call discussed
For developers and AI pipelines
Programmatic access to BioNTech SE earnings transcripts and 251,000+ others is available through the
EarningsCalls.dev REST API. Plans from $24.99/month — full transcripts, speaker segments,
full-text search, and the recently-added /api/v1/transcripts/recent polling endpoint for ETL pipelines.