Thermo Fisher Scientific Inc. (TMO) Earnings Call Transcript & Summary
November 18, 2020
Earnings Call Speaker Segments
Graziella Piras
executiveAll right. Good morning, everyone, and thank you for joining, and welcome to our webinar this morning. My name is Graziella Piras. I am Senior Technology Strategy Manager at Thermo Fisher Scientific, and I will be hosting the webinar today. We're very excited to talk today about Key Driver Identification, achieving consistent performance in bioprocessing through predictive modeling. We are live here. [Operator Instructions] And at the end, we will be answering your questions. So let's get started. Next slide. So at Thermo Fisher Scientific, we take pride in our mission by enabling our customers to make the world healthier, cleaner and safer. Next, and we're committed to the advancement of science by offering products and solutions that enable customers to push the boundaries of innovation through our deep customer focus, unmatched and great portfolio, industry-leading scale and our depth of capabilities. So it is my great pleasure today to introduce our speaker, Neel -- Dr. Neel Sengupta, who's a staff scientist at Thermo Fisher Scientific. Neel completed his PhD in chemical engineering at Purdue University, where he studied the role of cellular metabolism in protein expression production in total. He also hold a Master degree from IIT in Bombay with applications of mathematical modeling to cell signaling pathways. And he's been leading research to understand the impact of media and media components, such as components of media and supplements, on biological system through predictive modeling. And he leveraged this experience and knowledge to actually develop the Key Driver Identification approach, and he will gave you an overview today. So it's my great pleasure to introduce Neel. Over to you, Neel.
Neelanjan Sengupta
attendeeThanks, Graziella, for such a nice introduction. So welcome, everyone, to the talk today. And the agenda today would be, first, we'll go over bioprocessing drivers and sources of variability with particular focus on media and supplements. Next, I will introduce what are key drivers and our unique approach, the KDI or Key Driver Identification approach, to identify such sources of variability in media and supplements. Next, we will go over mathematical modeling or the mechanisms that we use to identify these key drivers. And finally, we will see some case studies where we have applied successfully this approach to achieve consistency in bioproduction and improve customer processes. And finally, we'll end with some conclusions. So moving on to bioprocessing drivers and sources of variability. Broadly speaking, the things which impact bioproduction can be divided into 3 categories. The cell line itself. So there can be specialized expression system, glyco-engineering, which can impact your both titers and protein quality. The other factors are process parameters, such as the pH, DO, cell culture conditions essentially, which will also greatly impact your bioproduction outcome. And finally, one of the things which are very critical for a successful bioproduction is the cell culture media and supplements and, as we all know, suboptimal formulations can impact product quality. And moreover, sometimes, these media can be cell line-specific as well, so there is not one universal media that will work across. So you might have special media for different cell scenarios. So our focus today on the talk will be on cell culture media and supplements and how the understanding -- or the deep understanding through our KDI approach can be used to leverage various favorable outcomes. So before we move forward, what I want to do is introduce the evolution of culture media and supplements over the years. So pre-1990s, in mammalian cell systems, serum supplementation was still being used. It's still being used today but mostly for vaccine applications. But as we all know, serum has some issues with lot-to-lot variability, supply issues, risk of BSEs, et cetera. And back then, supplements such as peptones, which are complex hydrolysates, such as yeast extract, soy peptones or other casein or animal-origin peptones were mostly being used for microbial cell culture. So in 1990s, there was a push to move away from serum for mammalian, especially to cell antibody applications. The many of the blockbuster drugs on the market today were very successful because they were using peptones in their formulations. The peptones were instrumental in giving a rapid solution to delivering high titers. So in mid-2000s, what was -- the focus was on consistency. The peptones, though they have a lot of advantages, they can lead to some sort of lot-to-lot variability, so there was a push for chemically defined media. And today, the focus is on product quality as well, and there is also a rekindled interest in peptones as supplements, particularly in biosimilars market as well, because peptones can deliver your rapid solution. CD media is also very useful, and what people have realized that chemically defined might not be chemically pure, so they also have their own unique challenges, so that -- in the space today, both peptones and chemically defined media and supplements are being used to get to your targets as they all have their unique challenges, which needs -- which we are addressing in this talk today. So before I move forward, I just wanted to focus on why is media so important and just a basic understanding why media and supplements could be so important. So here's an example where it's showing a CHO DHFR cell line, which was evaluated in 3 chemically defined media. And you can see the bars here define the growth, and these red dots are the production. So this is the same cell line, same process. And just changing out the media formulations can have a big impact on your bioproduction outcome. So here, media 3 is showing highest production, highest growth. So further looking into the protein quality, what was found that media 3, which is shown here on the plot on the right, has the lowest glycosylation. So in certain scenarios, yes, you're maybe getting highest production, but is that suitable for your needs? So the main take home point is here for the slide is just by changing the formulation of the media, you can have very drastic outcome. So this just shows the power of the media to get to your designed outcome. So going back to what could be potential sources of variability in media, and in particular chemically defined media, one source of variability, which people have reported in literature is the raw material purity. So again, going back to the concept, chemically defined might not be chemically pure, and it's well known that manganese salts -- manganese can be a trace contaminant in various vitamins, amino acids or even other trace metal salts. So here's a case where a customer was using a peptone containing media, and we evaluated this particular Thermo Fisher peptones for manganese content. And these are different lots of this particular peptone. And you can see that the levels of manganese in this particular peptone was very small, about 0.2 ppm, and the variation was also not that much. It's like 0.2. In one instance, that was 0.3 ppm. So is that enough to cause variation? So when we compare this peptone to a vitamin, which can be present in a commonly used vitamin, which is present in base media, what we found was this particular vitamin was bringing much higher order of magnitudes of manganese into the formulation. So this is an interesting scenario, like we often might think that peptone can be a source of variability. But overall, we have to have a holistic view of the media -- with -- in combination with the base media. So in certain instances, impurities in the basal media, which is chemically defined, can overshadow a natural variation between lots of peptones, so it's important to keep an eye on that. And to further add on to that, each process would be unique. So it's important to find the source of variability that -- for each particular process. So whether if you're using a peptone, is it coming from the peptone, is it from the base media components. So that's -- the holistic approach has to be taken. So if people who are using peptones, just to introduce what peptones are. Peptones are digest of protein sources of animal -- or animal origin free based materials, such as yeast extracts, soy peptones, so on. They do offer many advantages and development of a bioprocess. So as I've mentioned before, many of the blockbuster drugs are using peptone-based processes. A lot of new processes are still also using peptones because they have advantages of enhanced production, protein quality, better viability. But one disadvantage is peptones are derived from materials of biological origin, and they can show some inherent biological variability. So in our experience, this biological variation can vary between 10% to 20%, and that might or might not impact a customer's process. So here at Thermo Fisher, we have strict quality controls for critical manufacturing steps, which are the key to limit the variability in peptones. So what I'm showing here is a plot of a digestion pattern of a Thermo Fisher peptone. So this has different kilodaltons molecular weight profile, and what you can see is each bar is a different lot of the same peptone. Having strict quality controls, we can have a very consistent outcome from a peptone digestion process. So going further into a case study or scenario where the peptone and a combination of base media will always cause a variation, so let's take an example here and -- where we can assume, let's say, this component A is a critical driver for some processes. And again, assumption being the critical range is that it has to be above 20 ppm. So what I'm showing here is different lots of this peptone, which are some variation on this particular component A. You can see it varies from about 12 ppm to all the way to 20 ppm. So these are very tight ranges to begin with, about 10 ppm ranges, but is that enough to always cause variation? And the answer is it depends. So let's take up one process example where, let's say, one customer is using a base media, which also has this component A at 5 ppm. So in this scenario, this natural radiation coming from the peptone in combination with the base media might have an impact on your performance, especially if it's a critical factor for your process. The other scenario could be where a second customer or a second process has a base media, which also contains this component A but at 100 ppm. So in this scenario, this natural variation of this component A would not have any impact on the process because the base meter would overshadow the variation from this peptone. So one critical thing to keep in mind is, again, each process is unique. And any impact of lot-to-lot variation, either from the peptone or from the CD media, with -- would be dependent on the characteristics of that particular process. And we have to evaluate everything with a holistic approach here. So in this scenario, for the process one, component A would be something, which we might say is a key driver for that particular process. So moving on to what is a key driver and what is our unique KDI approach for identification of these key drivers. So key driver is a media component, which has a strong positive or a negative influence on the performance. There is an optimal range to achieve a target performance, and variation in this optimal range is causing variability in your cell culture performance. So for example, let's take a component A which have a variation shown by here on this X axis. And as you can see, the variation on this component, going back to the first example of scenario B with the component A, we can see that this variation is not having an impact on the yield. So this yield could be product quality. This could be titer. It could be whatever bioproduction outcome the customer is interested in. So in this scenario, component A is not a key driver. On the other hand, you can look at component B example here where this variation in this component range is causing a drastic impact on the yield of -- on the yield and the scenario. Moreover, there is a narrow range, shown by this gray bar, which the component B must be in to have a suitable target performance. So in this scenario, component B would be a critical driver or key driver component. So what is our Key Driver Identification approach? So it's a holistic approach where we leverage analytical data. And through our proprietary biostatistical models, we identify few key factors in your media or supplements out of many, which are driving your bioproduction outcomes. So the reason we are seeing this is holistic, it's not only leveraging chemistries and mathematical models. We also apply the biological knowledge for the system to discern what is the potential key driver for your process. And what are the advantages? So once we know what's driving your -- what are the components which are driving of your process, we can use that, leverage that to achieve consistency. We can improve our existing process, and we can even use this knowledge in scenarios where we are doing de novo media and supplement development for various projects. So once we have seen the approach, how do we identify the key drivers? So in general, it's the proprietary customizable biostatistical models are the workhorses of the KDI program. And it's a phase-gated approach, where we often work in collaboration with the customers to deliver these favorable outcomes. So the first step is data generation phase, where we have some lots of media or supplements where the customer has some performance data, and we generate analytical characterization on those media lots or supplements. So the next phase -- or the first phase is where we start this modeling process, where we do model development and discovery. And here, we're using our initial data analysis, our framework for our mathematical models. We start developing mathematical models, which tie the performance and the chemistry data together. So the output of the first phase is we identify a list of potential key drivers out of these many media components, and also we develop these initial predictive models. So the next phase is where we starting experimentation in collaboration with the customer, where we create some pilot materials with enhanced some of these potential key drivers. And once they're experimentally tested out, that gives further confirmation on which are the strongest out of this potential list identified. It further helps with readjusting the models, and the outputs would be confirmation of key drivers and updated models. And the Phase III is model finalization and validation, where we challenge the model with unseen lots and ask it to predict the outcome. And sometimes, we check the validation. And it depends on the customer whether they want to go through this. But regardless, there is a validated predictive model, which can be used for further optimization or screening. And finally, the implementation which -- where once we know what's driving a process through our micro-addition strategies or other adjusted component concentrations, what we can do is achieve optimal concentration of these key drivers in the media and supplements to drive very good bioproduction outcomes. And at the minimum, using the predictive models, we can also do raw material screening to, again, maintain biological -- sorry, bioproduction consistencies. So the first step is going back to this data generation and especially the analytical data generation, which is done by Thermo Fisher. So we here at Thermo Fisher have a very broad, very powerful analytical group. And for many of these scenarios, what we focus on is the small molecular analysis. So what I'm showing here is different media or lots of peptones are characterized by -- into using various small molecular analyzers. So there can be amino acids, vitamins, nucleosides, polyamines, carbohydrates, inorganic elements and other total carbohydrate quantifications, so on and so forth. So once we generate this large data set, what's important to understand is we have still 100-plus about analytical characteristics defining these media or peptone lots, but we don't know which are the key drivers for the customer's process. So in the past, what we have done is we often -- and the simplest analysis would be looking one variable at a time, which might not be always useful for finding these complex correlations because, in our experience, what we have found is often 2 or 3 factors can be acting together, and there are complex underlying interactions as well between these factors. So what we have done is we have developed proprietary mathematical modeling approaches to whittle down this 100-plus variables into critical few, and that can be leveraged to get the desired production outcome. So moving on to how are we implementing these modeling strategies for identification of these key drivers. So the first step is, as I mentioned before, we generate analytical data on multiple media or peptone lots. And we have some yield or product quality data from the customer on the same lot. So usually, they can be 10 to 15 lots or slightly less to begin with, depending on the scenario. So for the first step, as we -- sorry, the next step is we build multiple competing biostatistical models. And the unique thing about these models is through our codes. What we have done is we mimic biological behavior, and I will go into slightly more details in next couple of slides. And the other unique thing is through our proprietary codes, what we have done is we -- from -- starting from a large data set, we can reduce the models to potential top drivers. So from 100-ish, it can be 10. And the other thing is these models are predictive in nature. So as with any modeling activity, these could be a little bit iterative in the beginning. So as a base -- sorry, depending on scenario, there can be some rounds of further experimentation and data generation to make these models more robust. So once these initial models are built, we challenge them for both predictability and accuracy. So predictability refers to the models get challenged by blinded data sets and ask -- we ask them to predict. And if it's -- there's agreement, so that model can be further carried forward. So the model can undergo through some evolution because increasing data sets can impact the evolution in the mathematical structures as well as evolution in model parameters. So it essentially means like what component it's assigning more importance to. So once this exercise is completed, what we end up with potential key drivers list, so starting from 100-ish to about 10, 12, and we also get the preliminary models, which can be further evaluated down with experimentation. So as I mentioned before, the -- one of the key unique things is the biomimetic natures of these models, so we have used unique modeling strategies to mimic biological behavior. So at the simplest, you might have something like additive, which will -- you will find often in like standard statistical softwares. But through our proprietary codes, what we can also do is mimic other biological-like behaviors like -- which is shown here, enzyme kinetics, sigmoidals, other switch-type responses. And then there are more complex strategies to define more complex biological interactions. So essentially these equations, what they're doing is they're tying the relationship between the performance and the analytical chemistry data. So once we build these initial models, each of these models are subjected to a variable reduction, where we -- in step 1, we start with the -- each model is started with a full-on 100-plus variables. And through various iterative steps, each -- at each step, statistical testers turn to evaluate performance on significance of all coefficients on the models. And we eliminate one variable at a time. So this gets repeated in a loop and a code, and the process is stopped when all variables are significant or a smaller list of significant variables are obtained. So from 100, we end up with 10. And this, again, differentiates our models where we have capability from going from very large data sets to critical few, which can be worked upon further. So here's an example of where we are comparing 2 competing models. So what's shown here is an additive model where the Y axis as the experimented yield and the X axis as the simulated yield from the model. So for both the model 1 and model 2, what we can see is that the R-square values, which are measure of goodness of fit, gave a very good fit values. But when we challenge these models for blinded predictions, which is shown on the plots on your right here, so the blue bar here is the experimental yield and the red bar is the blind prediction from the model, meaning that the model was not built using these data sets, what we found was that additive model predictions were not in good alignment with experimental outcome. So in this scenario, the semilog model predictions were better. And even though fits were comparable, the semilog model was explaining the biological behavior better for this particular project. So just to summarize the overall Key Driver Identification approach here. So we start with large data sets, where we are trying to identify components that show an impact on performance across lots of media. We are looking at the components together, and the goal is to reduce the list to few key drivers using proprietary biostatistical models. So here, we again start with 100 plus. So the next stage is where we differentiate between parameters that caused the variability versus the ones which correlate to the variability experimentation. So meaning that the Phase II, where once we start with 100, we have a potential list of 10. And then we go for experimentation to whittle out these 10 drivers further. And the final outcome is we reduce these sets to about critical 2 to 3 drivers, and we determine the key driver optimal range to achieve their optimal performance through -- and we also end up with a predictive model, which could be very useful for any scenario. So some examples could be, for copper in a particular process, the range was defined to be 1 ppm to 2.5 ppm. A vitamin has to be in a range for -- from 0.5 grams per 100 grams or 0.5% to 1%. So you can see how tight these ranges are, and these ranges can have a big impact on a customer's process. So next, I will go over some case studies where we have successfully implemented this approach for a favorable outcome. So the first study is for a mammalian system using a peptone-containing process. So the process was using an animal-free peptone to produce a monoclonal antibody therapeutic. So the goal is to identify specific drivers in the peptone to enhance the yield and reduce the production variability. So using the framework, which was highlighted in the previous sections, applying all the chemistry, the modeling, we were able to hone in to 2 potential key driver, 1 and 2, for this particular customer scenario. So additionally, I want to highlight that what we also found was this particular behavior was pretty nonlinear, and you can see this particular green plot of us explaining the behavior for this particular key driver 1 and 2. And the surface plot below shows a simulated response, where both key driver 2 and 1 were acting together, and both have to be in optimum levels to give very good production. And this again highlights like the interaction -- not only the interaction between the 2 drivers, but also the nonlinear response we were observing. So once these potential drivers were identified, the next step was experimentation, where we had created some material where these key drivers 1 and 2 were enhanced and different pilot materials, and they were experimentally tested. So the plot here on the right, which shows the blue bar, is the base material and the outcome from the base material. And the red bar is the outcome from the enhanced material where we supplemented with these key driver 1 and for the lot 2, which was key driver 2. And both the supplementation showed an improvement in performance. So we were able to increase the yield over by 40% to 50%, hence confirming these were the key drivers. And this outcome was also used to finalize the predictive model. So the third phase was -- the model was verified further as a screening tool in lot selection. So the final lot model was used for selection of lots. And what you see here, the plot in the middle shows the green bars are the blinded predictions from the model. The red bars are the experimental outcome, and the red-dotted line is the desired yield from the customer. So using this modeling strategy, we were able to screen lots, which would be suitable for the customer, and we got very good success rate. So essentially, prior to implementing this KDI approach, if a lot was randomly selected, we would get about 40% to 50% success rate. But using this knowledge and this modeling approach, KDI approach, we were successful 100%. So that was a big improvement. And not only that, the customer was seeing consistent performance. So moving on to next example. Again, this is a peptone example, but this shows an example for negative drivers. So just to remind everyone, drivers could be positive and negative. So the background is this process was using 2 peptones in a pair, and the goal was again to identify specific drivers in the peptone pairs to enhance yield and reduce production variability. So what we found was these cations, cobalt, nickel, copper, were negative for the production. And the strategy was to screen out peptone pairs. So using this modeling strategy, again, a random selection of lots would have led to about 20%, 25% of success, whereas model-based selection led to 100% success rate. And the plot below, where we show peptone pair performance, the red line is the expected customer performance, and the blue is the performance we are getting from selected pairs, which was done using this modeling approach. And you can see that we, again, got very good success and consistent process for the customer using our knowledge as well as the KDI approach. So the final example, which I would go over is for chemically defined media. So the previous 2 examples were showing peptone media. So what I'm showing here is in-house screen of chemically defined media. So we have 42 in-house chemically defined library media, which was screened for a CHO DHFR line. What's shown here is the green bars are the protein quality as represented by percentage total glycosylation, and the red bars are the production values. So as expected and as we have gone before that changing the formulation can lead to various bioproduction outcome. And selection of your optimal media will require consideration on both production and glycosylation profile. So this was the scenario which was more early stage, so this just highlights, like, okay, you have to look into both production and glycosylation as -- to decide which media is best for you during screening processes itself. So although we understand that, okay, different media will have different outcomes, we also saw one interesting trend that as the production was increasing, the corresponding total glycosylation was decreasing. So what we still don't know is what's driving this behavior. So what we did was we applied the modeling approach to this screening data set to understand further impact of various media components contained in these 42 media formulations to understand better what's the impact on the production and glycosylation. So what I'm showing here is using our modeling approach, we were able to bifurcate or classify each media component. So each blue dot here represents a unique, let's say, amino acid or a trace metal or a vitamin which can be found in the media, into 4 different quadrants. So the first quadrant here are components, which are positively correlated to production and glycosylation, and this quadrant is where the production have a positive impact on production, but there's a negative impact on glycosylation. So as you can see, which -- this explains the previous data set or lot. There's a lot of components which have a very good impact. They are positive for production, but they are decreasing the glycosylation, which we were experiencing in that data set screen. So as a proof of concept, what we did was we selected a group 1 components, which were having a positive impact on production, but less impact on -- negative impact on glycosylation. And the proof of concept was, okay, if we test these group 1 components, can we enhance the production in a particular scenario without having a negative impact on the glycosylation. So we tested these group 1 components over a condition 42, which was a low-production condition. And the concept was like, "Okay. I want to increase the production, but not have a drastic negative impact on glycosylation." So what we -- what I'm showing here are the impact of addition of this group 1 component. These were added as like a bolus. The bars again represent the viable cell densities, and these diamonds are the production. As you can see, as increasing these group 1 components supplementing this to condition 42, we saw a 40% increase in production. Further evaluating the glycosylation data showed that we had a minimal impact on glycosylation. So this highlights that this type of modeling approaches can be also used for media development or -- and also be applied to chemically defined media scenarios, where just knowing what's driving your process in a chemically defined media scenario as well can help with further optimization. So in conclusion, variability can be caused by very small changes in specific components and peptones or impurities in chemically defined media. We have developed a unique approach to elucidate the performance drivers from this large data set through our mathematical modeling approaches. With the understanding of the drivers for your process, media can be leveraged to achieve your desired bioproduction goals. And as we have shown in these case studies, once we know -- once you know and understand your media, we can help with achieving a consistent process or even further optimize process. With that, I would like to acknowledge our various teams, math modeling and product support team; the chemistry team, who are very instrumental in generating large data sets; and our cell culture team, which was also helpful in generating some of the growth studies, which were shown here. With that, I would shift over to Graziella for questions and further conclusion of the presentation.
Graziella Piras
executiveThank you so much, Neel, for this very detailed and very interesting presentation on our Key Driver Identification approach. This approach is part of our bioprocessing analytics capabilities for early-phase solution, which allows us to provide media development and analytics as well as upstream and downstream process development, cell line development as well as media manufacturing capabilities. So we will now open up for questions. [Operator Instructions]
Graziella Piras
executiveSo Neel, there is a question in already for you. How do you identify the key drivers? Is it from high-throughput screen on well plate? And I think you should be able to see the question as well, Neel.
Neelanjan Sengupta
attendeeSo okay. I'm not seeing the question.
Graziella Piras
executiveI can repeat it.
Neelanjan Sengupta
attendeeYes. Can you repeat it? Thanks.
Graziella Piras
executiveYes. So how do you identify the key drivers? Is it from high-throughput screen on well plates?
Neelanjan Sengupta
attendeeSo the key drivers is primarily identified through our mathematical modeling strategy. So essentially, for the example which I have gone over, like, let's say you have 10, 15 media lots or a screening data, if you know the composition of your media or in scenarios where we are using peptone media, so we are analytically characterized those media or supplements, then that becomes a larger data set. That gets fed into the model where we start with like everything which is defining our process. And using our modeling strategy, we correlate that characterization to the performance. So the analytical data is correlated to something like, let's say, titer across different conditions. And then it gets spread to the model. And then it whittles down out of, let's say, 100 things in your media, what are the 4 or 5 things in the media, which is driving a process.
Graziella Piras
executiveGreat. Thank you, Neel. We have more questions coming in, and you should be able to see them under your Q&A, but I will read it out. So one of our attendees, "Thank you for the great talk. The question is, do you build models for each component separately? Or do you utilize multiple linear regressions, et cetera?"
Neelanjan Sengupta
attendeeSo let's say, we are doing production as an outcome. So what it's doing is they are in-built codes where we're not using linear regression. It's something else. And what it does is it looks at all the variables together, build these initial mathematical frameworks. And then it goes through these variable reduction processes, like it's on the code itself, where it will whittle down the various, let's say, 100 factors to 10. So it's the slide where I showed the -- we start with all the variables in the model and then it goes through iterative reduction.
Graziella Piras
executiveAnd you do look at all the components in the media together. Correct?
Neelanjan Sengupta
attendeeYes, yes, yes, because the assumption is everything is important. And then we -- as the model dictates, as the outcome dictates, it will use that data set to hone out, okay, what's the most important here.
Graziella Piras
executiveGreat. Thank you, Neel. We have another question. What software do you use to generate the model?
Neelanjan Sengupta
attendeeSo we have built the frameworks and math lab software because what we have is the customizable nature of these models. So standard softwares will have some capability similar to this. But again, they won't be able -- you won't be able to get like what's the few drivers. And the other unique thing is we have the biological mimetic equations in there, so we have written these proprietary codes in math lab.
Graziella Piras
executiveThank you, Neel. Another question about the modeling. So can the mathematical modeling predict beyond the typical or measure concentration of the KDI?
Neelanjan Sengupta
attendeeThat's a great question. So I would say the range would be something -- so when we the Phase II, we do enhance the driver sometimes deliberately outside what the model can see, and that's the fine tuning. But with any model fitting processes, the model will depend on what data set it sees. So yes, I would say, within 30%, 40% range, we should be -- it's pretty robust. But again, going beyond that might be little bit challenging for the models.
Graziella Piras
executiveThank you. Great. Another question about the type of components that we analyze. So is it the typical amino acid and vitamins? Or do you look beyond that? What are -- where do you focus? And what components do you look at in the medium to tease out these key drivers?
Neelanjan Sengupta
attendeeAgain, that's a great question. So for chemically defined media, again, we have a great, great analytical team here. We do look at amino acids, vitamins, nucleosides, trace metals, anions. They also have other small molecule analysis. Peptone containing media, again, some of these will be common, but other -- further enhancement can be through carbohydrate profiles, peptide profiles, fatty acids, so polyamines. So there's a lot of big -- a lot of small molecule analysis that's done.
Graziella Piras
executiveGreat. We have a lot more questions, so great input here from the audience. Another one, what criteria do you use for the factor exclusion, adjusted R2, Mallows Cp?
Neelanjan Sengupta
attendeeSo I cannot divulge the exact details. But essentially, what I can say is, like, what I'm checking for is the model fit, in particular analyte, if that coefficient is statistically significant or not when I'm doing one -- leave out one variable analysis. And from there, if it comes out to be not statistically significant, that particular variable gets removed from the model.
Graziella Piras
executiveGreat. Thank you, Neel. A question about the data. Any requirements on the data provided by the customer? Do you need data from multiple batches, which is also a very important question? So I'll let you answer that. So do you need data from multiple batches?
Neelanjan Sengupta
attendeeYes. That's again -- yes, so again, the more the data is better. But in our experience, 10-ish to 15-ish would be a good starting point, yes. Sometimes we -- it depends on how diverse the lots are. So essentially, what we are looking for, let's say, if you have 2 high performing lots, let's say, 3 average lots and 3 or 4, let's say, not high performing lots, so that will be really good data set to begin with because we are trying to demarcate, okay, what are the underlying hidden interactions and correlations here.
Graziella Piras
executiveThank you. Another question. How does the model performance correlates with number of training samples fed into the model?
Neelanjan Sengupta
attendeeSo yes, I mean, that's again in built into the code there. So when we are -- I have -- I don't remember the exact details here, what the -- how many training sets we are using. So there are some bifurcation early on. The challenge is that it's more on, like, how many samples we begin with. So usually, it's 10 to 15. So you don't have like a big luxury of having detailed training sets. So the way I'm doing is, like, one variable at a time, retrain the model, check it, retrain the model, check it. And then -- so it goes through multiple iterations, where each variable is left out at a time before anything is discarded as nonsignificant.
Graziella Piras
executiveGreat. All right. Maybe a couple more questions here. Can you use this approach to increase -- if working on a viral vaccine, increase viral titers? Can you use this approach to achieve higher viral titers?
Neelanjan Sengupta
attendeeYes. The answer would be yes because this approach, as you saw, was -- it can be applied to any outcome. So the concept is we are trying to tie the media or supplements to your final bioproduction outcome. So in our experience, we have done this for antibody titer, even like glycosylation, protein quality parameters as well. So the -- in terms of input, yes, viral titer can be done, provided that the media has an impact. The assumption will be the media has impact on your viral titer. So yes, that can be done as well. So in terms of input into the model, so when we are saying customer performance here does, so we are not looking for actual values. It's just a relative value. So given -- like, if you just scale it 200 and something could be 80, 90 or 120, so that's -- so normalized output is what we are looking into. So there is no issue of -- such that there is no issue of like proprietary sharing from the customer side.
Graziella Piras
executiveGreat. Thank you. Another question, what statistical test does your model uses in order to test for parameter significance?
Neelanjan Sengupta
attendeeI cannot discuss that. Sorry.
Graziella Piras
executiveThat may be -- yes, more of a -- part of your internal modeling that you've developed. Great. All right. So one question also came in. How long does it take to identify key drivers? Maybe I'll answer that one.
Neelanjan Sengupta
attendeeSure.
Graziella Piras
executiveSo as we've seen from Neel's presentation, there's multiple steps in this approach. There's data gathering, both on the performance, but also the analytical. So the time there would be depending on if you have previous data. And usually, our customers do. We then do the analytical testing to look at all the different components in the media or supplement. And then we move into the different phases of the model to build -- the approach to build the model basically. The time really depends also on the complexity of the system, the complexity of the medium and the end goal. But this is a collaborative approach, which involves both lab experimentation, but also more -- yes, modeling. So we're talking about more like months than days or weeks here, but it all depends on the complexity of the system. So Neel, if you want to add anything but...
Neelanjan Sengupta
attendeeYes. I mean, I would agree there. So one key part is the Phase II where, once we identify the potential list, once we generate this enhanced or pilot material, the customer collaboration would require that they test these drivers enhanced pilots in their small-scale systems. So that would be the collaboration piece from them. Depending on what system, so if it's a mammalian and you're looking for both product quality, production, so that would be taking slightly more time versus bacterial systems, which would be easier to execute on the customer end because the --just the time lines are small -- shorter. So they will be on a faster time scale. So the modeling itself -- generating the initial models might take just 1 to 2 months maybe and then it's testing. So that will be the -- again, the goal -- what's the goal here, what we are trying to achieve, that dictates the overall -- the service project time lines.
Graziella Piras
executiveGreat. Thank you, Neel. Another audience member would like to learn more about this approach. And you can see on the Thank You slide that you can either contact your local account manager or field application specialist or send an e-mail to gibcoservices@thermofisher.com, and we would be very happy to contact you and have a more in-depth discussion of your needs and how we can support you. Finally, a question about -- is there a way to review the slides in the presentation, and there will be. We will have a recording of this presentation that will be sent to all the attendees. All right. So I think there is one more question, maybe last one. Do your KDI identification model apply to different producer cell line, CHO-K1, DG44? Or is it cell line specific?
Neelanjan Sengupta
attendeeSo again, that's a great question. So this model is, I would say, it can be applied to anything. So as I went over, there were 2 examples. There was a bacterial cell example. So what's common is the framework, so -- because each process and cell line is unique. So if you're using a CHO-K1 with particular set of media, that can be fit into the model, and the model will customize -- will be customized to that particular process. And so there is no constraints on what it can be applied to. So it's generic in nature. Yes.
Graziella Piras
executiveThank you, Neel. Great. All right. So again, I would like to thank you, Neel, for this great in-depth presentation. And I would like to thank the audience for attending -- everyone, for attending and for all the great questions. Again, if you'd like to learn more, please contact your account manager or send us an e-mail at gibcoservices@thermofisher.com. And this concludes our webinar today. Thank you so much, everyone, and have a great rest of your day. Thank you.
Neelanjan Sengupta
attendeeThank you. Thank you, everyone.
Read the full transcript via the API
You're viewing the first half of this call. Get the complete Thermo Fisher Scientific Inc. transcript — plus 248,000+ transcripts from 12,000+ companies, speaker segments, AI summaries and full-text search — through the EarningsCalls.dev API.
Get the API View API docs →For developers and AI pipelines
Programmatic access to Thermo Fisher Scientific Inc. earnings transcripts and 248,000+ others is available through the
EarningsCalls.dev REST API. Plans from $24.99/month — full transcripts, speaker segments,
full-text search, and the recently-added /api/v1/transcripts/recent polling endpoint for ETL pipelines.