IQVIA Holdings Inc. (IQV) Earnings Call Transcript & Summary

January 26, 2023

New York Stock Exchange US Health Care Life Sciences Tools and Services special 56 min

Earnings Call Speaker Segments

Unknown Executive

executive
#1

Hello, and welcome to our webinar today: Overcoming Challenges in Rare Disease Research: Uncovering Hard-to-Find Patients Using Natural Language Processing in Electronic Medical Records. Today's speakers are Jane Reed, Director of Life Sciences; and Kimberley Jordan, Director of Artificial Engineering, Intelligence Engineering IQVIA Real World Data and Technology. You can find their full bios on the right side of your screen. At the bottom of your screen, you'll find resources relating to this webinar, which you can access and download. Time permitting, we will follow the presentation with our Q&A session, so please do submit any questions you have using the Q&A box on the left of your screen. Before the end of the webinar, we'd really appreciate it if you take part in our survey in the bottom right of your screen. And this helps us understand a little bit more about our audience and will help us improve our webinars in the future. This webinar is being recorded and will be e-mailed to you within 24 hours after the event. Jane, over to you.

Jane Reed

executive
#2

So thank you very much, [ Melanie ]. So yes. Thank you very much, everyone around the globe for joining our webinar today. So as Melanie said, my name is Jane Reed. I work for IQVIA NLP, and within -- I'm Director of Life Science. And I'm joined by my colleague, Kimberley. So we're going to be doing a kind of a joint presentation today. Kimberley is going to give an overview of what we mean when we talk about IQVIA Ambulatory EMR data, giving a brief overview. I'll talk a little bit about natural language processing. It's an innovative technology that IQVIA use to really help with unstructured data. And then the real meat of the webinar today will be Kimberley talking through some of the deep dive analytics we've been doing for some specific rare diseases, which I think will really bring this whole story to life. As Melanie said, please use the webinar Q&A form during the webinar, and we'll address your questions at the end. So I'm now going to hand over to Kimberley.

Kimberley Jordan

executive
#3

Thanks, Jane. Okay. Let's begin with an overview of IQVIA's Ambulatory EMR. This database contains records for over 800 medium to large ambulatory practices across the U.S. Approximately 40% of the 100,000 physicians in the database are general practitioners. There is an average of 3 years of observation for practices. And for new practices, we see about 5 years of history. Patient history and ambulatory EMR goes back to 2006. In addition to demographic and geographic information about the 87 million patients, the database contains lab results, vitals, prescriptions and procedures. Ambulatory EMR is updated monthly and is quite current with a 45- to 60-day lag. And now I think we're going to have a survey about whether people have used ambulatory EMR. We would love to understand more about your knowledge about EMR and our asset in particular, Ambulatory EMR. So we'd love to see your survey results. Give folks a couple more minutes to respond. Please do let us know about your experience with ambulatory EMR, IQVIA's Ambulatory EMR, with this survey. [Voting]

Kimberley Jordan

executive
#4

So it looks like, for the people who did respond to our survey about whether people have used ambulatory EMR, the majority have not. So I'm looking forward to sharing more about this EHR asset with you. And our second survey is about the ways in which you have used EMR if you've had that opportunity to use it. And this is whether you've used it in a stand-alone way or you've linked it to other assets. I'll give people a few more seconds to respond. [Voting]

Kimberley Jordan

executive
#5

Okay. All right. So we do see fairly equal distribution across the 3 areas of, I'm using it to find difficult-to-find patient cohorts, which is one of the things, of course, we're talking about today. Characterization of a cohort, which, again, you'll see some examples of that today. And it can also be used for stratifying a cohort by disease severity, it looks like we've got some folks who have done that. You can stratify a cohort by vitals or lab test results. And it looks like the majority of people have used it in other ways, which I'd actually, at some point, I'd love to hear more about. Okay. All right. Thanks very much. And we'll move -- we'll continue forward. So taking a closer look at the longitudinality of the data. We are able to observe 1 billion visits for 82 million patients back to 2006. In the last 5 years, there are 45 million patients with a recorded 381 million visits. And for the most recent year, we have 15 million patients with 67 million visits. So ambulatory EMR is rich in clinical detail. Recorded results for laboratory tests allow researchers to gain deeper insights into a patient's health, their prognosis and their health journey. Vitals measurements can help characterize disease severity or serve as predictor variables. EHR records offer additional clinical detail that are not available in health care claims. As a result, EHR can be used to identify hard-to-find populations and clinical subtypes or to characterize the severity of disease. Other patient characteristics or clinical details that ambulatory EMR offers include a patient's race or ethnicity, patient-reported family history and social behavior such as smoking and drinking. EMR -- Ambulatory EMR's longitudinality offers visibility to a patient's journey. 15.5 million patients in the database have at least 1 visit in 2021. And of these, 43% also visited their provider in each of the 2 previous years and 28% visited their provider in each of the 4 previous years. The patients in IQVIA's ambulatory EMR looks similar to the population of patients using ambulatory medical care services in the U.S. We can see that -- however, we can see that ambulatory EMR patients, our ambulatory EMR asset, has a little more representation in the South. And because ambulatory EMR has a greater proportion of GPs and internal medical practices, it's unsurprising that we see a somewhat greater proportion, a somewhat greater representation, in our ambulatory EMR asset of older patients. As mentioned previously, ambulatory EMR's value stems from its rich clinical detail. SNOMED codes and problem names as recorded by the HCP for greater granularity than ICD-10 codes. In addition to diagnosis, symptoms such as pain can be recorded, and SNOMED codes are in the problem description. Other clinical assessments of a patients' health status, such as measures of disability, things like wheelchair use or the use of other walking devices; as well as limitations in daily functioning, such as difficulty bathing, can also be found in our asset. Ambulatory EMR includes results and reference ranges for labs that have been sent for analysis as well as results for point-of-care testing and results from community labs. Prescriptions in Ambulatory EMR can be linked to recorded problems. Dosing instructions is available for some prescriptions and can be used to evaluate dose titration. Dosing for in-office therapy administration may also be observed, as may prescribed and patient reported use of OTC products. And as rich as the clinical insights that are offered by Ambulatory EMR are, our understanding of patients can be enriched by adding additional data. And here, you can see the many other data assets at IQVIA, including adjudicated and open claims that can be linked to ambulatory for even greater detail. And in fact, there's an overlap between our -- IQVIA's Ambulatory EMR asset and its closed claims data asset, which is called PharMetrics Plus. All right. So here, we're showing an example of how you can leverage Ambulatory EMR to find -- for hard-to-find populations. The hard-to-find population in this case, is primary progressive multiple sclerosis. So we started Ambulatory EMR, and we use both SNOMED and we used those problem name. In this example, we didn't have the advantage of leveraging NLP as we're talking about in this webinar. And so it was [indiscernible]

Unknown Executive

executive
#6

Kimberley, we've lost you momentarily. Kimberley, one moment. Your audio has gone a bit quiet. We can't hear you very well. [Technical Difficulty]

Jane Reed

executive
#7

So I'm going to -- apologies, audience while we try and see if we can resolve that. We are actually just moving from that overview of IQVIA Ambulatory EMR data into natural language processing is bringing value. So we'll just skip into my piece. And I'm hoping that hopefully, when I finished my half dozen slides, we can resume as we had hoped. So if you're familiar with any kind of electronic medical record, you know that there's a huge amount of structured information that's recorded within medical records. But there's also a lot of unstructured and semi-structured fields. But these can be -- if we can -- if you can find the right information within these semi-structured fields, you can really start drilling in to understand more about the patients of interest for your particular area. So the slide here just shows some of the areas where there's a lot of -- where the coding for the medical record is weak. So within the problem list, within procedures, particularly when, for example, CPT codes aren't sufficiently granular for the sort of subtypes that you're looking for. Obviously, there's often times when the medications aren't fully coded. And again, being able to find the information is really critical. And then also being able to think about the lab measurements, the vitals and so forth, and being able to, again, find the right information. But not just find it, but standardize it, normalizing this to enable that analytics on top, and that decision support. So what we've been doing over the last year is using IQVIA's natural language processing to really enhance those semi-structured fields within the Ambulatory EMR data set. And particularly, we've been looking at areas that we know are of value to pharma organizations, looking at rare disease and looking at social determinants of health. So social determinants of health, understanding the diversity across a patient population is really critical. And being able to drill into those kind of nuances of the social status is very valuable. So whether a particular patient has ambulatory issues, has smoking issues, there's isolated depression, drug abuse and so forth. Being able to highlight those gives a huge amount of value to the information. And as we're going to drill into in more detail, we can then think about rare disease. So again, over the last year, we've been working to normalize the top 50 rare diseases to really make the search for those much more effective. We can also pull out information on family history and carriers of those diseases, again, making sure that you're building rules to distinguish a patient with a family history of something from a patient with that. So this really enables us to maximize those cohorts of rare disease. And obviously, with so few patients, every patient is critical. And you can see using this NLP approach, we've been able to hugely increase our patient finding, and that's some of the details that Kimberley will be able to go into later. So just stepping back a little and giving a little bit more information on what we mean when we talk about NLP, or natural language processing. So it's an AI technology, an innovative technology that transform text into structured information. So the schematic here on the left-hand side, when people talk about text, not -- they often think about literature, publications, kind of big reports. But you'll find, as we've just discussed, unstructured text in a huge range of different areas. Not just the sort of literature such as you might do, but within safety reports, clinical reports, regulatory reviews. And particularly, as we're talking here in structured databases that still have unstructured fields. And what we can do is our natural language processing offering brings a tool set, a toolbox. So you can see in blue at the bottom, we have natural language processing. And this is the key computational linguistics that understands whether there's a noun or a verb and the association within those. But we also build rules around ontologies. So those are the big dictionaries of genes and diseases, statistical methods, machine learning and chemical understanding and all sorts of different technologies. And you can see from the schematic in the middle, we've been looking for -- the tool has been looking for ejection fraction and then the numerics, and then able to normalize the numerics on the left-hand side. Which then on the right-hand side of this little slide, enables you to drive your analytics and outcomes. So going back to the poll. I'd like to ask you if you're using any natural language processing in your current work. So it may be something that you've not heard of and you're not using it at all. It may be that you understand the value of this kind of AI technology, natural language processing, to really start looking at unstructured text and you're considering projects. So please start voting. It may be that you're already actively looking at vendors and tools and surveying the market, to open source and commercial vendors, to find the right fit for a natural language processing tool. It may be that you have internal tools who are -- and you're regularly using it in your current projects in your current work. Or again, that you're regularly using it with an external vendor. So if you just have a think, and I'll give us another 20, 30 seconds to really think about getting more of you to answer these questions. So we've nearly got quite a good representation. I'll just give it another 10 seconds, and then we can go and have a look at the polling answers. [Voting]

Jane Reed

executive
#8

It's been really interesting. We've been running these kind of polls in our webinars for the past couple of years, and just seeing the trend is fascinating. Okay. That's great. So let's go on and see what we have. So that's very interesting. There's 56% aren't using natural language at all. Hopefully, by the end of this webinar, you'll be convinced. A lot considering projects. And then a few are using kind of regularly with an internal technology. That's really, really good to see. So in the interest of time, because we've had a few hiccups with this, I'm just going to give 2 use cases, and then I'll hand back to Kimberley. So this is one where it's looking at the ability of NLP to find the right information from electronic medical records. This is working with a U.S. university medical center, so a big academic medical center that has a pediatric center and is looking at a lot of phenotyping to understand potential diagnostics for -- to see potentially rare disease patients. So they have a big microarray labs. They do a lot of genetic testing. And within the genetics notes, those are taken and free text. And being able to find the individual phenotypic characteristics within that to help with the diagnostic. Currently, it was being done manually and it was very slow. So they used, IQVIA natural language processing to build rules to find specific phenotypes that matched the human phenotype ontology in order to improve that diagnostic yield. And as you can see on the right-hand side in the results, they were able to increase their curation, speed it up by 200-fold. Find nearly 2.5x more terms. And with a very high level of accuracy. So that's running natural language processing over free text of EHRs within an academic medical center. I'm switching slightly differently. So this is thinking about, again, pulling out information around rare disease, but this is from literature to help inform a registry. Again, this is -- it's a rare disease registry. And here, on a regular basis -- monthly basis, the registry wants to be able to understand, has there been new information around this particular gene, around the severity of the phenotypes for Hunter's outcome within the literature? So again, it's a way of building rules using natural language processing to find the phenotypes, the genes, the gene variations involved. And then, again, being able to surface that information for downstream analytics. Okay. I'm now going to hand back to Kimberley to take you through some of the analyses that we've -- IQVIA been working on to really show you the value of this combined approach of using NLP to pull out much more detailed information from Ambulatory EMR. Kimberley.

Kimberley Jordan

executive
#9

Thanks, Jane. So we looked at -- specifically at 3 rare diseases to learn a little bit more about what we could understand about them after we've used these NLP labels within ambulatory EMR. And they are homozygous familial hypercholesterolemia, hereditary angioedema and ALS. I'm going to start with homozygous familial hypercholesterolemia. So this is a very rare, life-threatening disease, which is characterized by very high levels of LDL-C and resulting accelerated premature atherosclerotic cardiovascular disease. Untreated, most patients do not survive past 30. And so early diagnosis and initiation of dietary changes and lipid-lowering therapy are critical. So what makes finding these homozygous familial hypercholesterolemia patients difficult is that there's no ICD code specific to this disease. There are 2 SNOMED codes that are specific to it, and we do have the ability to search for SNOMED codes in EMR data because they are recorded by providers. And we -- so we can also use the problem name as it's been recorded by the provider. And so they -- in that problem, they may specify homozygous familial hypercholesterolemia. But again, it can come in, in a variety of ways, which can make finding it, right, finding all the cases of it in the problem name or the problem description difficult. And again, that was the benefit of having the NLP come in and do the labeling for us. And so when we used the NLP labels and the SNOMED codes, so we use the NLP labels from the -- that were affixed based on the problem description. And we also used the SNOMED codes. And we found over 2,000 patients with a label for homozygous. Now -- and if there was an indication that it was a family history or it was a screening for homozygous familial hypercholesterolemia, we dropped it. But after that, there were still 2,000 patients that indicated that they may have homozygous familial hyperclustemia. Now we know that we can capture patients who are being screened or getting rule-out diagnoses. And so there can also be inaccuracies in medical histories. So for our final cohort, we actually required that they have at least 2 records with an indication of homozygous familial hypercholesterolemia. And we also included them in the cohort if they had at least 1 prescription for Juxtapid or EVKEEZA because those are drugs that are specific to homozygous familial hypercholesterolemia. And so we wound up with 129 ambulatory patients. And given the prevalence of this condition in the United States, this is actually -- this is a good and reasonable number. When we profiled it, the clinical data in Ambulatory EMR was consistent with a cohort that has homozygous familial hypercholesterolemia. We see high proportions of the cohort that have lab tests for cholesterol and liver toxicity. We are pleased to see that there's high percentages of the cohort that have prescriptions for statin as well as the down statin cholesterol absorption inhibitor, Zetia. And we also see the Juxtapid prescriptions, right, which we're expecting to see. So these clinical results are -- and this clinical profile of this cohort, right, is positive and makes us feel good about this group of patients that we were able to identify. And you can see here as well, we have excellent linking to other assets at IQVIA, such as PharMetrics Plus, the closed claims database. So you can also, by linking out, you can get information on cost, real cost. Moving on to hereditary angioedema. So hereditary angioedema is a disorder that's characterized by recurrent episodes of severe swelling. This swelling often occurs without a known trigger. These episodes of swelling can occur in the airway, which can restrict breathing significantly and lead to a life-threatening obstruction of the airway. There can also be episodes involving the intestinal tract that cause severe abdominal pain, nausea and vomiting. Because their swelling attacks can often be mistaken for allergies or non-HAE angioedema, patients can go years without a diagnosis. So blood test can be used to help diagnose this condition. There are preventive treatments available for hereditary angioedema as well as treatments that are used that can -- that are used to treat acute episodes of hereditary angioedema. So once again, with hereditary angioedema, we have a small population that has a disease. And at the same time, there's no ICD code that's used -- that's specific to HAE. There are 5 SNOMED codes available. And again, we can use those in Ambulatory EMR to find those patients, as well as using the NLP label that was a affixed based on the problem name provided by the provider. So -- and then again, Ambulatory EMR offers some very nice, rich clinical detail. We may be able to further refine our patient cohort with C1 inhibitor levels or C4 levels. You can also get insights into symptoms, such as angioedema, abdominal pain, airway symptoms due to swelling. So here, we found 969 patients. We applied the same rule. We required a minimum of 2 indications of an HAE diagnosis at least 6 months apart. We started -- so we went from 2,000 patients with a label based on a SNOMED code or an NLP label. And then once we required at least 2 records that were 6 months apart, we wound up with a cohort of 969 patients. So again, given the prevalence level of HAE in the United States, we're unsurprised by this cohort size. And there is some nice clinical evidence in Ambulatory that supports, right, the -- that we have found a nice HAE cohort. We have -- we see testing, right, C4 and C1 testing in this cohort at high proportions. We see that they are being treated for pain as well as gastric disorders, which we expect to see. And we see them being treated for anxiety and depression, which is also unsurprising. And then when we look at specific prescription medications, we see they're taking Firazyr, right, which is an acute treatment for HAE. We see treatment for GERD. Again, unsurprising, given how this disease manifests, we see -- and we see some of the preventive HAE treatments in this cohort as well. These patients are seeing their doctor an average of 8 visits per year in Ambulatory EMR. Again, that is consistent with a difficult-to-manage disease and a difficult-to-diagnose disease. Okay. And our final disease that we looked at is ALS. So as many of you probably know, ALS is a progressive nervous system disease that causes loss of muscle control. Eventually, it affects the muscles that are needed to speak, eat and breathe. There is no cure, and the average life expectancy is 3 to 5 years following the onset of symptoms. What can make it difficult for both providers and patients is that early symptoms can be confused with other conditions. There are drugs available to treat ALS, and so we will look for those in our profile of these patients. So what is different about ALS is that there are ICD codes to use to help identify ALS patients. There are also SNOMED codes. And again, to make -- to collect as many of these patients as possible, we also used the NLP labels that have been affixed to the -- based on the problem description. Again, one of the things that EMR data brings to our examination of these ALS patients is the potential for greater insight into their symptomology as well as the severity of the disease, right? And we'll see -- we potentially can see evidence of supportive therapies and indications of their quality of life in EMR data. And then, of course, claims can be used to obtain information. Closed claims can be used to obtain information about cost. And both open and closed claims can be used to understand what prescriptions have been filled. So we began with nearly 14,000 patients who had a record indicating ALS. Again, we dropped out anybody who it was clear that they were being screened for ALS or they were indicating that they had a family history of ALS. And so we ended up with 14,000 patients after that. And then we -- again, we required a minimum of 2 records with an ALS indication at least 6 months apart and wound up with 5,539 patients. The clinical information in Ambulatory, again, supports that these are ALS patients and allows us to learn more about them. We see that -- right. We see that a lot of them are being treated for pain, which is, again, unsurprising. They are suffering from anxiety and depression and being treated for that. And of course, they are being treated for muscle spasticity and gastric disorders. So we do see specifically riluzole, right, which is used to treat ALS. 46% of the patients are taking that. And then we see other types of treatments being prescribed, again, for the mini side effects of ALS, right? Muscle. We see them being treated with baclofen for muscle spasticity. Gabapentin for -- which is an anticonvulsant. We see them that they are taking drugs for -- other drugs for the types of side effects that you can suffer when you have ALS. Okay. So at this point, I'm going to turn it back to Jane, who is going to summarize, take us back through what we've covered in this webinar.

Jane Reed

executive
#10

Brilliant. Thank you. So I hope you found this as interesting as I did. So in essence, to summarize, obviously, we've been talking about using IQVIA Ambulatory EMR data for rare disease. And it really shows you that, if you can get access to these real-world data sets, you can create analyses to provide insights to inform clinical trials, drug development, commercialization. And particularly focusing in on the AEMR data set. With that NLP enrichment, we've been able to show that you can find those insights for rare disease patients. But obviously, you can apply this technology across a broad range of diseases. Particularly, we've been talking about how a natural language processing can find the information around rare disease in order to surface that much more effectively. And really, just to call out. Please do reach out to us. We're about to go to the Q&A session. But if you want to bring the power of that AEMR data into your research, of IQVIA's NLP into your research, then please do reach out to us. So just this one, thank you for listening. Obviously, we have my and Kimberley's e-mail address on here. But I'll now hand back to Melanie for the Q&A session.

Unknown Executive

executive
#11

Thank you, Jane. And thank you, everyone, for your patience with some of our technical issues today, as expected these days. So I'm going to run through, we had a few questions come through, and we have 15 minutes to answer these. But whilst I run through these, I'll just also -- I'll just leave a survey or a poll up in the screen here. So I know that we didn't manage to cover everything we wanted to do today. So if anyone would like any kind of follow-up discussions with some of our experts, then do you just let us know. Whether it's you'd like more -- to know more about NLP, AEMR or both of them together. Kind of interested, that's also fine. But we'd just like to give you the opportunities to learn as much as you need to know. So I'll try and go through these questions here in order. So got one here. So are things like provider type, provider specialty, included from claims? That's one for Kimberley there.

Kimberley Jordan

executive
#12

So there is provider specialty in claims, and you can join ambulatory data to claims data, right? There's an overlapping set of patients, so you can join them and pull in the specialty from claims. You can also get some specialty information in ambulatory. But joining to claims is a really nice way to get provider specialty.

Unknown Executive

executive
#13

Thank you. Next one here. So how is Ambulatory EMR different from EHR data?

Kimberley Jordan

executive
#14

They are effectively the same thing, they're electronic health records. Ambulatory EMR is -- has mostly primary providers, family health and internal medicine. So it's a narrower set of providers that are in that system, but it is effectively an electronic health record system or database.

Unknown Executive

executive
#15

Thank you I'm not sure if this one's for you or Jane. Can you give an example of HPO?

Jane Reed

executive
#16

So I'll take that one. So the human phenotype ontology, it's a well-known -- I suppose, a publicly available ontology that you can license. And it's got things like -- I mean, a silly one is it's got phenotypes describing hair type. But obviously, it's got things like stature, conjoined fingers, webbed feet. It's a huge ontology, particularly focused on the phenotypes that can help with unusual diagnostics. And if you want, I can certainly send you a little bit more background and a link to some papers that talk about the use of the HPO, human phenotype ontology. So yes, please just add that into the chat or drop me an e-mail, and I can send more information.

Unknown Executive

executive
#17

Thank you, Jane. How many health systems does IQVIA ecosystem have? And does it contain the full EHR notes, or just limited sets?

Kimberley Jordan

executive
#18

So the Ambulatory EMR has over 800 medium to large practices. That's -- it has more than 100,000 physicians. They are mostly -- a majority of them are general practitioners, but they're -- 60% of them are specialists. And we do not have the provider notes, what we do have our semi-structured fields, which are -- they are text, but they're not as comprehensive as the nodes. They're much more limited. So an example that we talked a lot about on the webinar about and what we applied to NLP to was problem description. So doctors might be using a pick list to put in the problem or they may be typing it in theirselves. Different practices might be using different pick lists or have different rules around typing it in. So you see a wide range of information being entered into that field, but it is limited in scope. But all of the tables do have -- or a majority of these tables do have a field that does contain that semi structured text. So for example, you can...

Jane Reed

executive
#19

Sorry, go ahead.

Kimberley Jordan

executive
#20

I was -- sort of the meds table and the orders table, right, also have these types of semi-structured fields that provide additional information that could be very, very rich. In the vitals table, there's some really nice values in there in the semi-structured field. Go ahead, Jane. I'm sorry.

Jane Reed

executive
#21

Just to comment. I mean, obviously, at IQVIA, we have a huge number of partnerships as well. And we also partner with some vendors who -- actually means that we can get access to patient narratives. So -- and those are obviously much -- they're very -- a different sort of use case, but they're very rich in sort of comments. You may be would be able to use those much more around things like looking at switching, looking at patient behaviors and that kind of thing. So obviously, there's -- if you have particular challenges that you're trying to solve, then as Kimberley said, we have a broad range of data within IQVIA, and we can also access with our partners other types of medical record.

Unknown Executive

executive
#22

Thanks, Jane. Another one here for you, Jane. What is the data type where NLP was used on doctors and patient notes? I think you covered that in your use case. Is that right?

Jane Reed

executive
#23

If I -- well. So I mean, IQVIA have developed our own NLP. So IQVIA acquired an NLP company called Linguamatics about 3, 4 years ago. And Linguamatics is based in the U.K. and has been providing natural language processing for 2 decades to pharma companies, to health care companies. But we also use state-of-the-art models. So as I said, when we are looking about that kind of toolbox, we have the computational linguistics, but we also use machine learning-based NLP. So it's a whole range of different tools within the platform in order to, in essence, solve the right challenges that we need. I think that answered the question. But if not, please e-mail me and I can follow up. Melanie, you're on mute.

Unknown Executive

executive
#24

So another one here you, Jane, on NLP. So did you run a control analysis using only the structured data to obtain percentage of patients found thanks to the NLP method? And would this help assess the benefit of having NLP on unstructured text -- I mean, that's cut off there.

Jane Reed

executive
#25

Yes. I think probably I need to hand over to Kimberley for that. But as I showed in one of the slides, we really did -- for some of the rare disease, we weren't finding any patients at all. And using the natural language processing, we were able to kind of increase that by, in essence, 100%. But Kimberley, can you give any comments on how we did? Did we do a control?

Kimberley Jordan

executive
#26

Sure. What we did was, for each of the 62 rare diseases that we identified, we did actually look for them without the NLP label, and then determined the bump that we get from using the NLP label. And it ranged from a 5% bump to, as Jane said, 100%, where we weren't seeing SNOMED codes that were helping with identifying them. We weren't seeing -- either they weren't being entered in the database or they simply weren't -- they weren't available. And again, in the most frequent cases, right, there's no ICD-10 codes available. And so coming in, these are the low -- these are the -- it tends to happen with the hardest -- as you'd imagine, with the hardest-to-discover rare diseases, right? The more rare they are, the less likely we are to see ICD-10 codes for them and even SNOMED codes for them, or SNOMED codes being entered.

Unknown Executive

executive
#27

Thank you, Kimberley. It's actually one for you again here, Jane. How can medical scribes help NLP be more accepted?

Jane Reed

executive
#28

I suppose, I mean, in essence, with any innovative technology, it's -- to get that technology accepted, it's using it more, talking about it more. I mean, in essence, going back to that last question we just had, evangelizing about the value you get if you have that kind of tool. I mean, one of the critical things is -- and was talked about a lot kind of 3, 5 years ago, a clinician, you can't build an EMR system with enough drop-downs to capture the information. So we are always going to be needing some kind of text-mining tool if you're really going to drill into that context, that nuance, that deep value that's in medical records. So when medical scribes are thinking about how to bring that value, it's really knowing the value, understanding. And really -- I mean, you're absolutely right. A lot of people are like, "I don't trust the answers." But it's interesting. When you look at studies of human curation of almost any kind of record. Well, it's pulling information from literature. I mean, the FDA have done studies. And everyone sort of assumes people are going to be 100% accurate. But actually, when you do the studies, the inter-annotator agreement can be as low as 60%. So if we can build rules that are systematic and comprehensive and can pull out information at 85%, 95% accuracy. Then while, yes, if anyone tells you they have a tool that's 100% accurate over all data types, then either buy the company or they're possibly not being very truthful. So really, NLP can definitely bring huge value. And I think it's really understanding how to best use it and then sharing that information. Thanks for the question. Nice one.

Unknown Executive

executive
#29

Thanks. And just another one on NLP here. So what type of NLP pipelines were used in these studies?

Jane Reed

executive
#30

So actually, I'm not close enough to the actual studies. But I mean, we can certainly build NLP pipelines. So the way that we generally do this kind of work, is we'll get our experts looking at some of the data, building rules, testing rules. And once we've developed a set of rules, then we can then deploy that in a data factory, in the pipeline. Kimberley, do you have anything more on that?

Kimberley Jordan

executive
#31

I don't. I'm sorry, Jane, I do not.

Jane Reed

executive
#32

Yes. So in essence, I mean, we have our own -- IQVIA has our own way of deploying NLP pipelines, but we can also use -- we have a very rich set of APIs. We can embed provide SDKs for that. We can embed into workflow tools such as KNIME. We have our own data factory to build pipelines. But there's always that first piece of work, taking the out of the box NLP modules, reviewing those, developing those to make them accurate for your own data set, and then being able to deploy in a workflow or pipeline.

Unknown Executive

executive
#33

Thanks, Jane. But another one here, sort of a 2-parter for you, Kimberley. So what provides PharMetrics Plus claims data? And what kind of analysis can be made with PharMetrics? And there's just a question here surrounding social determinants data. So it says, I guess, Ambulatory EMR offers social determinants data differently. So we can repeat any of that if you need to.

Kimberley Jordan

executive
#34

Okay. So PharMetrics plus is a closed claims database. That means it has adjudicated health care claims in it. And it can be used for anything from targeting providers to do -- understanding the burden of the disease, to understanding the cost of the disease, understanding drug switches and that sort of thing. And we get -- the data comes in from a wide variety of health plans. Ambulatory EMR data does capture social determinants of health to some extent. And we -- in those -- again, in those semi-structured fields. We also can link to other sources to collect social determinants of health. So IQVIA has various options to offer people who are looking for social determinants of health information.

Unknown Executive

executive
#35

Thanks, Kimberley. So we are out of time, and we have a number of other questions that we haven't answered. So we will have to call it, but we will follow up separately with everyone else who has any outstanding questions that haven't been answered. And if anyone thinks of any other questions that they have, then please do feel free to get in touch with us. So I'll just put these e-mails back up on the screen for you. So there are Jane's and Kimberley's e-mails. And please do again touch with us. Thank you very much to our speakers today. Thank you for those of you who joined, and I hope you all have a lovely rest of your day.

Kimberley Jordan

executive
#36

Thank you very much, guys. Bye.

Read the full transcript via the API

You're viewing the first half of this call. Get the complete IQVIA Holdings Inc. transcript — plus 251,000+ transcripts from 12,000+ companies, speaker segments, AI summaries and full-text search — through the EarningsCalls.dev API.

Get the API View API docs →

This call discussed

For developers and AI pipelines

Programmatic access to IQVIA Holdings Inc. earnings transcripts and 251,000+ others is available through the EarningsCalls.dev REST API. Plans from $24.99/month — full transcripts, speaker segments, full-text search, and the recently-added /api/v1/transcripts/recent polling endpoint for ETL pipelines.