In this episode, I’ll discuss the performance evaluation of a large language model for medication management tasks.
With the growth of artificial intelligence (AI) tools based on large language models, there is uncertainty on how well they perform clinical tasks such as medication management. A group of researchers published in American Journal of Health-System Pharmacy a report evaluating the performance of a large language model for medication management tasks.
The authors gave OpenAI’s GPT-4o model tests on 3 medication tasks: identifying available formulations for a given generic drug name, identifying drug-drug interactions (DDIs) for a given medication regimen, and preparing a medication order for a given generic drug name. They scored results by clinical pharmacists or by comparing to the performance of LEXI-Comp drug interaction software on the same task.
For the first task of drug-formulation matching, GPT-4o had 49% accuracy for generic medications being matched to all available formulations, with an average of 1.23 omissions per medication and 1.14 hallucinations per medication.
For the second task of drug-drug interaction identification, the accuracy was 54.7% for identifying the DDI pair.
For the third task, GPT-4o generated order sentences containing no medication or abbreviation errors in 65.8% of the cases.
The authors concluded:
Model performance for basic medication tasks was consistently poor. This evaluation highlights the need for domain-specific training through clinician-annotated datasets and a comprehensive evaluation framework for benchmarking performance.
Leaving aside that the model used in this study is over 2 years old, which is an eternity in AI development, I think large language models, by their design, will fail tasks like these and should be considered for use differently. Because they work by predicting the next word that should logically follow the last words used, they cannot be expected to determine whether a drug interaction is present with 100% accuracy.
The weakness of using LEXI-Comp to check for drug interactions is the time it takes to put a medication list into a form the program can check by selecting one medication at a time from a dropdown list. Once the list is in, LEXI-Comp’s software will always perform the correct check for interactions, 100% of the time. Large language models’ strength is language, meaning that a messy list with some misspellings and missing spaces can easily be copied or dictated and ingested by the LLM and parsed correctly into a format that LEXI-Comp software can then run a check on. So, in my opinion, a better way to test a large language model would involve having the AI parse the list into LEXI-Comp’s preferred format, then compare the speed and accuracy of a pharmacist checking drug interactions manually with LEXI-Comp with one leveraging the LLM to quickly get the list into LEXI-Comp. This way, the pharmacist stays in the loop performing the medication management task while the AI model serves as leverage by taking on the tedious work of entering a medication list into the software.
To access my free download area with 20 different resources to help you in your practice, go to pharmacyjoe.com/free.
If you like this post, check out my book – A Pharmacist’s Guide to Inpatient Medical Emergencies: How to respond to code blue, rapid response calls, and other medical emergencies.
Leave a Reply