Caution urged in using artificial intelligence in medical expertise evaluation

Artificial intelligence (AI) holds great promise in various aspects of healthcare. Already, ChatGPT is able to produce clinical letters indistinguishable from those written by human doctors, while Google has developed an AI agent capable of clinical history-taking and diagnostic reasoning.

Judge Business school school

Yet use of AI poses risks, especially in relying on benchmarks to evaluate medical expertise and practice, according to research from the Psychometrics Centre at Cambridge Judge Business School.

Current general medical AI (GMAI) evaluations are based heavily on benchmarks – typically questions from medical licensing exams – and such evaluations are usually based on comparing the GMAI’s score to a passing score for aspiring human doctors. But such a comparison “is unable to inform the types of errors GMAI makes, identify their weaknesses, or provide insight into GMAI’s performance on tasks not within the benchmark assessment”, finds the research.

For a human doctor, doing well on a medical exam is a good predictor of performing well in a range of medical tasks that are not in the exam. But for GMAI, that is not necessarily the case. Despite passing the exam with flying colours, GMAI still makes unexpected mistakes that a human doctor with an equivalent score would never make. That’s a problem, because we need to be able to trust GMAI in the same way that we trust doctors.

“Who knows what the next patient will bring in: it may be something that no one has ever seen before – as both disease patterns and treatment trends change over time – and these generalised benchmarks may offer no predictive help in evaluating whether an AI doctor may deal with such a new situation effectively,” says Professor David Stillwell of Cambridge Judge Business School.

Read a longer article 



Looking for something specific?