Yet use of AI poses risks, especially in relying on benchmarks to evaluate medical expertise and practice, according to research from the Psychometrics Centre at Cambridge Judge Business School.
Current general medical AI (GMAI) evaluations are based heavily on benchmarks – typically questions from medical licensing exams – and such evaluations are usually based on comparing the GMAI’s score to a passing score for aspiring human doctors. But such a comparison “is unable to inform the types of errors GMAI makes, identify their weaknesses, or provide insight into GMAI’s performance on tasks not within the benchmark assessment”, finds the research.
For a human doctor, doing well on a medical exam is a good predictor of performing well in a range of medical tasks that are not in the exam. But for GMAI, that is not necessarily the case. Despite passing the exam with flying colours, GMAI still makes unexpected mistakes that a human doctor with an equivalent score would never make. That’s a problem, because we need to be able to trust GMAI in the same way that we trust doctors.
“Who knows what the next patient will bring in: it may be something that no one has ever seen before – as both disease patterns and treatment trends change over time – and these generalised benchmarks may offer no predictive help in evaluating whether an AI doctor may deal with such a new situation effectively,” says Professor David Stillwell of Cambridge Judge Business School.