Why Clinical AI Needs More Than Accuracy
As hospitals race to adopt predictive models, data scientist Mansi Goel argues that scoring well on a benchmark is the easy part and often the wrong thing to measure.
- Utility News
- 4 min read
As hospitals race to adopt predictive models, data scientist Mansi Goel argues that scoring well on a benchmark is the easy part, and often the wrong thing to measure.
Ask most people what makes a clinicalAI model good, and they will point to a number: its accuracy, or the area under some curve. Ask Mansi Goel, and she will tell you that number is where the real work begins, not where it ends.
“A model that looks accurate on paper but quietly fails for one subgroup of patients has not solved the problem,” says Goel, a data scientist at Lucem Health who builds machine learning models for early disease detection. “It has just moved the risk somewhere less visible.”
The problem with a good score
It is a distinction the healthcare industry is only beginning to reckon with. Predictive models are proliferating across hospitals and health systems, promising to flag disease earlier and to direct scarce clinical attention more wisely. But the same qualities that make these tools powerful, their scale, their speed, and their air of objectivity, also make their failures easy to overlook. A model that underperforms for a particular group of patients does not announce itself. It simply, and silently, gets those patients wrong.
A ‘clinical-first’ philosophy
Goel’s answer is an approach she describes as clinical-first. In her view, a model is only as useful as a clinician’s ability to understand and act on it. A prediction that cannot be interpreted, or that does not fit the way care is actually delivered, is not a breakthrough; it is a dead end. “Build with the end user in mind from day one,” she says. “In healthcare, that means understanding the clinical workflow before you write a single line of model code.”
That conviction shows up in how she works. Rather than treating a model as finished once it performs well on a benchmark, Goel takes ownership of the entire lifecycle, from the messy reality of raw electronic health record data through feature engineering, training, evaluation, validation, and the final and hardest mile of deployment into a live clinical setting. She has applied that approach across several conditions, including cardiac arrhythmia, liver disease, and type 1 diabetes, building models that run on the data health systems collect rather than on idealized research datasets.
That last point matters more than it sounds. Real-world health records are noisy, incomplete, and coded differently from one institution to the next. “Data quality is not a given, it is earned,” Goel says. “Investing in data engineering upfront saves enormous time later, and it is often the difference between a model that generalizes and one that only works in the lab.”
Accountability as a design requirement
Underneath the technical rigor is a concern that is fundamentally ethical: accountability. The patients most likely to be harmed by a poorly tested model are often the ones already underserved by the healthcare system, the groups least represented in the data a model learns from. Goel argues that checking how a model performs across those subgroups cannot be an afterthought bolted on at the end. “If we want clinicians to act on what these models say,” she says, “the accountability work has to be as rigorous as the modeling itself.”
It is a view increasingly shared at the frontier of the field, where responsible and explainable AI have moved from academic side-topics to central design requirements. For all the sophistication of the tools she builds, Goel keeps returning to a simple test. “Real-world impact is the only benchmark that matters,” she says. “In clinical AI, a model that cannot be used by a clinician is a model that does not help a patient.” The algorithms, in other words, are necessary. They are just not, on their own, enough.
Published By : Aniket Datta
Published On: 4 September 2026 at 20:14 IST