Medical AI Accuracy Depends on Consistent Results

New research suggests that the consistency of an AI’s diagnosis is a better indicator of accuracy than its internal confidence scores.
Researchers in Dresden have demonstrated that medical artificial intelligence systems can produce diagnoses that clinicians can trust, provided the technology operates within strict local boundaries. The study, published in Nature Medicine, highlights a critical shift in how these systems are evaluated. Instead of relying on the model's internal probability estimates, which often fail to reflect actual accuracy, the team found that the stability of the answer across multiple runs is the strongest predictor of a correct diagnosis.
This approach addresses two major hurdles in clinical adoption: data privacy and reliability. By keeping sensitive patient data on-premises, the system prevents external sharing of sensitive information. Simultaneously, it offers a practical method for doctors to distinguish between confident and uncertain AI outputs, ensuring that human oversight remains central to the diagnostic process.
Local Processing Ensures Data Privacy
A key challenge in deploying AI in healthcare is the protection of sensitive patient information. External cloud services often pose risks regarding data ownership and security. The system developed by the team at TU Dresden and Dresden University Hospital resolves this by running entirely on local hospital infrastructure. This on-premises architecture ensures that electronic patient data, including laboratory values and medical history, remains under the control of the institution rather than being transmitted to third-party servers.
According to GN technics/ai (en-US), this local setup is fundamental to the system's design. It allows the AI to interact with complex clinical cases without exposing private health records to external networks. This structural choice supports the broader goal of integrating AI into clinical workflows without compromising the ethical and legal obligations hospitals have to protect patient confidentiality.
Consistency Beats Internal Probability Scores
The study revealed a significant disconnect between an AI model's self-reported confidence and its actual diagnostic accuracy. When the researchers tested the system with standardized clinical cases, including conditions like pneumonia and appendicitis, they found that the model's internal likelihood scores were poor indicators of correctness. Even when the system was stressed by removing reliable information, its probability signals did not accurately reflect the drop in diagnostic performance.
In contrast, the stability of the diagnosis was highly informative. If the AI arrived at the same conclusion when processing the same case repeatedly, the diagnosis was likely correct. In benchmarks, the best locally operated model reached the correct diagnosis in approximately 90% of cases. Furthermore, when physicians reviewed a subset of these automated diagnoses, their consensus agreed with the AI in more than 90% of instances, validating the consistency-based reliability metric.
Selective Autonomy for Clinical Trust
The researchers propose a framework of selective autonomy to integrate these tools into practice. Rather than requiring clinicians to verify every single AI decision, or blindly trusting the system, this model allows for a tiered approach. Cases where the AI demonstrates high consistency are treated as reliable, while those with unstable outputs are flagged for mandatory human review. This trade-off prioritizes safety by ensuring that uncertainty is always handled by medical professionals.
Jakob N. Kather, the senior author, emphasizes that the goal is support, not replacement. The system is designed to be understandable to clinicians, clearly indicating when its outputs are uncertain. This transparency is essential for maintaining professional responsibility. By distinguishing between strong and weak reliability signals, the technology can assist in decision-making while keeping the final diagnostic authority firmly with the human doctor.






