AI Tool Audits Colonoscopy Quality at Scale

Researchers have developed an AI system that automatically reviews colonoscopy footage to measure procedural quality, offering a scalable alternative to manual audits.
Colonoscopies are a critical defense against colorectal cancer, yet the quality of these procedures varies significantly among physicians. Medical guidelines recommend regular performance reviews to ensure safety and effectiveness, but conducting these audits manually is time-consuming and difficult to implement across large healthcare systems. This gap leaves many practices without consistent feedback on their endoscopic techniques.
A team from Northwestern Medicine has introduced an artificial intelligence tool designed to close this gap. By analyzing video recordings of procedures, the software can rapidly assess key quality metrics. According to a study published in The American Journal of Gastroenterology, this represents the first instance of an AI system comprehensively measuring the quality of thousands of colonoscopies.
Automated Review of Procedural Footage
The study, reported by GN technics/ai (en-US), involved nearly 19,000 colonoscopies performed by 55 physicians over eleven months. The AI software reviewed the procedural videos to identify critical moments, such as when the scope reached the beginning of the colon, when it was withdrawn, and when polyps were removed. This automated detection allows for a standardized assessment that does not rely on human memory or manual data entry.
Key metrics include withdrawal time, which tracks how long the physician spends examining the colon as the tube is removed. The AI’s calculation of this time closely matched records kept by nurses, demonstrating high accuracy. The tool also tracked indicators that are difficult for humans to measure at scale, such as the specific techniques used for polyp removal.
Balancing Automation With Human Skill
While the tool offers a scalable solution for quality monitoring, experts caution that its role is currently limited to post-procedure assessment. Dr. Rajesh Keswani, the study’s lead author, noted that the system evaluates quality after the fact rather than guiding the physician in real-time. This distinction is important because the primary goal is to provide feedback for improvement, not to replace clinical judgment during the procedure itself.
Concerns About Physician Proficiency
There is a broader debate regarding how AI integration might affect doctor training and skill retention. Previous research suggested that reliance on AI for polyp detection could lead to a decline in manual proficiency over time. Keswani acknowledged this potential trade-off, describing it as a risk of "deskilling." However, he also suggested that AI could reveal blind spots in technique, potentially serving as a teaching tool for trainees.
The team is now studying how AI can be used in the education of new gastroenterologists. The ultimate aim is to create a feedback loop that enhances care quality without eroding the fundamental skills required to perform these delicate procedures safely and effectively.






