Validierung von künstliche Intelligenz-Algorithmen für die chirurgische Praxis
摘要
Artificial intelligence (AI) is increasingly being used in surgery; however, the validation of such systems is often methodologically insufficient.
ObjectiveWhich validation issues arise in surgical AI and what requirements can be derived for clinically meaningful validation strategies?
MethodsMetric-related pitfalls reported in the literature were analyzed, combined with insights from the interdisciplinary consensus process “metrics reloaded” and its ongoing extension to surgical applications.
ResultsRecurring weaknesses are observed at the levels of data, metrics and reporting. The lack of consideration of temporal structures and aggregation in video data is particularly critical.
DiscussionA structured, clinically grounded validation is essential for the safe use of surgical AI. The metrics reloaded procedure is currently being adapted to address surgery-specific requirements.