A Reproducibility Study on Consistent LLM Reasoning for Natural Language Inference over Clinical Trials
摘要
With the rapid expansion of AI in healthcare, ensuring that language models can reason accurately and consistently within the medical domain is essential for enhancing clinical decision-making. Consistent reasoning is particularly challenging, since once a model outputs a judgment for a given medical statement, it should retain that judgment when faced with a syntactically altered version of the statement. In addition, when the altered version of the statement reflects a semantic shift, the model should adjust its judgment accordingly. In this paper, we describe the process of reproducing state-of-the-art methods for safe biomedical Natural Language Inference for Clinical Trials (NLI4CT), emphasizing model vulnerability to small input variations and the inherent complexity of clinical trial reports (CTRs). We specifically evaluate the reasoning capabilities of Large Language Models (LLMs) in the SemEval-2024 NLI4CT dataset, thus considering a task that focused on robustness and which serves as a good proxy for real-world applications. To improve thoroughness, we extended this study and explore a broader set of techniques, establishing baseline scores for several widely used models. We conclude with an analysis of the results, highlighting key insights and empirical lessons that contribute to future research in this domain.