Background <p>The integration of generative artificial intelligence (AI) tools into nursing practice has accelerated documentation processes but it has also raised concerns regarding the completeness, accuracy, and clinical safety of AI-generated care plans. Despite the growing use of tools like ChatGPT, Gemini, and PopAI in clinical and academic settings, no validated instrument currently exists to assess the quality of such documentation across the nursing process.</p> Objective <p>This study aimed to develop and validate the Nursing Process Evaluation Tool (NPET), a multidimensional instrument designed to assess the quality of AI-generated nursing documentation within the ADPIE (Assessment, Diagnosis, Planning, Implementation, Evaluation) framework.</p> Methods <p>A two-phase cross-sectional study was conducted. Phase I focused on item development and content validation via two rounds of expert review (<i>n</i> = 23). Phase II evaluated the NPET’s psychometric properties by assessing 64 AI-generated nursing care plans based on eight clinical scenarios using eight AI models. A total of 368 individual expert ratings were yielded. Reliability (Cronbach’s α, ICC), content and construct validity (I-CVI, S-CVI/Ave, exploratory factor analysis), and comparative model performance (repeated-measures ANOVA with Tukey post hoc tests) were analyzed.</p> Results <p>The NPET demonstrated strong content validity (S-CVI/Ave = 0.88) and excellent internal consistency (α = 0.85–0.94 across domains). Inter-rater reliability was high (ICC_average = 0.85–0.94). Exploratory factor analysis supported the proposed structure: four domains were unidimensional, while the Assessment domain revealed two interpretable factors. Although the overall ANOVA did not reveal statistically significant differences among AI models (F (7, 360) = 1.57, <i>p</i> = 0.144, ω² = 0.01), descriptive trends and post hoc tests showed that paid models consistently outperformed free versions. PopAI Paid achieved the highest mean NPET score (M = 3.44 on a 4-point scale), followed by ChatGPT Paid (M = 3.37), while Microsoft Copilot scored the lowest (M = 2.99). The largest pairwise difference—between PopAI Paid and Copilot—yielded a moderate-to-large effect size (Cohen’s d = 0.60).</p> Conclusion <p>The NPET is a valid and reliable tool for evaluating the quality of AI-generated nursing care plans. While the overall ANOVA did not yield statistically significant differences across AI models, the consistently high performance across tools and meaningful differences observed in descriptive and post hoc comparisons support the tool’s utility in nursing education, clinical auditing, and AI benchmarking. Future research should explore its application in real-world documentation and monitor its adaptability to evolving AI technologies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development and validation of the Nursing Process Evaluation Tool (NPET): a multidimensional instrument for assessing the quality of AI-generated nursing documentation

  • Mohammad Othman Abudari,
  • Manar Abu-abbas,
  • Mohammad Al-Ma’ani,
  • Mutaz foad Alradaydeh,
  • Hamza Alduraidi

摘要

Background

The integration of generative artificial intelligence (AI) tools into nursing practice has accelerated documentation processes but it has also raised concerns regarding the completeness, accuracy, and clinical safety of AI-generated care plans. Despite the growing use of tools like ChatGPT, Gemini, and PopAI in clinical and academic settings, no validated instrument currently exists to assess the quality of such documentation across the nursing process.

Objective

This study aimed to develop and validate the Nursing Process Evaluation Tool (NPET), a multidimensional instrument designed to assess the quality of AI-generated nursing documentation within the ADPIE (Assessment, Diagnosis, Planning, Implementation, Evaluation) framework.

Methods

A two-phase cross-sectional study was conducted. Phase I focused on item development and content validation via two rounds of expert review (n = 23). Phase II evaluated the NPET’s psychometric properties by assessing 64 AI-generated nursing care plans based on eight clinical scenarios using eight AI models. A total of 368 individual expert ratings were yielded. Reliability (Cronbach’s α, ICC), content and construct validity (I-CVI, S-CVI/Ave, exploratory factor analysis), and comparative model performance (repeated-measures ANOVA with Tukey post hoc tests) were analyzed.

Results

The NPET demonstrated strong content validity (S-CVI/Ave = 0.88) and excellent internal consistency (α = 0.85–0.94 across domains). Inter-rater reliability was high (ICC_average = 0.85–0.94). Exploratory factor analysis supported the proposed structure: four domains were unidimensional, while the Assessment domain revealed two interpretable factors. Although the overall ANOVA did not reveal statistically significant differences among AI models (F (7, 360) = 1.57, p = 0.144, ω² = 0.01), descriptive trends and post hoc tests showed that paid models consistently outperformed free versions. PopAI Paid achieved the highest mean NPET score (M = 3.44 on a 4-point scale), followed by ChatGPT Paid (M = 3.37), while Microsoft Copilot scored the lowest (M = 2.99). The largest pairwise difference—between PopAI Paid and Copilot—yielded a moderate-to-large effect size (Cohen’s d = 0.60).

Conclusion

The NPET is a valid and reliable tool for evaluating the quality of AI-generated nursing care plans. While the overall ANOVA did not yield statistically significant differences across AI models, the consistently high performance across tools and meaningful differences observed in descriptive and post hoc comparisons support the tool’s utility in nursing education, clinical auditing, and AI benchmarking. Future research should explore its application in real-world documentation and monitor its adaptability to evolving AI technologies.