Epistemic injustice and phenomenological reductionism in psychiatric AI: an empirical-philosophical analysis
摘要
Large language models (LLMs) are increasingly deployed in clinical psychiatry. This paper examines three interrelated ethical risks: the epistemic marginalisation of patients (Fricker), the replacement of interpretive understanding with statistical pattern-matching (Jaspers), and the diffusion of moral responsibility in ways that may undermine clinician accountability (Dean et al.). I performed a secondary analysis of the HealthBench public dataset (57,237 rubric criteria, 5000 clinical prompts). A keyword filter identified 4752 criteria (8.3%) relating to psychiatry and mental health. Negative-scoring failure rates across five evaluation axes were mapped onto the three philosophical frameworks. Importantly, HealthBench is an evaluation framework of physician-derived rubric criteria; it does not contain recorded outputs from any specific LLM. Findings reflect expected scoring against these criteria, not measured performance of any particular model. Psychiatric prompts showed a nominal increase in context-awareness failures (26.8% vs. 25.0%), though no axis difference survived Bonferroni correction (adjusted α = 0.01). Qualitative analysis revealed clinically significant failure patterns: psychosis risk blindness, suppressed emergency referrals, and systematic under-hedging (mean 2.34 vs. 2.47), suggesting structural resonance with the deficits identified. LLMs deployed without adequate governance may marginalise patient narratives, miss clinically significant phenomenological signals, and erode accountability structures underpinning safe psychiatric care. For clinicians, uncritical deference risks delayed emergency referrals, diffused moral responsibility and attenuated epistemic agency. I propose three safeguards: phenomenological oversight, epistemic justice auditing, and explicit emergency referral accountability protocols.