From text to insight: a systematic literature review of keyword and keyphrase extraction techniques for healthcare text mining
摘要
The rapid digitization of healthcare has generated vast repositories of unstructured textual data—including electronic health records, clinical narratives, and patient-generated content—creating an urgent need for automated methods to distill salient information. Keyword and keyphrase extraction, the task of identifying terms and multi-word expressions that capture a document’s core content, has emerged as a fundamental natural language processing capability for unlocking clinical insights. This systematic literature review examines the state of research on extraction techniques applied to healthcare domains, synthesizing evidence from 110 peer-reviewed studies. We organize existing approaches into four methodological generations: statistical unsupervised methods, supervised sequence labeling, transformer-based methods (encompassing both fine-tuned sequence labeling and embedding-based unsupervised approaches), and emerging large language model approaches. Our analysis identifies key healthcare applications including systematic review automation, clinical note summarization, patient experience analysis, consumer health question answering, and coding assistance. We find that while unsupervised methods remain valuable for their computational efficiency and domain independence, transformer-based models demonstrate superior performance, particularly when fine-tuned on biomedical corpora, though at substantial computational cost. Emerging approaches include multi-task learning frameworks and hybrid systems that combine pattern-based precision with model flexibility. We conclude by identifying critical research gaps including the preservation of multi-word clinical concepts, scarce evaluation in real-world clinical workflows, and persistent hallucination risks in model-based extraction. This review provides researchers and practitioners with a structured framework for method selection based on their specific constraints and identifies six prioritized research directions for future investigation.