From Code to Insight: How LLMs Help and Hinder Qualitative Research
摘要
This study evaluates the performance of large language models (LLMs) in the qualitative analysis of an interview transcript. Four state-of-the-art LLMs—GPT-4, LLAMA 3.1, Gemini 2.0 flash, and DeepSeek-V3—were compared with human qualitative analysis using both inductive and deductive approaches. The deductive analysis employed the socio-ecological model as a framework to examine the complexities of supporting neurodivergent individuals in rural communities. Results indicate that Gemini achieved the highest cosine similarity to human coding in the inductive approach, while DeepSeek performed best in the deductive approach. Graph neural network (GNN) visualizations revealed that certain models struggled to capture the holistic context of the interview, demonstrating limitations in comprehension and contextual analysis compared to human coders. The models used in this study were freely available, and participant privacy was protected by anonymizing the transcript. The study was approved by the institutional review board. These findings highlight both the potential and the challenges of employing LLMs to augment qualitative research, particularly in nuanced and context-dependent data analyses. We discuss the implications of using LLMs to enhance qualitative research processes and propose future directions for improving their alignment with human analytical processes.