How to optimize large language model performance for human-centric post-occupancy evaluation of underground public space: From model selection to model tuning
摘要
Underground public space (UPS) is vital for three-dimensional urban development, and its post-occupancy evaluation (POE) forms a basis for human-centric planning, design and management. Although large language models (LLMs) provide powerful tools for POEs based on social media data (SMD), LLM application sare often constrained by lacking domain-specific knowledge. This study developed a framework to optimize LLM performance for SMD-based POEs. Model selection, parameter setting, prompt tuning and fine-tuning were examined for two specialized annotation tasks using 34 UPS cases and 14 LLMs. The methods proved satisfactory effects, achieving peak macro accuracy of approximately 0.95 and F1-scores over 0.83 for the two tasks. Reasoning models outperformed general-purpose models by up to 25.74% (F1-scores) in complex tasks, though with lower consistency. Prompt tuning was highly effective but task-dependent with a peak increment of 21.99% in macro F1-scores. Parameter settings and system role modifying had minimal effects. Excessively long prompts may reduce performance. Fine-tuning significantly improved performance and reduced prompt dependency. POEs identified spatial form, wayfinding and operational management as high-priority factors for future design and renovation of UPS. The findings offered a practical and cost-effective optimization pathway for applying LLMs to other similar textual analysis in underground space studies.