Benchmarking large language models against qualitative coding and natural language processing in decoding public sentiment on urban upzoning
摘要
Urban planners routinely engage with extensive textual materials, such as zoning codes, comprehensive plans, and public comments. The large volume of textual data being generated in cities offers new opportunities to use emerging computational technologies to process and analyze textual data. In this research, we compare how natural language processing (NLP) techniques and large language models (LLMs) compare to human qualitative coding techniques to identify public sentiment and topics contained in public comments about the Minneapolis 2040 upzoning. We use a custom rubric developed in collaboration with urban planners to assess outputs across these different methods, scoring outputs on factors such as accuracy, convergence, creativity, efficiency, and interpretability. Additionally, we conduct interviews with practicing urban planners to understand their perceptions of integrating these computational techniques into their existing workflows. We find that using NLP techniques are helpful in providing urban planners with an aerial view of their data, but require additional human interpretation. In contrast, using LLMs markedly improves efficiency, interpretability, and descriptiveness over traditional NLP techniques, but requires human validation to address concerns related to social biases and equity. We further find that urban planners are open to using new text processing technologies, but have reservations about entirely outsourcing decision-making to AI tools, viewing AI technologies more as “co-pilots” rather than autonomous agents. Our findings underscore the importance of integrating human judgment into using computational tools to develop a more informed, equitable, and reflective practice in an era of expanding urban data and computational technologies.