Beyond unimodal metrics: a multimodal evaluation framework for geo-referenced content on social media
摘要
Social media platforms generate large volumes of real-time, geo-referenced content during natural hazards, enabling multimodal topic models to uncover latent themes and their spatial organisation. However, the evaluation of such models typically relies on unimodal metrics, primarily semantic coherence, which overlook spatial structure, temporal dynamics, and the practical usefulness of discovered topics. This paper proposes a multidimensional evaluation framework for disaster-related social media analysis that integrates semantic, spatial, temporal, and operational indicators of topic quality. The framework distinguishes between Utility Information Value (UIV), capturing the informational usefulness of topic representations, and Diagnostic Information Value (DIV), capturing the structural and hazard-consistent validity of spatial and spatio-temporal topic patterns. We examine the framework through a comparative benchmarking analysis of multimodal topic models (MultiGraph and JSTTS) alongside a strong unimodal baseline, using eight geo-referenced datasets from X (Twitter) and Bluesky covering earthquakes, floods, hurricanes, and wildfires under a unified experimental setup. Results show that multimodal models achieve higher actionability (up to 0.75), while the unimodal baseline shows greater semantic diversity (up to 0.99). Spatio-temporal interaction varies across hazards (0.07-0.45), and spatial patterns exhibit hazard-dependent structure. UIV and DIV are strongly correlated (Pearson