Multi-domain Validation of LLM-Based Simulators via Interpretable and Latent Representations
摘要
The emergence of simulators powered by Large Language Models (LLMs) has enabled the realistic modeling of complex social phenomena while significantly reducing the costs and challenges of real-world data collection. Despite their promise, assessing the reliability of these simulators remains an open challenge. Existing validation methods often focus on isolated domains and operate at fixed levels of granularity, limiting their generalizability. In this work, we introduce simvale (simulator validation with latent embeddings), a generalizable, multi-domain framework for quantitatively assessing LLM-based social simulators. simvale leverages both interpretable features and latent representations to yield global and local assessments about the fidelity of a simulator to real-world dynamics, or about the effects of controlled interventions. We demonstrate the effectiveness of simvale through a case study evaluating a simulator’s ability to reproduce Online Social Network behavioral patterns and capture the impact of moderation interventions.