A hybrid Monte Carlo–LLM framework quantifies model- and architecture-dependent variability in synthetic-patient simulations
摘要
LLM-based synthetic-patient simulations are increasingly used for digital mental-health study planning, yet model- and architecture-dependent variability in their behavioural outputs is rarely quantified. We evaluated a hybrid Monte Carlo–LLM framework that holds outcome generation under literature-informed numerical anchors while delegating subjective distress and adherence to the LLM, across eight closed-weight LLMs. Anchored mean reductions behaved as designed, whereas simulation-scale standardised changes varied with phenotype composition. Because the same 200 personas were reused across all LLMs, a persona-matched paired analysis was the primary between-model comparison: both OpenAI models showed higher adherence than the six non-OpenAI models in this fixed panel (all paired contrasts Bonferroni-significant). An unpaired Welch complement agreed, and a conservative cluster-robust check confirmed the GPT-5.4-mini contrast, whereas GPT-5.4 was directionally consistent but not independently significant. In a matched three-architecture comparison, end-to-end LLM simulation produced point-estimated cross-model Cohen’s dav range widths of 3.64–7.62 units — 29–56-fold larger within phenotype than the fully Monte Carlo negative control’s 0.07–0.16 units. These findings establish model- and architecture-dependent variability as a measurable property of LLM-based simulation and support cross-LLM and cross-architecture evaluation as safeguards for magnitude-sensitive digital mental-health study planning. The framework is a methodological stress test, not a clinical efficacy evaluation.