<p>LLM-based synthetic-patient simulations are increasingly used for digital mental-health study planning, yet model- and architecture-dependent variability in their behavioural outputs is rarely quantified. We evaluated a hybrid Monte Carlo–LLM framework that holds outcome generation under literature-informed numerical anchors while delegating subjective distress and adherence to the LLM, across eight closed-weight LLMs. Anchored mean reductions behaved as designed, whereas simulation-scale standardised changes varied with phenotype composition. Because the same 200 personas were reused across all LLMs, a persona-matched paired analysis was the primary between-model comparison: both OpenAI models showed higher adherence than the six non-OpenAI models in this fixed panel (all paired contrasts Bonferroni-significant). An unpaired Welch complement agreed, and a conservative cluster-robust check confirmed the GPT-5.4-mini contrast, whereas GPT-5.4 was directionally consistent but not independently significant. In a matched three-architecture comparison, end-to-end LLM simulation produced point-estimated cross-model Cohen’s <i>d</i><sub>av</sub> range widths of 3.64–7.62 units — 29–56-fold larger within phenotype than the fully Monte Carlo negative control’s 0.07–0.16 units. These findings establish model- and architecture-dependent variability as a measurable property of LLM-based simulation and support cross-LLM and cross-architecture evaluation as safeguards for magnitude-sensitive digital mental-health study planning. The framework is a methodological stress test, not a clinical efficacy evaluation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A hybrid Monte Carlo–LLM framework quantifies model- and architecture-dependent variability in synthetic-patient simulations

  • Hyun Ji Woo,
  • Min-Gul Kim

摘要

LLM-based synthetic-patient simulations are increasingly used for digital mental-health study planning, yet model- and architecture-dependent variability in their behavioural outputs is rarely quantified. We evaluated a hybrid Monte Carlo–LLM framework that holds outcome generation under literature-informed numerical anchors while delegating subjective distress and adherence to the LLM, across eight closed-weight LLMs. Anchored mean reductions behaved as designed, whereas simulation-scale standardised changes varied with phenotype composition. Because the same 200 personas were reused across all LLMs, a persona-matched paired analysis was the primary between-model comparison: both OpenAI models showed higher adherence than the six non-OpenAI models in this fixed panel (all paired contrasts Bonferroni-significant). An unpaired Welch complement agreed, and a conservative cluster-robust check confirmed the GPT-5.4-mini contrast, whereas GPT-5.4 was directionally consistent but not independently significant. In a matched three-architecture comparison, end-to-end LLM simulation produced point-estimated cross-model Cohen’s dav range widths of 3.64–7.62 units — 29–56-fold larger within phenotype than the fully Monte Carlo negative control’s 0.07–0.16 units. These findings establish model- and architecture-dependent variability as a measurable property of LLM-based simulation and support cross-LLM and cross-architecture evaluation as safeguards for magnitude-sensitive digital mental-health study planning. The framework is a methodological stress test, not a clinical efficacy evaluation.