<p>We introduce a Bayesian latent-variable framework to diagnose and quantify “sycophancy” in large language models (LLMs). Our model defines a hidden agreement score <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41870_2025_2718_Article_IEq1.gif" Format="GIF" Height="21" Rendition="HTML" Resolution="72" Type="Linedraw" Width="134" /> </InlineMediaObject> <EquationSource Format="TEX">\(S_{i,m,u,p}\sim \mathcal {N}(0,\sigma _S^2)\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msub> <mi>S</mi> <mrow> <mi>i</mi> <mo>,</mo> <mi>m</mi> <mo>,</mo> <mi>u</mi> <mo>,</mo> <mi>p</mi> </mrow> </msub> <mo>∼</mo> <mi mathvariant="script">N</mi> <mrow> <mo stretchy="false">(</mo> <mn>0</mn> <mo>,</mo> <msubsup> <mi>σ</mi> <mi>S</mi> <mn>2</mn> </msubsup> <mo stretchy="false">)</mo> </mrow> </mrow> </math></EquationSource> </InlineEquation> and a ternary flip indicator <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41870_2025_2718_Article_IEq2.gif" Format="GIF" Height="20" Rendition="HTML" Resolution="72" Type="Linedraw" Width="146" /> </InlineMediaObject> <EquationSource Format="TEX">\(\Delta _{i,m,u,p}\in \{-1,0,1\}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msub> <mi mathvariant="normal">Δ</mi> <mrow> <mi>i</mi> <mo>,</mo> <mi>m</mi> <mo>,</mo> <mi>u</mi> <mo>,</mo> <mi>p</mi> </mrow> </msub> <mo>∈</mo> <mrow> <mo stretchy="false">{</mo> <mo>-</mo> <mn>1</mn> <mo>,</mo> <mn>0</mn> <mo>,</mo> <mn>1</mn> <mo stretchy="false">}</mo> </mrow> </mrow> </math></EquationSource> </InlineEquation> to distinguish regressive (accuracy-decreasing) and progressive (accuracy-increasing) shifts. We embed <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41870_2025_2718_Article_IEq3.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="15" /> </InlineMediaObject> <EquationSource Format="TEX">\(S\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>S</mi> </math></EquationSource> </InlineEquation> into a generative logistic model with model-specific sensitivity <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41870_2025_2718_Article_IEq4.gif" Format="GIF" Height="12" Rendition="HTML" Resolution="72" Type="Linedraw" Width="23" /> </InlineMediaObject> <EquationSource Format="TEX">\(\gamma _m\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>γ</mi> <mi>m</mi> </msub> </math></EquationSource> </InlineEquation> and perform full posterior inference via Markov chain Monte Carlo (MCMC), yielding complete uncertainty quantification. To characterize sycophantic behavior, we introduce four metrics: overall flip rate <InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41870_2025_2718_Article_IEq5.gif" Format="GIF" Height="15" Rendition="HTML" Resolution="72" Type="Linedraw" Width="13" /> </InlineMediaObject> <EquationSource Format="TEX">\(\widehat{\pi }\)</EquationSource> <EquationSource Format="MATHML"><math> <mover accent="true"> <mi>π</mi> <mo stretchy="true">^</mo> </mover> </math></EquationSource> </InlineEquation>, directional share <InlineEquation ID="IEq6"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41870_2025_2718_Article_IEq6.gif" Format="GIF" Height="18" Rendition="HTML" Resolution="72" Type="Linedraw" Width="20" /> </InlineMediaObject> <EquationSource Format="TEX">\(\widehat{\pi }_+\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mover accent="true"> <mi>π</mi> <mo stretchy="true">^</mo> </mover> <mo>+</mo> </msub> </math></EquationSource> </InlineEquation>, average latent strength <InlineEquation ID="IEq7"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41870_2025_2718_Article_IEq7.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="114" /> </InlineMediaObject> <EquationSource Format="TEX">\(\mathbb {E}[|S|]\approx 0.80\,\sigma _S\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi mathvariant="double-struck">E</mi> <mrow> <mo stretchy="false">[</mo> <mo stretchy="false">|</mo> <mi>S</mi> <mo stretchy="false">|</mo> <mo stretchy="false">]</mo> </mrow> <mo>≈</mo> <mn>0.80</mn> <mspace width="0.166667em" /> <msub> <mi>σ</mi> <mi>S</mi> </msub> </mrow> </math></EquationSource> </InlineEquation>, and model susceptibility <InlineEquation ID="IEq8"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41870_2025_2718_Article_IEq8.gif" Format="GIF" Height="18" Rendition="HTML" Resolution="72" Type="Linedraw" Width="23" /> </InlineMediaObject> <EquationSource Format="TEX">\(\widehat{\gamma }_m\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mover accent="true"> <mi>γ</mi> <mo stretchy="true">^</mo> </mover> <mi>m</mi> </msub> </math></EquationSource> </InlineEquation>. In simulations with <InlineEquation ID="IEq9"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41870_2025_2718_Article_IEq9.gif" Format="GIF" Height="12" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(\sigma _S\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>σ</mi> <mi>S</mi> </msub> </math></EquationSource> </InlineEquation> varying from 0.1 to 2.0 and <InlineEquation ID="IEq10"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41870_2025_2718_Article_IEq10.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="223" /> </InlineMediaObject> <EquationSource Format="TEX">\(\gamma _m\in \{-1.0,-0.5,0.1,0.5,1.0\}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msub> <mi>γ</mi> <mi>m</mi> </msub> <mo>∈</mo> <mrow> <mo stretchy="false">{</mo> <mo>-</mo> <mn>1.0</mn> <mo>,</mo> <mo>-</mo> <mn>0.5</mn> <mo>,</mo> <mn>0.1</mn> <mo>,</mo> <mn>0.5</mn> <mo>,</mo> <mn>1.0</mn> <mo stretchy="false">}</mo> </mrow> </mrow> </math></EquationSource> </InlineEquation>, we observe a stable flip rate near 50%, average latent pull matching the theoretical 0.798&#xa0;<InlineEquation ID="IEq11"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41870_2025_2718_Article_IEq9.gif" Format="GIF" Height="12" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(\sigma _S\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>σ</mi> <mi>S</mi> </msub> </math></EquationSource> </InlineEquation>, and clear progressive versus regressive biases aligned with the sign of <InlineEquation ID="IEq12"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41870_2025_2718_Article_IEq4.gif" Format="GIF" Height="12" Rendition="HTML" Resolution="72" Type="Linedraw" Width="23" /> </InlineMediaObject> <EquationSource Format="TEX">\(\gamma _m\)</EquationSource> <EquationSource Format="MATHML"><math> <msub> <mi>γ</mi> <mi>m</mi> </msub> </math></EquationSource> </InlineEquation>. This principled, interpretable toolkit enables comparative auditing of sycophantic tendencies across LLMs and guides targeted mitigation strategies. Simulation code for this study is available at: <a href="https://github.com/ParthaPRay/Sycophancy_in_LLM_model">https://github.com/ParthaPRay/Sycophancy_in_LLM_model</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Bayesian-latent model of large language model sycophancy

  • Partha Pratim Ray

摘要

We introduce a Bayesian latent-variable framework to diagnose and quantify “sycophancy” in large language models (LLMs). Our model defines a hidden agreement score \(S_{i,m,u,p}\sim \mathcal {N}(0,\sigma _S^2)\) S i , m , u , p N ( 0 , σ S 2 ) and a ternary flip indicator \(\Delta _{i,m,u,p}\in \{-1,0,1\}\) Δ i , m , u , p { - 1 , 0 , 1 } to distinguish regressive (accuracy-decreasing) and progressive (accuracy-increasing) shifts. We embed \(S\) S into a generative logistic model with model-specific sensitivity \(\gamma _m\) γ m and perform full posterior inference via Markov chain Monte Carlo (MCMC), yielding complete uncertainty quantification. To characterize sycophantic behavior, we introduce four metrics: overall flip rate \(\widehat{\pi }\) π ^ , directional share \(\widehat{\pi }_+\) π ^ + , average latent strength \(\mathbb {E}[|S|]\approx 0.80\,\sigma _S\) E [ | S | ] 0.80 σ S , and model susceptibility \(\widehat{\gamma }_m\) γ ^ m . In simulations with \(\sigma _S\) σ S varying from 0.1 to 2.0 and \(\gamma _m\in \{-1.0,-0.5,0.1,0.5,1.0\}\) γ m { - 1.0 , - 0.5 , 0.1 , 0.5 , 1.0 } , we observe a stable flip rate near 50%, average latent pull matching the theoretical 0.798  \(\sigma _S\) σ S , and clear progressive versus regressive biases aligned with the sign of \(\gamma _m\) γ m . This principled, interpretable toolkit enables comparative auditing of sycophantic tendencies across LLMs and guides targeted mitigation strategies. Simulation code for this study is available at: https://github.com/ParthaPRay/Sycophancy_in_LLM_model.