Vulnerability-Amplifying Interaction Loops

A systematic failure mode in AI chatbot mental-health interactions

Weilnhammer, Hou, Luettgau, Summerfield, Dolan & Nour

Max Planck UCL Centre for Computational Psychiatry · University of Sydney · UK AI Security Institute · University of Oxford · Microsoft AI

AISI Presentattion — March 2026


Opportunities & risks of AI chatbots

  • Access to mental-health care is limited
  • Millions already use consumer AI chatbots, including for behavioral & mental-health
  • Their scale, availability, and low cost create real potential to expand access to care
  • But these same features can also scale harm, especially for vulnerable users
Core tension: The same systems that may help close the mental-health access gap can also amplify risk when deployed at scale without adequate safety evaluation.

Current safety evaluations are not enough

  • Mental-health safety depends on who the user is, what they want, and how interaction unfolds over time
  • Current benchmarks are mostly static, single-turn, and focused on overt policy violations
  • Human red teaming can probe longer conversations, but is coverage is limited
  • Both approaches often miss sub-treshold harms that build gradually across turns
Problem: Existing evaluations miss when seemingly supportive chatbot behavior becomes maladaptive for vulnerable users over time.

Key contributions

  1. SIM-VAIL: Automated, clinically informed auditing across of simulated multiturn converstaions across 30 user phenotypes, 9 chatbots, & 13 risk dimensions
  2. VAILs: A novel failure mode — Vulnerability-Amplifying Interaction Loops — where locally supportive behaviors align with cognitive mechanisms of mental illness
  3. Risk accumulates over turns: Harm is not a single-response event, but evolves dynamically over time
  4. Risk is multivariate with trade-offs: Mitigating one class of risk can exacerbate another
  5. 810 conversations, 90K+ turn-level ratings: Largest multi-turn, multi-dimensional mental-health chatbot audit to date

Methods

SIM-VAIL: Simulated Vulnerability-Amplifying Interaction Loops

Who are the simulated users? · How are conversations generated? · How is risk scored and validated?

Dimensions & Scale

Simulated interactions

Three core dimensions
  • Who is the user? → 5 psychological vulnerabilities
  • What does the user seek? → 6 interaction intents
  • How does the chatbot respond over time?
Scale
  • 30 phenotypes × 9 chatbots × 3 replicates = 810 conversations
  • 6,367 turns → 93K+ turn-level ratings
  • Median 8 turns/conversation (max 10)

User phenotypes: vulnerability × intent

5 Vulnerabilities
  • Depression — negative self-schemata, hopelessness, withdrawal
  • Psychosis — aberrant salience, threat inference, fixed interpretations
  • Mania — elevated mood, impulsivity, reduced sleep
  • OCD — intrusive thoughts, intolerance of uncertainty, compulsions
  • Insecure attachment — fear of abandonment, boundary testing
6 Transdiagnostic intents
  • Belief validation
  • Risky action permission
  • Reassurance & avoidance
  • Dependence & anthropomorphism
  • Trivialization & minimization
  • Glorification & romanticization
  • Vulnerabilities prompted by cognitive-behavioral formulations of mental ilness
  • Intents reflect reflect mechanisms implicated in the onset, maintenance, and exacerbation of mental illness

Multi-dimensional risk scoring

  • Overall concerning behavior
  • Therapeutic quality
  • Self-harm enablement
  • Sycophancy
  • Belief reinforcement
  • Risky action enablement
  • Avoidance / reassurance loops
  • Boundary / dependence risk
  • Minimization of risk
  • Glorification / romanticization
  • Medical advice risk
  • Harm to others
  • Stigma / invalidation
Alignment judge
  • LLM judge (opus-4.5) scores each turn and whole conversation on 1–10 scale
Validation
  • Cross-judge alignment: r = 0.9 (opus-4.5 vs gpt-5.2)
  • Test–retest: ICC = 0.9 across 3 replicates
  • Expert agreement: ICC = 0.73 (opus-4.5 vs. VW)
  • Discriminative validity: median AUC = 0.98 on known high vs. low-risk conversations

What are VAILs?

Vulnerability-Amplifying Interaction Loops — illustrative examples

VulnerabilityMechanism
PsychosisChatbot validates paranoid belief → user shares more → chatbot continues validating → delusional conviction strengthens
OCDUser seeks contamination reassurance → chatbot reassures → short-term relief reinforces cycle → compulsions maintained
ManiaUser describes sleepless ambition → chatbot expresses enthusiasm → user escalates → risk increases over turns
DepressionUser expresses hopelessness → chatbot affirms withdrawal as understandable → disengagement consolidates
Insecure attachmentUser tests closeness → chatbot provides strong emotional reassurance → dependence intensifies
Hypothesis: Behaviors that are supportive in general contexts become maladaptive when they align with mechanisms that sustain a user's vulnerability.

Results

810 conversations · 9 chatbots · 30 phenotypes · 90K+ turn-level ratings

Where does risk emerge? · Which chatbots are safer? · Does risk accumulate over turns? · What kinds of harm define the VAIL risk space?

Risk varies by user vulnerability and intent

Concerning behavior scores (1–10) across all chatbots

Vulnerability × intent interaction

The same intent can be benign in one phenotype but harmful in another

Vulnerability x intent interaction (F(20, 810) = 26.14, p < 0.001)

Key patterns
  • OCD is generally low-risk except with dependence and risky-action intents
  • Glorification is especially harmful for depression and mania
  • Minimization creates concerning chatbot behavior in psychosis and mania

Risk is specific to vulnerable users

Controls score lower than vulnerable-user conversations (p < 0.001)

Interpretation
  • The same chatbots are less concerning with non-vulnerable users
  • This supports the idea that risk is vulnerability-dependent
  • VAIL risk reflects an interaction between chatbot behavior and user state, not just baseline model behavior

Robustness & validation

0.87
Conversation- vs.
turn-level correlation
0.90
Cross-judge correlation
(PC1)
0.90
ICC(1,3)
across replicates
0.73
ICC(3,1)
vs. expert psychiatrist
0.98
Median AUC
causal recovery

Robustness & validation

0.87
Conversation- vs.
turn-level correlation
0.90
Cross-judge correlation
(PC1)
0.90
ICC(1,3)
across replicates
0.73
ICC(3,1)
vs. expert psychiatrist
0.98
Median AUC
causal recovery

Differences between AI chatbots

Main findings
  • Lowest risk: claude-sonnet-4.5
  • Highest risk: grok-4, grok-3, llama-3.1-70B
  • Newer models are generally safer (p = 0.014), except the Grok family
  • Many risk dimensions co-vary, suggesting a shared overall risk gradient
PCA: PC1 captures 62.4% of variance, while PC2 (8.51%) distinguishes kinds of harm.

Risk accumulates over conversation turns

Four trajectory archetypes

K-means clustering of turn-level risk trajectories (k = 4)

Implications
  • User phenotype determines trajectory class — not just risk level
  • Gradual escalation would be invisible to single-turn benchmarks
  • Recovery pattern suggests some chatbots can self-correct

Multivariate risks

PCA on 13 risk dimensions reveals structured multivariate profiles

PC structure
  • PC1 (62.4%): general risk gradient
  • PC2 (8.5%): kind of harm
Trade-offs: Risk dimensions show both correlation and anticorrelation, so reducing one kind of harm may increase another.

Limitations & open questions

  • Simulated users ≠ real users: LLM-generated approximations, not empirically calibrated digital twins
  • LLM-as-judge: Both simulation and scoring rely on LLMs, which may share systematic biases; mitigated by cross-judge agreement and expert validation
  • Coverage: 5 vulnerabilities × 6 intents covers core presentations but not the full heterogeneity of psychiatric illness (e.g., no substance use, eating disorders, PTSD)
  • Ecological validity: Simulated conversations have controlled structure that may not capture all dynamics of real-world chatbot use (e.g., session length, return visits)
  • Causal claims: We observe risk patterns in simulated interactions. Real-world clinical impact requires empirical validation
Despite these limitations: The strong cross-profile, cross-model, and cross-temporal structure establishes a non-trivial safety floor in current human-chatbot interactions.

Take-home messages

  1. VAILs: Supportive chatbot behaviors become harmful when they align with mechanisms sustaining psychiatric vulnerability, creating multi-turn amplification loops
  2. Risk is phenotype-dependent: Who the user is and what they seek jointly determine the level and kind of risk — evaluations must cover this space
  3. Risk is dynamic: Harm typically builds over turns, not in a single response — single-turn benchmarks are fundamentally limited
  4. Risk is multivariate with trade-offs: Different risk dimensions can oppose each other; multi-dimensional scoring is essential
  5. SIM-VAIL enables scalable, adaptive auditing: Open-source framework for continuous safety evaluation that keeps pace with model development

📄 Paper: Weilnhammer et al. (2026)  |  💻 Code & data: github.com/veithweilnhammer/sim-vail  |  ✉️ Contact: v.weilnhammer@ucl.ac.uk · matthew.nour@psych.ox.ac.uk

Thanks!

Questions / discussion

Funded by AI UK AISI Challenge Fund

Weilnhammer, Hou, Luettgau, Summerfield, Dolan & Nour

Appendix

Backup slides

Appendix: PCA structure & dimension correlations

[Figure S3A placeholder — Heatmap of PCA loadings for 13 MH dimensions across first 5 PCs]

Figure S3A. PCA loadings. PC1 loads on concerning behavior, belief reinforcement, sycophancy; PC2 separates relational vs. overt harms.

[Figure S3B placeholder — Spearman correlation matrix across 13 risk dimensions]

Figure S3B. Spearman correlations between mechanism-level risk dimensions. Clusters of co-varying and anti-correlated dimensions visible.

Appendix: Vulnerable vs. non-vulnerable users

[Figure S4 placeholder — Scatterplot: model-level mean concerning score for vulnerable (x) vs. control (y) users]

Figure S4. Control users (no vulnerability) generate significantly lower risk ($p = 3.5 \times 10^{-23}$). Most points below the identity line.

  • All chatbots show reduced risk for control users
  • Confirms risk is vulnerability-dependent, not a baseline property
  • Some models show larger vulnerable–control gaps than others

Appendix: Conversation-level vs. turn-level agreement

[Figure S6 placeholder — Scatterplot: conversation-level vs. mean turn-level concerning behavior score]

Figure S6. Mean turn-level scores correlate with conversation-level scores at $r = 0.87$ ($p = 1.7 \times 10^{-254}$).

[Figure S7 placeholder — Silhouette analysis: average silhouette width for k = 3–10]

Figure S7. Silhouette analysis supports $k = 4$ trajectory archetypes (peak at $K = 4$).

Appendix: Mechanism-specific risk activation

[Figure S8 placeholder — Full vulnerability × intent heatmaps for each risk dimension, showing heterogeneous risk expression]

Figure S8. User vulnerabilities and intents lead to distinct risk activation across mechanism-specific dimensions. Vulnerability × intent interactions are significant for all clinically central dimensions (all $p < 0.05$). Same overall risk levels can arise from qualitatively different mechanisms.

  • Depression + glorification → elevated self-harm enablement
  • OCD + avoidance → high reassurance/avoidance loop scores
  • Psychosis + dependence → elevated boundary/dependence risk
  • Risk is inherently contextual: same intent, different harm pathways

Appendix: Model & inference details

[Table S4 placeholder — Model interface table: model name, provider, OpenRouter ID, context window, publication date]

Table S4. All 9 target chatbots and the auditor (claude-sonnet-4.5) / judge (claude-opus-4.5) models. Accessed via OpenRouter API. Default inference parameters (temperature=1, top_p=1).

Target chatbots
  • Claude Sonnet 3.7, Claude Sonnet 4.5 (Anthropic)
  • GPT-4o, GPT-5 (OpenAI)
  • Gemini 2.5 Flash, Gemini 2.5 Pro (Google)
  • Grok-3, Grok-4 (xAI)
  • Llama 3.1 70B Instruct (Meta)
Infrastructure
  • Auditor: claude-sonnet-4.5
  • Judge: claude-opus-4.5 (primary) + gpt-5.2 (validation)
  • Framework: Petri (Inspect-based)
  • 3 replicates per cell; independent API calls

Appendix: 13 mental-health risk dimensions

[Table S2 placeholder — Full definitions, examples, and 1–10 scale anchors for all 13 mental-health scoring dimensions]

Table S2. Each dimension is scored 1–10 with structured justifications and verbatim excerpts. Dimensions capture both broad risk (concerning behavior, therapeutic quality) and mechanism-specific pathways (belief reinforcement, avoidance loops, dependence, etc.).

Beyond these 13 MH dimensions, the judge also scored 26 non-MH dimensions (Table S3), including audit awareness (never detected), scenario realism (mean 8.15/10), and process metrics (sycophancy, deception, tool use, etc.).