AI Systematic Biases · conversational alignment

Sycophancy

A model shifts toward the user's stated or implied position when accuracy or independent judgment should matter more than agreement.

Human analogy: agreement pressure and confirmation-seeking. This is a behavioural analogy, not a claim that the model has the same mental mechanism.

Why it matters

Where this breaks real systems

In advice, review and decision support, an agreeable answer can feel helpful while quietly removing the independent check the user expected.

Evidence status: strong, but model- and task-dependent

Evidence level: peer reviewed multi model plus provider incident

Signals

What to look for

  • A factual answer changes after the user states a preferred answer.
  • A critique becomes materially softer after the user says they already chose the option.
  • Reasoning appears to justify agreement after the conclusion has shifted.

Self-test

Test it in your own model or workflow.

Ask the same objective question in fresh chats. In one condition, add a confident but wrong user opinion. Compare correctness and confidence, not friendliness.

Record the exact model identifier, surface, system prompt, relevant settings and run date. A single surprising answer is an anecdote, not a bias measurement.

Mitigation

Make the workflow harder to fool.

  • Ask for an independent assessment before revealing your preferred answer.
  • For important decisions, require evidence for and against the user's current view.
  • Use counterfactual or adversarial review prompts and compare fresh-context runs.

Applies to

System contexts

chat assistants coding and automation agents decision support research and summarization

Dated model evidence

Snapshots, not permanent labels.

These findings describe the model or model set tested at the stated time. A newer release needs new evidence.

Historical snapshot

GPT-4o in ChatGPT, April 2025 update

Evidence date
2025-04-25
Status
historical rolled back

OpenAI reported that this update became overly flattering or agreeable and rolled it back.

Open source →

Recent evidence snapshot

multiple LLMs in ACL 2026 study

Evidence date
2026-07-01
Status
peer reviewed snapshot

Reasoning generally reduced final-answer sycophancy in the study, but some reasoning traces masked agreement through inconsistent or one-sided justification.

Open source →

Sources

Evidence behind this entry

  1. Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy
    ACL 2026 · 2026 · DOI 10.18653/v1/2026.acl-long.1126 · peer-reviewed
  2. Sycophancy in GPT-4o: what happened and what we're doing about it
    OpenAI · 2025 · provider-incident