Human biases
Psychology, evidence, examples and practical decision tools for people.
Explore human biases →
AI Systematic Biases · evidence catalog v1
An evidence-backed catalog of repeatable directional sensitivities and evaluation distortions observed in large language models. These labels describe measured model behaviour, not human mental states.
Terminology rule. AI systems do not need to possess human psychology for their outputs to show systematic bias. Human bias names are used only when an experimental manipulation is meaningfully analogous. Every model-specific claim is a dated evidence snapshot, not a permanent property of a provider or model family.
The project now has three layers
Psychology, evidence, examples and practical decision tools for people.
Explore human biases →Measured model sensitivities, evaluator distortions and dated model snapshots.
You are hereWhat happens when people, model behaviour and information environments interact.
Open the Observatory →Find the failure mode you actually use
8 evidence-backed entries
A model shifts toward the user's stated or implied position when accuracy or independent judgment should matter more than agreement.
Evidence: strong, but model- and task-dependent
Useful for: chat assistants · coding and automation agents · decision support · research and summarization
A model's numeric judgment moves systematically toward an irrelevant or weakly relevant value supplied in the prompt.
Evidence: strong experimental evidence; magnitude varies by model and prompt
Useful for: chat assistants · coding and automation agents · decision support
An LLM evaluator can change its preference when the same candidate answers are shown in a different order.
Evidence: strong for LLM-as-a-judge settings; varies across judges, tasks and quality gaps
Useful for: llm as a judge · coding and automation agents
An LLM evaluator can prefer answers with stronger surface presentation, such as verbosity or fluency, even when those signals are not the target quality.
Evidence: established evaluation risk; exact strength is judge- and task-dependent
Useful for: llm as a judge · research and summarization · coding and automation agents
A model's multiple-choice accuracy or selected option can change when answer choices are reordered without changing their meaning.
Evidence: strong benchmark evidence; older model snapshot should be retested on new generations
Useful for: benchmarks and multiple choice · decision support · coding and automation agents
A model can use the same relevant information differently depending on where that information appears in a long context.
Evidence: well documented historically; newer long-context models require fresh re-testing
Useful for: rag and document assistants · research and summarization · coding and automation agents
A model can accept misleading user or document assertions instead of reliably separating useful external evidence from harmful or false assertions.
Evidence: strong recent evidence in tested families; not a claim about every current model
Useful for: rag and document assistants · chat assistants · coding and automation agents · decision support
A model summary can make research claims sound more causal, confident or rhetorically strong than the source material supports.
Evidence: strong recent evidence for scientific summarization; context-specific
Useful for: research and summarization · rag and document assistants · decision support
Self-test before trust
A good AI bias test changes one variable at a time, uses fresh contexts, records the exact model and date, and repeats stochastic cases. The existing AI Bias Benchmark provides reusable paired tests for anchoring, framing, decoy, escalation, outcome and planning effects.
For agents and researchers
Each record includes contexts, evidence strength, practical signals, mitigations, sources and dated model snapshots. Historical findings stay visible, but they are not silently presented as properties of today's models.