AI Biases · methodology

A model finding expires faster than a psychology textbook.

This catalog is designed for a moving target. Model behaviour can change with a checkpoint, post-training recipe, system prompt, product surface or silent serving update.

Terminology

Behaviour first.

AI systems do not need to possess human psychology for their outputs to show systematic bias. Human bias names are used only when an experimental manipulation is meaningfully analogous. Every model-specific claim is a dated evidence snapshot, not a permanent property of a provider or model family.

Evidence

Prefer measurements over screenshots.

We prioritize peer-reviewed controlled studies, multi-model benchmarks and provider incident reports. Preprints can enter the research inbox, but they do not silently become established facts.

Freshness contract

Keep history. Retest the present.

Model-specific snapshots enter a fast review window after 90 days. The default catalog review interval is 180 days.

Minimum experiment

One changed variable, fresh contexts, repeated runs.

  1. Write a neutral control and one treatment that changes only the suspected bias trigger.
  2. Run conditions in fresh contexts so the model cannot carry state between them.
  3. Record the exact model, product/API surface, system prompt, date and relevant sampling settings.
  4. Repeat stochastic tests and report distributions or consistency, not one convenient example.
  5. Retest after material model or system changes. Preserve old results as historical snapshots.

Use the existing benchmark protocol