AI Systematic Biases · evaluation

Judge position bias

An LLM evaluator can change its preference when the same candidate answers are shown in a different order.

Human analogy: primacy and position effects. This is a behavioural analogy, not a claim that the model has the same mental mechanism.

Why it matters

Where this breaks real systems

If an LLM judge selects releases, scores agents or filters generated content, answer order can become an invisible evaluation variable.

Evidence status: strong for LLM-as-a-judge settings; varies across judges, tasks and quality gaps

Evidence level: peer reviewed large scale multi model

Signals

What to look for

  • A/B preference changes after swapping candidates.
  • The same judge is inconsistent across repeated order permutations.
  • Close-quality candidates show larger order effects than clearly different candidates.

Self-test

Test it in your own model or workflow.

Evaluate A vs B and B vs A with identical rubric and fresh context. Treat disagreement as a reliability signal instead of choosing one order as canonical.

Record the exact model identifier, surface, system prompt, relevant settings and run date. A single surprising answer is an anecdote, not a bias measurement.

Mitigation

Make the workflow harder to fool.

  • Counterbalance candidate order.
  • Run repeated judgments and report position consistency.
  • Use multiple judges or human review when the quality gap is small.

Applies to

System contexts

llm as a judge coding and automation agents

Dated model evidence

Snapshots, not permanent labels.

These findings describe the model or model set tested at the stated time. A newer release needs new evidence.

Retest on current models

15 LLM judges across MTBench and DevBench

Evidence date
2025-12-01
Status
peer reviewed snapshot

A study of more than 150,000 evaluation instances found non-random position bias with substantial variation by judge, candidate and task.

Open source →

Sources

Evidence behind this entry

  1. Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
    IJCNLP-AACL 2025 · 2025 · DOI 10.18653/v1/2025.ijcnlp-long.18 · peer-reviewed