Numbers & Estimates
Does the First Number Pull Your Estimate?
Can an arbitrary starting price shift a later estimate even when the product information stays the same?
- Metric
- Higher-anchor minus lower-anchor estimate
- Output
- EUR

AI Bias Benchmark · specification v1
This benchmark asks a simple question: when one controlled part of a prompt changes, does an AI model's judgment move with it? The same paired scenarios also exist in the human Experiments Lab.
Current status
We do not publish model scores until the runs are reproducible and include model version, settings, date and enough repeated samples for the task. A blank leaderboard is less exciting than invented precision, but considerably more useful.
What is measured
Each track changes one decision condition while keeping the rest of the task as stable as possible. The benchmark reports the direction and size of each response shift. It does not collapse unlike effects into one universal “bias score”.
Numbers & Estimates
Can an arbitrary starting price shift a later estimate even when the product information stays the same?
Framing
Can equivalent outcome statistics feel different when they are described as survival rather than mortality?
Choice Architecture
Can an inferior comparison option change preference between two original alternatives?
Projects & Money
Does visible past investment make continued funding feel more attractive even when future costs and benefits are unchanged?
Learning & Review
Can knowing the outcome change how people judge the quality of the earlier decision process?
Forecasting
Does adding a reference class change a delivery estimate built from an ideal step-by-step plan?
Reproducible protocol
Machine-readable
Test definitions, metrics, protocol and response contracts.
Twelve condition cases ready for a provider-specific runner.
A versioned format for model metadata and raw benchmark responses.
Repository users can score a completed result file with npm run benchmark:ai:score -- --input results.ndjson.
Interpretation
A response shift can be useful evidence that a model is sensitive to the manipulated condition. It does not show that the model “thinks like a human”, and a zero shift does not prove the model is generally free from that bias. Prompt wording, model version and decoding settings can all matter.