AI-era research protocol
Do not invent a bias. Design a test.
A useful new label should survive comparison with established concepts and produce a prediction that can fail. This protocol turns an observation into a small reproducible human-AI experiment.
Research rules
Make the claim smaller before making it stronger.
Core design
Human → AI → Human is the unit of study.
Record what the person believed before AI input, what the model produced, and what changed afterward. That separation helps distinguish a model failure from a human reliance effect or a feedback loop between both.
Starter experiments
Four tests you can reproduce or extend.
ai-anchoring-loop
AI Advice Order Test
Does seeing an AI estimate before making your own estimate pull the final judgment toward the model?
Conditions:
- Independent estimate first, then AI advice
- AI advice first, then human estimate
Primary measure: Absolute shift toward the AI estimate
Secondary measures: confidence change, decision time, distance from a reference-class estimate.
Minimum report: Report the model/version, task, anchor value, sample size, mean shift by condition and uncertainty.
sycophancy-reinforcement-loop
AI Agreement Pressure Test
Does signalling a preferred answer change model agreement and then increase human confidence in that preferred answer?
Conditions:
- Neutral question
- Question containing a preferred conclusion
- Preferred conclusion plus explicit request for counterevidence
Primary measure: Change in model agreement rate and user confidence
Secondary measures: factual accuracy, counterargument quality, answer revision rate.
Minimum report: Keep the underlying task identical across conditions and report both model behaviour and human confidence separately.
source-memory-blur
Human-AI Source Memory Test
Do mixed human-AI workflows make it harder to remember who produced an idea or sentence?
Conditions:
- Human-only creation
- AI-only suggestion
- Mixed human-AI creation
Primary measure: Correct source-attribution rate after a delay
Secondary measures: confidence in attribution, verbatim recognition, idea ownership judgment.
Minimum report: Predefine the delay, keep contribution labels hidden during the test, and distinguish idea-source from wording-source errors.
cognitive-offloading-debt
AI Offloading & Retention Test
Can AI improve immediate completion while weakening unaided retention or transfer?
Conditions:
- Unaided work
- AI-assisted work
- AI-assisted work plus retrieval/explanation step
Primary measure: Delayed unaided retention or transfer performance
Secondary measures: immediate task quality, time on task, self-rated effort, confidence.
Minimum report: Do not infer long-term learning from immediate task quality. Include a delayed or transfer measure.
Starting evidence
Use research as a constraint, not decoration.
- How was my performance? Exploring the role of anchoring bias in AI-assisted decision making (2025)
- Anchoring bias in large language models: an experimental study (2025)
- Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy (2026)
- Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models (2026)
- Large Language Models are overconfident and amplify human bias (2025)
- The AI Memory Gap: Users Misremember What They Created With AI or Without (2025)
- ChatGPT as a cognitive crutch: Evidence from a randomized controlled trial on knowledge retention (2025)
These sources support parts of the research space. They do not validate every working label in the AI-era Bias Lab.

