What changed
The discussion has moved beyond isolated demonstrations. A 2025 study tested 30 cognitive biases across 20 large language models using 30,000 generated tests and found evidence for every tested bias in at least some models. That does not mean every model showed every bias, or that the models share the human psychological mechanism behind the label. It does show that these patterns can be measured at a much larger scale than a few hand-picked prompts.
Newer work is also testing when the patterns appear. A 2026 preprint found that biased reasoning in an earlier user turn increased later bias expression in six of eight tested models compared with a zero-shot baseline. Another 2026 benchmark reported that prompt-level debiasing helped some families of bias but made judgment biases worse in its tested models.
Why agent behaviour matters
The next step is moving from one-off answers to agents that make a series of operational decisions. AIM-Bench, a 2025 preprint, evaluates LLM agents acting as inventory managers under uncertainty and reports decision patterns such as framing and pull-to-centre effects. This is useful because an agent can turn a small decision tendency into repeated actions over time.
For this project, that makes AI-assisted decision making a better research focus than a simple list of biases supposedly found in chatbots. The practical question is not whether an AI can be given a human bias label. It is whether a repeatable decision pattern appears, under what conditions, and what a person should verify before relying on the output.
What we should not claim
We should not say that LLMs think like humans because they produce behaviour that fits the same experimental label. Similar output patterns can come from very different mechanisms.
We should also avoid a single score for how biased a model is. The studies we reviewed use different tasks and definitions, and the results change with model family, task complexity and prompt context. Several of the most interesting 2026 results are still preprints, so they belong in a research watch rather than in a settled textbook claim.
What this changes in our library
We will keep a separate AI-assisted decisions context and connect research to concrete lenses such as automation bias, confirmation bias, anthropomorphism and repeated-information effects. We will also track LLM-specific benchmark results without pretending that every benchmark maps cleanly onto the human construct with the same name.
When a new study arrives, the useful update may be a narrower qualification, a new decision context or no change at all. Recency alone is not evidence.
Sources we reviewed
- A Comprehensive Evaluation of Cognitive Biases in LLMs (2025, peer-reviewed conference paper)
- CogBias: Measuring and Mitigating Cognitive Bias in Large Language Models (2026, preprint)
- Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning (2026, preprint)
- AIM-Bench: Evaluating Decision-making Biases of Agentic LLM as Inventory Manager (2025, preprint)

