What is worth remembering
- Independent AI evaluation is not independent if an earlier score remains in the context.
- Repeated information can feel more true and more certain even when social endorsement cues are visible.
- A review can identify many plausible bias risks while still showing that direct evidence in the target domain is sparse.
- A named bias should be tested in the actual task. We should not assume that every model, person, or context will reproduce it.
- When LLMs simulate people, realistic cognitive limits may matter more than a persona prompt.
New context · provisional
A previous score can pull an LLM judge toward it
What the study found. Across 192,000 attempted evaluations of fixed texts, seven of eight tested models showed a task-bootstrap interval below zero for the total anchored-metadata effect. In a separate categorical dataset, anchored metadata blocked 48% of error corrections and changed 10.18% of correct judgments toward an assigned wrong label.
Why it matters. Many agent loops keep previous scores, revision numbers, or reviewer comments in context. If the goal is an independent second judgment, that history can become part of the decision instead of harmless metadata.
Try this. When you want an independent re-evaluation, hide the previous score first. Compare a blind judgment with a history-aware judgment instead of assuming they are equivalent.
What this changes here. Add an LLM-as-a-judge scenario to AI-assisted reasoning practice and treat prior-evaluation metadata as a testable anchoring pathway, not as a universal model trait.
Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence · 2026-08-26 · arXiv manuscript; paper states CIKM 2026 publication
Evidence update · strengthens · moderate
Repetition stayed powerful in an Instagram-like setting
What the study found. In a study of 165 participants, repeated statements were more likely to be judged true and were judged with greater confidence. Showing like counts did not meaningfully remove the repetition effects. Like magnitude mattered more for new statements, where repetition was not available as a cue.
Why it matters. The classic repetition effect is not limited to plain laboratory lists. In this experiment, a social-media-like interface added another cue, but repetition still strongly shaped truth and certainty judgments.
Try this. Before treating familiarity as evidence, count independent sources. Ten exposures to the same claim can still be one piece of evidence repeated ten times.
What this changes here. Strengthen the social-media boundary condition on the Illusory Truth Effect page and add a practice item that separates source independence from exposure count.
The robustness of repetition-based illusory truth and certainty effects in social media contexts · 2026-08-05 · peer-reviewed experiment · DOI 10.1038/s41598-026-61449-y
New context · moderate
AI-assisted medical decisions show many bias risks — and a domain-evidence gap
What the study found. A structured review synthesized 24 studies and identified 12 cognitive biases reported across AI-assisted medical decision making. Only one included study came directly from pathology, the review's target domain.
Why it matters. A long list of possible biases can look more settled than the underlying evidence. This review is useful partly because it exposes that gap: evidence from related medical settings can guide questions, but it should not be presented as direct pathology evidence.
Try this. Ask two questions separately: has this bias been observed in human-AI decisions, and has it been observed in this exact task and professional setting?
What this changes here. Add a domain-transfer caution to the AI-assisted decisions guide: related-domain evidence can motivate a check without proving the same effect in the target workflow.
Cognitive biases in AI-assisted medical decision making: A structured review as a primer for veterinary and human pathology · 2026-08-05 · peer-reviewed structured review · DOI 10.1177/03009858261472493
Evidence update · narrows · provisional
Not every bias manipulation works in every human-AI task
What the study found. A mixed-methods study compared 35 professionals with seven LLMs in neurodevelopmental support-allocation judgments. Neither group showed significant susceptibility to the study's anchoring or representativeness manipulations, while both showed inconsistencies between descriptive ratings and final decisions. LLMs reported higher intellectual humility, but that self-report was not related to decision consistency.
Why it matters. This is a useful negative result for a bias library. A familiar label should not be forced onto every decision problem. The task may reveal a different failure mode than the one the experiment expected.
Try this. Measure the decision change caused by the manipulation. Do not infer a bias from a plausible story about why the answer looks wrong.
What this changes here. Use this as a boundary-condition example in the AI research methodology: bias labels require task-level evidence, and null results are informative.
Language Models Reproduce Human Reductionist Bias and Decision Inconsistency in Neurodevelopmental Disorders Assessment · 2026-08-17 · preprint
Research watch · not a settled claim · provisional
Persona prompting can make simulated users unrealistically capable
What the study found. Using more than 71,000 reading-comprehension responses from 2,359 primary-school students, researchers found that standard persona prompting produced near-perfect, low-variance LLM performance. A simulator that imposed a working-memory-like episodic bottleneck reduced the gap between model behavior and the student data.
Why it matters. This is not evidence for a canonical cognitive bias. It is a research-method signal: an LLM asked to 'act like' a person may still use knowledge and memory that the simulated person would not have.
Try this. If an LLM is being used as a synthetic participant, test it against real human response distributions. A convincing persona is not validation.
What this changes here. Keep this in the research-method watchlist and connect it to future guidance on AI-based behavioral simulation rather than adding a new bias entry.
"Act Like a 5th Grader" is Not Enough: Bounding Knowledge in LLM-Based User Simulators · 2026-08-30 · preprint
What changed in the project
- Launched the monthly evidence-delta digest and added it to Research, the sitemap and the public data manifest.
- Updated the Illusory Truth evidence review with the 2026 social-media-like experiment and a clearer independent-source check.
- Added an evidence-distance check to the AI-assisted decisions guide so related-domain findings are not silently treated as direct evidence.
- Added practice scenarios for blind AI re-evaluation, repetition versus independent sources, and evidence transfer between professional workflows.
- Expanded the research scout beyond LLM benchmarks to debiasing, replication, forecasting and applied human decision research.
Questions we are carrying forward
- Can blind-versus-history-aware LLM judging be turned into a small reproducible experiment for the site?
- How often do repeated claims in AI answers come from genuinely independent sources rather than copied or circular evidence?
- Which debiasing interventions improve calibration in real decisions rather than only reducing a benchmark bias score?

