Research · claim history

Evidence should be allowed to change the page.

This log records research that actually changed how the library explains a concept or decision context. It is not a list of every paper we read.

What the labels mean

Strengthens · 1

A useful new result supports an existing evidence record without making it universal.

Narrows · 2

A boundary condition, null result, or competing finding makes an existing claim more precise.

New context · 2

Evidence moves a known decision problem into a practical setting that deserves explicit guidance.

A claim can be well supported overall and still be narrowed by a later boundary condition. That is not a contradiction; it is often what better evidence looks like.

August 2026 · New context · moderate

AI-assisted medical decisions show many bias risks — and a domain-evidence gap

Evidence delta. A structured review synthesized 24 studies and identified 12 cognitive biases reported across AI-assisted medical decision making. Only one included study came directly from pathology, the review's target domain.

Why we changed something. A long list of possible biases can look more settled than the underlying evidence. This review is useful partly because it exposes that gap: evidence from related medical settings can guide questions, but it should not be presented as direct pathology evidence.

Project change. Add a domain-transfer caution to the AI-assisted decisions guide: related-domain evidence can motivate a check without proving the same effect in the target workflow.

Practical check. Ask two questions separately: has this bias been observed in human-AI decisions, and has it been observed in this exact task and professional setting?

Cognitive biases in AI-assisted medical decision making: A structured review as a primer for veterinary and human pathology · monthly review

August 2026 · Evidence update · narrows · moderate

Repeating an opinion is not the same test as repeating a factual claim

Evidence delta. Two preregistered experiments with 457 participants found no reliable increase in truth ratings when social-political opinion statements were repeated. Equivalence tests indicated that any repetition effect in the tested conditions was smaller than the researchers' predefined meaningful threshold.

Why we changed something. The Illusory Truth Effect has a robust overall evidence base, but the popular version is often too broad. A checkable factual statement and an evaluative social-political opinion should not be assumed to respond to repetition in the same way.

Project change. Narrow the canonical Illusory Truth page, publish a boundary-condition research note, and keep both the positive social-media result and the opinion null result visible in the same digest.

Practical check. Before applying the Illusory Truth label, classify the statement. For factual claims, trace independent evidence. For evaluative opinions, also consider persuasion, identity, norms and preference rather than forcing a factual truth-judgment explanation.

Limits of the illusory truth effect for social-political opinions: Evidence from two experiments and a mini meta-analysis · monthly review

August 2026 · Evidence update · strengthens · moderate

Repetition stayed powerful in an Instagram-like setting

Evidence delta. In a study of 165 participants, repeated statements were more likely to be judged true and were judged with greater confidence. Showing like counts did not meaningfully remove the repetition effects. Like magnitude mattered more for new statements, where repetition was not available as a cue.

Why we changed something. The classic repetition effect is not limited to plain laboratory lists. In this experiment, a social-media-like interface added another cue, but repetition still strongly shaped truth and certainty judgments.

Project change. Strengthen the social-media boundary condition on the Illusory Truth Effect page and add a practice item that separates source independence from exposure count.

Practical check. Before treating familiarity as evidence, count independent sources. Ten exposures to the same claim can still be one piece of evidence repeated ten times.

The robustness of repetition-based illusory truth and certainty effects in social media contexts · monthly review

August 2026 · Evidence update · narrows · provisional

Not every bias manipulation works in every human-AI task

Evidence delta. A mixed-methods study compared 35 professionals with seven LLMs in neurodevelopmental support-allocation judgments. Neither group showed significant susceptibility to the study's anchoring or representativeness manipulations, while both showed inconsistencies between descriptive ratings and final decisions. LLMs reported higher intellectual humility, but that self-report was not related to decision consistency.

Why we changed something. This is a useful negative result for a bias library. A familiar label should not be forced onto every decision problem. The task may reveal a different failure mode than the one the experiment expected.

Project change. Use this as a boundary-condition example in the AI research methodology: bias labels require task-level evidence, and null results are informative.

Practical check. Measure the decision change caused by the manipulation. Do not infer a bias from a plausible story about why the answer looks wrong.

Language Models Reproduce Human Reductionist Bias and Decision Inconsistency in Neurodevelopmental Disorders Assessment · monthly review

August 2026 · New context · provisional

A previous score can pull an LLM judge toward it

Evidence delta. Across 192,000 attempted evaluations of fixed texts, seven of eight tested models showed a task-bootstrap interval below zero for the total anchored-metadata effect. In a separate categorical dataset, anchored metadata blocked 48% of error corrections and changed 10.18% of correct judgments toward an assigned wrong label.

Why we changed something. Many agent loops keep previous scores, revision numbers, or reviewer comments in context. If the goal is an independent second judgment, that history can become part of the decision instead of harmless metadata.

Project change. Add an LLM-as-a-judge scenario to AI-assisted reasoning practice and treat prior-evaluation metadata as a testable anchoring pathway, not as a universal model trait.

Practical check. When you want an independent re-evaluation, hide the previous score first. Compare a blind judgment with a history-aware judgment instead of assuming they are equivalent.

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence · monthly review

Machine-readable history

The same records are available as JSON so assistants and research tools can distinguish a current claim from the evidence changes that shaped it.

Evidence changes JSON Monthly digests