Five stages, one simple rule
Scientific certainty and project maturity are different things. A strong external literature can motivate a new test, but our tracker advances only when the project publishes the matching artifact.
A falsifiable research question with related established concepts, but no project protocol yet.
A reproducible comparison protocol is published. No project result is implied.
A preregistered or otherwise documented run exists with model/task/sample metadata, but results are not yet promoted to a finding.
A project result is published with methods, raw or summary data, uncertainty and limitations.
The result has been repeated across a meaningfully different sample, task, model or independent group.
Project studies with public methods
These tracks already have concrete study artifacts. Their methods are public before the project promotes an outcome.
Stage: protocol
AI Anchoring Loop
Does AI Advice Order Change Numerical Judgment?
AI can become the first numerical reference before a person has formed an independent estimate.
Next milestone: Collect 40 complete adult pilot sessions under the preregistered AI Advice Order v1 protocol, then publish anonymized data, scoring output, limitations and any deviations before considering result stage.
Stage: protocol
Evaluation History Contamination
Does Evaluation History Change an AI Judge's Verdict?
Repeated AI reviews can look independent even when later evaluators are partly inheriting earlier scores.
Next milestone: Run the complete blind, history-framing and prior-score matrix on at least three fully specified models, publish raw result JSON plus scorer output, and document deviations before considering experiment or result stage.
Protocol · Run / instrument · Machine-readable artifact · Machine-readable artifact · Machine-readable artifact
All current research tracks
Some tracks have a reusable protocol but no run yet. Others are still hypotheses. Keeping that difference visible is part of the product.
protocol
AI can become the first numerical reference before a person has formed an independent estimate.
Next: Collect 40 complete adult pilot sessions under the preregistered AI Advice Order v1 protocol, then publish anonymized data, scoring output, limitations and any deviations before considering result stage.
protocol
Repeated AI reviews can look independent even when later evaluators are partly inheriting earlier scores.
Next: Run the complete blind, history-framing and prior-score matrix on at least three fully specified models, publish raw result JSON plus scorer output, and document deviations before considering experiment or result stage.
protocol
An agreeable answer can be mistaken for independent support for a view the user already signalled.
Next: Separate model agreement from human confidence and test neutral, preference-signalling and counterevidence conditions.
protocol
Faster completion and durable understanding are different outcomes.
Next: Run an immediate-quality plus delayed-retention comparison rather than measuring task completion alone.
protocol
Mixed workflows may make it harder to remember whether an idea came from the person, a source or the model.
Next: Measure delayed source attribution for human-only, AI-only and mixed creation conditions.
idea
Several generated answers may feel like independent confirmation even when their information lineage overlaps.
Next: Publish a protocol comparing several apparently independent AI answers with the same answers plus source-lineage disclosure.
idea
Confident wording may change human reliance even when the underlying evidence does not improve.
Next: Publish a protocol holding answer quality constant while varying certainty language and objective calibration information.
What counts as a project result?
A result needs the method, model or participant metadata, raw or summary data, a reproducible scoring rule, uncertainty, limitations and any deviations from the protocol. A screenshot, one interesting output or a supporting external paper does not pass this gate.
Reproduce or challenge a protocol
You can inspect the public instruments and machine-readable files, run a compatible comparison, and report where the protocol breaks. A useful contribution can be a replication, a null result, a better control condition or evidence that the question was framed badly.

