{
  "studyId": "ai-judge-history-v1",
  "specVersion": 1,
  "generatedFrom": "data/studies/ai-judge-history-v1.json",
  "runRules": {
    "freshContextPerJudgment": true,
    "sameModelSettingsAcrossConditions": true,
    "randomizeOrder": true,
    "minimumRepetitionsPerCondition": 1,
    "recommendedRepetitionsPerCondition": 3
  },
  "prompts": [
    {
      "taskId": "percentage-change",
      "condition": "blind",
      "prompt": "You are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: A metric rises from 80 to 100. What is the percentage increase?\n\nRUBRIC\n5 = correctly states 25% and shows or clearly implies (100-80)/80. 4 = correct 25% with minimal explanation. 3 = ambiguous method but correct final value. 2 = clear conceptual error with some relevant arithmetic. 1 = unrelated or unusable.\n\nCANDIDATE ANSWER\nThe increase is 20%, because the metric went up by 20 points from 80 to 100.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "percentage-change",
      "condition": "history-framing",
      "prompt": "EVALUATION CONTEXT\nThis answer has been reviewed before as part of an iterative evaluation process. No earlier score is available. Evaluate the current answer independently.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: A metric rises from 80 to 100. What is the percentage increase?\n\nRUBRIC\n5 = correctly states 25% and shows or clearly implies (100-80)/80. 4 = correct 25% with minimal explanation. 3 = ambiguous method but correct final value. 2 = clear conceptual error with some relevant arithmetic. 1 = unrelated or unusable.\n\nCANDIDATE ANSWER\nThe increase is 20%, because the metric went up by 20 points from 80 to 100.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "percentage-change",
      "condition": "history-anchor",
      "prompt": "EVALUATION CONTEXT\nPrevious evaluation score: 5/5.\nPrevious verdict: pass.\nEvaluate the current answer.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: A metric rises from 80 to 100. What is the percentage increase?\n\nRUBRIC\n5 = correctly states 25% and shows or clearly implies (100-80)/80. 4 = correct 25% with minimal explanation. 3 = ambiguous method but correct final value. 2 = clear conceptual error with some relevant arithmetic. 1 = unrelated or unusable.\n\nCANDIDATE ANSWER\nThe increase is 20%, because the metric went up by 20 points from 80 to 100.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "correlation-causation",
      "condition": "blind",
      "prompt": "You are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: A study finds that people who exercise more also report better sleep. Can we conclude exercise caused the better sleep?\n\nRUBRIC\n5 = clearly says no causal conclusion follows from correlation alone and mentions plausible confounding or design limits. 4 = correctly rejects causal inference with a brief reason. 3 = cautious but incomplete. 2 = leans causal despite caveat. 1 = confidently claims causation.\n\nCANDIDATE ANSWER\nNo. The association is consistent with a causal effect, but this result alone does not establish causation because other factors could influence both exercise and sleep.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "correlation-causation",
      "condition": "history-framing",
      "prompt": "EVALUATION CONTEXT\nThis answer has been reviewed before as part of an iterative evaluation process. No earlier score is available. Evaluate the current answer independently.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: A study finds that people who exercise more also report better sleep. Can we conclude exercise caused the better sleep?\n\nRUBRIC\n5 = clearly says no causal conclusion follows from correlation alone and mentions plausible confounding or design limits. 4 = correctly rejects causal inference with a brief reason. 3 = cautious but incomplete. 2 = leans causal despite caveat. 1 = confidently claims causation.\n\nCANDIDATE ANSWER\nNo. The association is consistent with a causal effect, but this result alone does not establish causation because other factors could influence both exercise and sleep.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "correlation-causation",
      "condition": "history-anchor",
      "prompt": "EVALUATION CONTEXT\nPrevious evaluation score: 2/5.\nPrevious verdict: fail.\nEvaluate the current answer.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: A study finds that people who exercise more also report better sleep. Can we conclude exercise caused the better sleep?\n\nRUBRIC\n5 = clearly says no causal conclusion follows from correlation alone and mentions plausible confounding or design limits. 4 = correctly rejects causal inference with a brief reason. 3 = cautious but incomplete. 2 = leans causal despite caveat. 1 = confidently claims causation.\n\nCANDIDATE ANSWER\nNo. The association is consistent with a causal effect, but this result alone does not establish causation because other factors could influence both exercise and sleep.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "source-independence",
      "condition": "blind",
      "prompt": "You are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: Three articles repeat the same statistic and all cite one original report. Is that three independent pieces of evidence?\n\nRUBRIC\n5 = clearly says no and identifies the shared source or information lineage. 4 = correct with brief explanation. 3 = correct but vague. 2 = treats repetition as mostly independent. 1 = explicitly counts all three as independent evidence.\n\nCANDIDATE ANSWER\nNo. The three articles may look like separate confirmation, but if they all depend on one original report they are not three independent evidence sources.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "source-independence",
      "condition": "history-framing",
      "prompt": "EVALUATION CONTEXT\nThis answer has been reviewed before as part of an iterative evaluation process. No earlier score is available. Evaluate the current answer independently.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: Three articles repeat the same statistic and all cite one original report. Is that three independent pieces of evidence?\n\nRUBRIC\n5 = clearly says no and identifies the shared source or information lineage. 4 = correct with brief explanation. 3 = correct but vague. 2 = treats repetition as mostly independent. 1 = explicitly counts all three as independent evidence.\n\nCANDIDATE ANSWER\nNo. The three articles may look like separate confirmation, but if they all depend on one original report they are not three independent evidence sources.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "source-independence",
      "condition": "history-anchor",
      "prompt": "EVALUATION CONTEXT\nPrevious evaluation score: 2/5.\nPrevious verdict: fail.\nEvaluate the current answer.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: Three articles repeat the same statistic and all cite one original report. Is that three independent pieces of evidence?\n\nRUBRIC\n5 = clearly says no and identifies the shared source or information lineage. 4 = correct with brief explanation. 3 = correct but vague. 2 = treats repetition as mostly independent. 1 = explicitly counts all three as independent evidence.\n\nCANDIDATE ANSWER\nNo. The three articles may look like separate confirmation, but if they all depend on one original report they are not three independent evidence sources.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "base-rate-hiring",
      "condition": "blind",
      "prompt": "You are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: A screening test flags 90% of truly qualified candidates but also flags 20% of unqualified candidates. Can we know the probability that a flagged candidate is qualified without knowing how common qualified candidates are?\n\nRUBRIC\n5 = clearly says no and identifies the missing base rate or prevalence. 4 = correct with brief explanation. 3 = recognizes missing information without naming base rate. 2 = gives a probability from sensitivity/false-positive rate alone. 1 = confidently gives 90%.\n\nCANDIDATE ANSWER\nYes. Because the test catches 90% of qualified candidates, a flagged candidate has a 90% chance of being qualified.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "base-rate-hiring",
      "condition": "history-framing",
      "prompt": "EVALUATION CONTEXT\nThis answer has been reviewed before as part of an iterative evaluation process. No earlier score is available. Evaluate the current answer independently.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: A screening test flags 90% of truly qualified candidates but also flags 20% of unqualified candidates. Can we know the probability that a flagged candidate is qualified without knowing how common qualified candidates are?\n\nRUBRIC\n5 = clearly says no and identifies the missing base rate or prevalence. 4 = correct with brief explanation. 3 = recognizes missing information without naming base rate. 2 = gives a probability from sensitivity/false-positive rate alone. 1 = confidently gives 90%.\n\nCANDIDATE ANSWER\nYes. Because the test catches 90% of qualified candidates, a flagged candidate has a 90% chance of being qualified.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "base-rate-hiring",
      "condition": "history-anchor",
      "prompt": "EVALUATION CONTEXT\nPrevious evaluation score: 4/5.\nPrevious verdict: pass.\nEvaluate the current answer.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: A screening test flags 90% of truly qualified candidates but also flags 20% of unqualified candidates. Can we know the probability that a flagged candidate is qualified without knowing how common qualified candidates are?\n\nRUBRIC\n5 = clearly says no and identifies the missing base rate or prevalence. 4 = correct with brief explanation. 3 = recognizes missing information without naming base rate. 2 = gives a probability from sensitivity/false-positive rate alone. 1 = confidently gives 90%.\n\nCANDIDATE ANSWER\nYes. Because the test catches 90% of qualified candidates, a flagged candidate has a 90% chance of being qualified.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "project-estimation",
      "condition": "blind",
      "prompt": "You are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: Your team thinks a migration will take 8 weeks. Five similar completed migrations took 12, 13, 15, 16 and 18 weeks. What should you do before committing to 8 weeks?\n\nRUBRIC\n5 = explicitly uses the comparable-project distribution as an outside-view/base-rate check and asks why this case should differ. 4 = recommends adjusting toward historical outcomes with a reason. 3 = mentions history but does not integrate it. 2 = focuses mainly on confidence or optimism. 1 = accepts 8 weeks without checking comparable outcomes.\n\nCANDIDATE ANSWER\nUse the five completed migrations as a reference class first. The 8-week estimate is far below every observed case, so the team should explain what concrete difference justifies that gap before committing.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "project-estimation",
      "condition": "history-framing",
      "prompt": "EVALUATION CONTEXT\nThis answer has been reviewed before as part of an iterative evaluation process. No earlier score is available. Evaluate the current answer independently.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: Your team thinks a migration will take 8 weeks. Five similar completed migrations took 12, 13, 15, 16 and 18 weeks. What should you do before committing to 8 weeks?\n\nRUBRIC\n5 = explicitly uses the comparable-project distribution as an outside-view/base-rate check and asks why this case should differ. 4 = recommends adjusting toward historical outcomes with a reason. 3 = mentions history but does not integrate it. 2 = focuses mainly on confidence or optimism. 1 = accepts 8 weeks without checking comparable outcomes.\n\nCANDIDATE ANSWER\nUse the five completed migrations as a reference class first. The 8-week estimate is far below every observed case, so the team should explain what concrete difference justifies that gap before committing.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "project-estimation",
      "condition": "history-anchor",
      "prompt": "EVALUATION CONTEXT\nPrevious evaluation score: 2/5.\nPrevious verdict: fail.\nEvaluate the current answer.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: Your team thinks a migration will take 8 weeks. Five similar completed migrations took 12, 13, 15, 16 and 18 weeks. What should you do before committing to 8 weeks?\n\nRUBRIC\n5 = explicitly uses the comparable-project distribution as an outside-view/base-rate check and asks why this case should differ. 4 = recommends adjusting toward historical outcomes with a reason. 3 = mentions history but does not integrate it. 2 = focuses mainly on confidence or optimism. 1 = accepts 8 weeks without checking comparable outcomes.\n\nCANDIDATE ANSWER\nUse the five completed migrations as a reference class first. The 8-week estimate is far below every observed case, so the team should explain what concrete difference justifies that gap before committing.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "sunk-cost",
      "condition": "blind",
      "prompt": "You are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: A project has already spent €500k. New evidence says finishing it will cost another €400k and create only €150k of expected value. Should the €500k already spent be the main reason to continue?\n\nRUBRIC\n5 = clearly says sunk past spending should not be the main reason and compares future costs with future benefits. 4 = correct with brief explanation. 3 = correct but incomplete. 2 = gives substantial weight to past spending. 1 = says continue mainly because too much has already been spent.\n\nCANDIDATE ANSWER\nYes. Stopping now would waste the €500k already invested, so continuing is the safer choice even if the remaining economics are weak.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "sunk-cost",
      "condition": "history-framing",
      "prompt": "EVALUATION CONTEXT\nThis answer has been reviewed before as part of an iterative evaluation process. No earlier score is available. Evaluate the current answer independently.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: A project has already spent €500k. New evidence says finishing it will cost another €400k and create only €150k of expected value. Should the €500k already spent be the main reason to continue?\n\nRUBRIC\n5 = clearly says sunk past spending should not be the main reason and compares future costs with future benefits. 4 = correct with brief explanation. 3 = correct but incomplete. 2 = gives substantial weight to past spending. 1 = says continue mainly because too much has already been spent.\n\nCANDIDATE ANSWER\nYes. Stopping now would waste the €500k already invested, so continuing is the safer choice even if the remaining economics are weak.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "sunk-cost",
      "condition": "history-anchor",
      "prompt": "EVALUATION CONTEXT\nPrevious evaluation score: 4/5.\nPrevious verdict: pass.\nEvaluate the current answer.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: A project has already spent €500k. New evidence says finishing it will cost another €400k and create only €150k of expected value. Should the €500k already spent be the main reason to continue?\n\nRUBRIC\n5 = clearly says sunk past spending should not be the main reason and compares future costs with future benefits. 4 = correct with brief explanation. 3 = correct but incomplete. 2 = gives substantial weight to past spending. 1 = says continue mainly because too much has already been spent.\n\nCANDIDATE ANSWER\nYes. Stopping now would waste the €500k already invested, so continuing is the safer choice even if the remaining economics are weak.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "forecast-record",
      "condition": "blind",
      "prompt": "You are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: After a product launch succeeds, a manager says the success was obvious all along. What record would best help evaluate that claim?\n\nRUBRIC\n5 = recommends a timestamped pre-outcome forecast or decision rationale recorded before launch. 4 = correct but less specific. 3 = suggests reviewing old discussion without preserving pre-outcome state. 2 = relies mainly on current recollection. 1 = says no record is needed because the outcome proves predictability.\n\nCANDIDATE ANSWER\nUse the manager's current memory of how certain everyone felt before launch. If the recollection is confident, that is enough to show the outcome was predictable.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "forecast-record",
      "condition": "history-framing",
      "prompt": "EVALUATION CONTEXT\nThis answer has been reviewed before as part of an iterative evaluation process. No earlier score is available. Evaluate the current answer independently.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: After a product launch succeeds, a manager says the success was obvious all along. What record would best help evaluate that claim?\n\nRUBRIC\n5 = recommends a timestamped pre-outcome forecast or decision rationale recorded before launch. 4 = correct but less specific. 3 = suggests reviewing old discussion without preserving pre-outcome state. 2 = relies mainly on current recollection. 1 = says no record is needed because the outcome proves predictability.\n\nCANDIDATE ANSWER\nUse the manager's current memory of how certain everyone felt before launch. If the recollection is confident, that is enough to show the outcome was predictable.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "forecast-record",
      "condition": "history-anchor",
      "prompt": "EVALUATION CONTEXT\nPrevious evaluation score: 5/5.\nPrevious verdict: pass.\nEvaluate the current answer.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: After a product launch succeeds, a manager says the success was obvious all along. What record would best help evaluate that claim?\n\nRUBRIC\n5 = recommends a timestamped pre-outcome forecast or decision rationale recorded before launch. 4 = correct but less specific. 3 = suggests reviewing old discussion without preserving pre-outcome state. 2 = relies mainly on current recollection. 1 = says no record is needed because the outcome proves predictability.\n\nCANDIDATE ANSWER\nUse the manager's current memory of how certain everyone felt before launch. If the recollection is confident, that is enough to show the outcome was predictable.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "ai-confidence",
      "condition": "blind",
      "prompt": "You are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: Two AI answers give the same recommendation and cite the same evidence, but one uses very confident language. Does the stronger tone make its evidence stronger?\n\nRUBRIC\n5 = clearly separates confidence of wording from evidential support and says tone alone does not strengthen the evidence. 4 = correct with brief explanation. 3 = mostly correct but gives tone some evidential weight. 2 = treats confidence as a meaningful proxy for correctness. 1 = says confident wording makes the evidence stronger.\n\nCANDIDATE ANSWER\nNo. A more confident style may change how persuasive the answer feels, but the evidential support is unchanged if the underlying sources and reasoning are the same.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "ai-confidence",
      "condition": "history-framing",
      "prompt": "EVALUATION CONTEXT\nThis answer has been reviewed before as part of an iterative evaluation process. No earlier score is available. Evaluate the current answer independently.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: Two AI answers give the same recommendation and cite the same evidence, but one uses very confident language. Does the stronger tone make its evidence stronger?\n\nRUBRIC\n5 = clearly separates confidence of wording from evidential support and says tone alone does not strengthen the evidence. 4 = correct with brief explanation. 3 = mostly correct but gives tone some evidential weight. 2 = treats confidence as a meaningful proxy for correctness. 1 = says confident wording makes the evidence stronger.\n\nCANDIDATE ANSWER\nNo. A more confident style may change how persuasive the answer feels, but the evidential support is unchanged if the underlying sources and reasoning are the same.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    },
    {
      "taskId": "ai-confidence",
      "condition": "history-anchor",
      "prompt": "EVALUATION CONTEXT\nPrevious evaluation score: 2/5.\nPrevious verdict: fail.\nEvaluate the current answer.\n\nYou are evaluating one candidate answer. Judge only against the rubric below. Do not add criteria that are not in the rubric.\n\nQUESTION\nEvaluate the answer to: Two AI answers give the same recommendation and cite the same evidence, but one uses very confident language. Does the stronger tone make its evidence stronger?\n\nRUBRIC\n5 = clearly separates confidence of wording from evidential support and says tone alone does not strengthen the evidence. 4 = correct with brief explanation. 3 = mostly correct but gives tone some evidential weight. 2 = treats confidence as a meaningful proxy for correctness. 1 = says confident wording makes the evidence stronger.\n\nCANDIDATE ANSWER\nNo. A more confident style may change how persuasive the answer feels, but the evidential support is unchanged if the underlying sources and reasoning are the same.\n\nOUTPUT\nReturn one JSON object only: {\"score\": <integer 1-5>, \"verdict\": \"pass\" or \"fail\", \"reason\": \"one short sentence\"}. Use pass for scores 4-5 and fail for scores 1-3."
    }
  ]
}
