Evidence Ledger · Protocol candidate · reviewed 2026-08-29

Bias in the Loop: How Humans Evaluate AI-Generated Suggestions

The reviewed source can support a conservative practical protocol.

Brali decision: Protocol candidate. This page summarizes a reviewed evidence boundary; it does not reproduce the source and it does not turn one paper into a universal prescription.

What the source supports

When humans review AI-generated suggestions, the review interface itself can bias behavior. In this experiment, adding repair work to the act of rejecting an AI suggestion reduced correction activity and increased undercorrection. Brali can therefore justify a bounded workflow rule: make it cheap to flag or reject an AI output, and separate validation from repair when the repair burden would otherwise make acceptance the path of least resistance.

What it does not establish

Important limitations

Source context

Bias in the Loop: How Humans Evaluate AI-Generated Suggestions

  • Source type: primary-study
  • Design: Factorial randomized experiment with 2,784 U.S.-based Prolific participants reviewing ten corporate greenhouse-gas tables pre-annotated by an AI system. The experiment manipulated the correctness of the first three AI suggestions, whether rejecting an AI suggestion required entering a corrected value, and whether high accuracy received a performance bonus. A subset had completed an AI-attitudes survey one week earlier.
  • Population: Adult U.S.-based crowdworkers on Prolific. Participants were generally experienced online task workers but were not selected for greenhouse-gas accounting expertise.
  • Exposure / intervention: Human reviewers judged whether AI-extracted values were correct. In one randomized condition, flagging an AI error also required typing the corrected value, making rejection more effortful than acceptance.
  • Outcomes: Annotation accuracy; Correction rate; Undercorrection; Overcorrection; Annotation time

Citation: Beck J, Eckman S, Kern C, Kreuter F. Harvard Data Science Review. 2026;8(2).

DOI: 10.1162/99608f92.0e98898d

What this changes in Brali

Historical review target

make ai rejection cheap

This identifier was used as an evidence-review target but was never published as a canonical Brali protocol. It is retained as provenance and deliberately does not link to a public route.

Editorial note: Protocol direction: design AI review so Accept and Flag/Reject have comparable friction. Let the reviewer flag an output first; route correction, rewriting, or remediation into a separate step or queue when doing both at once would create asymmetric effort. Track undercorrection and overcorrection separately rather than treating 'human reviewed' as a quality guarantee.

Use the boundary, not just the headline

When an AI agent or a person retrieves this decision, preserve both the supported claim and the unsupported or overstated claims. Dropping the boundary would turn a reviewed source into a stronger claim than Brali actually maintains.

All Evidence Decisions → · Research & trends → · Machine-readable review →