Using AI · Growth Library

Benchmark AI on the Task Before You Trust the Workflow

Reviewed protocol

Test a small representative sample before turning assistance into a process

A person comparing notes and marking one useful insight
Protocol summary · reviewed

Try this

Choose a small sample containing easy, ordinary, and awkward cases. Define what counts as acceptable before running the model. Compare AI output with the source or a trusted human result, record failure patterns, and decide which cases can be assisted, which need review, and which should stay human-first.

Check-in: Do I know the cases where AI is reliable enough to help and the cases where its polished answer should not be trusted?

Evidence: Reviewed · source · Dell'Acqua et al. — Navigating the Jagged Technological Frontier

Review note: Reviewed against Dell'Acqua et al. on the jagged technological frontier. The evidence supports task-level evaluation and warns against assuming capability transfers to adjacent tasks. This protocol operationalizes that boundary through representative cases and acceptance criteria.

The useful question is rarely “Is this model smart?” It is “Does this model help on this exact step, with these inputs, under these failure costs?” The difference saves a great deal of automation theatre.

1. Build a representative sample

Include ordinary cases and the awkward cases that usually create rework. If you test only clean examples, you learn how the demo behaves, not how the workflow behaves.

2. Define acceptance before seeing output

Write the criteria first. Otherwise a fluent answer can quietly lower the standard after the fact.

3. Compare and classify failures

Mark errors by type: missing fact, invented fact, wrong calculation, bad judgment, format drift, privacy problem, or another failure that matters in your context.

4. Route work by risk

Use the result to separate safe assistance from mandatory review and human-only cases. Re-test after major model, prompt, data, or workflow changes.

Questions to consider

Why test the task instead of judging the model in general?

Because performance can change sharply between tasks that look similar. A model that is useful for one step in a workflow may be unreliable on the next step.

What should I score?

Use criteria tied to the real task: factual correctness, completeness, required format, safety, policy compliance, calculation accuracy, or another observable acceptance condition.

Sources

Evidence state: practical. Sources are provided so the underlying material can be inspected directly.

Maintenance record

Review history

Current status: Reviewed. An editorial or evidence review is recorded and no later event changes that conclusion.

  1. 2026-09-15 · reviewed · Brali editorial agent Source: review-registry-using-ai-starters.json

    Make the jagged-frontier result executable as a low-risk workflow evaluation protocol rather than a generic claim about AI productivity.

Browse the review ledger · Challenge or update this hack

Implementation observations can trigger a review, but they cannot change an evidence conclusion by themselves. Commercial relationships do not control review status or outcomes.

Continue from here

Choose the next useful path.

Article versions

Brali keeps substantive article history visible. Later reviews may refresh wording, sources, examples, or boundaries without silently replacing the record.

Canonical source record · Evidence state: reviewed.