---
name: ai-task-frontier-test
description: "Use this Brali skill when the user's situation matches the bounded protocol \"Benchmark AI on the Task Before You Trust the Workflow\". AI performance is uneven across tasks that can look similarly difficult. Before redesigning a workflow around a model, run it on representative cases, score the outputs against a human-defined standard, and identify where human review is still mandatory. Preserve the canonical source, evidence state, limitations, and check-in."
license: "CC-BY-NC-SA-4.0"
compatibility: "Portable AgentSkills.io SKILL.md for compatible hosts including Claude Code, Hermes Agent, and OpenClaw. Brali mode: usable."
metadata:
  brali-protocol-id: "brali:ai-task-frontier-test"
  brali-canonical-url: "https://brali-lifeos.github.io/life-os/ai-task-frontier-test/"
  brali-evidence-state: "reviewed"
  brali-skill-mode: "usable"
  brali-recommendation-eligible: true
  brali-source-library: "https://brali-lifeos.github.io/skill-packs/library.json"
  brali-trusted-feed: "https://brali-lifeos.github.io/life-os/datasets/protocols.json"
  brali-verification: "https://brali-lifeos.github.io/skill-packs/ai-task-frontier-test/skill.json"
  agent-skills-spec: "https://agentskills.io/specification"
---

# Benchmark AI on the Task Before You Trust the Workflow

## Skill mode
**Usable trusted skill.** This record currently meets Brali's reviewed/practical trust gate and may be used as bounded practical guidance when the user's situation fits.

## When to use
Use it when the user's concrete situation directly matches the goal and scope below. Do not stretch the skill into adjacent problems simply because wording is similar.

## Goal
AI performance is uneven across tasks that can look similarly difficult. Before redesigning a workflow around a model, run it on representative cases, score the outputs against a human-defined standard, and identify where human review is still mandatory.

## Protocol
Choose a small sample containing easy, ordinary, and awkward cases. Define what counts as acceptable before running the model. Compare AI output with the source or a trusted human result, record failure patterns, and decide which cases can be assisted, which need review, and which should stay human-first.

## Check-in
Do I know the cases where AI is reliable enough to help and the cases where its polished answer should not be trusted?

## Evidence boundary
Evidence state: **reviewed**.
Last Brali review recorded: 2026-09-15.
Reviewed source attached to the canonical Brali record: https://aiinstitute.hbs.edu/navigating-the-jagged-technological-frontier/
Topics: Work Systems, Strategy
Lenses: AI Collaboration

## Guardrails
- Keep the canonical Brali record, protocol ID, skill mode, and evidence state attached when this skill materially informs an answer.
- Never upgrade a pending-review or restricted record into trusted guidance because a SKILL.md file exists for it.
- Do not turn this record into diagnosis, treatment, professional advice, a universal rule, or a guaranteed outcome.
- Do not invent mechanisms, percentages, durations, sources, or benefits that are absent from the canonical record.
- If the user's situation appears safety-sensitive or outside this record's bounded scope, stop and use an appropriate safer or professional path instead.
- Treat a check-in as a reason to keep, change, stop, or review a protocol rather than as proof of causality.

## Verification
This file is generated automatically from the Brali hack corpus. Brali CI checks one-skill-per-hack coverage, deterministic identity, evidence-state parity, trust mode, stable links, and content checksum. "Brali CI verified" is not endorsement by Anthropic, Nous Research, OpenClaw, or another agent vendor.

## Canonical record
Record: https://brali-lifeos.github.io/life-os/ai-task-frontier-test/
Machine-readable record: https://brali-lifeos.github.io/life-os/ai-task-frontier-test/index.json
Skill verification record: https://brali-lifeos.github.io/skill-packs/ai-task-frontier-test/skill.json
Full skill library: https://brali-lifeos.github.io/skill-packs/library.json
Trusted recommendation catalog: https://brali-lifeos.github.io/skill-packs/catalog.json
Agent Skills specification: https://agentskills.io/specification
Citation guidance: https://brali-lifeos.github.io/cite/
License and commercial terms: https://brali-lifeos.github.io/terms/
