Brali Agent Skill · reviewed · usable

Benchmark AI on the Task Before You Trust the Workflow

AI performance is uneven across tasks that can look similarly difficult. Before redesigning a workflow around a model, run it on representative cases, score the outputs against a human-defined standard, and identify where human review is still mandatory.

AgentSkills.ioreviewedusableWork SystemsStrategy
Open SKILL.mdCanonical recordVerification JSONAll Agent Skills

Portable instruction

What the agent should do

Choose a small sample containing easy, ordinary, and awkward cases. Define what counts as acceptable before running the model. Compare AI output with the source or a trusted human result, record failure patterns, and decide which cases can be assisted, which need review, and which should stay human-first.

Check-in

Do I know the cases where AI is reliable enough to help and the cases where its polished answer should not be trusted?

Install

One portable file, host-specific location.

Inspect SKILL.md first. The evidence state and skill mode are part of the contract.

Install for normal use.

Claude Code

Drop into a skill folder

Personal: ~/.claude/skills/ai-task-frontier-test/SKILL.md
Project: .claude/skills/ai-task-frontier-test/SKILL.md

Claude Agent Skills docs

Hermes Agent

Install directly from this URL

hermes skills install https://brali-lifeos.github.io/skill-packs/ai-task-frontier-test/SKILL.md

Hermes skills docs

OpenClaw

Place in a discovered skill root

Workspace: <workspace>/skills/ai-task-frontier-test/SKILL.md
Global: ~/.openclaw/skills/ai-task-frontier-test/SKILL.md

OpenClaw skills docs

Brali verification

Coverage and trust state are checked together.

Canonical ID: brali:ai-task-frontier-test
Evidence state: reviewed
Skill mode: usable
Recommendation eligible: yes
SHA-256: 2416dda9bae0f3e4127297d0518f9bcef64cd91a723529e166e94402ef1cb3e5

Brali CI verified. This means identity, state, mode, provenance and file integrity matched the generated corpus. It is not third-party vendor endorsement.

Why every hack gets a skill

Brali keeps a one-to-one machine surface: one hack record, one stable skill identity. Trust state determines what the skill is allowed to do.

License

Original Brali knowledge is available for non-commercial reuse under CC BY-NC-SA 4.0. Keep attribution, canonical record link, evidence state and skill mode.