Drop into a skill folder
Personal: ~/.claude/skills/ai-task-frontier-test/SKILL.md
Project: .claude/skills/ai-task-frontier-test/SKILL.md
Brali Agent Skill · reviewed · usable
AI performance is uneven across tasks that can look similarly difficult. Before redesigning a workflow around a model, run it on representative cases, score the outputs against a human-defined standard, and identify where human review is still mandatory.
Portable instruction
Choose a small sample containing easy, ordinary, and awkward cases. Define what counts as acceptable before running the model. Compare AI output with the source or a trusted human result, record failure patterns, and decide which cases can be assisted, which need review, and which should stay human-first.
Do I know the cases where AI is reliable enough to help and the cases where its polished answer should not be trusted?
Install
Inspect SKILL.md first. The evidence state and skill mode are part of the contract.
Install for normal use.
Personal: ~/.claude/skills/ai-task-frontier-test/SKILL.md
Project: .claude/skills/ai-task-frontier-test/SKILL.md
hermes skills install https://brali-lifeos.github.io/skill-packs/ai-task-frontier-test/SKILL.mdWorkspace: <workspace>/skills/ai-task-frontier-test/SKILL.md
Global: ~/.openclaw/skills/ai-task-frontier-test/SKILL.md
Brali verification
Canonical ID: brali:ai-task-frontier-test
Evidence state: reviewed
Skill mode: usable
Recommendation eligible: yes
SHA-256: 2416dda9bae0f3e4127297d0518f9bcef64cd91a723529e166e94402ef1cb3e5
Brali CI verified. This means identity, state, mode, provenance and file integrity matched the generated corpus. It is not third-party vendor endorsement.
Brali keeps a one-to-one machine surface: one hack record, one stable skill identity. Trust state determines what the skill is allowed to do.
Original Brali knowledge is available for non-commercial reuse under CC BY-NC-SA 4.0. Keep attribution, canonical record link, evidence state and skill mode.