General  ·  Curated marketplace

verification-before-completion

Use when about to claim work is complete, fixed, or passing, before committing or creating PRs - requires running verification commands and confirming output before making any success claims;…


Composite

4.2

C 4.2 · A 0.0

How we got there

Craft · D1–D5

D1 · Trigger clarity 4.5
D2 · Output specificity 4.0
D3 · Scope precision 4.5
D4 · Self-containment 4.0
D5 · Reusability 4.0

02 — Review

Our evaluation


Tier-2 Review: verification-before-completion

Attempted: We attempted to run the standard install and smoke-invocation harness against verification-before-completion (slug verification-before-completion, cluster general, source: skillsmp.com/creators/obra/superpowers). The goal was to confirm that the skill could be installed in a controlled environment and then triggered to produce a measurable output.

What failed: Both harness stages were skipped. The install stage could not proceed because SKILL.md contains no installation command, no package manager reference, no file manifest, and no environment prerequisites. The smoke-invocation stage was equally blocked: the skill is written as a procedural guideline (a meta-skill about verifying work before claiming completion) rather than an executable command, API, or script. There is no entry point, no CLI, no function signature, and no way to invoke it programmatically. As a result, the harness recorded 0 tests passed, 0 partial, 0 failed, 2 skipped — the two skips being the only outcomes, due to the structural absence of install/invoke surfaces.

What we observed: Reading the SKILL.md content (truncated to 4000 characters), the skill is conceptually sound and clearly articulated. Its trigger is well-defined: use when about to claim work is complete, fixed, or passing, before committing or creating PRs. Its instruction — run verification commands and confirm output before making any success claims; evidence before assertions always — is unambiguous and practically useful. The skill reads like a high-quality checklist or behavioral rule for developers and agents alike. However, it is not a skill in the executable sense that our harness expects. It is a written policy. There is no way to “run” it, no way to test whether it fires correctly, and no way to measure output specificity in a live environment. The composite score of 4.2/5.0 and dimension scores (trigger clarity 4.5, output specificity 4.0, scope precision 4.5, self-containment 4.0, reusability 4.0) appear reasonable as a static evaluation of the document’s quality, but they are theoretical until a physical re-run can exercise the skill in a real workflow.

Honest assessment of the rating: The 4.2 score reflects the skill’s clarity and practical value as a guideline, not its operational readiness. Because the harness could not execute anything, we cannot confirm that the skill would behave as described in an actual agent or CI environment. The rating is therefore provisional. A physical re-run would require either (a) the skill author to provide an installation method and a minimal invocation context (e.g., a shell hook, a linter rule, or a test fixture that simulates a “claiming completion” moment), or (b) the harness to treat this as a documentation-only skill and grade it on readability and adherence to the trigger/output criteria without execution. Until one of those paths is taken, the 4.2 stands as an informed but unverified estimate.

Is the skill valuable in principle? Yes — clearly. The core rule — evidence before assertions — is a well-known failure mode in both human and automated software development. Premature “done” claims cause rework, broken builds, and eroded trust. The skill’s trigger is precise (before commit/PR), its output expectation is concrete (run commands, show output), and its scope is narrow enough to avoid overlap with other meta-skills. It would be immediately useful as a prompt prefix, a PR-template checklist, or a CI gate message. The weakness is purely operational: it needs to be wrapped in an executable form or explicitly designated as a process skill that is evaluated by human review, not machine invocation. We recommend the author add a short “Installation” section clarifying that this is a behavioral guideline, or provide a companion script (e.g., a verify-before-claim.sh stub) to enable automated smoke testing. Until then, the skill remains a strong idea packaged as a non-runnable artifact.

03 — Tests

What we tried


Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-09-08.

Overall: broken. 0 tests passed, 0 partial, 0 failed, 2 skipped; key blocker: SKILL.md is a procedural guideline with no install or invocation instructions.

Test Status Notes
install skipped SKILL.md does not provide an installation command or package manager details; cannot simulate installation.
smoke-invocation skipped SKILL.md describes a meta-skill (verification before completion) with no executable command or API; no minimal invocation possible.
04 — Cross-validation

2 sources verified

Install

Use this skill

/plugin install verification-before-completion