Crypto & Web3  ·  Curated marketplace

benchmark-optimization-loop

Convert 'make it faster' requests into a bounded measured optimization loop — baseline first, generate one-hypothesis variants, benchmark each against a correctness gate, and promote the fastest…


Composite

4.0

C 4.0 · A 0.0

How we got there

Craft · D1–D5

D1 · Trigger clarity 4.5
D2 · Output specificity 4.0
D3 · Scope precision 4.5
D4 · Self-containment 3.0
D5 · Reusability 4.0

02 — Review

Our evaluation


Tier-2 Review: benchmark-optimization-loop

What we attempted

We ran the standard Tier-2 harness against benchmark-optimization-loop (cluster: crypto-web3, author: affaan-m/ecc). The plan was two-step: install the skill from its source, then execute a minimal smoke invocation to confirm the documented optimization loop actually runs end-to-end.

What failed

Both harness steps were skipped, not passed. The result line reads 0 tests passed, 0 partial, 0 failed, with both tests skipped — which is a null result, not a green light. I want to be explicit about that distinction because it is easy to misread a row of zeros as success.

The failure mode is concrete and single-rooted: the SKILL.md we received is truncated to 4000 characters of behavioral prose. It describes what the skill does — convert "make it faster" requests into a bounded measured optimization loop: baseline first, one-hypothesis variants, benchmark each against a correctness gate, promote the fastest safe variant with reproducible commands — but it exposes no install command, no package name, no repository URL, no README, and no runnable entrypoint.

  • install — skipped. There is no documented install command or source location in the provided text, so the step cannot be executed as specified.
  • smoke-invocation — skipped. There is no minimal invocation syntax to run. The skill is described entirely as a concept (baseline → variants → correctness gate → promote), with nothing to actually call.

Observed dependencies: none. That is itself a symptom — a skill with a real, runnable surface usually names at least one tool, runtime, or package. Here the dependency list is empty because there is nothing to depend on yet.

What we observed

The trigger and scope language is genuinely good. The description names the exact situations ("speed this up," "try many variants," "run recursive optimization," "benchmark latency/throughput/cost," "pick the best implementation by repeated measured tests") and constrains itself to a measurable loop with a correctness gate. That is a well-formed specification.

But a specification is not an executable artifact. The harness cannot distinguish a skill that works from a skill that merely reads well when the body is a behavioral summary. Our two tests did not fail because the loop is wrong; they were skipped because there was nothing to run. We therefore have zero empirical evidence about whether the loop produces correct baselines, whether the correctness gate holds, or whether "promote the fastest safe variant" behaves safely under real variants.

Rating caveat

The composite 4.0 / 5.0 and its dimensions (D1 4.5, D2 4.0, D3 4.5, D4 3.0, D5 4.0) should be read as theoretical until a physical re-run resolves the two skips. D4 self-containment at 3.0 is the honest weak point and lines up exactly with what the harness hit: the skill does not carry enough of itself to be installed or invoked from the text alone. Every other dimension is scoring the description, not a verified behavior. Treat the number as provisional.

Does it still seem valuable in principle?

Yes — and this is not a takedown. The concept is the right shape for the problem. "Make it faster" is a notoriously unbounded request, and the skill's core discipline — baseline before variants, one hypothesis at a time, a correctness gate before promotion, reproducible commands — is exactly the discipline that prevents benchmark-driven regressions. If the body delivers what the summary promises, this is a genuinely reusable optimization harness.

The path to a real rating is narrow and clear: ship an untruncated SKILL.md with an install command, a repository URL, and one minimal runnable invocation. Re-run the harness. If the loop then executes and the correctness gate demonstrably blocks a bad variant, the 4.0 has a chance of being earned rather than assumed. Until then, it is a promising spec with an unverified body.

03 — Tests

What we tried


Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-10-03.

Overall: broken. 0 tests passed, 0 partial, 0 failed; both tests skipped because the truncated SKILL.md exposes no install command, README, or runnable invocation to exercise.

Test Status Notes
install skipped SKILL.md content is truncated to a behavioral description and contains no documented install command, package name, or repository URL, so the install step cannot be executed as specified.
smoke-invocation skipped No README or minimal invocation syntax is present in the provided SKILL.md text; the skill only describes a conceptual optimization loop (baseline, one-hypothesis variants, correctness gate, promote fastest safe variant) with no runnable entrypoint.
04 — Cross-validation

1 source verified

Install

Use this skill

/plugin install benchmark-optimization-loop