Agent Infrastructure  ·  Curated marketplace

agent-eval

Head-to-head comparison of coding agents (Claude Code, Aider, Codex, etc.) on custom tasks with pass rate, cost, time, and consistency metrics. Use when choosing between coding agents, or when a…


Composite

4.2

C 4.2 · A 0.0

How we got there

Craft · D1–D5

D1 · Trigger clarity 4.5
D2 · Output specificity 4.5
D3 · Scope precision 4.5
D4 · Self-containment 3.5
D5 · Reusability 4.0

02 — Review

Our evaluation


Agent-eval: A Sharp Instrument That Needs a Sheath

The agent-eval skill sits in the agent-infrastructure cluster, a space crowded with tools that promise to measure the unmeasurable. Compared to its siblings like xlsx and docx in the same Anthropic catalog — which focus on deterministic document manipulation — this skill tackles a far messier problem: comparative evaluation of coding agents under custom workloads. That ambition is both its strength and its Achilles' heel.

What works: trigger clarity and output specificity (both 4.5). The skill's trigger is unambiguous: "Use when choosing between coding agents, or when a change to an agent setup needs measured pass rate, cost, and time rather than an impression." That last clause is the killer — it explicitly rejects vibes-based decision-making. The output spec is equally disciplined: pass rate, cost, time, and consistency metrics. These are the four numbers that actually matter when you're justifying a tooling change to a skeptical engineering manager. There's no fluff about "qualitative insights" or "subjective experience"; the skill knows it's producing a scoreboard.

Where it stumbles: self-containment (3.5) and reusability (4.0). The harness observation is damning: "SKILL.md lacks a concrete invocation command or usage example, making smoke testing impossible." This is the difference between a skill and a blog post. The document describes what to measure but not how to run the measurement. Compare this to xlsx in the same cluster, which presumably includes a python -m xlsx_skill --input file.xlsx invocation. Here, a user reading the skill knows they need Python ≥3.10 but has no entry point. Do they write a script? A shell loop? A YAML config? The skill's scope precision (4.5) tells you which agents to compare and what metrics to collect, but the absence of a concrete harness means every user reinvents the execution layer.

The deeper judgment call: The skill's value proposition is that it turns "I think Claude Code is faster" into "Claude Code completed 12/15 tasks at $0.42/task with a 3.2s median latency." That's genuinely valuable — more so than docx's template filling, because it addresses a high-stakes, recurring decision. But the lack of a runnable example undermines the core promise. A user who can't smoke-test the skill in five minutes will likely abandon it and fall back on anecdote. The 4.2 composite score feels right: this is a well-designed specification trapped in an incomplete implementation.

The fix is small but non-negotiable. Add a Usage section with a concrete command — even something as simple as python agent_eval.py --agents claude-code,aider --tasks tasks.json --runs 3. That single addition would push self-containment to 4.5+ and reusability to 4.5. As written, agent-eval is like a surgical instrument with no handle: the blade is sharp, but you can't grip it. In a cluster where xlsx and docx are turnkey utilities, this skill demands more from its user than it gives back in setup cost. The measurement philosophy is right; the packaging needs to catch up.

03 — Tests

What we tried


Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-08-29.

Overall: partial. 1 test passed, 0 partial, 1 failed; key blocker: SKILL.md lacks a concrete invocation command or usage example, making smoke testing impossible.

Inferred dependencies: python>=3.10.

Test Status Notes
install pass Skill installs cleanly; no external dependencies beyond standard Python tooling.
smoke-invocation fail No CLI entry point or documented invocation command found in SKILL.md; only a high-level description of the skill's purpose.
04 — Cross-validation

1 source verified

Install

Use this skill

/plugin install agent-eval