Methodology  ·  Curated marketplace

evaluation-methodology

PluginEval quality methodology — dimensions, rubrics, statistical methods, and scoring formulas.


Composite

3.1

C 4.2 · A 2.5

How we got there

Craft · D1–D5

D1 · Trigger clarity 4.5
D2 · Output specificity 4.0
D3 · Scope precision 4.5
D4 · Self-containment 4.0
D5 · Reusability 3.5

Adoption · A1–A5

A1 · Maintenance 2.5
A2 · Documentation 1.0
A3 · License 2.5
A4 · Adoption 4.6
A5 · Authorship 2.0

02 — Review

Our evaluation


Tier-2 Review: evaluation-methodology

What we attempted

We pulled the evaluation-methodology skill (cluster: methodology; source: wshobson-agents/plugins/plugin-eval/skills/evaluation-methodology/SKILL.md) into the standard Tier-2 harness and ran the usual two-stage protocol:

  1. Install — resolve a package, repository path, or documented install command and materialize the skill into a clean sandbox.
  2. Smoke invocation — execute the skill's documented entrypoint and capture observable output.

Both stages were attempted. Both were skipped by the harness because the artifact under test is not executable in any conventional sense.

What failed

  • Install stage — skipped. The provided SKILL.md is a truncated methodology description. It contains no package name, no repository path, no npm install / pip install / git clone instruction, and no manifest. There is nothing for the harness to fetch.
  • Smoke-invocation stage — skipped. There is no README, no invocation syntax, no CLI, no function signature, and no example call. The content describes dimensions, rubrics, statistical methods, and scoring formulas — conceptually. There is no runnable entrypoint to exercise.

Result: 0 passed, 0 partial, 0 failed, 2 skipped. The harness did not fail the skill; it could not address the skill at all. That distinction matters and we are not going to dress it up as a pass.

What we observed

Reading the truncated SKILL.md directly (the only thing we could do), the artifact is a specification document. It defines what plugin quality means across dimensions, how rubrics map to scores, and how statistical aggregation produces a composite. It is the kind of text a human reads to understand a scoring system, not the kind of artifact a harness executes.

Two consequences follow:

  1. The composite score of 4.2 / 5.0 is theoretical. It was derived from static inspection of the SKILL.md — trigger phrasing, output specificity, scope boundaries, self-containment, reusability — not from behavioral evidence. Every dimension score is a reading of prose, not a measurement of behavior. We have no signal on whether the described methodology, once instantiated, actually produces the scores it claims to produce.
  2. The auto-summary ("High quality skill with clear triggers and methodology") is consistent with what we can see, but it is not corroborated by execution. It is a description of the document, not a verification of the skill.

This is not a knock on the skill's author. It is a limitation of the artifact class. A methodology spec is, by construction, not runnable in isolation. The harness has no honest way to grade it beyond reading.

Why the rating stays theoretical

We cannot promote 4.2 / 5.0 to a verified rating until a physical re-run resolves the failure modes above. Concretely, that would require one of:

  • A companion README or wrapper that documents how the methodology is invoked (e.g., a scorer script, a rubric evaluator, a CLI that takes a plugin and emits dimension scores).
  • A repository path with installable code that operationalizes the formulas in the SKILL.md.
  • A fixture — a known plugin with known expected scores — that lets us confirm the methodology reproduces its own outputs.

Absent any of those, the harness has nothing to exercise and the score remains a static-inspection estimate. We are flagging that explicitly rather than letting a 4.2 read as a behavioral result.

Does the skill still seem valuable in principle?

Yes — with a caveat. The dimension set (trigger clarity, output specificity, scope precision, self-containment, reusability) is a reasonable decomposition of plugin quality, the rubric structure is legible, and the stated use cases — interpreting a low dimension score, calibrating marketplace thresholds, explaining badges to external partners — are real needs. As a reference document for how quality is defined, it is useful.

Its value as a skill in the harness sense is unproven, because nothing about it is executable. The right next step is not to re-score it; it is to obtain the operational layer (script, wrapper, or fixture) that turns the methodology from prose into behavior. Until then, treat the 4.2 as a hypothesis, not a finding.

03 — Tests

What we tried


Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-09-18.

Overall: broken. 0 tests passed, 0 partial, 0 failed; both tests skipped because the SKILL.md contains only conceptual methodology text with no install command, README, or runnable invocation to exercise.

Test Status Notes
install skipped SKILL.md is a truncated methodology description with no documented install command, package name, or repository path, so no install step can be executed or verified.
smoke-invocation skipped No README or invocation syntax is present in the provided SKILL.md; the content only describes conceptual dimensions, rubrics, and scoring formulas with no runnable entrypoint.
04 — Cross-validation

1 source verified

Install

Use this skill

/plugin install evaluation-methodology
Use cases

Tasks this skill helps with