Agent Infrastructure · Curated marketplace
improve
Survey any codebase as a senior advisor and produce prioritized, self-contained implementation plans for OTHER models/agents to execute.
Composite
C 4.1 · A 0.0
How we got there
Our evaluation
Tier-2 Review: improve (motion/agents-skills-improve)
Cluster: agent-infrastructure Composite (theoretical): 4.1 / 5.0 Harness result: 0 passed, 0 partial, 1 failed, 1 skipped — not verified
What we attempted
We picked up improve because its dimension profile (trigger clarity 4.5, scope precision 4.5) suggested a well-bounded advisory skill, and the auto-summary described it as "high quality … with clear triggers and output." That is exactly the kind of skill a harness should be able to exercise cheaply: install it, invoke it against a fixture repo, and inspect the plan it emits.
Two probes were scheduled:
- install — resolve an installable artifact from the SKILL.md and the source URL, then install into a clean container.
- smoke-invocation — construct the minimal command that triggers the documented behavior and run it against a small fixture.
Neither completed. The results below are what actually happened, not what we hoped would happen.
What failed, and how
install (fail). The SKILL.md is a ~6000-character behavioral description. It contains no install command, no package name, no registry reference, and no artifact path. The source URL resolves to a web page (a skills-marketplace listing), not a CLI or package. In a clean container there was nothing to fetch and nothing to run — the install step had no target. This is not a transient failure or a network hiccup; the document simply does not describe a distributable unit.
smoke-invocation (skipped). With no install and no invocation syntax in the SKILL.md, we could not construct even a minimal command. The skill describes itself purely as read-only advisory behavior — "Survey any codebase … produce prioritized … plans" — with no entrypoint, no argument contract, and no example invocation. A skip here is a consequence of the install failure, not an independent data point.
What we observed
The failure mode is under-specification of the invocation surface, not a logic bug. The skill's intent is legible and internally consistent: read-only survey, prioritized output, self-contained plans for a downstream executor, explicit non-goals ("never implements, fixes, or refactors"). That is a coherent contract.
What is missing is the mechanical layer that lets any harness — human or automated — actually call it. There is no:
- install/distribution reference,
- invocation syntax or entrypoint,
- example input → output pair,
- declared runtime or dependency surface (the harness found none, which is itself ambiguous: zero dependencies could mean "truly standalone" or "unspecified").
D4 (self-containment, 3.5) is the dimension this failure most directly contradicts: a skill that cannot be resolved into a runnable artifact is not self-contained in the operational sense, regardless of how clean its prose is.
Why the rating is theoretical
The 4.1 composite is derived from the SKILL.md text and the auto-summary. No behavioral evidence supports it. We did not see the skill produce a plan, prioritize a finding, or hand off to another agent. Every dimension score is therefore an assessment of documentation quality, not observed behavior. Until a physical re-run resolves the install and invocation gaps, treat 4.1 as a hypothesis. The honest current status is unverified.
Does the skill still seem valuable in principle?
Yes — conditionally. The read-only-advisor pattern is genuinely useful in agent infrastructure: a survey-and-plan skill that deliberately refuses to mutate code is a clean separation of concerns, and the "handoff plans for another agent to implement" framing matches how multi-agent pipelines actually want to decompose work. The scope precision (4.5) reflects real discipline in the prompt.
But the value is latent until the invocation surface exists. The fix is small and concrete: add a package/registry reference or a runnable entrypoint, document the invocation syntax, and include one worked example. Do that, and this becomes testable — at which point the 4.1 can be confirmed, revised, or rejected on evidence rather than on the quality of its own description.
What we tried
Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-09-29.
Overall: broken. 0 tests passed, 0 partial, 1 failed, 1 skipped; key blocker: SKILL.md contains only a behavioral description with no install command, package reference, or invocation syntax, so neither install nor smoke-invocation could be executed.
| Test | Status | Notes |
|---|---|---|
| install | fail | SKILL.md is a truncated 6000-char description with no documented install command, package name, or registry reference; the source URL points to a web page rather than a CLI/package, so no installable artifact could be resolved in a clean container. |
| smoke-invocation | skipped | No README or invocation syntax is present in the provided SKILL.md; the skill is described only as a read-only advisory behavior ('Survey any codebase... produce prioritized... plans'), so no minimal command could be constructed or executed. |
1 source verified
- Best source
skillsmp.com - Authority tier Tier 2 — Curated marketplace
- Stars ★ 32,648
- Source link https://skillsmp.com/creators/motiondivision/motion/agents-skills-improve ↗
- First published 2026-07-12
- Last modified 2026-09-29
Use this skill
/plugin install improve Head-to-head pages featuring improve
More in Agent Infrastructure
skill-creator
Create, edit, improve, or audit AgentSkills.
loop-design-check
Design a goal-oriented agent loop, and review it for the ways loops go wrong — spinning and burning tokens, Goodhart-gaming the verifier, or running a wrong answer to completion.
exploring-scouts
How to explore and make sense of PostHog Signals scouts — the scheduled agents that scan a project and write reports into the Signals inbox.
blocks-network
Non-linear reference for managing Blocks Network agents — features, configuration, CLI, IO schemas, streaming, consumer SDK, publishing, invites, troubleshooting.