General  ·  Curated marketplace

idea-os

Five-phase pipeline (triage → clarify → research → PRD → plan) that turns a raw idea into four linked files: clarifying questions, deep research, a PRD with non-goals and metrics, and a phased…


Composite

4.2

C 4.2 · A 0.0

How we got there

Craft · D1–D5

D1 · Trigger clarity 4.5
D2 · Output specificity 4.5
D3 · Scope precision 4.5
D4 · Self-containment 3.5
D5 · Reusability 4.0

02 — Review

Our evaluation


Tier-2 Review: idea-os

Slug: idea-os
Cluster: general
Composite score (theoretical): 4.2 / 5.0
Dimension scores: trigger clarity 4.5 · output specificity 4.5 · scope precision 4.5 · self-containment 3.5 · reusability 4.0


What we attempted

We pulled the SKILL.md for idea-os from the source URL and ran it through our standard harness (D-048). The harness performs two checks: an install step (does the skill define how to install or load it?) and a smoke-invocation step (can we call it with a sample input and observe an output?). We also reviewed the full SKILL.md content for structure, dependencies, and internal consistency.

What failed

Both harness steps were skipped — not failed in the sense of an error, but skipped because the skill does not present itself as executable. The SKILL.md describes a five-phase conceptual pipeline (triage → clarify → research → PRD → plan) that transforms a raw idea into four linked artifacts: clarifying questions, deep research, a PRD with non-goals and metrics, and a phased execution plan with a mermaid user journey and kill criteria.

There is no install command, no CLI entry point, no API, no file layout, and no invocation syntax. The document reads like a process spec or a framework description, not like a runnable skill. Our harness correctly identified this as a blocker: we cannot execute a process description.

What we observed

The SKILL.md is well-written and internally coherent. It defines triggers clearly (e.g., when a user says “I have an idea” or “help me think this through”), and it specifies output artifacts with enough detail that a human or an LLM could follow the steps without additional context. The scope is precise — it does not wander into adjacent tasks like market sizing or coding — and the outputs are named and structured.

However, self-containment scores lower (3.5) because the skill assumes the reader knows what a “PRD” is, what “kill criteria” mean, and how to produce a mermaid user journey. There is no embedded template, no example output, and no fallback if the user lacks domain knowledge. The skill is more of a guide than a tool.

Reusability is moderate (4.0): the pipeline could apply to any idea domain, but the lack of any concrete interface means each user must manually re-implement the steps every time. There is no way to parameterize the input or chain the output into another tool.

Honest rating caveat

The composite score of 4.2 is theoretical. It reflects the quality of the written specification, not observed behavior. Until the skill is rewritten to include at least one of the following — a runnable script, a prompt template with placeholders, a JSON schema for inputs/outputs, or a clear “how to invoke” section — we cannot verify that it performs as described. The rating should be treated as an upper bound on potential, not a validated measurement.

Is the skill still valuable in principle?

Yes. Despite the execution gap, the underlying logic is sound and genuinely useful. The idea of forcing a structured triage before research, and a PRD with non-goals and kill criteria before planning, is a strong antidote to the common failure mode of brainstorming without rigor. The output artifacts are specific enough to be actionable, and the mermaid journey requirement adds a visualization that many teams lack.

If the author adds a minimal invocation layer — for example, a run.sh that accepts a one-line idea and prints the four files, or even just a prompt template that an LLM can follow step-by-step — this skill would likely score 4.5+ across the board. As written, it is a high-quality reference document masquerading as a skill. The content deserves a home in an executable wrapper.

We recommend a re-run after the author provides either an install command or an explicit “usage” section with sample input and expected output. Until then, treat the 4.2 as a design score, not a performance score.

03 — Tests

What we tried


Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-09-03.

Overall: broken. 0 tests passed, 0 partial, 0 failed, 2 skipped; key blocker: SKILL.md describes a conceptual pipeline with no install or invocation instructions.

Test Status Notes
install skipped SKILL.md does not document an install command; it describes a five-phase pipeline, not a CLI tool.
smoke-invocation skipped No executable or invocation syntax is provided in SKILL.md; the content is a process description, not a runnable command.
04 — Cross-validation

1 source verified

Install

Use this skill

/plugin install idea-os