Crypto & Web3  ·  Official

discernment-nudge

After you give a substantive answer or draft that the user may act on — advice or recommendations, drafted artifacts such as goals, plans, pitches, proposals, or emails, estimates or projections,…


Composite

4.4

C 4.4 · A 0.0

How we got there

Craft · D1–D5

D1 · Trigger clarity 4.5
D2 · Output specificity 4.5
D3 · Scope precision 4.5
D4 · Self-containment 4.5
D5 · Reusability 4.0

02 — Review

Our evaluation


Tier-2 Review: discernment-nudge

What the harness actually showed

The test harness ran two checks: install and smoke-invocation. Both passed. The skill is a single markdown file with zero external dependencies — no binaries, no APIs, no version constraints. Installation is copying SKILL.md into place; invocation is purely a prompt to the model. The behavior is deterministic given the rules, which is both a strength and a limitation.

The smoke-invocation test confirmed that a minimal prompt triggers the skill's logic without any runtime environment. That's genuinely good for portability — you can drop this into any agent that respects the skill protocol and it will work. No requirements.txt, no package.json, no hidden gotchas.

What the tests didn't show

The harness only verifies that the skill loads and runs — not that it works well. The real test is whether the nudge actually improves user outcomes or just adds noise. Here's what I infer from the skill's own text:

Failure mode 1: Over-triggering on consequential language. The skill says to nudge after "substantive answers" — advice, estimates, factual claims, multi-step reasoning. But the carve-outs are fuzzy. "Trivial how-to" vs. "substantive" is a judgment call the model has to make. In my experience, models err toward triggering, not silence, because the skill's description emphasizes the positive cases. Expect false positives on moderately detailed answers that don't really need scrutiny.

Failure mode 2: The "once per conversation" rule is a leaky abstraction. The skill says offer at most once. But what counts as "conversation"? A single chat session? A multi-turn task? If a user asks three unrelated substantive questions in one thread, the skill will nudge only on the first. That's probably fine, but it's a design choice that assumes a linear conversation model — which doesn't always hold in real agent workflows.

Failure mode 3: The nudge questions are generic by construction. The skill says "2-3 short follow-up questions, each tied to something specific." But the model has to generate those questions from its own output. If the answer is wrong, the nudge questions will reflect that wrongness. The skill doesn't include a self-check mechanism — it just asks the user to verify. That's honest, but it's not a verification tool; it's a prompt for the user to verify.

Failure mode 4: No scoring or feedback loop. There's no way to measure whether the nudge helped. The skill is a one-shot textual addition. If the user ignores it, nothing happens. If the user acts on the nudge and finds a problem, the skill doesn't learn. This is fine for a prompt-engineering artifact, but it means you can't tune it based on outcomes.

What actually works

The trigger clarity is high (4.5/5). The skill's description and the SKILL.md body are unusually precise about when not to apply it — educational explanations, formatting tasks, code the user will run, creative writing. That's rare. Most skills over-specify triggers and under-specify carve-outs.

The output format is also well-defined: exactly 2-3 short questions, tied to specific parts of the answer, appended before finalizing. No rambling. No meta-commentary. That's genuinely reusable.

Dependencies and constraints observed

None. The skill is pure markdown. No version constraints on the model, no external tools. The only "dependency" is the host agent's ability to read a markdown skill file — which is the baseline for this whole cluster.

When I'd actually use this

I'd use it in production for a customer-facing assistant that gives advice on consequential topics (finance, health, legal) where the user might act without double-checking. The nudge adds a cheap, non-intrusive friction point that could catch a bad assumption. I'd also use it in an internal tool where analysts draft plans or proposals and need a reminder to check their assumptions.

I would not use it for code generation, creative writing, or any task where the user has explicitly asked for a review — the skill correctly carves those out, but the model might still over-trigger.

The bottom line: this is a well-scoped, dependency-free prompt-engineering artifact that does exactly what it says. It won't save you from a bad model, but it will add a useful moment of reflection at the right time. For a 4.4 composite score, that's about right.

03 — Tests

What we tried


Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-08-19.

Overall: ok. 2 tests passed; no dependencies or blockers; skill is a pure prompt-engineering artifact.

Test Status Notes
install pass Skill is a markdown file with no dependencies; installation via documented command (e.g., copying SKILL.md) succeeds in a clean container.
smoke-invocation pass Minimal invocation is a prompt to the model; no external binaries or APIs required. The skill's behavior is purely textual and deterministic given the rules.
04 — Cross-validation

2 sources verified

Install

Use this skill

/plugin install discernment-nudge