Methodology · Curated marketplace
interview-me
Extracts what the user actually wants instead of what they think they should want.
Composite
C 4.1 · A 0.0
How we got there
Our evaluation
Tier-2 Review: interview-me (methodology)
What we attempted
We ran the standard two-stage harness against this skill: an install step and a smoke-invocation step, intended to confirm that the skill could be resolved from its source and exercised at least once in a minimal, non-interactive way.
The source is a skillsmp.com listing page. The SKILL.md we were able to retrieve is truncated to 6,000 characters, and the excerpt we inspected (a further 4,000-char cut) contains only the skill's descriptive prose — the trigger conditions and the intent-extraction rationale. There is no install command, no package name, no repository URL, and no invocation syntax anywhere in the retrieved text.
What failed
Both harness stages skipped, not passed:
- install — skipped. No install command, package identifier, or repository instructions were present in the truncated
SKILL.md. There was nothing to execute, so dependency resolution could not be attempted. Observed dependencies: none, but this is a statement about what we could see, not a verified absence. - smoke-invocation — skipped. No README, no CLI entrypoint, and no invocation syntax appeared in the excerpt. This skill is a prompt/methodology — a one-question-at-a-time interview loop — not a command-line tool, so a minimal invocation could not be constructed from the available material.
Net result: 0 passed, 0 partial, 0 failed, 2 skipped. The harness never reached a state where the skill's behavior could be observed. We are not reporting a failing skill; we are reporting a skill we could not exercise.
What we observed
The failure mode is specific and worth naming: the artifact is incomplete at the point of retrieval. The truncation is a property of how the content was served/captured, not necessarily of the skill itself, but the effect is the same — the machine-readable contract (how to install it, how to call it) is missing. For a pure-prompt methodology this is less damning than it would be for a CLI tool, because the "interface" is natural language. But it also means there is nothing for an automated harness to bind to. A methodology skill with no documented invocation surface can only be validated by a human reading it and choosing to apply it.
The auto-summary on the page ("High quality skill with clear triggers and focused scope") is consistent with the visible prose — the trigger conditions are genuinely well-drawn — but it is a reading of the description, not a result of execution. We could not corroborate it.
On the rating
The composite 4.1 / 5.0 and the per-dimension scores (D1 4.5, D2 3.5, D3 4.5, D4 4.0, D5 4.0) should be treated as theoretical until a physical re-run resolves the two skips. Specifically:
- D1 (trigger clarity) and D3 (scope precision) are scored from the descriptive text alone, which is the part most likely to be accurate even under truncation.
- D2 (output specificity) is the softest score at 3.5, and it is exactly the dimension a smoke test would have probed — what does the skill actually produce at the end of the interview? We have no evidence either way.
- D4 (self-containment) cannot be confirmed while the install path is unknown.
A re-run with the full, untruncated SKILL.md — or with a direct repository link — would let the harness attempt a real invocation and either raise or confirm these numbers.
Does the skill still seem valuable in principle?
Yes. The core idea — refusing to silently fill in ambiguous requirements, and instead running a one-question-at-a-time interview to ~95% confidence before any plan, spec, or code exists — addresses a real and common failure mode in agentic workflows. The trigger set is unusually well-chosen: it fires on underspecified asks, on explicit invocations ("interview me", "grill me", "stress-test my thinking"), and on the self-detection case where the agent notices itself guessing. That third trigger is the interesting one and is rarely articulated this cleanly.
The value is plausible on inspection. It is simply unverified by execution, and this review should be read as a description of an incomplete artifact rather than a judgment on the skill's merits.
What we tried
Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-09-26.
Overall: broken. 0 tests passed, 0 partial, 0 failed; both tests skipped because the truncated SKILL.md exposes no install command or invocation interface, so no dependencies could be verified.
| Test | Status | Notes |
|---|---|---|
| install | skipped | SKILL.md content is truncated to 6000 chars and contains no install command, package name, or repository instructions; the source is a skillsmp.com listing page, so no documented install command is present to execute. |
| smoke-invocation | skipped | No README or invocation syntax is provided in the SKILL.md excerpt; the skill is a prompt/methodology (one-question-at-a-time interview) with no CLI entrypoint, so a minimal invocation cannot be constructed. |
1 source verified
- Best source
skillsmp.com - Authority tier Tier 2 — Curated marketplace
- Stars ★ 42,726
- Source link https://skillsmp.com/skills/addyosmani-agent-skills-skills-interview-me-skill-md ↗
- First published 2026-05-24
- Last modified 2026-09-26
Use this skill
/plugin install interview-me Head-to-head pages featuring interview-me
More in Methodology
claude-api
Reference for the Claude API / Anthropic SDK — model ids, pricing, params, streaming, tool use, MCP, agents, caching, token counting, model migration.
prompt-engineering
Universal prompt engineering techniques for any LLM.
github-swyxio-ai-notes
notes for software engineers getting up to speed on new AI developments.
hatch-pet
Create, repair, validate, preview, and package Codex-compatible animated pet spritesheets from character art, screenshots, generated images, or visual references.