Agent Infrastructure  ·  Curated marketplace

exploring-ai-failures

Find where an AI/LLM application is failing in production and surface the failure patterns, working from real traces.


Composite

4.2

C 4.2 · A 0.0

How we got there

Craft · D1–D5

D1 · Trigger clarity 5.0
D2 · Output specificity 3.5
D3 · Scope precision 4.5
D4 · Self-containment 4.0
D5 · Reusability 3.5

02 — Review

Our evaluation


Tier-2 Review: exploring-ai-failures

Slug: exploring-ai-failures
Cluster: agent-infrastructure
Source: skillsmp.com/creators/posthog/posthog/products-ai-observability-skills-exploring-ai-failures
Composite score (theoretical): 4.2 / 5.0


What we attempted

We pulled the SKILL.md from the source URL and ran it through our standard harness (D-048). The harness checks for three things: (1) whether the skill has explicit install instructions, (2) whether it exposes a runnable CLI or executable, and (3) whether the skill’s trigger descriptions match its actual content. The first two checks require a package or script; the third is a text-contrast pass.

What failed

Both install and smoke-invocation checks were skipped — zero tests passed, zero partial, two skipped.

The blocker is structural: this skill is not an executable or a package. It is a conceptual methodology document. The SKILL.md contains no install command, no requirements.txt, no main.py, no CLI entry point, and no invocation syntax. There is nothing to run. The harness therefore could not validate that the skill “works” in any mechanical sense.

We did not observe a crash or a wrong output — because there was no output to observe. The skill is a guide for a human (or an agent) to manually investigate AI failures using traces. That is a legitimate artifact type, but it does not fit the harness’s assumption that a skill is a runnable unit.

What we observed (from content inspection)

The SKILL.md text is well-structured and internally consistent. Its trigger clarity (D1 = 5.0) is genuinely strong: the opening sentence names the use case explicitly (“Find where an AI/LLM application is failing in production”), and the trigger phrases are concrete (“what’s failing in my agent”, “surface error patterns”, “why are the responses bad”). Scope precision (D3 = 4.5) is also good — it narrows the investigation to one use case at a time and lists specific trace-selection signals (code errors, metric outliers, trace-type slices, manual review, eval spikes, clustering). The methodology is coherent: scope → find traces → rank into a failure taxonomy.

However, output specificity (D2 = 3.5) and reusability (D5 = 3.5) are lower, and that matches the content. The skill tells you what to do but not what to produce in a machine-readable or structurally enforced way. There is no schema for the failure taxonomy, no example output format, no checklist that could be mechanically verified. Reusability suffers because the skill depends heavily on the user already having PostHog-style trace tooling and domain context; it does not abstract beyond that.

Self-containment (D4 = 4.0) is decent — the skill does not reference external files or hidden dependencies — but it assumes the reader knows what “traces” and “eval spikes” mean in the PostHog ecosystem.

Honest rating caveat

The 4.2 composite score is theoretical. It reflects content quality only, because we could not execute the skill. Until the author either (a) packages this as a runnable CLI that ingests a trace export and emits a ranked taxonomy, or (b) explicitly marks it as a “human workflow” skill and the harness is updated to validate methodology documents against a rubric rather than an execution — the numerical score should not be treated as evidence of functional correctness. It is a content-quality estimate, nothing more.

Is the skill still valuable in principle?

Yes. The failure mode here is not a flaw in the skill’s thinking; it is a mismatch between artifact type and test harness. A methodology for turning raw traces into a ranked failure taxonomy is genuinely useful for anyone debugging LLM agents in production. The trigger phrases are realistic, the scoping advice is sound, and the trace-selection signals are practical. If the author later adds a concrete output schema (e.g., JSON lines with failure_class, evidence_trace_id, confidence), this could become both a strong guide and a testable artifact. As-is, it is a good read, not a good binary — and that is okay, as long as nobody mistakes the 4.2 score for a passing test result.

03 — Tests

What we tried


Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-09-09.

Overall: broken. 0 tests passed, 0 partial, 2 skipped; key blocker: skill is a non-executable guide with no install or invocation commands.

Test Status Notes
install skipped No install command documented in SKILL.md; skill is a conceptual guide, not a package.
smoke-invocation skipped No executable or CLI defined; skill is a methodology for manual investigation, no runnable code.
04 — Cross-validation

1 source verified

Install

Use this skill

/plugin install exploring-ai-failures