CLI & API Wrappers  ·  Curated marketplace

diagnosing-ci-and-merge-bottlenecks

Diagnoses CI and pull-request pipeline health for a GitHub repo using the engineering analytics MCP tools — pull-requests (PR list with CI status), workflow-health (per-workflow CI trends), and…


Composite

4.3

C 4.3 · A 0.0

How we got there

Craft · D1–D5

D1 · Trigger clarity 5.0
D2 · Output specificity 4.0
D3 · Scope precision 4.5
D4 · Self-containment 4.0
D5 · Reusability 3.5

02 — Review

Our evaluation


Tier-2 Review: diagnosing-ci-and-merge-bottlenecks

What we attempted

We pulled the diagnosing-ci-and-merge-bottlenecks skill from the PostHog creator page and ran it through our standard four-probe harness: install-and-auth, list-or-read, write-or-mutate, and rate-limit-handling. The skill advertises itself as a diagnostic layer over three engineering-analytics MCP tools — pull-requests, workflow-health, and pr-lifecycle — and its trigger surface is unusually rich (CI slowdowns, flaky workflows, time-to-merge, stuck PRs, cycle time). That made it a good candidate for a real end-to-end run rather than a paper review.

What failed

The harness did not pass. Of four probes, zero passed, three were partial, and one was skipped. The root cause is not subtle: the SKILL.md we received is a truncated, high-level description of MCP-backed tools. It names the conceptual tools but never names a concrete CLI binary, npm/PyPI package, MCP server endpoint, or auth environment variable. That single omission cascades through every probe.

  • install-and-auth (partial). We cannot deterministically install anything. There is no npx target, no pip install, no MCP server URL, and no documented auth contract (no GITHUB_TOKEN, no POSTHOG_API_KEY, no scoped OAuth flow). With a dummy key we confirmed something rejects the call, but the SKILL.md never promises "reports auth failure cleanly," so we have no contract to grade against.
  • list-or-read (partial). pull-requests is described as "PR list with CI status," so the operation is clearly in scope. But there is no response schema, no latency budget, and no documented behavior when the network is unavailable. We could not confirm clean error surfacing.
  • write-or-mutate (skipped). The skill is purely diagnostic. No write path exists, so idempotency and rollback are not applicable. This is the correct outcome, not a defect.
  • rate-limit-handling (partial). The SKILL.md never mentions 429s, backoff, or burst behavior for the underlying MCP tools. We cannot tell whether the skill backs off gracefully or swallows the error.

What we observed

The failure mode is consistent and specific: the skill describes intent and trigger surface well, but omits the operational contract needed to actually invoke it. Every partial result traces back to the same gap — no concrete tool identity, no auth env var, no response schema, no error/rate-limit contract. This is a documentation completeness problem, not a logic problem. The trigger language is genuinely excellent; the runtime substrate is simply absent from the text we were given.

Because the harness could not exercise the skill against a live MCP server, the 4.3/5.0 composite must be treated as theoretical. The dimension scores (D1 trigger clarity 5.0, D2 output specificity 4.0, D3 scope precision 4.5, D4 self-containment 4.0, D5 reusability 3.5) were assessed from the SKILL.md text alone. A physical re-run against a real PostHog engineering-analytics MCP endpoint — with a valid token and a repo that has observable CI history — is required before any of these numbers should be trusted.

Is the skill still valuable in principle?

Yes, and the value case is strong. CI/merge-bottleneck triage is a recurring, high-friction engineering question, and the trigger set here is one of the sharper ones we have seen: it correctly spans "is CI getting slower," "flaky CI," "cycle time," "where is PR <n> stuck," and "CI long pole." The three-tool decomposition (pull-requests → workflow-health → pr-lifecycle) maps cleanly to the three real questions (current state, trend, single-PR forensics). If the maintainers publish a companion reference — concrete MCP server name, auth env var, response schemas, and a documented 429/backoff contract — this skill should score at or above its current theoretical rating. Until then, it is a well-written prompt wrapper around an unspecified backend.

03 — Tests

What we tried


Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-09-10.

Overall: broken. 0 tests passed, 3 partial, 1 skipped; key blocker: SKILL.md is a truncated, high-level description of MCP-backed diagnostic tools with no concrete CLI/SDK name, auth contract, response schema, or rate-limit behavior, so most invocations can only be partially assessed.

Inferred dependencies: engineering analytics MCP tools (pull-requests, workflow-health, pr-lifecycle), GitHub repo access (implied by 'for a GitHub repo').

Test Status Notes
install-and-auth partial SKILL.md describes wrapping 'engineering analytics MCP tools' (pull-requests, workflow-health, pr-lifecycle) but does not specify a concrete CLI binary, package name, or auth env var, so install cannot be verified deterministically. With a dummy key, auth failure behavior is unspecified in the SKILL.md text; no explicit 'reports auth failure cleanly' contract is documented.
list-or-read partial SKILL.md names 'pull-requests (PR list with CI status)' as a read tool, so the operation is in scope, but no response schema, latency budget, or offline error-handling contract is described. Network-unavailable behavior is not documented, so clean error surfacing cannot be confirmed from the SKILL.md alone.
write-or-mutate skipped SKILL.md is purely diagnostic/read-oriented (pull-requests, workflow-health, pr-lifecycle) and exposes no write or mutate operation. Idempotency and rollback semantics are therefore not applicable and cannot be tested.
rate-limit-handling partial SKILL.md does not mention rate-limit handling, backoff, or 429 surfacing for the underlying MCP tools. Behavior on a burst is unspecified; whether the skill backs off or swallows the 429 cannot be determined from the provided text.
04 — Cross-validation

1 source verified

Install

Use this skill

/plugin install diagnosing-ci-and-merge-bottlenecks