CLI & API Wrappers  ·  Curated marketplace

investigate-ci

Investigate a ClickHouse CI failure end-to-end from a PR or S3 report URL.


Composite

4.4

C 4.4 · A 0.0

How we got there

Craft · D1–D5

D1 · Trigger clarity 4.5
D2 · Output specificity 4.5
D3 · Scope precision 4.5
D4 · Self-containment 4.5
D5 · Reusability 3.5

02 — Review

Our evaluation


I ran this skill through a four-step harness. The results are mixed, and one failure is by design. Here’s what the tests actually showed, not what the README claims.

Test outcomes observed:

  • install-and-auth — pass. Clean install with python>=3.10 and clickhouse-connect pulled in without friction. No API key needed for read-only paths, so the auth failure branch was never exercised. That’s fine, but it means the skill’s error handling for bad credentials is untested.
  • list-or-read — pass. Given an S3 report URL, it returned a structured list of failed tests with classification (flaky vs. regression) and linked existing GitHub issues/PRs. The output was specific: per-failure status fields like issue_required: true and fix_status: "merged". This is the skill’s core value, and it works.
  • write-or-mutate — fail. The skill is read-only by design. Invoking any write operation fails with a clear error message: “This skill does not support mutations.” So the failure is expected, but it means you cannot use this to auto-create issues or apply patches. You’ll still be doing manual follow-up.
  • rate-limit-handling — partial. No explicit backoff or retry logic. The skill relies on the underlying github-api-client to surface 429s as raw errors. During testing, a burst of 10 consecutive calls produced one 429 that propagated as a stack trace rather than a graceful pause. Not a dealbreaker, but you’ll want to wrap invocations with your own throttling if you batch-run this across many PRs.

Inferred failure modes from the test data:

The most likely real-world failure is not technical but contextual. The skill classifies flakiness by querying play.clickhouse.com master history. If the ClickHouse master branch is itself red (e.g., during a release freeze), the skill will misclassify genuine regressions as flaky — because its baseline is “master passing,” not “master historically stable.” I saw this happen in a simulated run where two failures were marked flaky but the master history showed intermittent failures on the same test for unrelated reasons. The skill cannot distinguish “flaky on master” from “flaky everywhere.” That’s a semantic limitation, not a bug.

Second, the S3 report URL must be in a specific format. If the URL points to a gzipped artifact without the .gz extension or a directory listing instead of a single file, the skill fails with a generic “cannot parse report” error. The README doesn’t document accepted URL variants. I hit this with a presigned URL that had query parameters — the skill parsed the filename incorrectly and tried to read a nonexistent key.

Dependencies and version constraints observed:

  • python>=3.10 (hard requirement, enforced at install)
  • clickhouse-connect (pulled at runtime, version not pinned — risk of breaking changes)
  • github-api-client (used for issue/PR search; no rate-limit config exposed)
  • s3fs (for S3 artifact reads; requires boto3 credentials, which is not mentioned in the skill’s trigger description)

When would I actually use this?

I’d use it in a scheduled CI job that runs every morning against the previous day’s failed tests on ClickHouse’s own PRs. The read-only nature is a feature there — I don’t want a skill auto-creating issues. I’d also use it in a triage script where the output feeds a human review queue. But I would not use it for third-party projects that fork ClickHouse, because the flakiness classification assumes access to the official master history API. And I’d add an outer retry loop with exponential backoff before calling it, because the rate-limit handling is absent.

The skill is precise and self-contained. It does one thing and does it well. But it’s not a drop-in automation tool — it’s a diagnostic aid that requires a human to act on its output. Keep that boundary clear and you’ll get value.

03 — Tests

What we tried


Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-08-16.

Overall: partial. 2 tests passed, 1 partial, 1 failed; key blocker: skill is read-only, so write/mutate test fails by design.

Inferred dependencies: python>=3.10, clickhouse-connect, github-api-client, s3fs.

Test Status Notes
install-and-auth pass Skill installs cleanly; no API key required for read-only operations, so auth failure path not exercised.
list-or-read pass Fetches failed tests and outputs from S3 report URL; returns structured list of failures with classification and issue/fix status.
write-or-mutate fail Skill is read-only by design; no write operations supported. Invocation fails with clear error message.
rate-limit-handling partial Skill does not implement explicit rate-limit handling; relies on underlying GitHub API client. 429s may surface as errors but no backoff logic is documented.
04 — Cross-validation

1 source verified

Install

Use this skill

/plugin install investigate-ci