CLI & API Wrappers  ·  Official

playwright

Use when the task requires automating a real browser from the terminal (navigation, form filling, snapshots, screenshots, data extraction, UI-flow debugging) via `playwright-cli` or the bundled…

  • playwright

Composite

3.4

C 4.3 · A 2.9

How we got there

Craft · D1–D5

D1 · Trigger clarity 4.5
D2 · Output specificity 4.0
D3 · Scope precision 4.5
D4 · Self-containment 4.5
D5 · Reusability 3.5

Adoption · A1–A5

A1 · Maintenance 2.5
A2 · Documentation 2.8
A3 · License 2.5
A4 · Adoption 4.3
A5 · Authorship 2.0

02 — Review

Our evaluation


Tier-2 Review: playwright (cli-and-api)

What we attempted

We ran the playwright skill through the standard Tier-2 harness to move its composite score (4.3 / 5.0) from a documentation-only estimate to something grounded in observed behavior. The intent was to exercise four surfaces the skill claims to cover — install/auth, read-only operations, mutation, and rate-limit handling — using the SKILL.md's own instructions as the sole runbook.

What failed

The harness reported 0 passed, 3 partial, 0 failed, 1 skipped. No test reached a clean pass. The three partials are partial for a single structural reason, not three separate bugs: this skill is a browser-automation CLI wrapper, and every real operation depends on npx plus a working browser binary plus network access to a live target URL. The harness container had none of the three in a guaranteed-available state, so each "operation" degraded to a prerequisite check rather than an execution.

The skipped test is the most diagnostic. The harness's rate-limit-handling probe assumes an HTTP API surface with 429/backoff semantics. This skill has no such surface — it drives a local browser session. The test is not applicable to the documented scope, which is itself a finding: the harness's generic cli-and-api rubric does not fit this skill's actual shape.

What we observed

  • install-and-auth (partial). SKILL.md contains no API-key or auth concept, correctly — it is not that kind of tool. The real gate is command -v npx. In a clean container without Node/npm, the prerequisite check fires and the skill instructs the user to install Node.js/npm rather than crashing. That graceful degradation is a genuine strength, but it means the test never got past the gate.
  • list-or-read (partial). There is no list/get/search API. The simplest read is open + snapshot, which requires network access and a browser binary. Offline, the open step fails; SKILL.md documents re-snapshotting on stale refs but does not describe offline error handling. The test could not distinguish "skill works" from "skill never ran."
  • write-or-mutate (partial). Mutation is fill/click against refs from the latest snapshot. SKILL.md explicitly warns refs go stale and that commands fail on missing refs, requiring a re-snapshot. It documents no idempotency or rollback semantics, so partial-failure behavior is undefined. This is a real gap independent of the harness environment.
  • rate-limit-handling (skipped). Not applicable; see above.

The common thread: the harness's simulated environment could not satisfy the skill's hard runtime dependencies, so we measured the prerequisite gate three times and the actual automation zero times.

Rating caveat

The 4.3 / 5.0 composite is therefore theoretical. It reflects the SKILL.md's documentation quality — and the dimension scores are credible on that basis (clear triggers at 4.5, precise scope at 4.5, strong self-containment at 4.5) — but it does not reflect observed execution. Until a physical re-run with Node/npm, a Playwright browser binary, and live network access resolves the three partials, treat the composite as an upper bound, not a measured result. The lowest dimension, reusability at 3.5, is the one most likely to move once real execution is possible — in either direction.

Does the skill still seem valuable in principle?

Yes. The failure mode here is environmental, not conceptual. SKILL.md is unusually disciplined: it front-loads the npx prerequisite check, tells the user how to recover, warns about stale refs, and explicitly steers away from @playwright/test unless asked. Those are the behaviors of a well-scoped skill. The gaps we can name honestly — undefined mutation idempotency, no offline error path — are fixable in the doc and worth fixing. A physical re-run is the only thing that will convert this from a promising spec into a verified one.

03 — Tests

What we tried


Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-09-15.

Overall: broken. 0 tests passed, 3 partial, 0 failed, 1 skipped; key blocker: the skill is a browser-automation CLI wrapper with no API-key auth or rate-limit surface, and all real operations depend on npx plus a working browser and network.

Inferred dependencies: npx (Node.js/npm), @playwright/cli (via npx --package @playwright/cli playwright-cli), playwright browser binaries, network access for target URLs, bash-compatible shell, CODEX_HOME env var (default ~/.codex).

Test Status Notes
install-and-auth partial SKILL.md has no API-key auth concept; it is a browser CLI wrapper. The prerequisite check for npx is the real gate, and the wrapper relies on npx --package @playwright/cli playwright-cli, so a clean container without Node/npm fails the prerequisite check and the skill correctly instructs the user to install Node.js/npm rather than crashing.
list-or-read partial The simplest read-only operation is open + snapshot (there is no list/get/search API). It requires network access to fetch the page and a browser binary; in an offline container the open step fails, and the SKILL.md does not describe explicit offline error handling beyond re-snapshotting on stale refs.
write-or-mutate partial Mutation is via fill/click on refs from the latest snapshot. SKILL.md explicitly warns refs can go stale and that commands fail on missing refs, requiring a re-snapshot; it does not document idempotency or rollback semantics, so partial-failure behavior is undefined by the skill.
rate-limit-handling skipped SKILL.md contains no rate-limit, 429, or backoff guidance; the CLI drives a local browser session rather than a rate-limited HTTP API, so this test is not applicable to the documented surface.
04 — Cross-validation

1 source verified

Install

Use this skill

/plugin install playwright