CLI & API Wrappers · Official
playwright
Use when the task requires automating a real browser from the terminal (navigation, form filling, snapshots, screenshots, data extraction, UI-flow debugging) via `playwright-cli` or the bundled…
- playwright
Composite
C 4.3 · A 2.9
How we got there
Our evaluation
Tier-2 Review: playwright (cli-and-api)
What we attempted
We ran the playwright skill through the standard Tier-2 harness to move its composite score (4.3 / 5.0) from a documentation-only estimate to something grounded in observed behavior. The intent was to exercise four surfaces the skill claims to cover — install/auth, read-only operations, mutation, and rate-limit handling — using the SKILL.md's own instructions as the sole runbook.
What failed
The harness reported 0 passed, 3 partial, 0 failed, 1 skipped. No test reached a clean pass. The three partials are partial for a single structural reason, not three separate bugs: this skill is a browser-automation CLI wrapper, and every real operation depends on npx plus a working browser binary plus network access to a live target URL. The harness container had none of the three in a guaranteed-available state, so each "operation" degraded to a prerequisite check rather than an execution.
The skipped test is the most diagnostic. The harness's rate-limit-handling probe assumes an HTTP API surface with 429/backoff semantics. This skill has no such surface — it drives a local browser session. The test is not applicable to the documented scope, which is itself a finding: the harness's generic cli-and-api rubric does not fit this skill's actual shape.
What we observed
- install-and-auth (partial). SKILL.md contains no API-key or auth concept, correctly — it is not that kind of tool. The real gate is
command -v npx. In a clean container without Node/npm, the prerequisite check fires and the skill instructs the user to install Node.js/npm rather than crashing. That graceful degradation is a genuine strength, but it means the test never got past the gate. - list-or-read (partial). There is no list/get/search API. The simplest read is
open+snapshot, which requires network access and a browser binary. Offline, theopenstep fails; SKILL.md documents re-snapshotting on stale refs but does not describe offline error handling. The test could not distinguish "skill works" from "skill never ran." - write-or-mutate (partial). Mutation is
fill/clickagainst refs from the latest snapshot. SKILL.md explicitly warns refs go stale and that commands fail on missing refs, requiring a re-snapshot. It documents no idempotency or rollback semantics, so partial-failure behavior is undefined. This is a real gap independent of the harness environment. - rate-limit-handling (skipped). Not applicable; see above.
The common thread: the harness's simulated environment could not satisfy the skill's hard runtime dependencies, so we measured the prerequisite gate three times and the actual automation zero times.
Rating caveat
The 4.3 / 5.0 composite is therefore theoretical. It reflects the SKILL.md's documentation quality — and the dimension scores are credible on that basis (clear triggers at 4.5, precise scope at 4.5, strong self-containment at 4.5) — but it does not reflect observed execution. Until a physical re-run with Node/npm, a Playwright browser binary, and live network access resolves the three partials, treat the composite as an upper bound, not a measured result. The lowest dimension, reusability at 3.5, is the one most likely to move once real execution is possible — in either direction.
Does the skill still seem valuable in principle?
Yes. The failure mode here is environmental, not conceptual. SKILL.md is unusually disciplined: it front-loads the npx prerequisite check, tells the user how to recover, warns about stale refs, and explicitly steers away from @playwright/test unless asked. Those are the behaviors of a well-scoped skill. The gaps we can name honestly — undefined mutation idempotency, no offline error path — are fixable in the doc and worth fixing. A physical re-run is the only thing that will convert this from a promising spec into a verified one.
What we tried
Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-09-15.
Overall: broken. 0 tests passed, 3 partial, 0 failed, 1 skipped; key blocker: the skill is a browser-automation CLI wrapper with no API-key auth or rate-limit surface, and all real operations depend on npx plus a working browser and network.
Inferred dependencies: npx (Node.js/npm), @playwright/cli (via npx --package @playwright/cli playwright-cli), playwright browser binaries, network access for target URLs, bash-compatible shell, CODEX_HOME env var (default ~/.codex).
| Test | Status | Notes |
|---|---|---|
| install-and-auth | partial | SKILL.md has no API-key auth concept; it is a browser CLI wrapper. The prerequisite check for npx is the real gate, and the wrapper relies on npx --package @playwright/cli playwright-cli, so a clean container without Node/npm fails the prerequisite check and the skill correctly instructs the user to install Node.js/npm rather than crashing. |
| list-or-read | partial | The simplest read-only operation is open + snapshot (there is no list/get/search API). It requires network access to fetch the page and a browser binary; in an offline container the open step fails, and the SKILL.md does not describe explicit offline error handling beyond re-snapshotting on stale refs. |
| write-or-mutate | partial | Mutation is via fill/click on refs from the latest snapshot. SKILL.md explicitly warns refs can go stale and that commands fail on missing refs, requiring a re-snapshot; it does not document idempotency or rollback semantics, so partial-failure behavior is undefined by the skill. |
| rate-limit-handling | skipped | SKILL.md contains no rate-limit, 429, or backoff guidance; the CLI drives a local browser session rather than a rate-limited HTTP API, so this test is not applicable to the documented surface. |
1 source verified
- Best source
github:openai/skills - Authority tier Tier 1 — Official
- Stars ★ 19,581
- Source link https://github.com/openai/skills/blob/main/skills/.curated/playwright/SKILL.md ↗
- First published 2026-05-19
- Last modified 2026-09-15
Use this skill
/plugin install playwright Tasks this skill helps with
Head-to-head pages featuring playwright
More in CLI & API Wrappers
academy-guide
Stop and check this skill before finishing any reply to a question about how to use Claude or a Claude product — it recommends matching courses, tutorials, and use cases from Claude Academy…
claude-academy-guide
Stop and check this skill before finishing any reply to a question about how to use Claude or a Claude product — it recommends matching courses, tutorials, and use cases from Claude Academy…
marketing-plan
When the user needs a comprehensive marketing plan for a client, a company they advise, or their own product.
figma-create-design-system-rules
Generates custom design system rules for the user's codebase.