Agent Infrastructure · Curated marketplace
page-agent
Embed alibaba/page-agent into your own web application — a pure-JavaScript in-page GUI agent that ships as a single <script tag or npm package and lets end-users of your site drive the UI with…
Composite
C 4.2 · A 0.0
How we got there
Our evaluation
Tier-2 Review: page-agent
What we attempted
We tried to take page-agent from a documented capability to a verified one: install the package via npm, wire it into a minimal page, and run the canonical smoke invocation from the SKILL.md ("click login, fill username as John") against a local Ollama backend. The goal was to confirm three things — that the package resolves and installs, that the JS API surface is callable as documented, and that the natural-language example actually drives a DOM interaction end to end.
What failed
Both of the two test steps we could attempt came back partial, not passing. Zero tests passed cleanly.
The blocker is the same for both: the SKILL.md content we have is truncated, and the truncation cuts exactly where the operational detail lives. Specifically:
- Install path (partial). The skill states page-agent "ships as a single
<script>tag or npm package," which tells us npm is a supported route but not how to take it. There is no explicitnpm install <package>line, no version pin, and no registry confirmation in the available text. We could not verify that the package name resolves, because the package name is never stated in full. "alibaba/page-agent" appears as a repo reference, which is not the same as a published npm identifier. - Smoke invocation (partial). The only invocation guidance is the natural-language example. The actual JS surface — constructor name, init method, run/dispatch method, how you point it at an LLM backend — is not present in the truncated text. We cannot confirm a minimal call verbatim, so we cannot confirm the example runs.
This is not a case of the skill being wrong. It is a case of the skill being under-specified in the artifact we could read, and the missing pieces are precisely the ones a harness needs to execute anything. A natural-language example is a great demonstration and a poor test fixture when the programmatic entry point is undocumented.
What we observed
- The dependency list is coherent and self-consistent: node/npm, a browser with DOM (the agent runs client-side), and a pluggable LLM backend (local Ollama or cloud Qwen/OpenAI/OpenRouter). Nothing here is exotic, and the "no Python, no headless browser, no extension" framing is a genuine differentiator versus server-side automation stacks.
- The scope boundary in the SKILL.md is sharp and correct: it explicitly routes server-side browser automation elsewhere. That is exactly the kind of negative scoping that D3 rewards, and it reads as deliberate.
- The failure mode is documentation truncation, not a broken concept. We cannot distinguish "the full SKILL.md contains the install command and API surface" from "it does not," because we only received 4000 characters. That uncertainty is the honest ceiling on this review.
On the ratings
The composite of 4.2/5.0 and the per-dimension scores (D1 4.5, D2 4.0, D3 4.5, D4 4.0, D5 3.5) should be read as theoretical until a physical re-run resolves the two partials. D4 self-containment at 4.0 in particular is optimistic given that the install command and API surface are absent from the text we could inspect — if the truncation is representative, D4 is the dimension most likely to move down on re-test. D5 reusability at 3.5 already reflects that a skill you cannot invoke programmatically is hard to reuse in a pipeline.
Does the skill still seem valuable in principle?
Yes. The underlying capability — a pure-JS, in-page GUI agent that lets end-users of a SaaS or admin panel drive the UI in natural language, with a swappable local/cloud LLM backend — is genuinely useful and not well-served by server-side automation tools. The trigger clarity and scope precision are real strengths. The problem is packaging, not premise: the skill needs to expose the exact npm package name, an install command with a version, and the minimal JS call sequence before it can graduate from "promising" to "verified." Until then, treat the 4.2 as a strong hypothesis, not a measured result.
What we tried
Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-09-19.
Overall: broken. 0 tests passed, 2 partial, 0 failed; key blocker: SKILL.md is truncated and does not specify the exact npm package name, install command, or JS API surface, so both install and smoke-invocation can only be partially validated.
Inferred dependencies: node/npm (for npm package install path), modern browser with DOM (in-page agent runs client-side), LLM backend: local Ollama OR cloud Qwen/OpenAI/OpenRouter (per SKILL.md), API key for chosen cloud LLM provider (OpenAI/OpenRouter/Qwen).
| Test | Status | Notes |
|---|---|---|
| install | partial | SKILL.md describes page-agent as shipping 'as a single |