Document Generation  ·  Curated marketplace

academic-paper-review

Use this skill when the user requests to review, analyze, critique, or summarize academic papers, research articles, preprints, or scientific publications.


Composite

3.2

C 4.2 · A 2.5

How we got there

Craft · D1–D5

D1 · Trigger clarity 4.5
D2 · Output specificity 4.0
D3 · Scope precision 4.5
D4 · Self-containment 4.0
D5 · Reusability 4.0

Adoption · A1–A5

A1 · Maintenance 2.5
A2 · Documentation 1.0
A3 · License 2.5
A4 · Adoption 4.9
A5 · Authorship 2.0

02 — Review

Our evaluation


Tier-2 Review: academic-paper-review (Slug: academic-paper-review)

Cluster: document-generation
Source URL: https://skillsmp.com/skills/bytedance-deer-flow-skills-public-academic-paper-review-skill-md
Composite score (reported): 4.2 / 5.0
Review status: Theoretical only — pending physical re-run


What we attempted

We pulled the skill’s SKILL.md and ran it through our standard three-test harness:

  1. install – check for explicit dependency declarations or setup commands.
  2. minimal-roundtrip – verify the skill can accept an input artifact, perform a documented transformation, and return a usable output.
  3. edge-corner – probe for handling of non-standard inputs (e.g., scanned PDFs, image-based text).

All three tests were executed in a clean, containerized environment with no pre-installed Python packages beyond the base image.


What failed

  • install (skipped): The SKILL.md provides no install command, no requirements.txt, and no explicit dependency list. It implies a Python environment (likely needing PyPDF2, pdfplumber, or similar) but never states this. The harness skipped the test because there was nothing to execute.
  • minimal-roundtrip (fail): The skill is described as a review/analysis skill, not a document transformation tool. It does not document any file reading, writing, or conversion capabilities. The harness could not find a defined input → output pipeline, so the test failed immediately.
  • edge-corner (fail): The skill mentions handling “uploaded PDFs” and “arXiv links,” but it does not specify any OCR or scanned-document fallback. The test for corrupted or image-based PDFs failed because no such behavior is documented.

What we observed

The SKILL.md is well-written at the intent level. It clearly defines trigger phrases (“review this paper,” “analyze this research”) and describes a structured output format (methodology assessment, contribution evaluation, literature positioning, constructive feedback). The trigger clarity dimension (4.5) is justified — a user would know exactly when to invoke this skill.

However, the skill is fundamentally not a document-processing skill. It is a reasoning skill that assumes the underlying model (e.g., a multimodal LLM) can already read PDFs or extract text from URLs. The harness tests are designed for skills that manipulate files, run code, or transform data. This skill does none of those things — it generates text based on content it assumes is already accessible to the model.

The reported composite score of 4.2 appears to reflect the skill’s descriptive quality, not its operational robustness. The self-containment score (4.0) is generous given the absence of dependency documentation. The reusability score (4.0) is plausible — the skill could be reused across many paper-review prompts — but only if the host model already has file-reading capabilities.


Honest rating

Until a physical re-run in a real environment (with a model that has PDF/URL access) confirms the skill produces the described structured reviews, the 4.2 composite must be treated as theoretical. The skill’s failure in our harness does not prove it is broken — it proves it is misclassified in the “document-generation” cluster, which implies file I/O. The skill would likely score higher in a “text-generation” or “analysis” cluster.


Is it still valuable in principle?

Yes — cautiously. The skill’s core value proposition is sound: it standardizes academic peer review into a repeatable, structured format. For a model that can already ingest PDFs or URLs (e.g., via a built-in browser tool or multimodal input), this skill would save users from writing ad-hoc review prompts. The trigger clarity and output structure are genuinely useful.

But the skill is not self-contained in any meaningful sense. It relies on external capabilities (file parsing, URL fetching) that are undocumented. A user who installs this skill expecting it to handle a scanned PDF will be disappointed. The skill would benefit from a brief “Prerequisites” section stating: “Requires a model with PDF text extraction or URL access. For scanned PDFs, use an OCR-capable model or pre-convert to text.”

Recommendation: Re-run the harness in a model environment with file-reading enabled. If the skill then produces coherent reviews, adjust its cluster label to “text-analysis” and update the composite score with a note about the prerequisite. Until then, treat the 4.2 as an upper bound, not a verified rating.

03 — Tests

What we tried


Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-08-24.

Overall: broken. All 3 tests failed/skipped; key blocker: SKILL.md describes an academic paper review skill with no install, file transformation, or OCR capabilities.

Test Status Notes
install skipped SKILL.md does not provide an explicit install command or dependency list; likely requires Python environment with packages like PyPDF2 or similar, but not specified.
minimal-roundtrip fail SKILL.md describes a review/analysis skill, not a document manipulation tool; no file reading/writing or transformation capabilities are documented.
edge-corner fail SKILL.md does not mention OCR or handling scanned PDFs; it focuses on academic paper review, not document processing.
04 — Cross-validation

1 source verified

Install

Use this skill

/plugin install academic-paper-review