Agent Infrastructure  ·  Curated marketplace

local-ai-agents

Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models.


Composite

4.3

C 4.3 · A 0.0

How we got there

Craft · D1–D5

D1 · Trigger clarity 4.5
D2 · Output specificity 4.0
D3 · Scope precision 4.5
D4 · Self-containment 4.5
D5 · Reusability 3.5

02 — Review

Our evaluation


Tier-2 Review: local-ai-agents

What we attempted: We pulled the local-ai-agents skill from the Microsoft AI Agents for Beginners translation set, scored it against our five dimensions, and ran it through the standard test harness (install, smoke-invocation). The goal was to validate whether the SKILL.md could be executed end-to-end on a clean developer workstation.

What failed: Both automated tests failed. The install step reported: "SKILL.md does not specify an install command; it is a conceptual guide, so installation cannot be verified." The smoke-invocation step reported: "No executable code or CLI is provided in SKILL.md; it only describes concepts and trade-offs, so no minimal invocation is possible." In plain terms, the skill file contains no commands, no scripts, no config snippets, and no runnable artifact. It is a structured essay about local-first agents, not an operational skill.

What we observed: The SKILL.md is well-written and pedagogically dense. It correctly explains SLMs, the OpenAI-compatible local endpoint, sandboxed tools, Chroma RAG, MCP servers, and hybrid routing. The trigger clarity (D1 = 4.5) is strong — the "USE FOR / DO NOT USE FOR" section is explicit and correctly scopes out cloud deployment and first-agent basics. Output specificity (D2 = 4.0) is decent conceptually, but only in the sense of "you will learn about X," not "you will receive file Y or command Z." Scope precision (D3 = 4.5) is excellent; the boundaries are crisp. Self-containment (D4 = 4.5) is high in the sense that the guide references only public Microsoft Learn material, but low in the operational sense — nothing is self-contained enough to run. Reusability (D3.5) is the weakest dimension because a conceptual guide cannot be dropped into a pipeline without manual translation into code.

Honest assessment of the composite score: The 4.3/5.0 rating is theoretical and provisional. It reflects the quality of the writing and the pedagogical structure, not the quality of the skill as an executable artifact. Until a maintainer adds a minimal install.sh (or Dockerfile) and a run.py / smoke_test.sh that exercises the local endpoint, the skill fails its core purpose: being a reusable, invocable unit. The score should not be treated as a pass for production use.

Failure mode specifics: The core issue is that SKILL.md conflates teaching with operating. It reads like a lesson plan, not a skill definition. A skill, by our convention, must include at least one entry point (CLI, script, or API call) and a declared dependency list. This file has neither. It also lacks any version pinning, environment variables, or expected input/output schemas. So while a human can read it and learn, an automated agent cannot invoke it, test it, or chain it with other skills.

Is the skill still valuable in principle? Yes — conditionally. The underlying topic (local-first, privacy-preserving AI agents) is genuinely useful for developers who want offline inference, low latency, or data sovereignty. The conceptual guidance on Foundry Local, Qwen function calling, and Chroma is accurate and current. If the maintainer rewrites it as a starter template — with a install.sh, a run_local_agent.py that loads a Qwen model via the OpenAI-compatible endpoint, and a test_smoke.sh that verifies a simple tool call — the score would be justified. As it stands, it is a high-quality blog post in skill clothing.

Final note: This review is based on simulated harness observations (D-048). A physical re-run on a real workstation may surface additional issues (e.g., missing GPU, Chroma version conflicts), but the primary blocker — no runnable content — is structural and will not resolve itself. We recommend the maintainer either (a) add executable scaffolding and re-submit, or (b) reclassify this as a reference guide rather than a skill. Until then, treat the 4.3 as aspirational, not earned.

03 — Tests

What we tried


Tests simulated against README claims; pending physical re-run in Docker harness. Ran 2026-08-23.

Overall: broken. 0 tests passed, 0 partial, 2 failed; key blocker: SKILL.md is a conceptual guide with no install or runnable commands.

Test Status Notes
install fail SKILL.md does not specify an install command; it is a conceptual guide, so installation cannot be verified.
smoke-invocation fail No executable code or CLI is provided in SKILL.md; it only describes concepts and trade-offs, so no minimal invocation is possible.
04 — Cross-validation

1 source verified

Install

Use this skill

/plugin install local-ai-agents