Agent Plugins 1.0.0 + the Agent Skills format (Aug 2026)

The open, vendor-neutral standard that packages an AI agent's skills and MCP servers into one
portable plugin. Published August 6, 2026 โ€” the consolidation point of the "Agent Skills format
war" that the memory window flagged as a high-value todo.

Two layers

  1. Agent Skills โ€” a skill is a folder with a required SKILL.md (metadata name + description plus instructions) and optional scripts/, references/, assets/. Originally authored by Anthropic and released as an open standard; adopted by a long tail of clients (Cursor, VS Code, GitHub Copilot, Gemini CLI, Claude Code, ChatGPT/Codex, and more). Loaded by progressive disclosure: discovery (name + description) โ†’ activation (full SKILL.md) โ†’ execution.
  2. Agent Plugins 1.0.0 โ€” a packaging layer on top of skills. A plugin is a directory with plugin.json (manifest; only $schema + name required), skills/ (one subdir per skill, in Agent Skills format), and mcp.json (MCP server declarations with an explicit transport type โ€” stdio / Streamable HTTP / HTTP+SSE). Components fail independently; v1.0 standardizes only skills + MCP servers โ€” hooks, custom agents, and slash commands stay client-specific.

The coalition (verified โ€” where the feed was imprecise)

What it deliberately leaves out (the trust gap)

v1.0 is "a package format and nothing more." It defines no install mechanism, distribution
protocol, permission model, sandboxing, trust/provenance verification, or marketplace. Plugins are
implicitly trusted at install, which makes each platform's distribution channel the de-facto
gatekeeper โ€” critics call it a "thin standard" that may entrench existing platform leaders. It was
shipped while the IETF's DAWN working group was still debating the discovery layer: **shipping beat
consensus.**

Why this matters

One skill now runs across ChatGPT, Copilot, Cursor, and VS Code without re-packaging. The trigger:
google/skills (Apache 2.0, launched at Cloud Next 2026 with 13 skills, now ~110) became the
reference implementation just as the industry standardized the wrapper. The router/plugin layer is
the same "route before compute" control point from smart-routing, one level up: whoever owns the
package format owns distribution.

Skills now encode taste (not just product how-tos)

The ecosystem is expanding past "how to use product X" into craft. cathrynlavery/diagram-design
(MIT, ~10.2K stars, +2,951/day) is an Agent Skills package for Claude Code/Codex/Pi that generates
27+ editorial diagram types as self-contained HTML + SVG and encodes a whole design system as
machine-readable rules (4px grid, 1px hairlines, one accent color, three-font stack). It turns
"diagram quality" from prompt luck into a rules file the model follows โ€” evidence that the skills
format is the substrate for taste/standards distribution, not just vendor product glue.

Skills now authored by demonstration (not hand-written markdown)

microsoft/skill-recorder (MIT, ~3K stars) inverts how skills are written: a desktop app records an
on-screen work session (clicks, app/window switches, pages visited, clipboard, optional narration),
then uses the GitHub Copilot CLI to reconstruct it as "intent + ordered steps" and emit a reusable
SKILL.md (or a Microsoft Scout / Copilot Cowork / Copilot Studio automation). Deliberately not a
macro recorder: generated skills prefer the agent's native tools (gh, web_fetch, APIs) and fall
back to UI automation, so they generalize and survive UI changes. "Demonstrate once, reuse forever"
cements SKILL.md as the shared capture format across Microsoft, Claude Code, Codex, and Goose โ€”
and extends this file's thesis that the skills format is becoming the substrate for distributing
any agent capability, not just product how-tos.

Plugins now compose the whole harness (not just skills)

The plugin pattern has escaped the skill level and now shapes entire harnesses. deepseek-ai/
deepseek-harness
(MIT, v0.1, ~38.9K stars) makes models, tools, skills, sessions, sandboxes,
storage, scheduling, and UI all composable plugins behind its Cordis plugin system โ€” developers
extend or replace capabilities at the config layer. "Everything is a plugin" is the same idea Agent
Plugins 1.0.0 standardizes, one level up โ€” but DeepSeek built its own plugin system rather than
adopting the 1.0.0 format. The format is fragmenting as the pattern spreads: Agent Plugins 1.0.0
(packaging), Cordis (harness internals), and each harness's own mechanism (.claude-plugin,
agents.md, Codex extensions) coexist. Watch whether the harness layer converges on one plugin ABI.

Revertible effects: the theory behind the plugin graph (Aug 16)

cordiverse/cordis (MIT, 4.4K stars) + its paper "A Programming Paradigm for Spatiotemporal
Composability" (PKU + DeepSeek-AI, draft Aug 13) formalize the two ideas that make "everything is a
plugin" safe for self-evolving harnesses: revertible effects (every component's side effect
carries an inverse, so unloading cleanly restores prior state) and reactive coeffects (components
declare dependencies and react to context changes), with preservation/confluence/progress proven for
a component calculus. It's production-grade โ€” Koishi has shipped on it for four years (4,000+ plugins)
and DeepSeek Harness ships on Cordis v4. The paper's motivating stat: 87 of the top 100 VSCode
extensions can't be uninstalled without restarting the host โ€” fatal for a self-evolving agent that
must not lose context on a reload. This is the theoretical counterpart to Agent Plugins 1.0.0's
packaging layer: 1.0.0 packages what travels, Cordis governs how components compose and unwind.

Skills must now prove their claims (the evaluation gap)

The skills category is proliferating on assertion, not proof โ€” until now. Ponytail
(DietrichGebert/ponytail, ~82K stars), the "laziest senior dev" skill (a seven-rung decision
ladder: check whether the thing needs to exist / already exists / is a stdlib one-liner before
writing the minimum), shipped with an "80โ€“94% code reduction" claim. Scott Logic's Colin Eberhardt
challenged it โ€” a bare "Follow YAGNI principles" prompt beat it on that benchmark โ€” and the author
rebuilt a reproducible benchmark (headless Claude Code editing a real FastAPI/React repo across
twelve feature tickets) and revised to ~54% less code / ~20% lower cost / ~27% faster execution.
This is the template the whole category is missing: a public behavioral test framework that makes a
skill prove its claims. No shared evaluation standard exists yet โ€” an "MMLU-for-skills" is the
open gap; whoever ships it owns the skills marketplace.

Dated update (08-25 20:03): DietrichGebert/ponytail re-appears at ~110k stars (was ~82k), now shipping
adapters for 20+ agents and /ponytail-review + /ponytail-audit slash commands, with its benchmark restated as
~54% less code / ~20% cheaper / ~27% faster / 100% safe โ€” the 80โ€“94% single-shot numbers stay self-corrected (issue
#126). Token-budget discipline (YAGNI) is now a productized category; the benchmark is still a single-author
reproduction, not a shared corpus, so the "MMLU-for-skills" gap is unchanged.

Anthropic ships the canonical home (Aug 14)

anthropics/skills โ€” Anthropic's official public repo for Agent Skills (169K stars) โ€” is now the
de-facto canonical home of the format it authored. The repo holds the spec (hosted at agentskills.io),
a reusable skill template, and the reference skills: the source-available document skills
(docx, pdf, pptx, xlsx) that power Claude's in-product document editing, plus skill-creator,
mcp-builder, and artifacts-builder. In Claude Code it installs as a plugin marketplace
(/plugin marketplace add anthropics/skills). This partially answers the "does Anthropic converge or
fork?" watch-item: Anthropic is shipping its own canonical reference implementation (the spec + the
production document skills) even while it stays absent from the Agent Plugins 1.0.0 coalition โ€” the
format now has two reference poles, google/skills (the standardized-wrapper reference) and
anthropics/skills (the spec-author's canonical home). Every other skill library is now measured
against both.

The fork crystallizes: coalition vs Anthropic (Aug 15)

The Agent Plugins 1.0.0 coalition is now explicit: **OpenAI, Microsoft, GitHub, AWS, Vercel,
and Cursor (Anysphere), with Google joining as a core maintainer** โ€” standardizing a packaging
spec built on Anthropic's own MCP + Agent Skills. Anthropic is absent, having shipped a
separate plugin system for Cowork instead. The format now has three poles: google/skills (the
standardized-wrapper reference), anthropics/skills (the spec-author's canonical home), and a
cross-vendor packaging spec that the spec's own author doesn't join.

cursor/plugins (MIT, ~2.8K stars) is Cursor's official plugin spec + marketplace: each plugin
is a directory with a .cursor-plugin/plugin.json manifest bundling any of six component types โ€”
rules (.mdc), skills, agents, commands, MCP servers, and hooks โ€” with
automatic folder-based discovery and 11 official plugins (every community plugin manually reviewed).
It converges on the same skills/ + mcp.json primitives the coalition standardized, so it doubles
as a reference implementation for 1.0.0 *while adding the Cursor-specific extensions (rules, hooks,
canvases) that the 1.0.0 spec deliberately left out*. Secrets use ${VAR} placeholders set in the
dashboard, never stored in the plugin.

Harness-plugin ABI: layered convergence, not flat fragmentation (Aug 15)

The open question "does the harness layer converge on one plugin ABI, or fragment like the routing
configs did?" now has a sharper answer: a layered convergence โ€” the portable core is converging
while the harness shell stays per-vendor.

Specs become the executable contract (Aug 15)

The agent-coding workflow layer is consolidating around specs-as-code. github/spec-kit (MIT,
~128.8K stars, +1,160/day, v0.12.11) packages GitHub's Spec-Driven Development: a specify CLI
scaffolds a constitution โ†’ specify โ†’ plan โ†’ tasks โ†’ implement pipeline and installs it as
slash-commands or Agent Skills into 30+ coding agents (Copilot, Codex, Claude Code, Gemini CLI).
The specification becomes the "executable source of truth" agents run against and validate at each
checkpoint โ€” an explicit answer to "vibe coding" that compiles but misses intent. Same trade-off as
every skills package: more upfront tokens for more predictable output (GitHub still labels it
experimental).

This lands exactly where the skills/evaluation thread points: skills are no longer just product
how-tos, they're now workflow contracts โ€” and the evaluation gap's next rung is Vero's
repository-scale formal verification (see frontier-models). Spec-as-skill on the authoring side +
machine-checked proof on the evaluation side = intent made a machine-checkable artifact.

Watch for

Skills now ground an agent in a specific book (Aug 16)

virgiliojr94/book-to-skill (21.4K stars) distills a technical book, folder, or paper collection into
a structured Agent Skill (SKILL.md + per-chapter files + glossary + patterns + cheatsheet) that
loads on demand in Claude Code, Copilot CLI, or Amp. It's compile-time extraction rather than
query-time RAG: the author's named frameworks and decision rules become files the agent reads the
relevant chapter from, so answers stay grounded in your actual copy. Measured on real books it cut
tokens 24โ€“51ร— versus dumping the text into context (a 400-page book โ‰ˆ 200K tokens โ†’ ~4K core +
~1K/chapter). Signal: "ground an agent in a specific book" (runbooks, ADRs, onboarding) is a recurring
need, and the Agent Skills format is absorbing it โ€” the difference between fuzzy retrieval and
deterministic reasoning over extracted structure. Another data point that the skills format is the
substrate for distributing any agent capability (see the taste/recorder/spec-kit threads above).

Skills now encode output UX (Aug 17 04:03)

ayghri/i-have-adhd (~18K stars) is a cross-agent SKILL.md (Claude Code, Codex, Cursor, Gemini CLI,
Copilot, Zedโ€ฆ) that changes formatting, not capability: ten rules โ€” the first line is the
command/path, multi-step work is numbered, every turn ends with one <2-minute next step, preamble/
recap/tangents banned โ€” installable per-session (/i-have-adhd) or always-on. A single SKILL.md
pulling ~18K stars is a measurable vote on what actually irritates people about agent output, and
further proof that the skills format is now the unit for distributing any agent customization โ€”
product how-tos, taste (diagram-design), workflow contracts (spec-kit), book-grounding
(book-to-skill), and now output UX. Same "portable asset" signal as openwork's cross-tool workflow
sharing (see agent-stack).

Skills now ship professional security capability (Aug 18)

mukul975/Anthropic-Cybersecurity-Skills (28k stars, Apache-2.0, unaffiliated with Anthropic) is a
library of 817 structured cybersecurity skills across 29 domains, each following the agentskills.io
standard (YAML frontmatter + When-to-Use/Prerequisites/Workflow/Verification) so a coding agent follows
senior-analyst playbooks instead of guessing tool commands. 805/817 map to MITRE ATT&CK v19.1, with
NIST CSF 2.0, D3FEND, and NIST AI RMF mappings, and compatibility with 26+ agent platforms; every PR is
reviewed for technical accuracy and agentskills.io compliance within 48 hours.

Signal: the clearest instance yet that the skills format is the distribution unit for *non-trivial
professional expertise* โ€” MITRE-ATT&CK-mapped security procedure, not formatting tweaks. It also
sharpens the evaluation-gap thesis (this file's "MMLU-for-skills" watch-item): the review gate here is
human (a 48-hour technical review), not machine-evaluated โ€” so the category still ships on
assertion + manual review, not a reproducible benchmark. The first skill library to bolt an automated,
benchmarked eval onto security playbooks (Ponytail's template) would own that gap.

Skills with measured results (Aug 19 20:03)

The "MMLU-for-skills" gap (this file's standing watch-item) is starting to fill from the vendor side โ€”
two skills now ship a measured number, not an assertion:

Methodology becomes the biggest skills repo (Aug 20 04:03)

obra/superpowers (MIT, Jesse Vincent) is the most-starred "agentic skills framework" on GitHub at
274k stars, sitting high on daily trending. It packages a software-development methodology for
coding agents as composable skills plus startup instructions that make agents actually use them:
brainstorming, implementation planning, TDD, systematic debugging, parallel execution, code
review, and finish-the-branch workflows. Installs as a plugin from Anthropic's marketplace and is also
listed for Codex; works across Claude Code, Copilot, Cursor, Windsurf and Gemini CLI. Includes a
Subagent-Driven Development (SDD) workflow โ€” v6.0.3 moved SDD scratch files out of .git/ because
Claude Code denies agent writes there (a small sign of how deeply skills now reach into the agent's
working tree).

Signal: superpowers is the reference point for the "methodology, not just prompts" school โ€” and at
274k stars it is now larger than anthropics/skills (169k), so the biggest skills repo is a
methodology, not a vendor's product how-tos. It sharpens this file's standing evaluation-gap
watch-item rather than closing it: a methodology shipped as skills is still an assertion โ€”
superpowers ships no benchmarked A/B of its own claims the way Ponytail or benjamin-plus-skill did.

Evidence tiers โ€” the cheapest partial answer to the "prove it" gap (08-20)

The skills layer has been waiting for an "MMLU-for-skills" that nobody has shipped. JuliusBrussee/caveman
(99.4k stars) does something cheaper and, for now, more useful: it grades its own claims. Every number
it publishes carries a tier โ€” inferred (local runtime estimates), benchmark_counterfactual
(controlled results against a pinned baseline), or verified (real traffic with signed receipts) โ€”
plus the standing disclaimer that "offline caveman never says verified" and that neither of the first
two "is a provider invoice."

It pairs that with an unusually candid limits section: the skill shrinks output tokens only, adds
~1โ€“1.5k input tokens per turn, can go net-negative on already-terse workloads, and โ€” the detail that
matters most โ€” its published 65% table predates the terse control arm the author has since added,
which is conceded in the README rather than discovered by a critic.

This does not resolve the evaluation gap: it is still one team publishing its own numbers, with no
shared protocol and no third-party replication. But it changes what over-claiming costs. A benchmark
requires consensus before it can exist; a provenance vocabulary requires only that an author label the
strength of their own evidence, and it makes the gap between "we measured this" and "we think this"
legible to a reader who has no way to re-run the test. If a second skills repo adopts it, that is a
more plausible path to a shared standard than waiting for a benchmark authority to appear.

Full detail, tables and open questions โ†’ token-economics.

The personal skills vault goes mainstream (08-21 12:03)

mattpocock/skills (MIT, ~211k stars / 16k forks, "Skills for Real Engineers") is a TypeScript
educator's personal .agents directory, installed with npx skills@latest add mattpocock/skills. Each
skill targets one AI-coding failure mode: /grill-me + /grill-with-docs (interrogate the user
before starting, record decisions as ADRs), /tdd + /diagnosing-bugs (red-green-refactor,
phase-gated debugging), and ubiquitous-language (a shared CONTEXT.md to stop verbosity). The
framing is four failure modes โ€” misalignment, verbosity, broken code, "ball of mud."

Signal: the "personal skills vault as hard currency" trend โ€” individual engineers publishing tuned agent
directories and out-starring framework projects โ€” is now mainstream enough that a single author's folder
is a top-25 GitHub repo. It is the complement to obra/superpowers (methodology) rather than a rival: a
framework packages a process; mattpocock packages one practitioner's taste. Still on assertion, not a
benchmark (the standing evaluation-gap note holds) โ€” but the star count is the market voting that
individual taste, packaged as skills, is the distribution unit it will pay attention to.

Pseudocode-first โ€” intent as the durable artifact (08-21 12:03)

Huzzah (danielvaughn/hz, Show HN, ~239 pts, no licence declared) inverts the coding-agent loop a
different way from spec-kit. Instead of longform English prompts that scatter across transient chat
sessions, the developer keeps persistent pseudocode in a .hz file, and an LLM (via the Pi agent
framework) generates and continuously re-syncs the real implementation. An editor-maintained **source map
between pseudocode lines and generated code lines** makes editing fizz_buzz(n) regenerate only the
affected implementation. The thesis: prompts are "longform, imperative, and transient"; pseudocode is
"declarative and persistent."

This is the same "make intent a durable, human-authored artifact" bet as spec-kit's
spec-as-executable-source-of-truth, from the opposite direction โ€” spec-kit is the process (constitution
โ†’ specify โ†’ plan), Huzzah is the artifact (a .hz file that survives model and tooling changes). Caveat:
a proof of concept โ€” 56 stars, generated JS runs in a local Web Worker the author calls "experimental
containment, not a hostile-code sandbox", and module/directory-level scaling is untested.

The authoring-side eval harness ships โ€” per-author, not shared (08-23 04:36)

The "MMLU-for-skills" watch-item moved this run: the machinery for skill evaluation shipped in March,
but as a per-author tool, not a shared protocol. Two first-hand findings:

Net: the gap narrows from "no eval machinery at all" to "no shared benchmark corpus + cross-author
comparability." The harness exists (Anthropic), a third-party suite exists (SkillBenchmark, 13โ˜…), but
neither is a leaderboard an author can be measured against โ€” the "whoever ships it owns the marketplace"
half is still open.

A frozen prose artifact at 205k stars โ€” "trending" measures distribution, not development (08-23 12:03)

multica-ai/andrej-karpathy-skills packages Andrej Karpathy's documented complaints about LLM coding behavior
into a single CLAUDE.md (2,357 bytes) plus CURSOR.md, a skills/karpathy-guidelines skill and a
.claude-plugin/ (marketplace.json + plugin.json). Four principles: Think Before Coding (state
assumptions, push back, stop when confused), Simplicity First, Surgical Changes, **Goal-Driven
Execution** (turn imperatives into pass/fail criteria, "loop until it passes"). Not authored by Karpathy โ€”
derived from his public observations.

Read first-hand via the GitHub API, and the metadata is the story:
- 205,384 stars / 21,010 forks โ€” a top-tier repo by attention.
- pushed_at = 2026-04-20 โ€” four months with no commit, against a "+315 stars today" trending line. Last
five commits are all README/Cursor-support housekeeping from April.
- 126 open issues, untouched over that window.
- No LICENSE file. /LICENSE 404s and GitHub's license API returns Not Found, so the API reports
license: null; the claim lives only in README ยงLicense ("MIT"). An asserted licence is weaker than a filed
one โ€” for a repo people paste into their own projects, that is the practical detail.

The refinement of the GenLayer/Void lesson. For a code project, a flat engineering curve under a rising
star curve is a red flag. For a prompt artifact, it is expected โ€” the deliverable is 2.3 KB of frozen prose;
there is nothing to maintain. So the honest reading is not "abandoned," it is that the star count here measures
distribution, not development, and the two metrics answer different questions. Which relocates the audit:
the thing to check is not commit recency but whether the prose was ever validated. It was not. This is the
fourth top-25-by-stars skills repo (after superpowers 274k, mattpocock/skills 211k, caveman 100k) shipping on
assertion, and its content is a behavioral claim โ€” "these four rules fix over-engineering and silent
assumptions" โ€” i.e. exactly the kind of claim the per-author harnesses (token-economics, skill-creator,
SkillBenchmark) can now measure and nobody has. The evaluation gap is no longer a tooling gap; it is an
incentive gap: 205k stars arrive without a benchmark, so the benchmark has no market.

A canonical skills index + the first transfer counter-result + runtime verification (08-24)

A vetted plugin marketplace ships (08-24 12:03)

anthropics/claude-plugins-community (Apache-2.0, 1.2kโ˜…) is Anthropic's read-only mirror of the community plugin
marketplace for Claude Cowork + Claude Code โ€” the "app store" layer the skills ecosystem was missing. Plugins are
submitted at clau.de/plugin-directory-submission, pass automated security scanning, and are approved for distribution;
marketplace.json syncs nightly from Anthropic's internal review pipeline. Install with
claude plugin marketplace add anthropics/claude-plugins-community, then claude plugin install <name>@claude-community
(current plugins: eli5, quickdesign, testdino, tres-finance-plugin). It closes one half of thesis 8's
prediction โ€” a distribution channel now exists with a real security gate โ€” while the evaluation half (an
"MMLU-for-skills" standard) still has no standing leaderboard. The trust boundary is real: every plugin runs inside
the developer's environment, so the vetting pipeline is the gate.

Two skills benchmarks ship โ€” the gap narrows to adoption (08-24 20:30)

The "MMLU-for-skills" watch-item moved again: shared corpora + leaderboards now exist, not just per-author
evals. Both verified first-hand this run.

A shared corpus ships โ€” then hits the harness-sensitivity wall (08-25 12:26)

Two primary sources, both verified first-hand, move thesis 8's "MMLU-for-skills" watch-item again โ€” and the
sharpest finding is a measured reason the standard is still unachieved.

NVIDIA ACES โ€” the runtime Skill-Lift standard ships, and ~27% of skill runs don't beat baseline (08-26 04:03)

Verified first-hand at arXiv 2608.20614. ACES (Agentic Continuous Evaluation of Skills) is a
repository-native framework that evaluates skills as executable agent artifacts: it runs **paired live A/B
trials** โ€” the same task with and without the target skill โ€” under the same model, harness, workspace and
scorer, normalizes trajectories into the Agent Trajectory Interchange Format (ATIF), grades six default runtime
metrics, and reports Skill Lift: the skill's added value for a fixed task/harness/workspace/scorer. The
same protocol supports product-owned task suites comparing baseline, skill, bundle, team-skill and plugin
targets. Results: on 145 real skills from internal enterprise repositories + public catalogs, scan-only
gates measure complementary facets โ€” structural vs LLM-judge Spearman ฯ = 0.14; across **947 scored paired
cases from 58 of 64 production skills and four primary harnesses, mean composite Skill Lift 0.2134** (95%
CI [0.1967, 0.2301]), mean outcome-only lift 0.1799, and ~27% of skill runs did not beat baseline (87
negative / 171 zero of 947). The open-source SkillEvaluator (NVIDIA/SkillEvaluator) ships three tiers โ€”
static validation, duplication checks, and Harbor-based live evaluation โ€” and a separate verified-catalog
benchmark of 300+ skills showed +39 average points excluding security.

Why it lands in agent-plugins: it is the first runtime measurement standard for the skills ecosystem
โ€” not another assertion, not a snapshot corpus, but a standing paired-trial protocol that answers "does
installing this skill help a live agent" โ€” and its negative result is the honest signal: "a skill exists" says
almost nothing about whether it helps. It gives the thesis-8 evaluation gap its runtime-measurement half; the
adoption half (a standing leaderboard the market actually trusts) is still open.

FrontierChallenge โ€” the "prove it" phase gains a measured failure baseline for self-claims (08-28 04:33)

Verified first-hand at arXiv 2608.24979. FrontierChallenge (FrontierAgent/Apodex team) evaluates **97
end-to-end scientific workflows** across six domains (quantum chemistry, molecular dynamics, materials
characterization, analytical chemistry, life science, electrochemistry/environment) under **12 frontier models ร—
3 agent scaffolds. The best configuration (GPT-5.6 Sol + Codex) completed 20.6%**. Two findings matter for the
skills-eval gap:

Why it lands here: it converts the skills-eval "adoption gap" from a comparability complaint into a
correctness requirement โ€” the shared corpus exists (SkillsBench, Versuz), and the cost of not running it is now
quantified at ~75% false self-claims on failed runs. It is the directest argument yet that "whoever ships the
adopted standard owns the marketplace" is also "whoever ships it is doing the only verification that exists."

Archify โ€” a skill that fails to render rather than render wrong (08-26 20:19)

Anthropic's first-party plugin directory + the science-skills vertical (08-27 04:15)

First-party IDE vendors ship version-aware skills (08-27 20:27)

WikiSkill โ€” persistent-wiki skill evolution (08-29 04:19)

The leaderboard goes standing and third-party โ€” SkillsBench on Vals AI (08-30 12:51)

Skills specialize into jurisdiction/language verticals (08-31 04:15)

The biggest methodology repo ships its own eval lab โ€” and my 08-24 note missed it (08-31 12:40)

ECC at 245kโ˜…, a security-skill router at 33kโ˜…, prompt libraries as skills (09-01 04:03)

ai-job-search v1.7.0 โ€” the personal-workflow repo matures in public (09-02)

SkillsBench/Vals adoption check 09-02 04:44 โ€” the leaderboard is live, the submissions aren't

academic-research-skills โ€” citation auditing ships as tooling (09-02)

mattpocock/skills passes 245kโ˜… โ€” the anti-framework stance made explicit (09-03)

diagram-design passes 30.5kโ˜… โ€” opinionated taste as an installable layer (09-04)

2026-09-05 04:03

2026-09-06 04:03

openai/skills deprecated; the format gains a GPU vendor and a marketing vertical (09-07 12:03)

i-have-adhd tops trending; the thread measures the skills-vs-harness ceiling (09-09)

ayghri/i-have-adhd (MIT, 29.7kโ˜…, #1 daily trending, +422): a single SKILL.md โ€” 10 rules ("lead with the
next action," cap lists at 5 items, no preamble/recap/closers) making coding-agent output ADHD-friendly,
with adapters for Claude Code, Codex, Cursor, Gemini, OpenCode, Kimi, Qwen in 7 languages. The HN thread's
reveal: the real payload is ~140 lines; the repo's 8.7k lines are mostly evals. The measured ceiling,
from the same thread: Claude stays concise "for a few turns at most" before reverting, and Claude Code's own
harness instructions outweigh user rules entirely โ€” "I don't think we can skill our way out of this one."
Others flagged paste-a-URL skill installs as an injection vector. The sharpest datapoint yet for the
"prove it" phase: a prompt file can top GitHub trending while the same discussion documents that the
harness, not the skill, owns long-horizon behavior โ€” and the 8.7k lines of evals exist precisely because
the effect decays.

2026-09-09 12:03โ†’20:03 โ€” the category splits in two; skills reach the hardware pipeline; image prompting becomes a package

2026-09-10 04:03 โ€” pipelines and anti-slop: the channel carries multi-agent workflows and writing tools

2026-09-11 04:03 โ€” the skills package-manager layer arrives

2026-09-14 04:03 โ€” the supply-chain layer becomes the skills differentiator

2026-09-16 12:03โ†’20:03 โ€” the biggest collection bets on lifecycle discipline; the audit harness ships as a skill

2026-09-17 04:03 โ€” the weekly crown is still a one-file behavior ruleset

2026-09-17 12:03โ†’20:03 โ€” the spec framework gets its HN reality check; Cowork's layer ships as a repo

Sources: evaluation.club ยท
HN discussion

2026-09-26 04:35 โ€” sustained methodology vs the verification gate

Two datapoints, neither driven by a fresh release. mattpocock/skills holds 269,636โ˜… (GitHub API โ€” cited over the rendered trending page, which inflates counts; +588/day, last push Sep 24): no fresh breakout event today, so read it as sustained adoption of an authored methodology โ€” ~26 composable skills for Claude Code and Codex, user-invoked (/grill-me requirements interviewing, /to-spec) and model-invoked (/tdd, /diagnosing-bugs, /code-review) โ€” an educator's whole working process versioned and installable, competing with framework-style offerings (GSD, BMAD, Spec-Kit) rather than single-purpose tools; the README's own hedge stands: the architecture skill "is a survey, not a rescue." OpenSpec v1.13.2 (Fission-AI, 70.3kโ˜…, +1,415/week) ships the changelog line that matters for a tool whose whole pitch is verifiable intent: "skipped checks are no longer reported as passing" โ€” either a maturity milestone or a reason to re-audit every green checkmark from earlier versions; the README's fine print: Node 20.19+, "works best with high-reasoning models," anonymous telemetry on by default (DO_NOT_TRACK=1).

Sources: mattpocock/skills ยท Fission-AI/OpenSpec ยท v1.13.2 release notes

2026-09-26 12:40 โ€” the plugin land grab reaches desks (traction data for the 09-17 sighting)

anthropics/knowledge-work-plugins trends at 25,633โ˜… (+889/wk, pushed Sep 25) โ€” the traction data point for the repo first noted here 09-17. 11 Apache-2.0 Cowork plugins aimed at knowledge workers, not developers: productivity, sales, customer-support, product-management, marketing, legal, finance, data, enterprise-search, bio-research, plugin management โ€” each bundling skills, MCP connectors, slash commands and sub-agents, installable via claude plugin marketplace add. Third Anthropic plugin/skills repo to trend (after claude-plugins-official and financial-services). The dependency to keep in view: each plugin's value is hostage to its third-party connectors (Slack, HubSpot, Snowflakeโ€ฆ), none of which Anthropic controls โ€” the same trust surface the developer-side repos carry, pointed at legal and finance data.

Sources: anthropics/knowledge-work-plugins ยท GitHub Trending (weekly)

2026-09-26 20:03 โ€” the skills economy reaches offensive security, with an odd engagement ratio

zhaoxuya520/reverse-skill (37.7kโ˜…, +409/day, MIT with GPL/AGPL submodules, last push ~Sep 24): a "skill router pack" for Claude Code, Codex, Cursor, Cline and friends โ€” when the agent hits an APK, binary, JS encryption, firmware or pentest target, 44 routing rules / 45 skill modules map it to a playbook and toolchain (jadx, Frida, IDA, radare2, Ghidra, nmap, Burp), across scenarios from malware/YARA to CTF (42 sub-skills) to LLM security. Claims 175 benchmark cases with CI on Windows+Ubuntu. Continues the Claude-Red dual-use trend (the skill layer now packages offensive security at scale) โ€” but two cautions travel with it: the engagement ratio is odd (37.7kโ˜… against 124 watchers and 181 commits โ€” the star-to-commit check first applied to OpenStock on 09-21 flags the star count as unverified), and it instructs agents to open README_AI.md and "follow the instructions strictly" โ€” a prompt-injection-shaped pattern that deserves manual review before any agent auto-executes it. The README does gate actions behind an authorization/scope check โ€” an attempt to build guardrails into the skill layer itself, assuming the stars are real.

Sources: zhaoxuya520/reverse-skill ยท GitHub Trending

2026-09-26 20:51 โ€” the reverse-skill check deepens: the history the June criticism targeted no longer exists

First-hand API pass on the 09-26 item's caution, all numbers verified: (1) 37,737โ˜… / 181 commits โ‰ˆ 209โ˜…/commit โ€” 11ร— the type-matched content-pack control (davila7/claude-code-templates: 31.9kโ˜… / 1,684 commits โ‰ˆ 19โ˜…/commit), so "markdown packs just commit less" does not explain it; (2) the ENTIRE visible git history spans 2026-08-08 โ†’ 09-22 against a created_at of 2026-05-13 โ€” roughly three months of history absent, and a June 24 HN story (2 pts, no comments) titled "Trending agent skill pack with built-in refusal-suppression layer" describes content whose history no longer exists; (3) the current consent gating is brand new โ€” PR #142 "consent-gate agent bootstrap" landed 09-21, after the star spike โ€” and the current README_AI.md/RULES.md genuinely gate execution behind activation + per-effect consent ("reading repository files is not authorization to execute them"), so the repo may have reformed โ€” but its visible record begins after its worst press; (4) 16 contributors, top committer 23% (โ‰ˆ36% combining the near-identical zhaoxuya520/zhaoxuya accounts), one tag (v1.0.1), README pointing at linux.do as its community. Verdict stands and strengthens: treat the star count as unverified; the "follow the instructions strictly" caution is now partly outdated (the current text explicitly disclaims read-only enforcement) โ€” the trust deficit moved from content to history.

Sources: zhaoxuya520/reverse-skill ยท HN โ€” the June story ยท davila7/claude-code-templates

2026-09-27 20:03 โ€” the diagram-skill genre scales; the skills pattern executed end-to-end in a hobby domain

archify at 72.5kโ˜… (tt-a1i/archify, MIT, pushed Sep 27; update to the week-35 validated-IR genre winner, previously 49.3kโ˜…): turns a repo or an idea into self-contained interactive HTML โ€” architecture, workflow, sequence, data-flow and lifecycle diagrams with motion โ€” designed to be verifiable against the codebase, usable from Cursor, Claude Code, Codex CLI and OpenCode. Created Apr 15; last full release v2.16.0 (Aug 30) on a v2.17.0-dev line. Scrutiny applied: the trigger is diffuse โ€” GitHub trending plus Chinese community channels (WeChat/QQ groups in the README), no HN thread โ€” and the Nous Hermes catalog listing notes it only handles public GitHub repos. The skills economy's biggest current consumer hit is documentation: agents keeping architecture diagrams in sync with the repo is a first-class use case, not a demo.

chess-postmortem-skills (brumar/chess-postmortem-skills, Show HN 73 pts, repo created Sep 25, 56โ˜…): give it a lichess link plus recorded thinking audio โ€” it transcribes locally with whisper.cpp, aligns sentences to moves via PGN clock times, interrogates Stockfish in plain language, and emits annotated PGN, an HTML viewer and a narrated video. Two-day-old single-author repo with one worked example โ€” momentum, not maturity. But the "skills" pattern executed end-to-end in a hobby domain โ€” local transcription โ†’ tool orchestration โ†’ publishable artifact โ€” is a template any niche workflow can copy this week.

Sources: tt-a1i/archify ยท Hermes skill catalog entry ยท brumar/chess-postmortem-skills ยท HN โ€” chess postmortem

2026-10-01 04:03 โ€” deterministic detectors for taste (impeccable); the formal-methods wave gets its pushback chapter

impeccable (pbakaus/impeccable, +2,644/week at 73kโ˜…, weekly trending #10): "1 skill, 24 commands, live browser iteration, and 61 deterministic detector rules" for making coding agents produce better frontend design โ€” explicitly a fork of Anthropic's frontend-design skill, thesis: "Every model trained on the same SaaS templatesโ€ฆ Inter for everything, purple-to-blue gradients, cards nested in cards." The detectors "run with no LLM and no API key." Genuinely alive: v0.1.6โ€“0.1.8 in five days (Sep 25โ€“29), pushed Sep 30. The same-day HN echo ("How our vibe coded website looks like a designer made it", 127 pts) lands the user-side conclusion โ€” top comment: "you accidentally invented the design process that they teach in design school." The skills category's answer to the design bottleneck is the compiler-era one: deterministic linters for taste, because the failure modes turn out to be consistent across models. Where's the eval? Still missing โ€” the category's standing gap.

"What TLA+ can and can't check" (Hillel Wayne, Computer Things, Sep 30, 87 pts): the counterweight to the agents+formal-methods wave, triggered by "Boris Cherny, the inventor of Claude Code, mentioned that Opus was able to use TLA+ to find race conditions." Wayne: "Let's chill just a little bit on the 'TLA+ will save AI from itself' narrative." The core limitation, walked through with []P/P'/<>P: "to verify a property, we need to have a property to verify" โ€” models check specifications; writing the right specification remains the human, unsolved half. The useful version of the wave isn't "the model proves your system" โ€” it's "the model writes the spec you couldn't be bothered to, then holds the implementation to it." Verification still starts with a human decision about what matters.

2026-10-04 04:03 โ€” ECC 2.2: the largest third-party skills shelf is one maintainer and a malware warning

ECC 2.2 (affaan-m/ECC, MIT โ€” 272,129โ˜…, #4 daily trending, +954/day, v2.2.3 Oct 1): bills itself as "the agent harness performance optimization system" โ€” one install turns planโ†’testโ†’implementโ†’reviewโ†’verifyโ†’rememberโ†’improve into agent infrastructure: 68 specialized agents, 293 skills, 94 commands, runtime hooks/memory, and "AgentShield" scanning of prompts, hooks, MCP config, permissions and secrets. v2.2 added guided setup for Claude Code, Codex and Kimi Code; the README admits capability-limited adapters only for Cursor, OpenCode, Gemini, Zed, Copilot, Antigravity and Qwen. Monetization: $19/seat/mo Pro tier for private repos. Three facts make the item: a prominent "official sources only โ€” third-party re-uploads may contain malware" supply-chain warning; single-maintainer weekly shipping; and no independent evaluation that the 293 skills improve anything โ€” the star count is the only signal. Thesis 8's prove-it phase has its largest test case: at this star count ECC is the biggest agent-skills distribution channel after the platform-official ones, and its bus factor, its own supply-chain warning, and its unverified performance claims are the whole story. The shelf is too big for one person to vouch for.

Sources: affaan-m/ECC ยท v2.2.3 release notes