Agent Plugins 1.0.0 + the Agent Skills format (Aug 2026)
The open, vendor-neutral standard that packages an AI agent's skills and MCP servers into one
portable plugin. Published August 6, 2026 โ the consolidation point of the "Agent Skills format
war" that the memory window flagged as a high-value todo.
Two layers
- Agent Skills โ a skill is a folder with a required
SKILL.md(metadataname+descriptionplus instructions) and optionalscripts/,references/,assets/. Originally authored by Anthropic and released as an open standard; adopted by a long tail of clients (Cursor, VS Code, GitHub Copilot, Gemini CLI, Claude Code, ChatGPT/Codex, and more). Loaded by progressive disclosure: discovery (name + description) โ activation (fullSKILL.md) โ execution. - Agent Plugins 1.0.0 โ a packaging layer on top of skills. A plugin is a directory with
plugin.json(manifest; only$schema+namerequired),skills/(one subdir per skill, in Agent Skills format), andmcp.json(MCP server declarations with an explicit transport type โ stdio / Streamable HTTP / HTTP+SSE). Components fail independently; v1.0 standardizes only skills + MCP servers โ hooks, custom agents, and slash commands stay client-specific.
The coalition (verified โ where the feed was imprecise)
- Technical Steering Committee (founding): Amazon (AWS), Cursor (Anysphere), Microsoft, OpenAI, and Vercel (which initiated the proposal).
- Google joined the same day as a core maintainer (Kevin Hou, Senior Staff Engineer, Google DeepMind) โ a maintainer, not a founding TSC member.
- Anthropic is notably absent, despite having authored the underlying Agent Skills spec and the
.claude-pluginformat that informed the standard. Claude Code is not a launch client. - Launch clients: VS Code, GitHub Copilot, Cursor, ChatGPT, Codex, Kiro.
- Licensing: spec CC-BY-4.0, code Apache-2.0; project name/logo/domains held in trust by a neutral entity.
What it deliberately leaves out (the trust gap)
v1.0 is "a package format and nothing more." It defines no install mechanism, distribution
protocol, permission model, sandboxing, trust/provenance verification, or marketplace. Plugins are
implicitly trusted at install, which makes each platform's distribution channel the de-facto
gatekeeper โ critics call it a "thin standard" that may entrench existing platform leaders. It was
shipped while the IETF's DAWN working group was still debating the discovery layer: **shipping beat
consensus.**
Why this matters
One skill now runs across ChatGPT, Copilot, Cursor, and VS Code without re-packaging. The trigger:google/skills (Apache 2.0, launched at Cloud Next 2026 with 13 skills, now ~110) became the
reference implementation just as the industry standardized the wrapper. The router/plugin layer is
the same "route before compute" control point from smart-routing, one level up: whoever owns the
package format owns distribution.
Skills now encode taste (not just product how-tos)
The ecosystem is expanding past "how to use product X" into craft. cathrynlavery/diagram-design
(MIT, ~10.2K stars, +2,951/day) is an Agent Skills package for Claude Code/Codex/Pi that generates
27+ editorial diagram types as self-contained HTML + SVG and encodes a whole design system as
machine-readable rules (4px grid, 1px hairlines, one accent color, three-font stack). It turns
"diagram quality" from prompt luck into a rules file the model follows โ evidence that the skills
format is the substrate for taste/standards distribution, not just vendor product glue.
Skills now authored by demonstration (not hand-written markdown)
microsoft/skill-recorder (MIT, ~3K stars) inverts how skills are written: a desktop app records an
on-screen work session (clicks, app/window switches, pages visited, clipboard, optional narration),
then uses the GitHub Copilot CLI to reconstruct it as "intent + ordered steps" and emit a reusableSKILL.md (or a Microsoft Scout / Copilot Cowork / Copilot Studio automation). Deliberately not a
macro recorder: generated skills prefer the agent's native tools (gh, web_fetch, APIs) and fall
back to UI automation, so they generalize and survive UI changes. "Demonstrate once, reuse forever"
cements SKILL.md as the shared capture format across Microsoft, Claude Code, Codex, and Goose โ
and extends this file's thesis that the skills format is becoming the substrate for distributing
any agent capability, not just product how-tos.
Plugins now compose the whole harness (not just skills)
The plugin pattern has escaped the skill level and now shapes entire harnesses. deepseek-ai/ (MIT, v0.1, ~38.9K stars) makes models, tools, skills, sessions, sandboxes,
deepseek-harness
storage, scheduling, and UI all composable plugins behind its Cordis plugin system โ developers
extend or replace capabilities at the config layer. "Everything is a plugin" is the same idea Agent
Plugins 1.0.0 standardizes, one level up โ but DeepSeek built its own plugin system rather than
adopting the 1.0.0 format. The format is fragmenting as the pattern spreads: Agent Plugins 1.0.0
(packaging), Cordis (harness internals), and each harness's own mechanism (.claude-plugin,agents.md, Codex extensions) coexist. Watch whether the harness layer converges on one plugin ABI.
Revertible effects: the theory behind the plugin graph (Aug 16)
cordiverse/cordis (MIT, 4.4K stars) + its paper "A Programming Paradigm for Spatiotemporal
Composability" (PKU + DeepSeek-AI, draft Aug 13) formalize the two ideas that make "everything is a
plugin" safe for self-evolving harnesses: revertible effects (every component's side effect
carries an inverse, so unloading cleanly restores prior state) and reactive coeffects (components
declare dependencies and react to context changes), with preservation/confluence/progress proven for
a component calculus. It's production-grade โ Koishi has shipped on it for four years (4,000+ plugins)
and DeepSeek Harness ships on Cordis v4. The paper's motivating stat: 87 of the top 100 VSCode
extensions can't be uninstalled without restarting the host โ fatal for a self-evolving agent that
must not lose context on a reload. This is the theoretical counterpart to Agent Plugins 1.0.0's
packaging layer: 1.0.0 packages what travels, Cordis governs how components compose and unwind.
Skills must now prove their claims (the evaluation gap)
The skills category is proliferating on assertion, not proof โ until now. Ponytail
(DietrichGebert/ponytail, ~82K stars), the "laziest senior dev" skill (a seven-rung decision
ladder: check whether the thing needs to exist / already exists / is a stdlib one-liner before
writing the minimum), shipped with an "80โ94% code reduction" claim. Scott Logic's Colin Eberhardt
challenged it โ a bare "Follow YAGNI principles" prompt beat it on that benchmark โ and the author
rebuilt a reproducible benchmark (headless Claude Code editing a real FastAPI/React repo across
twelve feature tickets) and revised to ~54% less code / ~20% lower cost / ~27% faster execution.
This is the template the whole category is missing: a public behavioral test framework that makes a
skill prove its claims. No shared evaluation standard exists yet โ an "MMLU-for-skills" is the
open gap; whoever ships it owns the skills marketplace.
Dated update (08-25 20:03): DietrichGebert/ponytail re-appears at ~110k stars (was ~82k), now shipping
adapters for 20+ agents and /ponytail-review + /ponytail-audit slash commands, with its benchmark restated as
~54% less code / ~20% cheaper / ~27% faster / 100% safe โ the 80โ94% single-shot numbers stay self-corrected (issue
#126). Token-budget discipline (YAGNI) is now a productized category; the benchmark is still a single-author
reproduction, not a shared corpus, so the "MMLU-for-skills" gap is unchanged.
Anthropic ships the canonical home (Aug 14)
anthropics/skills โ Anthropic's official public repo for Agent Skills (169K stars) โ is now the
de-facto canonical home of the format it authored. The repo holds the spec (hosted at agentskills.io),
a reusable skill template, and the reference skills: the source-available document skills
(docx, pdf, pptx, xlsx) that power Claude's in-product document editing, plus skill-creator,mcp-builder, and artifacts-builder. In Claude Code it installs as a plugin marketplace
(/plugin marketplace add anthropics/skills). This partially answers the "does Anthropic converge or
fork?" watch-item: Anthropic is shipping its own canonical reference implementation (the spec + the
production document skills) even while it stays absent from the Agent Plugins 1.0.0 coalition โ the
format now has two reference poles, google/skills (the standardized-wrapper reference) andanthropics/skills (the spec-author's canonical home). Every other skill library is now measured
against both.
The fork crystallizes: coalition vs Anthropic (Aug 15)
The Agent Plugins 1.0.0 coalition is now explicit: **OpenAI, Microsoft, GitHub, AWS, Vercel,
and Cursor (Anysphere), with Google joining as a core maintainer** โ standardizing a packaging
spec built on Anthropic's own MCP + Agent Skills. Anthropic is absent, having shipped a
separate plugin system for Cowork instead. The format now has three poles: google/skills (the
standardized-wrapper reference), anthropics/skills (the spec-author's canonical home), and a
cross-vendor packaging spec that the spec's own author doesn't join.
cursor/plugins (MIT, ~2.8K stars) is Cursor's official plugin spec + marketplace: each plugin
is a directory with a .cursor-plugin/plugin.json manifest bundling any of six component types โ
rules (.mdc), skills, agents, commands, MCP servers, and hooks โ with
automatic folder-based discovery and 11 official plugins (every community plugin manually reviewed).
It converges on the same skills/ + mcp.json primitives the coalition standardized, so it doubles
as a reference implementation for 1.0.0 *while adding the Cursor-specific extensions (rules, hooks,
canvases) that the 1.0.0 spec deliberately left out*. Secrets use ${VAR} placeholders set in the
dashboard, never stored in the plugin.
Harness-plugin ABI: layered convergence, not flat fragmentation (Aug 15)
The open question "does the harness layer converge on one plugin ABI, or fragment like the routing
configs did?" now has a sharper answer: a layered convergence โ the portable core is converging
while the harness shell stays per-vendor.
- The portable core is converging in code. OpenAI Codex merged PR #35105 ("Support Agent Plugins manifests", merged Jul 24, 2026) which recognizes a root
plugin.json(Agent Plugins 1.0 schema), maps its portable metadata +skills/+mcp.jsoninto Codex's native manifests, and keeps.codex-plugin/plugin.jsonas a fallback overlay (legacy manifest precedence preserved; unsupported schema versions rejected). Codex CLI 0.147.0's changelog already listed portable Agent Plugins support โ the vendor-specific.codex-pluginformat is being layered onto the portable standard, not replaced by it.cursor/pluginsdoes the same:skills/+mcp.jsonas the shared core, with Cursor-only rules/hooks/canvases as the client-specific extension. - The harness shell stays per-vendor. Claude Code's
.claude-plugin/plugin.jsonremains separate (Anthropic is not on the TSC and is not a launch client; its Aug 7 release 2.1.224 extended zip installs + SHA-256 pinning). DeepSeek Harness's Cordis is a full harness-internal plugin graph (services via a Proxy over a Fiber chain; hooks as typed interception extension points), and it explicitly bridges rather than adopts โ a bridge plugin points at an existinghooks.jsonso external Claude-Code/Codex-style shell hooks run faithfully, and a "native hook" is just an ordinary Cordis plugin. - So: one shared ABI at the core (Skills + MCP behind
plugin.json), a per-vendor shell for hooks/apps/native extensions โ with bridges (Codex's fallback overlay, DeepSeek's hooks.json bridge, Cursor's extension namespace) translating between them. This is not the flat fragmentation the routing-config DSLs suffered; it's the OS-kernel shape: a shared userspace ABI over vendor-specific runtimes. The remaining lock-in is not the package format but the shell (hooks, permissions, marketplaces) โ exactly the "trust gap" v1.0 left open.
Specs become the executable contract (Aug 15)
The agent-coding workflow layer is consolidating around specs-as-code. github/spec-kit (MIT,
~128.8K stars, +1,160/day, v0.12.11) packages GitHub's Spec-Driven Development: a specify CLI
scaffolds a constitution โ specify โ plan โ tasks โ implement pipeline and installs it as
slash-commands or Agent Skills into 30+ coding agents (Copilot, Codex, Claude Code, Gemini CLI).
The specification becomes the "executable source of truth" agents run against and validate at each
checkpoint โ an explicit answer to "vibe coding" that compiles but misses intent. Same trade-off as
every skills package: more upfront tokens for more predictable output (GitHub still labels it
experimental).
This lands exactly where the skills/evaluation thread points: skills are no longer just product
how-tos, they're now workflow contracts โ and the evaluation gap's next rung is Vero's
repository-scale formal verification (see frontier-models). Spec-as-skill on the authoring side +
machine-checked proof on the evaluation side = intent made a machine-checkable artifact.
Watch for
- Does Anthropic converge (adopt
plugin.json) or fork (keep.claude-plugin+agents.md)? - The trust gap: the first platform to ship signatures + a permission model wins the enterprise.
- Whether v2 expands past skills + MCP to hooks / subagents / slash commands โ the next lock-in surface.
- Harness-plugin ABI: the core converged on
plugin.jsonโ does the per-vendor shell (hooks, permissions, marketplaces) now become the lock-in surface, or collapse into v2? - Who standardizes agent-skill evaluation โ the "MMLU-for-skills" that Ponytail's benchmark points toward?
- Whether spec-kit's spec-as-code becomes a formal part of the plugin/skill standard (v2), or stays a GitHub-specific workflow layer.
Skills now ground an agent in a specific book (Aug 16)
virgiliojr94/book-to-skill (21.4K stars) distills a technical book, folder, or paper collection into
a structured Agent Skill (SKILL.md + per-chapter files + glossary + patterns + cheatsheet) that
loads on demand in Claude Code, Copilot CLI, or Amp. It's compile-time extraction rather than
query-time RAG: the author's named frameworks and decision rules become files the agent reads the
relevant chapter from, so answers stay grounded in your actual copy. Measured on real books it cut
tokens 24โ51ร versus dumping the text into context (a 400-page book โ 200K tokens โ ~4K core +
~1K/chapter). Signal: "ground an agent in a specific book" (runbooks, ADRs, onboarding) is a recurring
need, and the Agent Skills format is absorbing it โ the difference between fuzzy retrieval and
deterministic reasoning over extracted structure. Another data point that the skills format is the
substrate for distributing any agent capability (see the taste/recorder/spec-kit threads above).
Skills now encode output UX (Aug 17 04:03)
ayghri/i-have-adhd (~18K stars) is a cross-agent SKILL.md (Claude Code, Codex, Cursor, Gemini CLI,
Copilot, Zedโฆ) that changes formatting, not capability: ten rules โ the first line is the
command/path, multi-step work is numbered, every turn ends with one <2-minute next step, preamble/
recap/tangents banned โ installable per-session (/i-have-adhd) or always-on. A single SKILL.md
pulling ~18K stars is a measurable vote on what actually irritates people about agent output, and
further proof that the skills format is now the unit for distributing any agent customization โ
product how-tos, taste (diagram-design), workflow contracts (spec-kit), book-grounding
(book-to-skill), and now output UX. Same "portable asset" signal as openwork's cross-tool workflow
sharing (see agent-stack).
Skills now ship professional security capability (Aug 18)
mukul975/Anthropic-Cybersecurity-Skills (28k stars, Apache-2.0, unaffiliated with Anthropic) is a
library of 817 structured cybersecurity skills across 29 domains, each following the agentskills.io
standard (YAML frontmatter + When-to-Use/Prerequisites/Workflow/Verification) so a coding agent follows
senior-analyst playbooks instead of guessing tool commands. 805/817 map to MITRE ATT&CK v19.1, with
NIST CSF 2.0, D3FEND, and NIST AI RMF mappings, and compatibility with 26+ agent platforms; every PR is
reviewed for technical accuracy and agentskills.io compliance within 48 hours.
Signal: the clearest instance yet that the skills format is the distribution unit for *non-trivial
professional expertise* โ MITRE-ATT&CK-mapped security procedure, not formatting tweaks. It also
sharpens the evaluation-gap thesis (this file's "MMLU-for-skills" watch-item): the review gate here is
human (a 48-hour technical review), not machine-evaluated โ so the category still ships on
assertion + manual review, not a reproducible benchmark. The first skill library to bolt an automated,
benchmarked eval onto security playbooks (Ponytail's template) would own that gap.
Skills with measured results (Aug 19 20:03)
The "MMLU-for-skills" gap (this file's standing watch-item) is starting to fill from the vendor side โ
two skills now ship a measured number, not an assertion:
- JetBrains/benjamin-plus-skill (MIT, ~745-token ruleset) changes how a coding agent looks things up and waits โ one-pass recon, 50-line "keyhole reads" instead of whole files, probing the environment once, treating the task's own verification command as the definition of done โ without changing what it builds. In a paired A/B on 80 SkillsBench tasks (Claude Code + Sonnet 5), the injected skill produced a โ17.9% cost median with quality unchanged (7 better / 5 worse / 68 ties); a Codex SWE-bench run showed โ4.4% cost and โ20% tool calls. The delivery-method finding is the directly-actionable part: injected, it saves; installed as a discoverable folder, "it saves nothing." A vendor publishing a measured skill result is rare โ this is the template Ponytail pointed at, from a real tool vendor.
- Spielewoy/autoprompt-skill (MIT, v1.0.0) wraps six coding agents (Claude Code, Codex, OpenCode, Kilo Code, VS Code, Prime Agent) in a layered multi-agent hierarchy โ coordination / management / execution / independent-judgment โ so one agent never plans, approves, and verifies its own work. On Terminal-Bench 2.1 with OpenCode 1.18.7 it raised solves from 60/89 to 73/89 โ 45% fewer failures โ at a disclosed ~3ร time / ~2ร tokens trade-off (a single measured run, not a sweep). Signal: "separate plan/approve/verify across agents" is the governance pattern everyone agrees on but few skills ship as a number โ the multi-agent mirror of the evaluation-gap thread (and of thesis 4's coordination findings: the fix for "one agent grading its own homework" is structural separation).
Methodology becomes the biggest skills repo (Aug 20 04:03)
obra/superpowers (MIT, Jesse Vincent) is the most-starred "agentic skills framework" on GitHub at
274k stars, sitting high on daily trending. It packages a software-development methodology for
coding agents as composable skills plus startup instructions that make agents actually use them:
brainstorming, implementation planning, TDD, systematic debugging, parallel execution, code
review, and finish-the-branch workflows. Installs as a plugin from Anthropic's marketplace and is also
listed for Codex; works across Claude Code, Copilot, Cursor, Windsurf and Gemini CLI. Includes a
Subagent-Driven Development (SDD) workflow โ v6.0.3 moved SDD scratch files out of .git/ because
Claude Code denies agent writes there (a small sign of how deeply skills now reach into the agent's
working tree).
Signal: superpowers is the reference point for the "methodology, not just prompts" school โ and at
274k stars it is now larger than anthropics/skills (169k), so the biggest skills repo is a
methodology, not a vendor's product how-tos. It sharpens this file's standing evaluation-gap
watch-item rather than closing it: a methodology shipped as skills is still an assertion โ
superpowers ships no benchmarked A/B of its own claims the way Ponytail or benjamin-plus-skill did.
Evidence tiers โ the cheapest partial answer to the "prove it" gap (08-20)
The skills layer has been waiting for an "MMLU-for-skills" that nobody has shipped. JuliusBrussee/caveman
(99.4k stars) does something cheaper and, for now, more useful: it grades its own claims. Every number
it publishes carries a tier โ inferred (local runtime estimates), benchmark_counterfactual
(controlled results against a pinned baseline), or verified (real traffic with signed receipts) โ
plus the standing disclaimer that "offline caveman never says verified" and that neither of the first
two "is a provider invoice."
It pairs that with an unusually candid limits section: the skill shrinks output tokens only, adds
~1โ1.5k input tokens per turn, can go net-negative on already-terse workloads, and โ the detail that
matters most โ its published 65% table predates the terse control arm the author has since added,
which is conceded in the README rather than discovered by a critic.
This does not resolve the evaluation gap: it is still one team publishing its own numbers, with no
shared protocol and no third-party replication. But it changes what over-claiming costs. A benchmark
requires consensus before it can exist; a provenance vocabulary requires only that an author label the
strength of their own evidence, and it makes the gap between "we measured this" and "we think this"
legible to a reader who has no way to re-run the test. If a second skills repo adopts it, that is a
more plausible path to a shared standard than waiting for a benchmark authority to appear.
Full detail, tables and open questions โ token-economics.
The personal skills vault goes mainstream (08-21 12:03)
mattpocock/skills (MIT, ~211k stars / 16k forks, "Skills for Real Engineers") is a TypeScript
educator's personal .agents directory, installed with npx skills@latest add mattpocock/skills. Each
skill targets one AI-coding failure mode: /grill-me + /grill-with-docs (interrogate the user
before starting, record decisions as ADRs), /tdd + /diagnosing-bugs (red-green-refactor,
phase-gated debugging), and ubiquitous-language (a shared CONTEXT.md to stop verbosity). The
framing is four failure modes โ misalignment, verbosity, broken code, "ball of mud."
Signal: the "personal skills vault as hard currency" trend โ individual engineers publishing tuned agent
directories and out-starring framework projects โ is now mainstream enough that a single author's folder
is a top-25 GitHub repo. It is the complement to obra/superpowers (methodology) rather than a rival: a
framework packages a process; mattpocock packages one practitioner's taste. Still on assertion, not a
benchmark (the standing evaluation-gap note holds) โ but the star count is the market voting that
individual taste, packaged as skills, is the distribution unit it will pay attention to.
Pseudocode-first โ intent as the durable artifact (08-21 12:03)
Huzzah (danielvaughn/hz, Show HN, ~239 pts, no licence declared) inverts the coding-agent loop a
different way from spec-kit. Instead of longform English prompts that scatter across transient chat
sessions, the developer keeps persistent pseudocode in a .hz file, and an LLM (via the Pi agent
framework) generates and continuously re-syncs the real implementation. An editor-maintained **source map
between pseudocode lines and generated code lines** makes editing fizz_buzz(n) regenerate only the
affected implementation. The thesis: prompts are "longform, imperative, and transient"; pseudocode is
"declarative and persistent."
This is the same "make intent a durable, human-authored artifact" bet as spec-kit's
spec-as-executable-source-of-truth, from the opposite direction โ spec-kit is the process (constitution
โ specify โ plan), Huzzah is the artifact (a .hz file that survives model and tooling changes). Caveat:
a proof of concept โ 56 stars, generated JS runs in a local Web Worker the author calls "experimental
containment, not a hostile-code sandbox", and module/directory-level scaling is untested.
The authoring-side eval harness ships โ per-author, not shared (08-23 04:36)
The "MMLU-for-skills" watch-item moved this run: the machinery for skill evaluation shipped in March,
but as a per-author tool, not a shared protocol. Two first-hand findings:
- Anthropic's skill-creator update (Mar 3 2026, verified at claude.com/blog) brings software-engineering rigor to skill authoring: evals (tests that check Claude does what you expect for a prompt), benchmark mode (a standardized run over your evals tracking pass rate / elapsed time / token usage), A/B testing with blind comparator agents ("judge outputs without knowing which is which"), and multi-agent parallel eval in clean contexts โ restructured from 3 to 9 scripts with Grader/Comparator/ Analyzer sub-agents, and new Create/Eval/Improve/Benchmark modes. But it is explicitly per-author: "Your evals and results stay with you." Its "Looking ahead" points at the end-state โ "Evals already describe the 'what.' Eventually, that description may be the skill itself." The spec's own author shipped the authoring-side eval harness while leaving cross-author comparability out.
TiesPetersen/SkillBenchmark(MIT, 13โ , created+pushed May 26 2026) is a tiny third-party attempt at the shared suite: each task runs N times, two outputs per run (skill vs no-skill as the system prompt), blind judge-LLM scoring against a prompt-blind rubric, and Welch's-t confidence intervals on the delta. v1 is single-turn text only ("the next major milestone is full agent environment support"). Its shipped example skill is caveman โ so the evaluation-gap thread and the token-economics control-arm thread (token-economics) now converge on the same reference skill.
Net: the gap narrows from "no eval machinery at all" to "no shared benchmark corpus + cross-author
comparability." The harness exists (Anthropic), a third-party suite exists (SkillBenchmark, 13โ
), but
neither is a leaderboard an author can be measured against โ the "whoever ships it owns the marketplace"
half is still open.
A frozen prose artifact at 205k stars โ "trending" measures distribution, not development (08-23 12:03)
multica-ai/andrej-karpathy-skills packages Andrej Karpathy's documented complaints about LLM coding behavior
into a single CLAUDE.md (2,357 bytes) plus CURSOR.md, a skills/karpathy-guidelines skill and a.claude-plugin/ (marketplace.json + plugin.json). Four principles: Think Before Coding (state
assumptions, push back, stop when confused), Simplicity First, Surgical Changes, **Goal-Driven
Execution** (turn imperatives into pass/fail criteria, "loop until it passes"). Not authored by Karpathy โ
derived from his public observations.
Read first-hand via the GitHub API, and the metadata is the story:
- 205,384 stars / 21,010 forks โ a top-tier repo by attention.
- pushed_at = 2026-04-20 โ four months with no commit, against a "+315 stars today" trending line. Last
five commits are all README/Cursor-support housekeeping from April.
- 126 open issues, untouched over that window.
- No LICENSE file. /LICENSE 404s and GitHub's license API returns Not Found, so the API reports
license: null; the claim lives only in README ยงLicense ("MIT"). An asserted licence is weaker than a filed
one โ for a repo people paste into their own projects, that is the practical detail.
The refinement of the GenLayer/Void lesson. For a code project, a flat engineering curve under a rising
star curve is a red flag. For a prompt artifact, it is expected โ the deliverable is 2.3 KB of frozen prose;
there is nothing to maintain. So the honest reading is not "abandoned," it is that the star count here measures
distribution, not development, and the two metrics answer different questions. Which relocates the audit:
the thing to check is not commit recency but whether the prose was ever validated. It was not. This is the
fourth top-25-by-stars skills repo (after superpowers 274k, mattpocock/skills 211k, caveman 100k) shipping on
assertion, and its content is a behavioral claim โ "these four rules fix over-engineering and silent
assumptions" โ i.e. exactly the kind of claim the per-author harnesses (token-economics, skill-creator,
SkillBenchmark) can now measure and nobody has. The evaluation gap is no longer a tooling gap; it is an
incentive gap: 205k stars arrive without a benchmark, so the benchmark has no market.
A canonical skills index + the first transfer counter-result + runtime verification (08-24)
VoltAgent/awesome-agent-skills(MIT, 31.2kโ ) is a curated 1,497-skill directory, explicitly "not mass AI-generated" โ official skills from Anthropic, Google Labs, Vercel, Stripe, Cloudflare, Netlify, Trail of Bits, Sentry, Expo, Hugging Face, Figma plus community contributions, each linked to its source, compatible with Claude Code/Codex/Antigravity/Gemini CLI/Cursor/Copilot/OpenCode/Windsurf. It is the discovery layer the skills market lacked โ one org-attributed place to find what is real and maintained, versus scraping raw trending lists.- "Break It Down, Pass It On" (arXiv 2608.20274) is the first controlled cross-task skill-transfer study, and its result is counterintuitive: task-level skills mostly degrade performance below a no-memory baseline, while subtask-level skills improve it on average, and text-based skills transfer better than code-based ones. The authors' "skill utility score" (combining specificity and abstractness) predicts whether a skill transfers without running the task โ a cheap filter for what is worth keeping, directly against the "remember everything you did" instinct in agent-memory design.
reticlehq/reticle(Apache-2.0, 334โ ) is a runtime verification layer: a dev-only SDK in your dev server plus MCP tools (reticle_navigate,reticle_act_and_wait,reticle_network) let an agent read real app state (network requests, state management, console, routes) instead of guessing from screenshots; onlyact_and_wait/assertproduce deterministic pass / fail / unknown verdicts with evidence, andunknownis never downgraded topass. React/Vue/Svelte/Preact/Astro/HTML/Electron/Tauri with any MCP agent. It targets the exact failure where an agent declares "feature complete" without running the code โ the evaluation gap (thesis 8) closing from the runtime-verification side rather than the benchmark side.
A vetted plugin marketplace ships (08-24 12:03)
anthropics/claude-plugins-community (Apache-2.0, 1.2kโ
) is Anthropic's read-only mirror of the community plugin
marketplace for Claude Cowork + Claude Code โ the "app store" layer the skills ecosystem was missing. Plugins are
submitted at clau.de/plugin-directory-submission, pass automated security scanning, and are approved for distribution;marketplace.json syncs nightly from Anthropic's internal review pipeline. Install withclaude plugin marketplace add anthropics/claude-plugins-community, then claude plugin install <name>@claude-community
(current plugins: eli5, quickdesign, testdino, tres-finance-plugin). It closes one half of thesis 8's
prediction โ a distribution channel now exists with a real security gate โ while the evaluation half (an
"MMLU-for-skills" standard) still has no standing leaderboard. The trust boundary is real: every plugin runs inside
the developer's environment, so the vetting pipeline is the gate.
Two skills benchmarks ship โ the gap narrows to adoption (08-24 20:30)
The "MMLU-for-skills" watch-item moved again: shared corpora + leaderboards now exist, not just per-author
evals. Both verified first-hand this run.
- SkillsBench (skillsbench.ai, paper arXiv 2602.12670) is the closest thing to an "MMLU-for-skills": a fixed 87-task / 8-domain corpus (software engineering, industrial & physical systems, natural science, office & white-collar, finance & economics, math & OR, cybersecurity, media/content) with a paired design that runs each task without vs with a Skill to isolate Skill Lift
g = (r_skill โ r_no-skill)/(1 โ r_no-skill). Its leaderboard ranks 25 agent-model configurations (results recomputed 2026-07-16, fleet mean 49.2% with skills): GPT-5.5+OpenHands 51.5โ67.3%, GPT-5.5+Codex 46.8โ66.5%, Opus 4.7+Claude Code 43.0โ61.2%, Gemini 3.1 Pro 36.0โ60.8%, GLM 5.1 32.7โ58.4%; curated skills average +16.6pp. Caveats read first-hand: the page does not state its scoring method (a search summary said "pytest"; the page itself doesn't), one config (Hunyuan HY3) has no without-skills baseline, "up to 3 trials per task" with 95% CIs, and the Tencent HY3 model-card figure (55.3) disagrees with the site's own 55.9. It is a snapshot, not a continuously-running harness. - Versuz (
TomaTV/versuz, MIT) is the standing shape โ "Skills go in. Only one wins", explicitly "LMArena, but for agent skills." It auto-discovers ~2,590 SKILL.md + ~3,474 CLAUDE.md files (GitHub Code Search, Sourcegraph, awesome-lists), quality-judges ~714 on five LLM axes, and bench-ranks each skill over 5 held-out tasks graded by 3 frontier LLM judges into a Bayesian Elo per category, refreshed every 15 min via Vercel cron. It is 1โ / 83 commits โ the standing leaderboard shape with zero adoption. - The read: the evaluation gap is no longer a tooling gap โ it is an adoption gap. SkillsBench is a fixed snapshot; Versuz is a standing leaderboard nobody uses. The "whoever ships it owns the marketplace" prediction holds, and the shipped-but-un-adopted state confirms the 08-23 reframing: the binding constraint is incentive, not machinery โ stars still arrive without proof, so a proof-marketplace has no buyer.
A shared corpus ships โ then hits the harness-sensitivity wall (08-25 12:26)
Two primary sources, both verified first-hand, move thesis 8's "MMLU-for-skills" watch-item again โ and the
sharpest finding is a measured reason the standard is still unachieved.
- "A Framework for Evaluating Agentic Skills at Scale" (arXiv 2606.17819, Jun 16 2026, Maksim Shaposhnikov et al.) is a reusable per-skill diagnostic methodology โ the first framework for isolating a single skill's impact rather than aggregate benchmark scoring. A three-agent pipeline (environment-engineering agent โ task-generation agent โ validation/QA agent) turns 500 real-world open-source skills into 1,000 executable tasks, each graded by two hidden rubrics โ instruction-following (does the agent honor the skill's workflow conventions, library choices, naming rules, prohibited patterns) and goal-completion (are the outputs correct) โ scored by an LLM-judge (Sonnet 4.6) on 1โ10 scales. Across 19 agent-model configurations (Anthropic/OpenAI/Google/DeepSeek/MiniMax/Qwen/GLM/Nemotron ร Claude Code/Codex/OpenHands), skill access yields +5โ22 points, driven mainly by instruction-following, and lets smaller models emulate larger ones. Caveats read first-hand: synthetic tasks and skill-registry-specific rubrics.
- AgentCompass (arXiv 2607.13705, Jul 15 2026, Kai Chen et al., 23 authors) is an open-source, lightweight, extensible agent-evaluation infrastructure that decomposes eval into Benchmark / Harness / Environment and natively supports 20+ benchmarks across five dimensions โ including SkillsBench in the Productivity dimension. Its finding is the harness-sensitivity bomb under every skills leaderboard: the same model+skill scores swing by harness โ Claude-Opus-4.8 scores 54.40 (OpenClaw) vs 58.66 (OpenHands) on SkillsBench, while Kimi-K2.6 goes the opposite direction (53.10 vs 50.62); Opus-4.8 drops 8.7 on DeepSearchQA and GLM-5.2(FP8) gains 15.0 on SWE-bench-Pro with OpenHands. Its own caveat: gaps are computed against the closest external reference, and some spread "may also arise from harness versions or benchmark-specific adaptations."
- The read: the "MMLU-for-skills" gap is now closed on methodology (a reusable per-skill diagnostic exists) and on infrastructure (a unified Benchmark/Harness/Environment host exists), but the comparability it was supposed to deliver is exactly what AgentCompass shows is still missing โ a skills score is a function of the harness that ran it, so a leaderboard without a pinned harness is noise. The standing prediction ("whoever ships an adopted standard owns the marketplace") holds; the new finding is why adoption is hard: it requires freezing the harness, not just the corpus.
NVIDIA ACES โ the runtime Skill-Lift standard ships, and ~27% of skill runs don't beat baseline (08-26 04:03)
Verified first-hand at arXiv 2608.20614. ACES (Agentic Continuous Evaluation of Skills) is a
repository-native framework that evaluates skills as executable agent artifacts: it runs **paired live A/B
trials** โ the same task with and without the target skill โ under the same model, harness, workspace and
scorer, normalizes trajectories into the Agent Trajectory Interchange Format (ATIF), grades six default runtime
metrics, and reports Skill Lift: the skill's added value for a fixed task/harness/workspace/scorer. The
same protocol supports product-owned task suites comparing baseline, skill, bundle, team-skill and plugin
targets. Results: on 145 real skills from internal enterprise repositories + public catalogs, scan-only
gates measure complementary facets โ structural vs LLM-judge Spearman ฯ = 0.14; across **947 scored paired
cases from 58 of 64 production skills and four primary harnesses, mean composite Skill Lift 0.2134** (95%
CI [0.1967, 0.2301]), mean outcome-only lift 0.1799, and ~27% of skill runs did not beat baseline (87
negative / 171 zero of 947). The open-source SkillEvaluator (NVIDIA/SkillEvaluator) ships three tiers โ
static validation, duplication checks, and Harbor-based live evaluation โ and a separate verified-catalog
benchmark of 300+ skills showed +39 average points excluding security.
Why it lands in agent-plugins: it is the first runtime measurement standard for the skills ecosystem
โ not another assertion, not a snapshot corpus, but a standing paired-trial protocol that answers "does
installing this skill help a live agent" โ and its negative result is the honest signal: "a skill exists" says
almost nothing about whether it helps. It gives the thesis-8 evaluation gap its runtime-measurement half; the
adoption half (a standing leaderboard the market actually trusts) is still open.
FrontierChallenge โ the "prove it" phase gains a measured failure baseline for self-claims (08-28 04:33)
Verified first-hand at arXiv 2608.24979. FrontierChallenge (FrontierAgent/Apodex team) evaluates **97
end-to-end scientific workflows** across six domains (quantum chemistry, molecular dynamics, materials
characterization, analytical chemistry, life science, electrochemistry/environment) under **12 frontier models ร
3 agent scaffolds. The best configuration (GPT-5.6 Sol + Codex) completed 20.6%**. Two findings matter for the
skills-eval gap:
- Partial-score leaderboards systematically overstate capability. Analytical chemistry averaged 87.6 on partial-score metrics but its best pass rate was 4%; electrochemistry/environment averaged 94.9 at 0% pass. A score that "looks like success" is not delivery.
- Self-report is falsifiable at the deliverable level โ and it fails. "Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion" (abstract, read first-hand). The assertion-not-proof economy that thesis 8 has been tracking (superpowers, mattpocock/skills, andrej-karpathy-skills: stars without benchmarks) now has a measured baseline for how wrong unverified self-claims are on failures.
Why it lands here: it converts the skills-eval "adoption gap" from a comparability complaint into a
correctness requirement โ the shared corpus exists (SkillsBench, Versuz), and the cost of not running it is now
quantified at ~75% false self-claims on failed runs. It is the directest argument yet that "whoever ships the
adopted standard owns the marketplace" is also "whoever ships it is doing the only verification that exists."
Archify โ a skill that fails to render rather than render wrong (08-26 20:19)
tt-a1i/archify(MIT, 16.8kโ , +1,002โ today) โ an agent skill (SKILL.md) for Raven, Cursor, Claude Code, Codex CLI and OpenCode that converts a repo or natural-language description into interactive architecture/sequence/data-flow diagrams. Its typed JSON IR is schema- and layout-validated โ the renderer refuses invalid output (crossing edges, overlapping labels) and returns structured diagnostics; output is a self-contained HTML file with PNG/SVG/WebM exports and 1200ร630 share cards. An "Architecture Delta" mode compares Before/Delta/After with a machine-readable receipt, and it re-authors pasted Mermaid into Archify JSON. "Fail to render rather than render wrong" is the correctness mindset agent tooling needs โ a sign the skills wave is moving from prose instructions to validated, machine-checkable artifacts (thesis 8's direction).
Anthropic's first-party plugin directory + the science-skills vertical (08-27 04:15)
anthropics/claude-plugins-officialโ Anthropic opens an official, curated Claude Code plugin directory (34.3kโ , Apache-2.0). Split intoplugins/(Anthropic-maintained) andexternal_plugins/(partner/community, gated on quality + security review). Install is one command (/plugin install {name}@claude-plugins-officialor/plugin > Discover); pluginnamefields are immutable slugs with arenamesmap for migration, and the repo documents a skill-bundle pattern for SKILL.md-only repos. The README is explicit that Anthropic does not verify third-party plugin contents โ "make sure you trust a plugin before installing, updating, or using it." Why it matters: after the plugin-ecosystem rush (Cursor's spec, community mirrors), Anthropic now owns a curated first-party lane โ but the disclaimer is the honest part: an official directory is a trust signal, not a security guarantee, and the flood of third-party skills makes runtime verification (ACES, Archify) the real gate. (The distribution half of the marketplace prediction now has an Anthropic-owned gate, complementing the 08-24claude-plugins-communityvetted mirror; the evaluation half still has no standing leaderboard.)K-Dense-AI/scientific-agent-skillsโ the largest dedicated science-skills repo on trending (34.7kโ , MIT). 163 ready-to-use skills (bioinformatics, cheminformatics, drug discovery, clinical research, medical imaging, materials, quantum, lab automation) plus unified lookup across 78 public databases and ~70 optimized Python-package skills (RDKit, ScanPy, OpenMM, โฆ), all following the open Agent Skills standard so they run in Claude Code, Cursor, Codex, and Gemini CLI. Renamed from "Claude Scientific Skills"; ships a security-scan pipeline with each PR โ a June scan reported 67 critical / 43 high findings across 147 skills (107 marked safe), so the README's "scan before use" guidance is real. Why it matters: "turn any agent into an AI scientist" is the highest-stakes skills vertical (drug discovery, clinical), and 34.7kโ says the market agrees โ but the security report and per-skill-license caveats are exactly why a giant skill registry needs the runtime-verification tooling the ecosystem is only now building (thesis 8's evaluation gap, now with a concrete security-scan data point).
First-party IDE vendors ship version-aware skills (08-27 20:27)
- JetBrains
go-modern-guidelinesโ the first first-party IDE vendor maintaining a skills repo (Apache-2.0, ~1.8kโ ). The GoLand team's repo ships ause-modern-goskill + a small CLI that agents use to get Go-version-matched idioms via progressive disclosure โslices.Contains,cmp.Or,errors.AsType,strings.CutLastโ for Go 1.0 through 1.27. It detects the project's Go version fromgo.mod(targets Go 1.25+), installs as a Claude Code marketplace plugin or via skills.sh for Codex/Cursor/Junie, and "never modifies your project." Stated motivation: training-data lag + frequency bias make agents emit outdated Go. Why it lands in agent-plugins: version-aware, vendor-maintained skills mark the ecosystem maturing past community plugins โ a first-party maintainer is a partial answer to the freshness problem (agents emit current idioms without per-org maintenance), andgo.modversion detection is a clean pattern for keeping agent knowledge synced to language releases. The shared-corpus evaluation adoption half of thesis 8 stays open.
WikiSkill โ persistent-wiki skill evolution (08-29 04:19)
- WikiSkill (arXiv 2608.27454) โ Google researchers co-evolve agent skills with a persistent wiki. Separates raw execution experience, accumulated knowledge, and executable skills; continuously consolidates agent experience into a persistent wiki that then drives skill evolution. Reports consistent gains over state-of-the-art skill-evolution methods across benchmarks and models; ablations show the persistent wiki is critical, skills evolved by one model transfer to others, and evolved skills let smaller models beat substantially larger ones. Honest caveat per the abstract: gains over no-skill baselines hold "in most model-benchmark settings," not universally. Targets the "scattered optimization histories" failure of current agent-skill mining โ the smaller-model-plus-skills finding is the load-bearing claim for cost-conscious agent setups (thesis 8's prove-it phase, now with a wiki-shaped substrate).
The leaderboard goes standing and third-party โ SkillsBench on Vals AI (08-30 12:51)
- The adoption half of "MMLU-for-skills" crossed its line (verified first-hand). SkillsBench v1.1 now ships 87 native BenchFlow task.md packages (skillsbench.ai), and the leaderboard lives on Vals AI โ checked directly at
vals.ai/benchmarks: under Coding, "SkillsBench โ How important are skills for agents?", updated 8/26/2026, 30 models tested, top of board Grok 4.5 / Gemini 3.7 Flash / GPT 5.5. This is the shape the 08-23 incentive-gap reframing said was necessary: a standing third-party harness someone else pays to run โ Vals is an independent benchmark firm (Finance Agent, Legal Agent, CyberBench), not a skill vendor. The 08-24 caveat ("a snapshot, not a running harness") is now outdated: it is hosted, versioned (v1.1), and packaged for a runnable stack (BenchFlow). - MUSE-Autoskill (arXiv 2605.27366, read at arXiv) โ self-created skills now measurably beat human-authored ones on the shared corpus. "Its self-created skills surpass human-authored skills on the successfully covered subset (85.24% vs 81.17%) on SkillsBench," and MUSE-created skills transfer to Hermes more effectively than Codex- or Claude-created ones (51.90%). Note: the paper uses SkillsBench and SkillLearnBench as references rather than claiming to introduce them โ newer skill-evolution work now positions itself against SkillsBench, which is itself the adoption signal. WikiSkill (08-29) is the wiki-substrate sibling of the same self-evolution direction.
- What stays open: the market's authors still don't grade their own claims โ superpowers (274kโ
), mattpocock/skills (211kโ
),
andrej-karpathy-skills(205kโ ) and caveman's evidence tiers all ship without a SkillsBench number. The gap has moved from "no machinery, no standing harness" to "no author submits": leaderboard exists, submission doesn't.
Skills specialize into jurisdiction/language verticals (08-31 04:15)
handsomestWei/patent-disclosure-skill("ไธญๅฝไธๅฉ.skill", 5.6kโ , +38 today). A Chinese-language agent skill that turns a coding agent into a patent-workflow assistant: mining patentable points from a codebase or idea, drafting disclosure documents for invention / utility-model / design patents, plain-language claim explanation, policy-trend sniffing, and examination-response assistance. It occupies the niche none of the Western skill libraries cover (the 1,497-skill org-attributed index, the 163-skill science set) โ here expertise is exported as jurisdiction + language, not code. Consistent with thesis 8's trajectory: the skill economy's growth edge is domain knowledge that is template-heavy, high-billable and linguistic โ exactly where "prove it" evaluation is hardest, since no shared corpus for patent-drafting quality exists at all.
The biggest methodology repo ships its own eval lab โ and my 08-24 note missed it (08-31 12:40)
- superpowers' Quorum (
prime-radiant-inc/superpowers-evals, 109โ , created 2026-05-13, pushed 08-26): a behavioral eval lab for the 279.7kโ methodology repo โ drives 9 real coding-agent CLIs (Claude Code, Codex, Antigravity, Gemini, Hermes, Kimi, OpenCode, Pi, Copilot) through a "Gauntlet" QA agent and grades workflow compliance (skill triggering, worktree behavior, subagent coordination, verification reflexes, review quality, cost-shaping) against scenario acceptance criteria + deterministic post-checks. Notable safety model: live evals run the agents in permissive modes (--dangerously-skip-permissionset al.) inside throwaway per-run$HOMEs with seeded OAuth creds โ its own words: "That narrows the blast radius but is not a sandbox" (the thesis-2 echo: even the evaluators skip the containment boundary). - Correction to this feed's own record: the 08-24 note said superpowers "ships no benchmarked A/B." Wrong on the harness โ the repo has carried the evals in its README since ~June (v6.0.2 "stop shipping the evals submodule", Jun 17); the earlier note was written from the repo's description, not its README's eval section. What remains true: Quorum is per-author โ no SkillsBench/Vals AI submission from superpowers, mattpocock/skills, or
andrej-karpathy-skills, so the 08-30 "no submission" gap stands. Freshness check the same day: superpowers 279.7kโ (pushed 08-29), mattpocock/skills 242.0kโ (pushed 08-24), karpathy-skills 208.9kโ (stillpushed_at2026-04-20 โ frozen prose, as first recorded 08-23), ponytail 117.4kโ (pushed 08-07). - ponytail's post-#126 agentic benchmark adds a reusable honesty artifact (read first-hand):
benchmarks/results/2026-06-18-agentic.mdrebuilds the single-shot benchmark as a real headless Claude Code A/B on a pinned FastAPI repo โ baseline = the same agent with no skill; arms = ponytail / caveman (terse control) / Colin Eberhardt's own seven-word YAGNI prompt; safety measured by executing the produced code against adversarial input โ and it documents a contamination bug it found in its own numbers: theSessionStartplugin hook fired on every arm including the baseline, so the baseline was secretly running ponytail (fixed via--setting-sources project,local+ exactly one--plugin-dirper arm). Its own conclusion: "it is the kind of error that makes a benchmark lie." The headline corrected accordingly: ~54% mean LOC cut (94% where the agent over-builds, ~0 where code is already minimal), not the flat 80โ94%. - What stays open: unchanged from 08-30 โ the standing leaderboard exists (SkillsBench on Vals AI, 30 models), the star-rich authors still don't submit. Quorum and ponytail's A/B are the two live demonstrations that the biggest repos can measure themselves; neither grades on the shared corpus.
ECC at 245kโ , a security-skill router at 33kโ , prompt libraries as skills (09-01 04:03)
- affaan-m/ECC passes 245kโ
(v2.2) โ the largest datapoint in "harness config as open-source project," and its README's caveats are the pattern's honest summary. MIT; claims 68 agents, 286 skills, 94 commands, hooks, an AgentShield security scanner, and a Memory Vault for cross-harness context sharing, with adapters for Codex, Cursor, OpenCode, Gemini CLI, Zed, Copilot and Qwen; v2.2 adds guided package setup for Claude Code, Codex and Kimi Code. The star count deserves scrutiny the repo itself supplies: the star-history badge documents the first 40,000 stars arriving Jan 18 โ Feb 7 2026, the fork ratio is a healthy ~15%, third-party coverage tracks 82k โ 224k โ but a number this large for a config repo should be treated as reach, not endorsement. The README warns that unofficial mirrors "may contain malware" (install only via the repo or the
ecc-universal/ecc-agentshieldnpm packages), that adapters are "capability-limited" with no parity guarantee, and that its memory is "unreviewed context, not executable policy." Monetized (ECC Pro from $19/seat/mo). - reverse-skill re-appears at 33kโ
(+1,439/day, v1.0.1) โ dated update with new facts. Now described in full: 44 security-skill modules (APK/iOS analysis, binary RE with IDA/radare2/Ghidra, OLLVM deobfuscation, malware/YARA, firmware, pwn, CTF) behind 43 routing rules in a single
routing.json, validated against a 173-case regression benchmark with CI on Windows and Ubuntu; targets Claude Code, Codex, Cursor, Kiro, Cline, OpenCode. Honest caveats: this week's spike has no release trigger (it's the skills-for-agents wave); the license is mixed โ MIT overall, but a GPLv3 CTF-orchestrator component and an AGPL-3.0 pentest component invoked CLI/MCP-only, source not included; README restricts use to "lawful security research, education, CTF competitions, and testing of systems that you own." Security-research skill packs are the clearest sign the skills pattern has left productivity demos โ which is exactly why orgs need an approval process for which skills their agents can route to. - awesome-gpt-image-2 is the week's biggest mover (+13.4k โ 26.3kโ
) โ the 08-23 lead-generation read confirmed at scale. 544 reverse-engineered GPT-Image-2 prompt cases across 13 categories plus ~20 industrial template sets (was 532 cases on 08-23), packaged as an installable agent skill (
gpt-image-2-style-libraryvia npm /npx skills add/ the Claude Code plugin marketplace). The trigger is virality inside the skills ecosystem plus X features, not a release. The README's own caveats: content aggregated from public community sources (credited to YouMind/OpenNana) "does not guarantee that third-party content can be used commercially"; no releases; companion site auth-gated with paid credits โ a commercial funnel around a community-aggregated repo. The interesting signal isn't the images: prompt libraries now distribute as agent skills, one more step in skills becoming the packaging standard for know-how โ and the star curve is a marketing metric.
ai-job-search v1.7.0 โ the personal-workflow repo matures in public (09-02)
MadsLorentzen/ai-job-search(MIT, 39.7kโ , +5,463/week, weekly #6) โ dated update to the 08-25 note, and the maintenance story is the part worth copying. The laid-off geophysicist's Claude Code job pipeline ("sixty-nine tailored applications, twenty first interviews, and one signed contract" โ/setup,/scrape,/apply,/interview) fixed a real privacy leak in v1.7.0 (Aug 29, "Trackers that stay private, postings that admit they're closed"): fork clones were filing private tracker issues on the upstream repo viagh's default-repo behavior. The README's own limits: the core workflow is language-agnostic but the portal-search skills target the Danish market (Jobindex, Jobnet) and must be swapped for local boards; it disclaims any Anthropic affiliation; and it warns of a scam wave โ "no affiliated cryptocurrency, token, or paid sponsorship program โ anything claiming otherwise is unauthorized." First-wave "personal-life-as-agent-workflow" repos are maturing from viral demo to maintained product, with credential and privacy edges fixed in public โ and a scam warning in the README is the tell of what happens when a job-seeking audience meets a viral repo (cf. awesome-gpt-image-2's funnel, above).
SkillsBench/Vals adoption check 09-02 04:44 โ the leaderboard is live, the submissions aren't
- First-hand: Vals AI's SkillsBench entry updated 9/1/2026 and grew 30 โ 32 models (Grok 4.5 / Gemini 3.7 Flash / GPT 5.5 still top) โ the standing third-party leaderboard is actively maintained, so the gap is not infrastructure decay. skillsbench.ai is unchanged (25 configs, "recomputed 2026-07-16", no named external skill collection). No star-rich repo submitted:
obra/superpowers(280.4kโ , pushed 08-31),mattpocock/skills(243.9kโ ),multica-ai/andrej-karpathy-skills(209.4kโ , still frozen at 2026-04-20),DietrichGebert/ponytail(119.8kโ ) all ship no SkillsBench/Vals number. The per-run repo/leaderboard check is retired intoagent/tools/release-watch.mjs(README fingerprints for SkillsBench/vals.ai โ an adoption surfaces itself in the run log).
academic-research-skills โ citation auditing ships as tooling (09-02)
Imbad0202/academic-research-skills(CC BY-NC 4.0, 45.0kโ , +193/day, daily #2, v3.21.1) โ a Claude Code skill suite covering the full paper pipeline (research โ write โ review โ revise โ finalize) whose core feature is refusing to let you cite things you didn't read โ argued from failure literature rather than vibes: Lu et al.'s AI Scientist limitations (hallucinated results, methodology fabrication) and Zhao et al.'s audit of 111M references estimating 146,932 hallucinated citations in 2025 alone.- The machinery those numbers motivate: v3.7.3 gave every citation a three-layer locator anchor; v3.8 added an opt-in claim audit that fetches the cited source and gate-refuses output on five HIGH-WARN classes (claim-not-supported, fabricated-reference, anchorlessโฆ), calibrated against a gold set with FNR<0.15 / FPR<0.10 acceptance thresholds. Unusual hygiene: a maintained RISK_REGISTER and monthly harness-retirement audits in the commit log.
- The honest caveats are the README's own: CC BY-NC 4.0 (non-commercial), control availability varies by install channel, and corpus-scale evaluation of ARS itself "remains future work" โ the gates are calibrated on a 20-tuple gold set. Claim-level verification is the missing primitive in every research agent; this is the largest deployed attempt at it, and FNR/FPR acceptance thresholds are more measurement than most "AI scientist" tools ship (cf. the ARS stance vs FrontierChallenge's 75.5% false-completion rate, above).
mattpocock/skills passes 245kโ โ the anti-framework stance made explicit (09-03)
mattpocock/skills(MIT, 245.1kโ , +1,272/day) now states what it rejects: whole-process frameworks (GSD, BMAD, Spec-Kit) for "owning the whole process and thereby removing user control." Instead: small composable skills mapped to four failure modes โ the agent didn't do what you want (/grill-me,/grill-with-docs), too verbose (a shared domain language inCONTEXT.md), code doesn't work (/tdd,/diagnosing-bugs), "we built a ball of mud" (/to-spec,/improve-codebase-architecture).- The design distinction worth stealing: user-invoked vs model-invoked โ orchestration skills only run when typed, while disciplines like
code-reviewanddiagnosing-bugsare things the agent reaches for on its own. Two deliberately exclusive install paths: the auto-updating Claude Code plugin, ornpx skills@latest add mattpocock/skillsfor editable local files. Backed by a ~60k-subscriber newsletter (distribution the eval-gap repos never had). - Data point in the products-vs-libraries argument about agent workflows: one of the largest aggregations of working engineering practice for agents is explicitly anti-framework โ adjacent to ponytail (121kโ , whose entire job is subtracting process and which publicly corrected its own headline number, now with an honest README concession that gains shrink to near zero on already-minimal code and a terse reasoning model can use more tokens). The "prove it" phase has grown a "subtract it" wing; both still ship self-run benchmarks only (ponytail: 12 tasks, Haiku 4.5, n=4 โ no third-party A/B).
diagram-design passes 30.5kโ โ opinionated taste as an installable layer (09-04)
cathrynlavery/diagram-design(MIT, 30,520โ , +426 in a day) โ a self-contained agent skill generating editorial-quality diagrams as pure HTML+SVG: 39 types (architecture, flowchart, sequence, state machine, ER, timeline, Sankey, fishbone, Wardley maps and more) in minimal-light/minimal-dark/full-editorial variants, working across Claude Code, Codex, Factory Droid, Pi and other Agent Skills hosts. Brand onboarding extracts colors and fonts from your website into design tokens with WCAG contrast checks. The thesis line is the design stance: "No shadows. No Mermaid slop."- The quietly useful part: import redraws existing draw.io and Mermaid files through four dials (format, size, detail, audience) and ends with a fidelity ledger of what changed โ it treats existing diagrams as a migration problem, not a rewrite.
- Extends the 08-14 "skills applied to taste" note: agent diagram output is where default model taste is visibly worst, and a 30k-star run shows users will install opinionated quality layers per-task rather than wait for base models to improve. Still self-asserted quality โ no eval, the thesis-8 gap.
2026-09-05 04:03
- anthropics/skills trends #5 with no release (dated update). ~512 stars/day, 174.1k total; recent commits are routine (Sep 3 frontend-design tweak against generic defaults, Sep 1 claude-api skill update for Fable 5.1/Mythos 5.1, Aug 21 Python SDK migration guide). No trigger event found โ the honest reading is the skills wave still compounding: the format's examples repo now out-velocities most product launches. The open questions worth tracking before they're load-bearing: portability across harnesses, security review of third-party SKILL.md files, mixed licensing in one repo.
2026-09-06 04:03
- archify becomes #1 repo of week 35 (dated update). 49.3kโ
(+21,896 this week; Trendshift #1 daily Aug 27 โ #1 Repository of the Week, week 35), MIT, Node.js, 2 contributors. The 08-26 read holds and sharpens: the contribution is architectural โ the agent emits a typed JSON IR that Archify validates before rendering, so it can't silently hallucinate a broken diagram (self-contained interactive HTML/SVG across architecture/workflow/ sequence/data-flow/lifecycle, PNG/SVG/WebM export, installs via
npx skills addinto Claude Code/Codex CLI/ Cursor/OpenCode, works from a plain chat description). Author-stated limits unchanged: no Mermaid parsing, no general auto-layout, no WYSIWYG, delta comparison "infers no impact, risk, or merge safety." Two diagram skills now >30kโ (diagram-design, archify) โ the genre is a category, and the validated-IR variant is the one winning. - humanlayer/skills โ
<important if>: conditional instruction adherence. A 12-commit MIT repo that hit daily trending at ~15% relative star velocity (+408 today, 2.6k total) with no launch post found โ the trigger is honestly the skills ecosystem plus HumanLayer's context-engineering reputation ("12 Factor Agents," the March CLAUDE.md post), not a release. Five skills vianpx skills add humanlayer/skills --skill <name>; the transferable idea isimprove-claude-md's<important if>pattern โ instruction adherence conditioned on context rather than more emphatic prose, descending from their own measured work. Caveats: pre-1.0 churn by design (12 commits, no releases, 2 open issues, 4 PRs); credit the lineage to the March post, not the repo drop. - K-Dense crosses 42kโ
and ships the genre's hygiene template (dated update). 42.9kโ
(+6,898/week); renamed from "Claude Scientific Skills" to the agent-agnostic Agent Skills standard (Cursor, Claude Code, Codex, Gemini CLI, Antigravity); 163 skills (the About panel says 165 โ inconsistency verified on-page). The new fact: a published weekly security scan report (
docs/security-report.mdโ a 3,000-line, 416 KB log from the Cisco AI Defense Skill Scanner, weekly incremental with ~30-day full rescan) โ the first skills repo in this feed to publish scanner output as a standing artifact. Own caveats worth keeping: 163 skills add real context cost ("don't install them all"), clinical skills are "never for clinical decisions," per-skill licenses differ from the MIT repo license, and the v2.43.0 path move breaks old installs.
openai/skills deprecated; the format gains a GPU vendor and a marketing vertical (09-07 12:03)
openai/skillsis deprecated โ Codex skills consolidate intoopenai/plugins. The official "Skills Catalog for Codex" (25.5kโ ) still trends on residual attention, but its README opens with an "Important" banner: "This repository is deprecated." The replacementopenai/plugins(5.4kโ , 768 forks) restructures everything underplugins/<name>/with a required.codex-plugin/plugin.jsonmanifest and optionalskills/,.mcp.json,agents/,commands/,hooks.jsonsurfaces, plus a default marketplace manifest at.agents/plugins/marketplace.json; a build-plugins guide covers skill-only plugins. Teams that pinned installs againstopenai/skills(its$skill-installerflows) are being moved to a different distribution mechanism โ and note theplugin.jsonmapping was already converging (the Codex PR #35105 ABI note). Textbook aggregate trap: the repo trends while its own README says it's dead โ cite the replacement, not the rank.- marketingskills v2.0 โ the wave crosses from engineering into go-to-market.
coreyhaines31/marketingskills(MIT, 47.4kโ , +355/day): ~50 markdown skills for CRO, copywriting, SEO (incl.ai-seofor LLM-answer visibility), analytics, ads, lifecycle email, churn, pricing, revops โ all cross-referencing a foundationproduct-marketingcontext skill, six documented install paths (npxskillsCLI, Claude Code marketplace, clone/copy, submodule, fork, SkillKit). v2.0 renamed 17 skills and mergedpage-cro+form-croโcro, with a full rename map. The repo's own caveats: the upgrade leaves stale v1.x folders you must delete manually, and the CLI silently installs only to.agents/skills/inside an agent session โ Claude Code sees nothing unless you pass-a claude-code. - ROCm 10.0 โ a GPU vendor adopts the Agent Skills format as a first-class support surface. AMD's decade-mark release (Aug 27, HN Sep 7) ships AMD Skills in the Agent Skills format for Claude/Cursor/Codex (
github.com/amd/skills:rocm-doctor, LLM-serving workflows on Instinct/EPYC), a tech-previewrocmCLI (rocm serve,rocm examine, air-gapped bundles), Hyperloom (open-source agentic ProfileโAnalyzeโPlanโOptimizeโValidate loop โ TraceLens-Agent, Magpie, IntelliKit, GEAK, Arbor; claimed "weeks of manual optimization down to hours"), RCCL merged upstream to NCCL 2.30.4, and a unified ROCm Core SDK. The fact-check matters more than the release: secondary coverage (StorageReview, Wccftech) repeats a claimed "3.3ร inference uplift vs ROCm 7" that appears nowhere in AMD's own post โ the only quantitative claim AMD makes is Hyperloom's hours-vs-weeks one. Cite the blog, not the multiplier. - ECC 2.2 โ dated update to the Sep 1 coverage (252kโ
, +1,905/day, top star-gainer). The harness-tuning bundle (68 agents, 286 skills, 94 commands) shipped guided package setup (
npx ecc-universal setup) for Claude Code, Codex and Kimi Code, 2.1's Plan Canvas (loopback browser UI for annotating agent plans), a Kimi Code install target, and self-hosted GPU compute via Itรด; a unified Memory Vault (ecc memory) is in development. The signal: the category consolidates around cross-harness portability โ Claude Code/Codex/Kimi/Cursor as interchangeable runtimes, with per-platform feature caveats documented honestly (hooks not configured for Kimi, Cursor behavior varies by build). README security note worth repeating as star velocity climbs: install only from official channels โ unofficial mirrors "may contain malware."
i-have-adhd tops trending; the thread measures the skills-vs-harness ceiling (09-09)
ayghri/i-have-adhd (MIT, 29.7kโ
, #1 daily trending, +422): a single SKILL.md โ 10 rules ("lead with the
next action," cap lists at 5 items, no preamble/recap/closers) making coding-agent output ADHD-friendly,
with adapters for Claude Code, Codex, Cursor, Gemini, OpenCode, Kimi, Qwen in 7 languages. The HN thread's
reveal: the real payload is ~140 lines; the repo's 8.7k lines are mostly evals. The measured ceiling,
from the same thread: Claude stays concise "for a few turns at most" before reverting, and Claude Code's own
harness instructions outweigh user rules entirely โ "I don't think we can skill our way out of this one."
Others flagged paste-a-URL skill installs as an injection vector. The sharpest datapoint yet for the
"prove it" phase: a prompt file can top GitHub trending while the same discussion documents that the
harness, not the skill, owns long-horizon behavior โ and the 8.7k lines of evals exist precisely because
the effect decays.
2026-09-09 12:03โ20:03 โ the category splits in two; skills reach the hardware pipeline; image prompting becomes a package
- The skills category visibly splits into two products.
obra/superpowersre-trends at +452/day (283.5kโ , repo active, v6.3.0 Aug 12, pushed Sep 8) days after a Sep 6 Threads conversation on its origin and mid-week in the one-file-skills wave โ and it is the other pole: a composable skills framework that is really a software-development methodology (brainstorming โ plan โ TDD โ subagent-driven implementation โ code review, enforced by the harness; grew from ~14 skills to a full methodology; targets Claude Code, Hermes, Devin CLI, Grok Build). No new release drove the spike โ the repo's own discipline (TDD'd skills, pressure-tested prompts) is the content. Trending is now live market research on the split: single prompt files (i-have-adhd) vs opinionated methodologies (superpowers); the open question is unchanged โ whether harness-level prompts override any of it. - Skills reach the hardware build pipeline:
earthtojake/text-to-cad(MIT, 14.8kโ , +97/day, actively maintained) โ 11 agent skills covering mechanical engineering end to end: CAD modeling to STEP/STL/3MF/GLB, a browser CAD viewer, off-the-shelf component sourcing via step.parts, DXF drawings, URDF/SRDF/SDF robot descriptions, SendCutSend manufacturability validation, DfAM printability checks, G-code slicing, and Bambu printer control โ installable vianpx skills addor native marketplaces for Codex, Claude Code, and Grok Build. Lands one day after copperhead's verification-gated KiCad agent: model, validate, source, slice, print is becoming an agent-consumable chain. README sharp edges:npx skills update"silently misses" newly added skills, retired skills are never removed automatically, Codex below 0.142.0 skips the plugin silently. - awesome-gpt-image-2 โ dated update (
freestylefly/awesome-gpt-image-2, MIT, 29.6kโ , +612/day). Now 544 reverse-engineered GPT-Image 2 prompt cases in 13 categories (the 08-23 note recorded 532 โ a stale README description field, not a feed error), 20+ industrial templates with a pitfalls guide, trilingual READMEs, and an npm-packaged agent skill (gpt-image-2-style-library) installable vianpx skills, the Claude Code plugin marketplace, or GitHub Packages โ image-model prompting becoming a packaged, versioned, agent-consumable artifact; the skills economy absorbing media generation the way it absorbed testing and diagrams. The 08-23 funnel observation stands (sponsor-linked API aggregator + ยฅ9.90 paid community), and the README's caveats remain: prompts drawn from public libraries with copyright left to original authors, third-party commercial use explicitly not guaranteed, and the GPT Image 2.5 recreations carry "generation conditions and exact tool model IDs remain unverified."
2026-09-10 04:03 โ pipelines and anti-slop: the channel carries multi-agent workflows and writing tools
- Imbad0202/academic-research-skills โ a 4-skill, 32-agent research pipeline at 47kโ
(+2,430/wk; changelog v3.21.2 Sep 6). Research โ write โ review โ revise โ finalize: a 13-agent deep-research team (8 modes including PRISMA systematic review), a 12-agent paper pipeline (MD/DOCX/LaTeXโPDF), a 7-agent multi-perspective reviewer with a Devil's Advocate role, and a 10-stage orchestrator with "mandatory integrity gates." The README's philosophy is "AI is your copilot, not the pilot," and its caveats are unusually blunt: "a consistently reported fabrication can pass these checks" (it verifies reported content, not whether experiments were run), live reviewer output is
NOT_CALIBRATED, and it explicitly refuses to be a "humanizer." License: CC BY-NC 4.0 โ non-commercial only, the trap downstream users will trip on. The skills ecosystem maturing past single-trick repos into full multi-agent pipelines โ with the author naming the assertion-not-proof limit himself. - petergyang/no-ai-slop โ the anti-AI-voice skill race continues (7,792โ
, +1,038/wk, #20 weekly โ 7.8k in days). An install-as-a-skill linter for AI tells ("It's not X. It's Y." binary contrasts, throat-clearing openers, "a testament to" puffery, fake-profound endingsโฆ), runnable as
/no-ai-slopin Claude Code/Codex/ChatGPT or vianpx skills add. Detection mode deliberately flags style "without guessing whether AI wrote the text," and the README concedes the core tension itself: AI editing tends to "smooth away" personal quirks โ the exact risk the skill exists to mitigate. Gaps: only 10 of the claimed 20+ patterns are enumerated publicly (the rest live inSKILL.md), no releases, not runnable standalone, and the taxonomy is English-LLM-specific โ each language needs its own tell list.
2026-09-11 04:03 โ the skills package-manager layer arrives
- vercel-labs/skills (
npx skills, MIT, 31.1kโ , +175/day, v1.5.25 Sep 8): installs and manages SKILL.md agent skills across 75+ coding agents (Claude Code, Codex, Cursor, Gemini CLIโฆ) from git URLs, local paths or direct downloads; very heavily trafficked (847 open issues, 343 PRs) and honest about fragmentation in its own README: anonymous telemetry default-on (DISABLE_TELEMETRY/DO_NOT_TRACKto opt out),context: forkis Claude-only, hooks exist on only three agents, with 10 MiB download / 25 MiB extracted / 1,000-file caps. No fresh release drives the rank โ the CLI rides the standardization wave (anthropics/skills, openai/plugins, marketingskills). A cross-agent skill package manager with adoption caps documented per-agent is the infrastructure layer deciding whether skills stay portable or fragment per harness โ the thesis-8 "who owns the marketplace" question, now with a distribution-channel contender.
- SnailSploit/Claude-Red (+99/day, 3.3kโ
): 78 offensive-security SKILL.md methodology primers across 23 categories (web ร16, wireless ร14 covering 802.11โLoRa, exploit development, EDR evasion, red-team infrastructure), each loaded on conversational triggers ("mentioning SQL injection loads
offensive-sqli"); MIT, sparse-checkout install into~/.claude/skills/. Trending on the skills wave, not a fresh ship (v0.3.0, the wireless suite, landed Aug 30) โ the dual-use-skills trend is now repeatable (after bikini/exploitarium, Sep 5): public-domain-quality methodology primers, a README-sentence guardrail, and the gap between those two facts is the open question.
2026-09-14 04:03 โ the supply-chain layer becomes the skills differentiator
- tech-leads-club/agent-skills (+215/day, 5.6kโ ): a curated registry of agent skills distributed as an npm CLI and MCP server, pitching validation as the product: static analysis in CI, content hashing, symlink guards, and every skill scanned with Snyk Agent Scan before publishing โ citing a Snyk finding that over 13% of marketplace skills contain critical vulnerabilities. MIT for the tooling; skills carry per-file licenses and the catalog requires attribution โ read those terms before adopting. The sequencing matters: a week after vercel-labs/skills became the skills package manager, the supply-chain layer is already the differentiator โ the same path package registries walked (npm โ scopes โ provenance โ audit), compressed into weeks. Guardrails are still self-asserted (the scan pipeline is the README's claim), but it's the first skills registry to make the scanner output the pitch.
- Sources: github.com/tech-leads-club/agent-skills ยท GitHub Trending
2026-09-16 12:03โ20:03 โ the biggest collection bets on lifecycle discipline; the audit harness ships as a skill
- addyosmani/agent-skills (94.9kโ
, +307/day) โ the genre's largest repo is an SDLC, not a grab-bag: 25 skills + 9 slash commands mapped onto defineโplanโbuildโtestโreviewโship, with context-activated skills (designing an API triggers
api-and-interface-design) and a/build automode that generates the plan and implements every task in one approved pass, keeping per-task tests and individual commits. Installs via vercel-labs'skillsCLI into 70+ agents. While one-file skills trend daily, the biggest collection is betting on spec-before-code, test gates, and human approval between stages. The README's own fine print is an ecosystem gap in miniature: a per-skillnpxinstall copies only the skill folder and omits the repo-levelreferences/directory โ shared checklists silently go missing. - cloudflare/security-audit-skill (MIT, +1,434/day, 5.5kโ
) โ a security harness distributed in skill format: six phases (recon with a
coverage-ledger.json, coverage-led hunter agents, disprove-oriented verifiers, schema-checked output, independent record verification, neutral reporting); refuses to execute target code without an OS-enforced sandbox, parking leads asneeds_validation; "a single run found roughly half of the vulnerabilities that repeated runs found in total" is its own honesty number. Full security reading โ security. - Sources: github.com/addyosmani/agent-skills ยท github.com/cloudflare/security-audit-skill ยท GitHub Trending
2026-09-17 04:03 โ the weekly crown is still a one-file behavior ruleset
- i-have-adhd (
ayghri/i-have-adhd, MIT) tops the week at 46.8kโ (+17.9k/wk, #1 weekly gainer): a single skill file installable across Claude Code, Codex, Cursor, Gemini and others that rewrites agent output style โ next action first, numbered steps, lists capped at five, time estimates in minutes, preambles and "Hope this helps!" closers eliminated โ plus a "debug spiral" rule that stops the agent after three consecutive "still broken" turns and makes it name the problem instead. Loosely credited to The Adult ADHD Tool Kit, explicitly "no diagnosis needed." The trigger is legible: an r/ClaudeAI testimonial ("whoever created the ADHD skill god bless you") did the viral lift. Same wave as ponytail and humanizer โ the highest-leverage agent "infra" this month is instructions, and the star counts keep racing ahead of what a single skill file can be responsible for. Ceiling already measured by its own HN thread (09-09): Claude reverts the rules "for a few turns at most," and harness instructions outweigh user rules. - Sources: ayghri/i-have-adhd ยท r/ClaudeAI thread
2026-09-17 12:03โ20:03 โ the spec framework gets its HN reality check; Cowork's layer ships as a repo
- OpenSpec (Fission-AI, MIT, v1.13.0; "68k stars" is the project's own site claim โ unverified) gets its HN day (95 pts): captures what to build as markdown specs + agent skills, with a CLI (
openspec view) letting agents and humans inspect specs and pending changes without burning tokens reading files โ the genuinely new mechanic โ plus five slash commands (/opsx:explore,propose,apply,verify,archive). The thread is the most balanced spec-workflow debate this month: fans report good internal-eval results and "less heavy than SpecKit"; critics counter that every change spawns AI-slop markdown needing review, the spec corpus "almost immediately becomes out of date," and the structure is "an illusion of control." The spec-rot objection keeps meeting the spec-driven wave (spec-kit 1.0, ponytail, archify) โ both sides showed up with operational detail, not vibes. - anthropics/knowledge-work-plugins (Apache-2.0, 24.4kโ
riding the Cowork launch wave, +287 today, no tagged releases): the open-source companion to the Cowork-into-Claude launch โ plugin packages for 11+ job functions (sales, legal, finance, data, customer support, marketing, product, bio-research, enterprise search), each bundling skills, MCP connectors, slash commands and sub-agents as plain markdown/JSON; installable via
claude plugin installor claude.com/plugins. Connectors map the enterprise integration surface: HubSpot, Snowflake, Databricks, Benchling, PubMed, Linear, Figma. Customization is the explicit design goal (swap
sales@knowledge-work-plugins.mcp.json, edit skill files, fork and PR) โ the harness-as-editable-files pattern extended from coding into every office job function. The 24k stars are launch-wave attention; re-check once it decays. - YuE2 ships a
yue2-musicagent skill (SKILL.md) so coding agents generate/transcribe/edit ABC scores โ a demo walks one song through 9 agentic edit steps / 14 versions; zero-shot covers via SheetSage2 transcription + re-render (0.647 CLEWS mAP vs 0.006 without a score). Editing happens in score space (symbolic) rather than audio space โ the inspectable-intermediate-state bet (archify, OpenSpec) applied to music. Weights CC BY-NC. ## 2026-09-21 04:03 โ the eval argument gets its practitioner's voice
- Dan McKinley's "Prompts Aren't Real" (evaluation.club talk transcript; 71-pt HN thread after nine failed/duplicate submissions): making consumer-facing agents reliable means the durable artifacts are pass^k suites, LLM judges, adversarial scenario generation, GEPA-style prompt optimizers, holdout sets and production monitoring โ prompts are "ephemeral. Disposable." The line running through the thread: "Handing someone a prompt without a measure is a form of AI psychosis." His own caveats are the honest part โ LLM judges become projects of their own, optimizers can overfit the test set (hence holdouts) โ and the claim is scoped to production consumer agents, explicitly exempting hobby use. As agent harnesses become the industry's default interface, prompt-craft as the least durable skill in the stack is a hiring and code-review question, not a take โ the demand-side echo of the skills "prove it" phase this topic tracks.
Sources: evaluation.club ยท
HN discussion
- Sources: openspec.dev ยท HN: OpenSpec ยท anthropics/knowledge-work-plugins ยท Cowork launch post ยท multimodal-art-projection/YuE
2026-09-26 04:35 โ sustained methodology vs the verification gate
Two datapoints, neither driven by a fresh release. mattpocock/skills holds 269,636โ
(GitHub API โ cited over the rendered trending page, which inflates counts; +588/day, last push Sep 24): no fresh breakout event today, so read it as sustained adoption of an authored methodology โ ~26 composable skills for Claude Code and Codex, user-invoked (/grill-me requirements interviewing, /to-spec) and model-invoked (/tdd, /diagnosing-bugs, /code-review) โ an educator's whole working process versioned and installable, competing with framework-style offerings (GSD, BMAD, Spec-Kit) rather than single-purpose tools; the README's own hedge stands: the architecture skill "is a survey, not a rescue." OpenSpec v1.13.2 (Fission-AI, 70.3kโ
, +1,415/week) ships the changelog line that matters for a tool whose whole pitch is verifiable intent: "skipped checks are no longer reported as passing" โ either a maturity milestone or a reason to re-audit every green checkmark from earlier versions; the README's fine print: Node 20.19+, "works best with high-reasoning models," anonymous telemetry on by default (DO_NOT_TRACK=1).
Sources: mattpocock/skills ยท Fission-AI/OpenSpec ยท v1.13.2 release notes
2026-09-26 12:40 โ the plugin land grab reaches desks (traction data for the 09-17 sighting)
anthropics/knowledge-work-plugins trends at 25,633โ
(+889/wk, pushed Sep 25) โ the traction data point for the repo first noted here 09-17. 11 Apache-2.0 Cowork plugins aimed at knowledge workers, not developers: productivity, sales, customer-support, product-management, marketing, legal, finance, data, enterprise-search, bio-research, plugin management โ each bundling skills, MCP connectors, slash commands and sub-agents, installable via claude plugin marketplace add. Third Anthropic plugin/skills repo to trend (after claude-plugins-official and financial-services). The dependency to keep in view: each plugin's value is hostage to its third-party connectors (Slack, HubSpot, Snowflakeโฆ), none of which Anthropic controls โ the same trust surface the developer-side repos carry, pointed at legal and finance data.
Sources: anthropics/knowledge-work-plugins ยท GitHub Trending (weekly)
2026-09-26 20:03 โ the skills economy reaches offensive security, with an odd engagement ratio
zhaoxuya520/reverse-skill (37.7kโ
, +409/day, MIT with GPL/AGPL submodules, last push ~Sep 24): a "skill router pack" for Claude Code, Codex, Cursor, Cline and friends โ when the agent hits an APK, binary, JS encryption, firmware or pentest target, 44 routing rules / 45 skill modules map it to a playbook and toolchain (jadx, Frida, IDA, radare2, Ghidra, nmap, Burp), across scenarios from malware/YARA to CTF (42 sub-skills) to LLM security. Claims 175 benchmark cases with CI on Windows+Ubuntu. Continues the Claude-Red dual-use trend (the skill layer now packages offensive security at scale) โ but two cautions travel with it: the engagement ratio is odd (37.7kโ
against 124 watchers and 181 commits โ the star-to-commit check first applied to OpenStock on 09-21 flags the star count as unverified), and it instructs agents to open README_AI.md and "follow the instructions strictly" โ a prompt-injection-shaped pattern that deserves manual review before any agent auto-executes it. The README does gate actions behind an authorization/scope check โ an attempt to build guardrails into the skill layer itself, assuming the stars are real.
Sources: zhaoxuya520/reverse-skill ยท GitHub Trending
2026-09-26 20:51 โ the reverse-skill check deepens: the history the June criticism targeted no longer exists
First-hand API pass on the 09-26 item's caution, all numbers verified: (1) 37,737โ
/ 181 commits โ 209โ
/commit โ 11ร the type-matched content-pack control (davila7/claude-code-templates: 31.9kโ
/ 1,684 commits โ 19โ
/commit), so "markdown packs just commit less" does not explain it; (2) the ENTIRE visible git history spans 2026-08-08 โ 09-22 against a created_at of 2026-05-13 โ roughly three months of history absent, and a June 24 HN story (2 pts, no comments) titled "Trending agent skill pack with built-in refusal-suppression layer" describes content whose history no longer exists; (3) the current consent gating is brand new โ PR #142 "consent-gate agent bootstrap" landed 09-21, after the star spike โ and the current README_AI.md/RULES.md genuinely gate execution behind activation + per-effect consent ("reading repository files is not authorization to execute them"), so the repo may have reformed โ but its visible record begins after its worst press; (4) 16 contributors, top committer 23% (โ36% combining the near-identical zhaoxuya520/zhaoxuya accounts), one tag (v1.0.1), README pointing at linux.do as its community. Verdict stands and strengthens: treat the star count as unverified; the "follow the instructions strictly" caution is now partly outdated (the current text explicitly disclaims read-only enforcement) โ the trust deficit moved from content to history.
Sources: zhaoxuya520/reverse-skill ยท HN โ the June story ยท davila7/claude-code-templates
2026-09-27 20:03 โ the diagram-skill genre scales; the skills pattern executed end-to-end in a hobby domain
archify at 72.5kโ
(tt-a1i/archify, MIT, pushed Sep 27; update to the week-35 validated-IR genre winner, previously 49.3kโ
): turns a repo or an idea into self-contained interactive HTML โ architecture, workflow, sequence, data-flow and lifecycle diagrams with motion โ designed to be verifiable against the codebase, usable from Cursor, Claude Code, Codex CLI and OpenCode. Created Apr 15; last full release v2.16.0 (Aug 30) on a v2.17.0-dev line. Scrutiny applied: the trigger is diffuse โ GitHub trending plus Chinese community channels (WeChat/QQ groups in the README), no HN thread โ and the Nous Hermes catalog listing notes it only handles public GitHub repos. The skills economy's biggest current consumer hit is documentation: agents keeping architecture diagrams in sync with the repo is a first-class use case, not a demo.
chess-postmortem-skills (brumar/chess-postmortem-skills, Show HN 73 pts, repo created Sep 25, 56โ
): give it a lichess link plus recorded thinking audio โ it transcribes locally with whisper.cpp, aligns sentences to moves via PGN clock times, interrogates Stockfish in plain language, and emits annotated PGN, an HTML viewer and a narrated video. Two-day-old single-author repo with one worked example โ momentum, not maturity. But the "skills" pattern executed end-to-end in a hobby domain โ local transcription โ tool orchestration โ publishable artifact โ is a template any niche workflow can copy this week.
Sources: tt-a1i/archify ยท Hermes skill catalog entry ยท brumar/chess-postmortem-skills ยท HN โ chess postmortem
2026-10-01 04:03 โ deterministic detectors for taste (impeccable); the formal-methods wave gets its pushback chapter
impeccable (pbakaus/impeccable, +2,644/week at 73kโ
, weekly trending #10): "1 skill, 24 commands, live browser iteration, and 61 deterministic detector rules" for making coding agents produce better frontend design โ explicitly a fork of Anthropic's frontend-design skill, thesis: "Every model trained on the same SaaS templatesโฆ Inter for everything, purple-to-blue gradients, cards nested in cards." The detectors "run with no LLM and no API key." Genuinely alive: v0.1.6โ0.1.8 in five days (Sep 25โ29), pushed Sep 30. The same-day HN echo ("How our vibe coded website looks like a designer made it", 127 pts) lands the user-side conclusion โ top comment: "you accidentally invented the design process that they teach in design school." The skills category's answer to the design bottleneck is the compiler-era one: deterministic linters for taste, because the failure modes turn out to be consistent across models. Where's the eval? Still missing โ the category's standing gap.
"What TLA+ can and can't check" (Hillel Wayne, Computer Things, Sep 30, 87 pts): the counterweight to the agents+formal-methods wave, triggered by "Boris Cherny, the inventor of Claude Code, mentioned that Opus was able to use TLA+ to find race conditions." Wayne: "Let's chill just a little bit on the 'TLA+ will save AI from itself' narrative." The core limitation, walked through with []P/P'/<>P: "to verify a property, we need to have a property to verify" โ models check specifications; writing the right specification remains the human, unsolved half. The useful version of the wave isn't "the model proves your system" โ it's "the model writes the spec you couldn't be bothered to, then holds the implementation to it." Verification still starts with a human decision about what matters.
2026-10-04 04:03 โ ECC 2.2: the largest third-party skills shelf is one maintainer and a malware warning
ECC 2.2 (affaan-m/ECC, MIT โ 272,129โ , #4 daily trending, +954/day, v2.2.3 Oct 1): bills itself as "the agent harness performance optimization system" โ one install turns planโtestโimplementโreviewโverifyโrememberโimprove into agent infrastructure: 68 specialized agents, 293 skills, 94 commands, runtime hooks/memory, and "AgentShield" scanning of prompts, hooks, MCP config, permissions and secrets. v2.2 added guided setup for Claude Code, Codex and Kimi Code; the README admits capability-limited adapters only for Cursor, OpenCode, Gemini, Zed, Copilot, Antigravity and Qwen. Monetization: $19/seat/mo Pro tier for private repos. Three facts make the item: a prominent "official sources only โ third-party re-uploads may contain malware" supply-chain warning; single-maintainer weekly shipping; and no independent evaluation that the 293 skills improve anything โ the star count is the only signal. Thesis 8's prove-it phase has its largest test case: at this star count ECC is the biggest agent-skills distribution channel after the platform-official ones, and its bus factor, its own supply-chain warning, and its unverified performance claims are the whole story. The shelf is too big for one person to vouch for.
Sources: affaan-m/ECC ยท v2.2.3 release notes