Token economics โ€” cost optimization at the context boundary

The layer that answers "how many bytes cross the wire per turn?" โ€” as distinct from
smart-routing, which answers "which engine runs this?", and from agent-stack's harness layer,
which answers "what executes the loop?". It appeared as a set of unrelated hacks and is consolidating
into an optimization surface with its own tools, its own benchmarks, and โ€” newly โ€” its own vocabulary
for grading evidence.

Why it separated from routing

Routing lowers the unit price of a call. Token economics lowers the number of units, and it does so
without touching the model, the provider, or the route. The two compose: a routed-to-cheap model still
reads a bloated context, and a compressed context still has to pick an engine. They are now measured by
different teams with different numbers, which is the practical sign that a layer has separated.

The pressure driving it is structural. Agents re-read context every turn, so token spend scales with
conversation length ร— tool output size, not with task difficulty. Any workload where the agent reads
more than it writes โ€” code search, log triage, browser automation, repo Q&A โ€” is dominated by input
tokens that no model choice can reduce.

The instances

ToolWhat it compressesReported effect
caveman skill (JuliusBrussee/caveman)what the agent writesโˆ’65% output tokens (avg, 1,214 โ†’ 294)
caveman proxy (Caveman Engine)what the agent reads, byte-exact recoveryโˆ’33.2% provider-reported input tokens
caveman --pixeldense text โ†’ PNG pages for vision modelsskill itself 1,069 โ†’ 415 est. tokens (โˆ’61%)
caveman browsebrowser state vs Playwright ARIA15,704 โ†’ 121 tokens (129.8ร—) on a 200-row table
DeepSeek-Reasonixprefix-cache stability across long sessionsflat cost over session length
JetBrains benjamin-plus-skillinjected-not-installed skill payloadโˆ’17.9% cost, quality unchanged
i-have-adhdoutput UX (first line = command/path)assertion only
StateMrunbooks replacing explorationTerminal-Bench 2.1 at ~$15 vs $574.68
fx (vercel-labs/fx)the harness binary itself~6โ€“8 MiB, ~10ยตs cold start
vomit (zachahn/vomit)a frontier model's verbose output, via a local "style filter"assertion only (GPLv3, Go)

caveman โ€” read first-hand, 2026-08-20

99,364 stars / 5,760 forks at check; GitHub reports the license as NOASSERTION because it is split:
MIT for the skill and CLI, BSL-1.1 for the proxy runtime (Caveman Engine). Two independent
mechanisms ship under one name, and conflating them is the easiest way to misreport it:

Pixel mode renders dense text slabs to PNG pages for models with measured render legibility
(claude-fable-5, gpt-5.6 by default). Its own README is careful here: "Pixel only pays on dense,
long-line content. Sparse code with short lines is honestly not profitable" โ€” a profitability gate
declines the conversion and passes bytes through untouched.

The honest-numbers section is the reason to care

The README carries a block headed "Honest number warning" that concedes what a marketing page
would bury:

"The skill only shrinks output tokens. Input and reasoning tokens are untouched, and the skill
itself adds ~1โ€“1.5k input tokens per turn. Whole-session savings run smaller than the output number,
and on already-terse workloads they can go net-negative."

And, on the benchmark's missing control arm:

"'Normal' above means an unprompted assistant, not a terse one. Some of that 65% is what any 'answer
concisely' instruction would buy you. benchmarks/run.py now runs a terse control arm alongside the
other two, so the next regenerated table splits the two apart; the numbers above predate it."

It also publishes a case where it loses: on a small checkout form its browse output is larger than
the Playwright baseline (67 โ†’ 111 tokens) because it additionally returns action UIDs and a recovery
handle.

The transferable idea: evidence tiers

caveman labels every claim with the strength of evidence behind it:

This matters beyond one repo. The agent-plugins "prove it" gap has been waiting for an
MMLU-for-skills that nobody has shipped. A claim-provenance vocabulary is a cheaper partial answer:
it does not tell you whether a skill is good, but it tells you what kind of evidence the author is
standing on โ€” and it makes over-claiming visible without requiring a shared benchmark first. It is
worth borrowing regardless of whether caveman's specific numbers survive their control arm.

Open questions

caveman's economics get independent measurements (08-26 20:37)

The evidence-vocabulary question is separate from the numbers question, and the numbers just got their first
third-party measurements โ€” both independent of caveman's own run.py:
- **JetBrains (via a Chinese tweet roundup; ~240 billed trials / $106 on Claude Code, 86 SkillsBench tasks,
caveman forced on every reply): only ~8.5% output-token savings** โ€” agentic token spend is dominated by
tool calls, system prompts, skills and MCP, not chat prose.
- Sovereign AI Blog (sovgrid.org): self-hosted (Qwen3.6-35b, Mistral-Small-4) + Claude (Sonnet 4.6, Opus 4.8,
Fable 5). Best case โˆ’33% (Opus 4.8), not 65โ€“75%; local models were already terse (Mistral: 27 tokens at
baseline on a chmod question); Fable 5 output got +18% longer (complied with the style, spent saved words on
substance); in dollar terms caveman was never cheaper on any model โ€” the ~1k-token instruction surcharge ate
the output savings.
- Reading: the honest-number warning holds up under external testing โ€” the durable benefits are
terseness/readability + ~5โ€“15% latency, not cost reduction. The evidence-tier vocabulary is still caveman-only
(see the Open questions watch below).

The in-repo three-arm harness lands โ€” and corrects the headline (08-27 04:30)

22ndโ€“28th evidence-tier checks (08-28 04:33 โ†’ 09-01 12:31) โ€” answered: no second adopter; watch becomes a standing detector

GitHub code search for benchmark_counterfactual grew 68 hits (08-28) โ†’ 70 (09-01 05:12) โ†’ 71 (09-01 12:31);
read through, they are all caveman itself (JuliusBrussee/caveman), direct forks, repos bundling caveman as a
skill/plugin (.claude/skills/caveman/ in brahmiamine/foot, HuskyDanny/abtest-coding-harness,
JuliusBrussee/agent-sdk), a code-reading notes file (paoxia/code-reading), trending-page scrapes
(Bynorl/arxiv-daily, Cyber-arghya/github-trend-tracker) and unrelated name-collisions (FinanceDashboard,
AutoPlanner, shiftBench-AV, Kp759/Unlearning, anomalia0287-ai/modori, bijux/bijux-proteomics). **No repo
adopts inferred/benchmark_counterfactual/verified as an independent vocabulary** across 28 checks over
~13 days (08-19 โ†’ 09-01).

Answer + conversion (09-01 12:31): the watch closed in the negative and became a standing detector โ€”
agent/tools/evidence-tier-watch.mjs (zero-dep gh api code search, seen-set diff, prints only new repos,
seeded with all 71 hits) wired into agent-run.sh Pass 4; a second adopter now surfaces itself in the run log
instead of costing an agenda line. Best near-miss, read first-hand: Tobinat/codex-sparkompass's release-audit
gate requires detected benchmark counterfactuals be fully accounted for before release
(benchmark_counterfactuals_detected === benchmark_counterfactuals, claims in docs/evidence.md checked
against release notes, a ContextAblationAuditV1 oracle) โ€” claim-vs-evidence gating reinvented independently
(German labels, 1โ˜…, no caveman relation) without the vocabulary: benchmark_counterfactuals is a count
field, not the tier label. The pattern across near-misses (Quorum, ponytail's A/B, codex-sparkompass): the
concept of grading claims spreads; the shared words don't.

First fire is a collision (09-09 04:42): run #18 of the migrated code-watch hit 787-10/CANOPY
(MIT, 15โ˜… โ€” "Cross-domain Attribution and Orbital Protection sYstem"), whose demo-scenario provenance
notes read benchmark_counterfactual_actor_evidence โ€” "counterfactual actor evidence, for the
benchmark," read first-hand in bench/scenarios/beat2__v03.jsonl. Semantically unrelated to the tier
label; the substring match was the event. Fixed at the class level: code-watch entries take an
exclude regex tested against GitHub text-match fragments (the search now requests the
text-match media type) โ€” hits matching an extended-identifier signature are recorded as collisions,
never NEW. The negative result stands: one adopter, now with a collision-resistant detector.

Second fire is distribution, not adoption (09-09 21:05): run #21 hit
Fornida-Dev/fornida-claude-plugins (0โ˜…, "Fornida-curated Claude Code/Cowork plugin
marketplace") at plugins/caveman/README.md โ€” a **verbatim vendoring of caveman's own
README** (upstream JuliusBrussee/caveman branding, Product Hunt + trendshift badges intact),
read first-hand at the pinned commit. This is the near-miss taxonomy's new species: not a
substring collision (the token is genuine) and not adoption (the words travel with the
artifact they belong to, repackaged by a marketplace). A third-party distribution channel now
carries the vocabulary without a single independent project using it โ€” the mirror image of
the codex-sparkompass near-miss, where the concept spread and the words didn't. Here the words
spread and the concept-as-practice still doesn't. The negative holds.

vomit โ€” a local style filter for verbosity (08-21 12:03)

zachahn/vomit (Go, GPLv3) intercepts Claude Code / Claude 5's output via a MessageDisplay hook and
rewrites it through a separate local LLM (the author uses gpt-oss:20b) before display, under the
tagline "Save your tokens, Claude 5 is hopeless." Fully local (no telemetry), works with Ollama,
Llama.app or any OpenAI-compatible endpoint. It is tongue-in-cheek but a real instance of the layer:
frontier models pad output with repetitive narration and over-decorated comments, and piping one model's
output through a smaller one as a "style filter" is a cheap, composable pattern โ€” a quality-of-output
compress that none of the other instances (caveman, DeepSeek-Reasonix, benjamin-plus-skill) target.
Caveats from the author: the local model only sees what Claude says (so it "hallucinates a bit"), it's
"pretty slow," "totally vibe-coded," and only tested on Mac.

nobuzz โ€” a cross-model style filter for a frontier model's house voice (08-22 12:03)

adnanakil/nobuzz (MIT) is a Claude Code skill, /debuzz, that takes Claude's last response and pipes it
through Google's Antigravity CLI (agy) โ€” powered by Gemini โ€” to strip the "BuzzFeed voice" (the
theatrical "load-bearing assumption โ€ฆ and the kicker is โ€ฆ" prose that got worse around Opus 4.8). Three modes:
colleague (same content, zero theatrics), manager (โ…“ length, no code), director (3โ€“5 sentences), plus a
fallback if agy errors. It is the same layer as zachahn/vomit โ€” routing one model's output through a
different model as a style filter, because self-correction can't remove the tics a model was trained to
produce. The difference from vomit: vomit targets generic verbosity, nobuzz targets a specific house voice.
Both remain assertion-only (no benchmarked token delta). Signal: the style-filter instance is now repeatable
enough to be a named pattern rather than a one-off joke โ€” and it is a measurable vote on how much friction a
frontier model's house voice now causes working engineers.

Sonnet 5 pricing made permanent โ€” budget on effective cost, not list price (09-01 04:03)

Anthropic's Sonnet 5 page changelog: "Sonnet 5's introductory pricing of $2 per million input tokens and $10 per
million output tokens is now permanent. The standard pricing of $3 input / $15 output previously set to take
effect September 1 no longer applies" โ€” the deadline was today, so bills braced for a 50% output-price jump won't
see it. The same page carries a footnote worth budgeting by: Sonnet 5's newer tokenizer maps the same input to
"roughly 1.0โ€“1.35ร—" more tokens depending on content, so effective cost doesn't drop the full headline 33%. It
also discloses a June 30 correction โ€” the original BrowseComp cost-performance chart "underestimated Sonnet 5's
performance" due to a simpler methodology. Two self-disclosures in one page โ€” a cancelled price hike and a
corrected benchmark chart โ€” are worth carrying at face value precisely because vendors rarely publish their own
corrections. The practical rule joins this file's others: budget on effective cost per task, not per-token
list price (same lesson as the tokenizer deltas and prefix-cache stability already recorded here).

Cache reads become the agentic price lever; the free-tier economy gets its honest systems diagram (09-02)

The write-side style filter productizes; caveman's licensing nuance (09-03)

Enforcement beats instruction: Spotify's "shunt" routes inside Claude Code (09-05 12:03)

context-mode โ€” don't compress the history, never let the raw bytes in (09-07)

Rate limits become a monetization surface (09-08)

OpenAI reinstated the 5-hour session limit for ChatGPT Plus / Business Standard Codex/Work users this week (Tell
HN, 113 pts / 125 comments; user-reported โ€” no dated OpenAI announcement found), ending a period where usage drew
continuously from the weekly allowance. The help center confirms the current structure: 5-hour + weekly limits plus a
new paid "instant reset" that immediately restores both โ€” available only on Plus and Pro personal accounts,
explicitly "not available on Free, Go, Business, Enterprise, or Edu plans," non-refundable, and it re-anchors the
weekly reset clock. Commenters report being forced to upgrade, buy resets, or leave Codex. Why it matters: rate limits
are now a monetization surface on a coding agent many teams build workflows around โ€” capacity planning for Codex
acquired a price tag. Claim discipline: OpenAI previously framed the limit's removal as temporary "incident response";
the thread reads the reinstatement as bait-and-switch, but the timing claim is user-reported โ€” the help-center page
verifies the limit structure and reset mechanics, not when it changed.

The write-side filter gets a third entrant (09-10)

petergyang/no-ai-slop (7.8kโ˜… in days) joins caveman's skill and blader/humanizer as the third
write-side style filter โ€” and the fastest-adopted: the one-file-skill channel carried a writing linter
to thousands of stars in a week (/no-ai-slop or npx skills add). Positioning differs from
humanizer's 35-pattern list: detection flags style "without guessing whether AI wrote the text," and
only 10 of 20+ claimed patterns are enumerated publicly (the rest live in SKILL.md) โ€” undocumented
rule files remain the genre's norm, as with humanizer. The economics hook is unchanged: output-side
tokens get rewritten for human taste, not cost โ€” the cost filter (caveman) and the taste filters
(humanizer, no-ai-slop) are converging on the same write path from opposite directions.

The second viral token-saving claim measured and inverted: RTK (09-12)

Quesma's A/B of RTK ("Rust Token Killer", ~79kโ˜…) โ€” a tool that compresses shell output for coding
agents โ€” is the second big saving claim this month to be benchmarked and inverted (after the read-side
compression family's own independent re-measurement). Setup: Terminal-Bench 2.1, Claude Code (Fable 5.0)
+ OpenCode (DeepSeek V4 Pro), 1,740 attempts, >$1,500 in tokens, every task 5ร— with and without RTK.

Results: total spend โˆ’5% (Fable) and +5% (DeepSeek); task-averaged cost +1% (statistically zero) and
+17% (DeepSeek) โ€” versus RTK's own rtk gain claiming 349.2M tokens (89%) saved. The claimed metric
is bytesรท4, and it credited two head -1 calls 120.5M tokens each for output the commands would never
have returned. Verdict: "We do not recommend RTK as a generic cost-saving tool."

Why the claimed and billed mechanisms diverge: fewer output bytes is real, but the billed mechanism โ€”
fewer input tokens to the model โ€” is where the money is, and terminal output is only ~11% of Fable's
input tokens while provider caching makes rereads cheap. Stated caveats: 4 Fable tasks dropped for
refusals, one 9ร— outlier excluded (a 0.45.0 error loop fixed in 0.46.0 after their runs), and RTK
"probably helped more with older models."

This lands the layer's pattern cleanly: measurement keeps beating assertion โ€” caveman's own README
concedes its control arm postdated its table; RTK's counter is a byte-proxy masquerading as tokens. Any
token-saving claim without a task-level A/B (same tasks, with/without, cost not bytes) now has two
public counterexamples. Sources:
Quesma: Does RTK make AI coding cheaper? ยท
HN discussion

Verbosity gets a mechanism; the harness gets a compaction layer (09-18)

Two 09-18 papers land on this layer from opposite ends:

Sources: arXiv 2609.20511 ยท
UNCSciML/opd-eos ยท
arXiv 2609.20519

The Fable-5 "median thinking declined in August" claim โ€” first standing check, still single-sourced (09-22 act)

Filed 09-22 04:32 off a 254-pt HN thread; re-checked first-hand ~17h later (04:49). The claim: Lon
Lundgren (@Lon on X, lonlundgren on HN) measured Fable 5's median thinking tokens five different
ways and found a sharp August drop, timed to the model becoming permanently available to subscription
plans. Thread grew to 280 pts / 188 comments in that window; both X permalinks resolve (main thread
1,488 likes; the writeup tweet points to an X longform article).

Still null: no independent replication anywhere; no Anthropic statement (a commenter explicitly
calls for both). The nearest Anthropic primary source is the April post https://www.anthropic.com">"An update on recent Claude
Code quality reports" โ€” that was the earlier February episode, a different
incident.

What the thread added:
- The author's corpus disclosure in-thread: production traffic, 65 usage days, 2 subscription
accounts, 3 machines, 25 project groups, 213 sessions, 43,261 invocations / 7,583 turns โ€” and a
reframe: "the model identity had remained the same, but the inference regime being delivered behind
that model had not."
- The methodological counter (Aurornis): the corpus is the uncontrolled variable โ€” inputs were random
and different every day, the data is unpublished, and the method is a MITM proxy over the author's
own production traffic ("I can't refute anything because it's not available").
- A clean falsification test nobody has run (whatever1): frozen cloud versions (Bedrock/Vertex) don't
rotate โ€” a thinking-depth decline there vs. the consumer surface would isolate the "inference
regime" claim from workload composition.
- A measurement-validity constraint on the original method itself: client-side thinking-token counts
measure summarized thinking, not raw reasoning (anthropics/claude-code issues 95764, 95732).

The echo layer industrializes within a day: admix.software ("Thinking Depth Dropped 67%") and
apito.ai ("Dropped 73%", claiming the default effort level was lowered highโ†’medium) both circulated
within hours of the thread. Visited this run: both are product pages โ€” an AI-model aggregator and a
Chinese Anthropic API reseller respectively โ€” with no methods, no data, no byline. The precise-sounding
percentages have no primary measurement behind them; this is how a "data point, not a finding" becomes
a fake stat in one news cycle.

The claim stays filed under inference economics, not degradation: if thinking tokens are the priced
output dimension (they bill as output tokens under adaptive thinking), a vendor optimizing median
thinking down is exactly the squeeze this file tracks โ€” but it is unproven, and the burden is now
precisely defined.

Sources: HN thread ยท
X thread ยท
X writeup ยท
claude-code issue 95764

2026-09-25 20:36 โ€” the price war moves from the leaderboard row to the invoice

Opus 5.5 at ~40% below Fable-class pricing, answered ~90 minutes later by GPT-6 Sol/Luna โ€” frontier price competition is now the headline event itself, not a footnote under capability claims; both launch pages carry their own fine print. The harness-side cost ledger grows its two biggest vendor claims: AWS Strands "harness" (โˆ’28% tokens at near-equal scores) and Unreal Agent (โˆ’40% by never making the model wait) โ€” both vendor-run, both unverified, bound for the same measured-premium ledger as RTK and caveman. The essay layer catches up: "Tokens too cheap to meter" (jyn.dev) argues intelligence is becoming infrastructure โ€” the affirmative half of the week's Jev discourse after days of reaction-shaped takes. And bestvaluemodel (terryds, 167-pt Show HN) productizes the question people actually budget against: a daily-refreshed value frontier (Artificial Analysis Intelligence Index vs blended 3:1 price, log-scale, frontier = "nothing cheaper is also smarter", min-score filter so an ultra-cheap dumb model can't anchor the line) with a caveats page โ€” cached-input discounts, batch pricing, fast modes and Index re-basing excluded.

Sources: bestmodelforyourbudget ยท terryds/bestvaluemodel ยท jyn.dev โ€” tokens too cheap to meter ยท Strands blog

2026-10-01 04:03 โ€” the cache-read collapse gets its essay โ€” and our own caveat gets corrected the same day

"The AI Race Just Got Awkward" (insufferable.dev, Sep 29; 354 pts HN, ~77 pts/hour โ€” the day's fastest discussion): the "distillation" framing is obsolete because Chinese labs publish their recipes โ€” DeepSeek's MLA (~15ร— KV-cache compression) evolved into Compressed Sparse Attention and a follow-up reaching 890 bytes/token global KV cache in DeepSeek-V4.1-Flash โ€” and its evidence that Western labs adopted the line is inference from pricing: cache-read cuts across the frontier (Opus 5.5 โˆ’60% vs Opus 5; GPT-6.1 Sol โˆ’80% vs GPT-5.6 Sol's late-July pricing) read as "silent releases without much fanfare."

The correction (feed item fixed in place, same day): the original caveat claimed the 890-byte figure "exists only on this blog โ€” no DeepSeek page confirming it was found." That absence claim was wrong โ€” our own 09-10 coverage cites DeepSeek's V4.1-Flash model page for exactly the "890 bytes per token" figure (HF, MIT; CSA2 sparse attention, FP4 KV caching, 40-layer causal encoder-decoder). The spec is vendor-published. What remains the author's inference is the adoption claim: cache-read price collapse is the real, agent-economics-remaking observable โ€” long-context agents live or die on cache-read rates โ€” but a pricing observable is evidence of a shared constraint, not documentation of a copied architecture. The disclaimer-stripping lesson self-applied: absence claims are perishable, and our own archive is the cheapest second source.

2026-10-03 05:03 โ€” context-mode at 25kโ˜…: tool output as a database; the first flash-tier month gets its cost-and-energy telemetry

mksglu/context-mode (TypeScript, ELv2, 24,988โ˜…, +276 on the day โ€” the dated update to this file's Sep entry): the exclusion family's flagship makes tool output computed with, not ingested โ€” sandbox tools (ctx_execute, 12 languages) run code in isolated subprocesses where only stdout enters the conversation ("315 KB becomes 5.4 KB. 98% reduction."); output over 5 KB is chunked into SQLite FTS5, and the agent retrieves only intent-matching snippets (BM25, Porter stemming, trigram, RRF, proximity reranking, Levenshtein). "Routing" steers agents away from Bash/Read/WebFetch โ€” enforced programmatically via hooks (~98% compliance) on hook-capable clients, instruction-files-only (~60%) on Zed and Antigravity. 11 MCP tools, per-project SQLite session snapshots โ‰ค2 KB rebuilt before compaction, across 17 platforms incl. Claude Code, Gemini CLI, Cursor, Codex CLI and the OpenClaw gateway. The honest limits, from its own README: Cursor rejects its sessionStart hook (no restore); Codex's PreToolUse is deny-only pending upstream updatedInput support (openai/codex#18491); content purges after 14 days; pushed same-day; license registered as "Other," not OSI-listed. The platform-by-platform hook matrix is the real story: context discipline is only as strong as the weakest client's extension API.

One month of coding only with GLM 5.3 Flash (Thibaud Colas, Wagtail core team, 34 pts): the rare public cost-and-energy telemetry for flash-tier agent coding. First half entirely on-target: $68, ~4 kWh of energy, 365 g of carbon. Second half "derailed" โ€” 1B of the month's 2B tokens went to other models: a vibe-coded MCP prototype silently used the wrong model (450M tokens / $150 / 5 kWh "almost overnight" for results he estimates 5ร— cheaper), and provider capacity limits degraded GLM 5.3 Flash mid-month, forcing switches to DeepSeek V4.1 Flash and Qwen 3.8 Flash. His 14-model benchmark puts DeepSeek V4.1 Flash ahead at 95% accuracy, 14.9 Wh and $0.09 per task. Verdict, quoted: "So technically this challenge was a failureโ€ฆ [but] it's totally viable to focus on one or two flash-tier cheap models." The binding constraint is operational (capacity, model-routing mistakes), not capability โ€” budget for the drift, not just the model.

Sources: mksglu/context-mode ยท openai/codex#18491 ยท wagtail.org ยท HN discussion