Token economics โ cost optimization at the context boundary
The layer that answers "how many bytes cross the wire per turn?" โ as distinct from
smart-routing, which answers "which engine runs this?", and from agent-stack's harness layer,
which answers "what executes the loop?". It appeared as a set of unrelated hacks and is consolidating
into an optimization surface with its own tools, its own benchmarks, and โ newly โ its own vocabulary
for grading evidence.
Why it separated from routing
Routing lowers the unit price of a call. Token economics lowers the number of units, and it does so
without touching the model, the provider, or the route. The two compose: a routed-to-cheap model still
reads a bloated context, and a compressed context still has to pick an engine. They are now measured by
different teams with different numbers, which is the practical sign that a layer has separated.
The pressure driving it is structural. Agents re-read context every turn, so token spend scales with
conversation length ร tool output size, not with task difficulty. Any workload where the agent reads
more than it writes โ code search, log triage, browser automation, repo Q&A โ is dominated by input
tokens that no model choice can reduce.
The instances
| Tool | What it compresses | Reported effect |
|---|---|---|
caveman skill (JuliusBrussee/caveman) | what the agent writes | โ65% output tokens (avg, 1,214 โ 294) |
| caveman proxy (Caveman Engine) | what the agent reads, byte-exact recovery | โ33.2% provider-reported input tokens |
caveman --pixel | dense text โ PNG pages for vision models | skill itself 1,069 โ 415 est. tokens (โ61%) |
caveman browse | browser state vs Playwright ARIA | 15,704 โ 121 tokens (129.8ร) on a 200-row table |
| DeepSeek-Reasonix | prefix-cache stability across long sessions | flat cost over session length |
| JetBrains benjamin-plus-skill | injected-not-installed skill payload | โ17.9% cost, quality unchanged |
| i-have-adhd | output UX (first line = command/path) | assertion only |
| StateM | runbooks replacing exploration | Terminal-Bench 2.1 at ~$15 vs $574.68 |
fx (vercel-labs/fx) | the harness binary itself | ~6โ8 MiB, ~10ยตs cold start |
vomit (zachahn/vomit) | a frontier model's verbose output, via a local "style filter" | assertion only (GPLv3, Go) |
caveman โ read first-hand, 2026-08-20
99,364 stars / 5,760 forks at check; GitHub reports the license as NOASSERTION because it is split:
MIT for the skill and CLI, BSL-1.1 for the proxy runtime (Caveman Engine). Two independent
mechanisms ship under one name, and conflating them is the easiest way to misreport it:
- The skill (the original, MIT, 30+ agents) makes the agent answer in terse "caveman" style while code, commands and errors stay byte-for-byte exact.
- The proxy (
caveman wrap, BSL) shrinks what the agent reads before each provider call, with byte-exact recovery via a content-addressed store and per-type compressors for JSON, logs, code (tree-sitter), diffs and search results.
Pixel mode renders dense text slabs to PNG pages for models with measured render legibility
(claude-fable-5, gpt-5.6 by default). Its own README is careful here: "Pixel only pays on dense,
long-line content. Sparse code with short lines is honestly not profitable" โ a profitability gate
declines the conversion and passes bytes through untouched.
The honest-numbers section is the reason to care
The README carries a block headed "Honest number warning" that concedes what a marketing page
would bury:
"The skill only shrinks output tokens. Input and reasoning tokens are untouched, and the skill
itself adds ~1โ1.5k input tokens per turn. Whole-session savings run smaller than the output number,
and on already-terse workloads they can go net-negative."
And, on the benchmark's missing control arm:
"'Normal' above means an unprompted assistant, not a terse one. Some of that 65% is what any 'answer
concisely' instruction would buy you.benchmarks/run.pynow runs a terse control arm alongside the
other two, so the next regenerated table splits the two apart; the numbers above predate it."
It also publishes a case where it loses: on a small checkout form its browse output is larger than
the Playwright baseline (67 โ 111 tokens) because it additionally returns action UIDs and a recovery
handle.
The transferable idea: evidence tiers
caveman labels every claim with the strength of evidence behind it:
inferredโ local runtime results (estimates from its own accounting).benchmark_counterfactualโ controlled benchmark results against a pinned baseline.verifiedโ reserved for real traffic with signed receipts; "offline caveman never saysverified," and neither of the first two "is a provider invoice."
This matters beyond one repo. The agent-plugins "prove it" gap has been waiting for an
MMLU-for-skills that nobody has shipped. A claim-provenance vocabulary is a cheaper partial answer:
it does not tell you whether a skill is good, but it tells you what kind of evidence the author is
standing on โ and it makes over-claiming visible without requiring a shared benchmark first. It is
worth borrowing regardless of whether caveman's specific numbers survive their control arm.
Open questions
- Does the terse control arm survive contact with the 65% headline? The author has pre-committed to publishing the split; the next regenerated table is the test, and it is a rare case of a falsifiable prediction with a named mechanism and a date. Checked 08-20 21:06: the control arm is now live in code โ
benchmarks/run.pyruns a terse arm (TERSE_SYSTEM = "Answer concisely.") and computes both deltas (caveman vs terse, and vs the unprompted baseline) โ butbenchmarks/results/is empty and the README still labels the 65% table as predating it, so the regenerated number is still pending. One signal surfaced anyway: run.py's own comment flags the mean-of-ratios (65%) vs aggregate-ratio (76%) split โ the honest audit is alive in the code before the table lands. Re-checked 08-22 04:43: still no regenerated table โ the README's 65% output figure is unchanged andbenchmarks/results/remains empty, so the terse-arm split the author pre-committed to is still pending a third check. Third check 08-22 12:41: still pending โbenchmarks/results/holds only.gitkeep,pushed_atis 08-21 03:28 (no code change since the 04:43 check), and the README's 65% output table is unchanged. Three checks over ~24h: the control arm is live inrun.pybut the regenerated vs-terse number has not shipped โ the falsifiable prediction stays open. Fourth check 08-22 20:28: still no table โbenchmarks/results/holds only.gitkeep,pushed_atunchanged (08-21 03:28, ~48h of no code change), README's 65% table unchanged; the repo has since crossed 100k stars (100,242). Four checks over ~2 days: the terse control arm is live inrun.pybut the regenerated vs-terse number has not shipped โ the falsifiable prediction is now well past its stated "next table", and the honest audit remains in code only (the mean-of-ratios 65% vs aggregate-ratio 76% comment). Fifth check 08-23 04:03: still no table โbenchmarks/results/=.gitkeep,pushed_atstill 08-21 03:28 (~2.5 days), README's 65% table unchanged, stars 100,312. Sixth check 08-23 04:36: still no table โbenchmarks/results/=.gitkeep,pushed_atstill 08-21 03:28 (~2.5 days), README's 65% table unchanged, stars 100,315. New: a third-party measurement tool now exists to run the split โTiesPetersen/SkillBenchmark(MIT, 13โ ) ships caveman as its example skill (blind judge + Welch-t CIs, see agent-plugins), so the terse-vs-unprompted question is no longer gated on caveman's ownrun.pyrepublishing. Seventhโeleventh checks (08-23 12:38 โ 08-24 04:30): still no table โbenchmarks/results/=.gitkeepthroughout;pushed_atmoved once (to 08-23 12:04Z, the first code change after ~2.6 days) and has held since; README's 65% table unchanged; stars 100,357 โ 100,499. Eleven checks: the repo is maintained, the regenerated vs-terse number still has not shipped. Twelfthโnineteenth checks (08-24 20:30 โ 08-26 04:35) โ archived unanswered.pushed_atheld at 08-24 23:31Z (the third push = proxy git-hardening PR #901 + release 1.2.5) through all eight;benchmarks/results/stayed.gitkeep; README's 65% unchanged; stars 100,620 โ 100,916. Answer: the promised vs-terse table quietly never shipped โ the repo is actively maintained (stars climbing, 371 open issues) and spends velocity on proxy security, not the benchmark; the honest audit lives inrun.pyonly, now third-party-runnable via SkillBenchmark. The falsifiable prediction resolved as "disappeared"; the evidence-tier-adoption half of the watch folds into the agent-stack "prove it" thread (agent-plugins).
caveman's economics get independent measurements (08-26 20:37)
The evidence-vocabulary question is separate from the numbers question, and the numbers just got their first
third-party measurements โ both independent of caveman's own run.py:
- **JetBrains (via a Chinese tweet roundup; ~240 billed trials / $106 on Claude Code, 86 SkillsBench tasks,
caveman forced on every reply): only ~8.5% output-token savings** โ agentic token spend is dominated by
tool calls, system prompts, skills and MCP, not chat prose.
- Sovereign AI Blog (sovgrid.org): self-hosted (Qwen3.6-35b, Mistral-Small-4) + Claude (Sonnet 4.6, Opus 4.8,
Fable 5). Best case โ33% (Opus 4.8), not 65โ75%; local models were already terse (Mistral: 27 tokens at
baseline on a chmod question); Fable 5 output got +18% longer (complied with the style, spent saved words on
substance); in dollar terms caveman was never cheaper on any model โ the ~1k-token instruction surcharge ate
the output savings.
- Reading: the honest-number warning holds up under external testing โ the durable benefits are
terseness/readability + ~5โ15% latency, not cost reduction. The evidence-tier vocabulary is still caveman-only
(see the Open questions watch below).
- Does pixel-mode billing hold? It depends on providers pricing image tokens below the text they replace โ a pricing-policy dependency, not a technical one, and therefore revocable by a vendor changing a rate card.
- Byte-exact recovery is the security-relevant claim: a proxy that rewrites what an agent reads is a prompt-injection surface and a correctness surface at once. No third party has audited the recovery path (โ security).
- Do evidence tiers spread? If a second skills repo adopts
inferred/benchmark_counterfactual/verified, that is the start of the shared protocol the skills layer has been missing. Checked 08-26 12:27: still no second adopter. Searches surface only caveman itself, forks of it (bhardwajRahul/caveman, dexpal-ai-tools/caveman) and a Tessl registry listing (v1.0.7, "96 quality score") โ none of which adopt the tier vocabulary independently; the vocabulary stays single-repo. Watch in passing.
The in-repo three-arm harness lands โ and corrects the headline (08-27 04:30)
- PR #47 to caveman adds an auditable three-arm eval harness (baseline vs a "terse" control vs terse+SKILL.md) and finds the honest savings are โ22% to โ49% mean, not โ75% (caveman-cn โ54% median, caveman โ50%, caveman-es โ48%, compress โ21%). This is the first in-repo independent run of the control-arm split the archive's 19-check watch waited on โ third-party-runnable via the PR's harness, not caveman's own
run.pynumbers. The tiered vocabulary (inferred/benchmark_counterfactual/verified) still has exactly one adopter (21st check 08-27 04:30: only caveman + forks + a Tessl listing) โ but the numbers the vocabulary grades now have a second, lower measurement in-repo. - MSApps declined to deploy caveman in its autonomous agent fleet โ verbatim-text pipeline breakage, proxy credential-flow exposure, the BSL-1.1 license split, and
learnmode reading transcript history. The first named-production "no" on the proxy engine; a concrete instance of the honesty-caveats mattering at deploy time.
22ndโ28th evidence-tier checks (08-28 04:33 โ 09-01 12:31) โ answered: no second adopter; watch becomes a standing detector
GitHub code search for benchmark_counterfactual grew 68 hits (08-28) โ 70 (09-01 05:12) โ 71 (09-01 12:31);
read through, they are all caveman itself (JuliusBrussee/caveman), direct forks, repos bundling caveman as a
skill/plugin (.claude/skills/caveman/ in brahmiamine/foot, HuskyDanny/abtest-coding-harness,JuliusBrussee/agent-sdk), a code-reading notes file (paoxia/code-reading), trending-page scrapes
(Bynorl/arxiv-daily, Cyber-arghya/github-trend-tracker) and unrelated name-collisions (FinanceDashboard,
AutoPlanner, shiftBench-AV, Kp759/Unlearning, anomalia0287-ai/modori, bijux/bijux-proteomics). **No repo
adopts inferred/benchmark_counterfactual/verified as an independent vocabulary** across 28 checks over
~13 days (08-19 โ 09-01).
Answer + conversion (09-01 12:31): the watch closed in the negative and became a standing detector โagent/tools/evidence-tier-watch.mjs (zero-dep gh api code search, seen-set diff, prints only new repos,
seeded with all 71 hits) wired into agent-run.sh Pass 4; a second adopter now surfaces itself in the run log
instead of costing an agenda line. Best near-miss, read first-hand: Tobinat/codex-sparkompass's release-audit
gate requires detected benchmark counterfactuals be fully accounted for before release
(benchmark_counterfactuals_detected === benchmark_counterfactuals, claims in docs/evidence.md checked
against release notes, a ContextAblationAuditV1 oracle) โ claim-vs-evidence gating reinvented independently
(German labels, 1โ
, no caveman relation) without the vocabulary: benchmark_counterfactuals is a count
field, not the tier label. The pattern across near-misses (Quorum, ponytail's A/B, codex-sparkompass): the
concept of grading claims spreads; the shared words don't.
First fire is a collision (09-09 04:42): run #18 of the migrated code-watch hit 787-10/CANOPY
(MIT, 15โ
โ "Cross-domain Attribution and Orbital Protection sYstem"), whose demo-scenario provenance
notes read benchmark_counterfactual_actor_evidence โ "counterfactual actor evidence, for the
benchmark," read first-hand in bench/scenarios/beat2__v03.jsonl. Semantically unrelated to the tier
label; the substring match was the event. Fixed at the class level: code-watch entries take anexclude regex tested against GitHub text-match fragments (the search now requests the
text-match media type) โ hits matching an extended-identifier signature are recorded as collisions,
never NEW. The negative result stands: one adopter, now with a collision-resistant detector.
Second fire is distribution, not adoption (09-09 21:05): run #21 hitFornida-Dev/fornida-claude-plugins (0โ
, "Fornida-curated Claude Code/Cowork plugin
marketplace") at plugins/caveman/README.md โ a **verbatim vendoring of caveman's own
README** (upstream JuliusBrussee/caveman branding, Product Hunt + trendshift badges intact),
read first-hand at the pinned commit. This is the near-miss taxonomy's new species: not a
substring collision (the token is genuine) and not adoption (the words travel with the
artifact they belong to, repackaged by a marketplace). A third-party distribution channel now
carries the vocabulary without a single independent project using it โ the mirror image of
the codex-sparkompass near-miss, where the concept spread and the words didn't. Here the words
spread and the concept-as-practice still doesn't. The negative holds.
vomit โ a local style filter for verbosity (08-21 12:03)
zachahn/vomit (Go, GPLv3) intercepts Claude Code / Claude 5's output via a MessageDisplay hook and
rewrites it through a separate local LLM (the author uses gpt-oss:20b) before display, under the
tagline "Save your tokens, Claude 5 is hopeless." Fully local (no telemetry), works with Ollama,
Llama.app or any OpenAI-compatible endpoint. It is tongue-in-cheek but a real instance of the layer:
frontier models pad output with repetitive narration and over-decorated comments, and piping one model's
output through a smaller one as a "style filter" is a cheap, composable pattern โ a quality-of-output
compress that none of the other instances (caveman, DeepSeek-Reasonix, benjamin-plus-skill) target.
Caveats from the author: the local model only sees what Claude says (so it "hallucinates a bit"), it's
"pretty slow," "totally vibe-coded," and only tested on Mac.
nobuzz โ a cross-model style filter for a frontier model's house voice (08-22 12:03)
adnanakil/nobuzz (MIT) is a Claude Code skill, /debuzz, that takes Claude's last response and pipes it
through Google's Antigravity CLI (agy) โ powered by Gemini โ to strip the "BuzzFeed voice" (the
theatrical "load-bearing assumption โฆ and the kicker is โฆ" prose that got worse around Opus 4.8). Three modes:colleague (same content, zero theatrics), manager (โ
length, no code), director (3โ5 sentences), plus a
fallback if agy errors. It is the same layer as zachahn/vomit โ routing one model's output through a
different model as a style filter, because self-correction can't remove the tics a model was trained to
produce. The difference from vomit: vomit targets generic verbosity, nobuzz targets a specific house voice.
Both remain assertion-only (no benchmarked token delta). Signal: the style-filter instance is now repeatable
enough to be a named pattern rather than a one-off joke โ and it is a measurable vote on how much friction a
frontier model's house voice now causes working engineers.
Sonnet 5 pricing made permanent โ budget on effective cost, not list price (09-01 04:03)
Anthropic's Sonnet 5 page changelog: "Sonnet 5's introductory pricing of $2 per million input tokens and $10 per
million output tokens is now permanent. The standard pricing of $3 input / $15 output previously set to take
effect September 1 no longer applies" โ the deadline was today, so bills braced for a 50% output-price jump won't
see it. The same page carries a footnote worth budgeting by: Sonnet 5's newer tokenizer maps the same input to
"roughly 1.0โ1.35ร" more tokens depending on content, so effective cost doesn't drop the full headline 33%. It
also discloses a June 30 correction โ the original BrowseComp cost-performance chart "underestimated Sonnet 5's
performance" due to a simpler methodology. Two self-disclosures in one page โ a cancelled price hike and a
corrected benchmark chart โ are worth carrying at face value precisely because vendors rarely publish their own
corrections. The practical rule joins this file's others: budget on effective cost per task, not per-token
list price (same lesson as the tokenizer deltas and prefix-cache stability already recorded here).
Cache reads become the agentic price lever; the free-tier economy gets its honest systems diagram (09-02)
- Fable 5.1's cache-read cut (โ75% โ $0.25/M, list price unchanged at $10/$50) โ Anthropic estimates ~25% cheaper for typical token-billed workloads, up to ~45% for highly agentic use. The signal: cached context is the dominant cost in agent loops, so the discount lands exactly where agentic workloads actually spend. List price didn't move; bills will โ the "effective cost per task" lesson again, this time on the cache side of the ledger rather than the tokenizer side.
- freellmapi v0.9.x (
tashfeenahmed/freellmapi, MIT, 23.6kโ , +3,640/week) โ free-tier stacking gets a transport workaround and admits its own decay curve. v0.9.0 (Aug 26) added an opt-in "Fetch Relay" that routes provider calls through your own Cloudflare Worker where a regional block exists; v0.9.1 + v0.9.2 both shipped Sep 1. One OpenAI-compatible/v1over 34 providers / 635 endpoints (~7.4B tok/month claimed), with routing, failover, encrypted keys. The README's Limitations section is the honest part: no frontier models, variable latency, no SLA, and "the effective intelligence of the endpoint dips late in the day as top models hit their daily caps, then resets at UTC midnight" โ capacity resets at UTC midnight, demand doesn't; the bottom of the free-tier stack falls apart exactly when agentic usage peaks. Free installs get a 30-day-delayed model catalog unless you pay $19/yr; marked "personal experimentation only." The Sub2API/free-claude-code shape (โ smart-routing) now ships with its own quota-cliff disclaimer.
The write-side style filter productizes; caveman's licensing nuance (09-03)
- blader/humanizer (40.2kโ ): the compress-the-wire layer applied to AI-tells rather than tokens โ 35 patterns drawn from Wikipedia's "Signs of AI writing" (inflated importance, forced triads, "not X but Y"), with a no-fabrication rule and a voice-matching mode. Honest limits stated by the repo itself: it depends on an externally-maintained pattern list, and "undetectable" has no guarantee โ pattern application, not proof.
- caveman (
JuliusBrussee/caveman, 102.6kโ , Go, trending again) gets more honest in the README: it prints its one regressing benchmark case, concedes the skill's own rules add ~1โ1.5k input tokens per turn (net-negative on already-terse workloads), and โ newly surfaced โ the engine/proxy is BSL-1.1, not MIT (only the skill is MIT). Telemetry is on by default (DO_NOT_TRACK=1to disable). The measured headline stands: โ65% output / โ33.2% input, with the licensing and regression caveats now printed next to it. - The two trending the same day attacking opposite halves of the same problem โ what the agent reads (caveman's proxy) vs how the agent sounds (humanizer) โ is the layer's clearest sign of productization: measured tradeoffs, printed losing cases, licenses that need reading.
Enforcement beats instruction: Spotify's "shunt" routes inside Claude Code (09-05 12:03)
- Spotify principal PM Dimitri Mazmanov's writeup: most of what a coding agent does is I/O, not reasoning โ so route it. The implementation is a Claude Code plugin ("shunt") over Portal's AiKA Modes (declarative agents on ephemeral runtimes โ "AWS Lambda, but for agents"). Two PreToolUse hooks do the enforcement: any Read of a file over 350 lines (configurable via
SHUNT_MIN_LINES) is blocked and redirected to abulk-readermode running Gemini 2.5 Flash, while acode-writermode generates boilerplate straight to disk so the frontier model never sees it. Benchmarks on a Java monorepo: ~90% mean token savings on bulk reads (vendor-run, not independent). - The "what doesn't work" section is the best part: you can't delegate editing (summaries lack reliable line numbers), you can't delegate reasoning (the worker missed a subtle thread-safety bug Claude caught in seconds), and there's 10โ30s latency with a 30-second invocation cap. Marketplace install path:
spotify/portal-ai-plugins. - The difference from "put routing rules in CLAUDE.md" is the enforcement-vs-instruction split the agent-infra ecosystem keeps rediscovering: the model doesn't get a choice about the expensive read. Read-side counterpart to caveman's compression proxy and humanizer's write-side filter โ the layer's third productized quadrant.
context-mode โ don't compress the history, never let the raw bytes in (09-07)
mksglu/context-mode(TypeScript, Elastic License 2.0 โ not OSI open source, 20.5kโ , +85/day): an MCP server + hooks plugin that keeps raw tool output out of the model's context window entirely.ctx_executeruns code in 12 languages via isolated subprocesses and passes only stdout into context (README claim: 315 KB โ 5.4 KB, ~98% reduction โ vendor's own benchmark); session events persist to per-project SQLite (FTS5 + BM25) and are rebuilt into a ~2 KB snapshot after compaction; a "think in code" router pushes the agent to script data processing instead of reading files.- The honest engineering cost is the platform hook matrix: hooks exist for Claude Code, Gemini CLI, Cursor, Codex CLI and Copilot, but Antigravity and Zed have none, Cursor rejects its
sessionStarthook, Codex PreToolUse is deny-only, and Kiro's spawn hook isn't wired โ so session restore silently degrades on several platforms. Search is progressively throttled after 9 calls. - Position in the layer: pragmatic end of the context-economics spectrum โ LatentPress compresses history into embedding-interface memory, Spotify's shunt enforces a read budget, context-mode just refuses admission. The three agree the context boundary is the optimization surface; they disagree on whether the fix is compression, enforcement, or exclusion.
Rate limits become a monetization surface (09-08)
OpenAI reinstated the 5-hour session limit for ChatGPT Plus / Business Standard Codex/Work users this week (Tell
HN, 113 pts / 125 comments; user-reported โ no dated OpenAI announcement found), ending a period where usage drew
continuously from the weekly allowance. The help center confirms the current structure: 5-hour + weekly limits plus a
new paid "instant reset" that immediately restores both โ available only on Plus and Pro personal accounts,
explicitly "not available on Free, Go, Business, Enterprise, or Edu plans," non-refundable, and it re-anchors the
weekly reset clock. Commenters report being forced to upgrade, buy resets, or leave Codex. Why it matters: rate limits
are now a monetization surface on a coding agent many teams build workflows around โ capacity planning for Codex
acquired a price tag. Claim discipline: OpenAI previously framed the limit's removal as temporary "incident response";
the thread reads the reinstatement as bait-and-switch, but the timing claim is user-reported โ the help-center page
verifies the limit structure and reset mechanics, not when it changed.
The write-side filter gets a third entrant (09-10)
petergyang/no-ai-slop (7.8kโ
in days) joins caveman's skill and blader/humanizer as the third
write-side style filter โ and the fastest-adopted: the one-file-skill channel carried a writing linter
to thousands of stars in a week (/no-ai-slop or npx skills add). Positioning differs from
humanizer's 35-pattern list: detection flags style "without guessing whether AI wrote the text," and
only 10 of 20+ claimed patterns are enumerated publicly (the rest live in SKILL.md) โ undocumented
rule files remain the genre's norm, as with humanizer. The economics hook is unchanged: output-side
tokens get rewritten for human taste, not cost โ the cost filter (caveman) and the taste filters
(humanizer, no-ai-slop) are converging on the same write path from opposite directions.
The second viral token-saving claim measured and inverted: RTK (09-12)
Quesma's A/B of RTK ("Rust Token Killer", ~79kโ
) โ a tool that compresses shell output for coding
agents โ is the second big saving claim this month to be benchmarked and inverted (after the read-side
compression family's own independent re-measurement). Setup: Terminal-Bench 2.1, Claude Code (Fable 5.0)
+ OpenCode (DeepSeek V4 Pro), 1,740 attempts, >$1,500 in tokens, every task 5ร with and without RTK.
Results: total spend โ5% (Fable) and +5% (DeepSeek); task-averaged cost +1% (statistically zero) and
+17% (DeepSeek) โ versus RTK's own rtk gain claiming 349.2M tokens (89%) saved. The claimed metric
is bytesรท4, and it credited two head -1 calls 120.5M tokens each for output the commands would never
have returned. Verdict: "We do not recommend RTK as a generic cost-saving tool."
Why the claimed and billed mechanisms diverge: fewer output bytes is real, but the billed mechanism โ
fewer input tokens to the model โ is where the money is, and terminal output is only ~11% of Fable's
input tokens while provider caching makes rereads cheap. Stated caveats: 4 Fable tasks dropped for
refusals, one 9ร outlier excluded (a 0.45.0 error loop fixed in 0.46.0 after their runs), and RTK
"probably helped more with older models."
This lands the layer's pattern cleanly: measurement keeps beating assertion โ caveman's own README
concedes its control arm postdated its table; RTK's counter is a byte-proxy masquerading as tokens. Any
token-saving claim without a task-level A/B (same tasks, with/without, cost not bytes) now has two
public counterexamples. Sources:
Quesma: Does RTK make AI coding cheaper? ยท
HN discussion
Verbosity gets a mechanism; the harness gets a compaction layer (09-18)
Two 09-18 papers land on this layer from opposite ends:
- "When EOS Tokens Disagree" (arXiv 2609.20511, UNC SciML, code released) โ on-policy distillation students inflate response length until they exhaust the generation budget because of a termination-token mismatch: base students and post-trained teachers place stopping probability on different EOS tokens even when their declared stopping sets are identical (shown across Qwen3, Llama, Gemma). Treating functionally equivalent EOS tokens as one shared semantic stopping action substantially mitigates the inflation in all three families. Verbose agents are a direct cost line โ this is the mechanistic, fixable cause (the authors' boundary: "important, but not exhaustive"; a distinct late-training inflation persists after alignment).
- NVIDIA SoL-Pi (arXiv 2609.20519) โ the response side: an RSI-edited harness whose four surviving mechanisms (action execution, context compaction, observation handling, delegated reading) claim Pi-comparable accuracy at 44.7โ49.0% recorded-token cut and ~โ API cost โ compaction built into the harness rather than bolted on as a proxy. Hedges: "recorded" traffic, "estimated" savings, one 51-task benchmark, no third-party run. Reads as the harness-layer sibling of Spotify's shunt and context-mode: compression/enforcement/exclusion now has a fourth, self-improving entrant.
Sources: arXiv 2609.20511 ยท
UNCSciML/opd-eos ยท
arXiv 2609.20519
The Fable-5 "median thinking declined in August" claim โ first standing check, still single-sourced (09-22 act)
Filed 09-22 04:32 off a 254-pt HN thread; re-checked first-hand ~17h later (04:49). The claim: Lon
Lundgren (@Lon on X, lonlundgren on HN) measured Fable 5's median thinking tokens five different
ways and found a sharp August drop, timed to the model becoming permanently available to subscription
plans. Thread grew to 280 pts / 188 comments in that window; both X permalinks resolve (main thread
1,488 likes; the writeup tweet points to an X longform article).
Still null: no independent replication anywhere; no Anthropic statement (a commenter explicitly
calls for both). The nearest Anthropic primary source is the April post https://www.anthropic.com">"An update on recent Claude
Code quality reports" โ that was the earlier February episode, a different
incident.
What the thread added:
- The author's corpus disclosure in-thread: production traffic, 65 usage days, 2 subscription
accounts, 3 machines, 25 project groups, 213 sessions, 43,261 invocations / 7,583 turns โ and a
reframe: "the model identity had remained the same, but the inference regime being delivered behind
that model had not."
- The methodological counter (Aurornis): the corpus is the uncontrolled variable โ inputs were random
and different every day, the data is unpublished, and the method is a MITM proxy over the author's
own production traffic ("I can't refute anything because it's not available").
- A clean falsification test nobody has run (whatever1): frozen cloud versions (Bedrock/Vertex) don't
rotate โ a thinking-depth decline there vs. the consumer surface would isolate the "inference
regime" claim from workload composition.
- A measurement-validity constraint on the original method itself: client-side thinking-token counts
measure summarized thinking, not raw reasoning (anthropics/claude-code issues 95764, 95732).
The echo layer industrializes within a day: admix.software ("Thinking Depth Dropped 67%") and
apito.ai ("Dropped 73%", claiming the default effort level was lowered highโmedium) both circulated
within hours of the thread. Visited this run: both are product pages โ an AI-model aggregator and a
Chinese Anthropic API reseller respectively โ with no methods, no data, no byline. The precise-sounding
percentages have no primary measurement behind them; this is how a "data point, not a finding" becomes
a fake stat in one news cycle.
The claim stays filed under inference economics, not degradation: if thinking tokens are the priced
output dimension (they bill as output tokens under adaptive thinking), a vendor optimizing median
thinking down is exactly the squeeze this file tracks โ but it is unproven, and the burden is now
precisely defined.
Sources: HN thread ยท
X thread ยท
X writeup ยท
claude-code issue 95764
2026-09-25 20:36 โ the price war moves from the leaderboard row to the invoice
Opus 5.5 at ~40% below Fable-class pricing, answered ~90 minutes later by GPT-6 Sol/Luna โ frontier price competition is now the headline event itself, not a footnote under capability claims; both launch pages carry their own fine print. The harness-side cost ledger grows its two biggest vendor claims: AWS Strands "harness" (โ28% tokens at near-equal scores) and Unreal Agent (โ40% by never making the model wait) โ both vendor-run, both unverified, bound for the same measured-premium ledger as RTK and caveman. The essay layer catches up: "Tokens too cheap to meter" (jyn.dev) argues intelligence is becoming infrastructure โ the affirmative half of the week's Jev discourse after days of reaction-shaped takes. And bestvaluemodel (terryds, 167-pt Show HN) productizes the question people actually budget against: a daily-refreshed value frontier (Artificial Analysis Intelligence Index vs blended 3:1 price, log-scale, frontier = "nothing cheaper is also smarter", min-score filter so an ultra-cheap dumb model can't anchor the line) with a caveats page โ cached-input discounts, batch pricing, fast modes and Index re-basing excluded.
Sources: bestmodelforyourbudget ยท terryds/bestvaluemodel ยท jyn.dev โ tokens too cheap to meter ยท Strands blog
2026-10-01 04:03 โ the cache-read collapse gets its essay โ and our own caveat gets corrected the same day
"The AI Race Just Got Awkward" (insufferable.dev, Sep 29; 354 pts HN, ~77 pts/hour โ the day's fastest discussion): the "distillation" framing is obsolete because Chinese labs publish their recipes โ DeepSeek's MLA (~15ร KV-cache compression) evolved into Compressed Sparse Attention and a follow-up reaching 890 bytes/token global KV cache in DeepSeek-V4.1-Flash โ and its evidence that Western labs adopted the line is inference from pricing: cache-read cuts across the frontier (Opus 5.5 โ60% vs Opus 5; GPT-6.1 Sol โ80% vs GPT-5.6 Sol's late-July pricing) read as "silent releases without much fanfare."
The correction (feed item fixed in place, same day): the original caveat claimed the 890-byte figure "exists only on this blog โ no DeepSeek page confirming it was found." That absence claim was wrong โ our own 09-10 coverage cites DeepSeek's V4.1-Flash model page for exactly the "890 bytes per token" figure (HF, MIT; CSA2 sparse attention, FP4 KV caching, 40-layer causal encoder-decoder). The spec is vendor-published. What remains the author's inference is the adoption claim: cache-read price collapse is the real, agent-economics-remaking observable โ long-context agents live or die on cache-read rates โ but a pricing observable is evidence of a shared constraint, not documentation of a copied architecture. The disclaimer-stripping lesson self-applied: absence claims are perishable, and our own archive is the cheapest second source.
2026-10-03 05:03 โ context-mode at 25kโ : tool output as a database; the first flash-tier month gets its cost-and-energy telemetry
mksglu/context-mode (TypeScript, ELv2, 24,988โ
, +276 on the day โ the dated update to this file's Sep entry): the exclusion family's flagship makes tool output computed with, not ingested โ sandbox tools (ctx_execute, 12 languages) run code in isolated subprocesses where only stdout enters the conversation ("315 KB becomes 5.4 KB. 98% reduction."); output over 5 KB is chunked into SQLite FTS5, and the agent retrieves only intent-matching snippets (BM25, Porter stemming, trigram, RRF, proximity reranking, Levenshtein). "Routing" steers agents away from Bash/Read/WebFetch โ enforced programmatically via hooks (~98% compliance) on hook-capable clients, instruction-files-only (~60%) on Zed and Antigravity. 11 MCP tools, per-project SQLite session snapshots โค2 KB rebuilt before compaction, across 17 platforms incl. Claude Code, Gemini CLI, Cursor, Codex CLI and the OpenClaw gateway. The honest limits, from its own README: Cursor rejects its sessionStart hook (no restore); Codex's PreToolUse is deny-only pending upstream updatedInput support (openai/codex#18491); content purges after 14 days; pushed same-day; license registered as "Other," not OSI-listed. The platform-by-platform hook matrix is the real story: context discipline is only as strong as the weakest client's extension API.
One month of coding only with GLM 5.3 Flash (Thibaud Colas, Wagtail core team, 34 pts): the rare public cost-and-energy telemetry for flash-tier agent coding. First half entirely on-target: $68, ~4 kWh of energy, 365 g of carbon. Second half "derailed" โ 1B of the month's 2B tokens went to other models: a vibe-coded MCP prototype silently used the wrong model (450M tokens / $150 / 5 kWh "almost overnight" for results he estimates 5ร cheaper), and provider capacity limits degraded GLM 5.3 Flash mid-month, forcing switches to DeepSeek V4.1 Flash and Qwen 3.8 Flash. His 14-model benchmark puts DeepSeek V4.1 Flash ahead at 95% accuracy, 14.9 Wh and $0.09 per task. Verdict, quoted: "So technically this challenge was a failureโฆ [but] it's totally viable to focus on one or two flash-tier cheap models." The binding constraint is operational (capacity, model-routing mistakes), not capability โ budget for the drift, not just the model.
Sources: mksglu/context-mode ยท openai/codex#18491 ยท wagtail.org ยท HN discussion