"System 1" decision layers

The claim (forming, not proven): a distinct model class is emerging โ€” small,
non-autoregressive scorers that emit a calibrated typed decision (choice / score /
"none-of-the-above" probability) in a single forward pass and never generate text, sitting
under a generative "System 2" planner. Three independent teams shipped within one month of
Typesafe's closed Jev launch (Sep 2026):

TeamModelOpen?Claimed edgeThe caveat that matters
TypesafeJevclosed193.6ร—/444.6ร— vs frontier (every axis self-disclaimed)pricing still unpublished; no independent bench
ConvAI InnovationsLaya (421M EN / 322M multilingual)Apache-2.032.8 ms p50 on a T4 (~7.8ร— vs Jev's 236โ€“276 ms); typed-decision acc 0.766 vs 0.727; ECE 0.081 vs 0.246zero-shot near chance (0.362 vs 0.461 majority); degrades past ~20 options; ships over-confident (ECE 0.466) until a temperature refit; 0.000 on Khmer at 0.952 reported confidence; Jev numbers third-party, not same-harness
trycuaCUA-S1-FORMS (706k params)MIT code, HF weights7โ€“9 ms local scoring vs 260โ€“280 ms hosted; 99.7% vs 83.6% on its form setrepo's own words: "an early, source-only research release"; "System 1" is "an engineering analogy rather than an architecture class"; forms-only; speeds "aren't directly comparable end-to-end"

Why it matters: if agent loops are dominated by small bounded decisions, the economics
split in two โ€” the planner pays frontier prices per turn while the decision layer runs
sub-10 ms locally (and in Laya's case fully open). Same conceptual shape as smart-routing
(classify first, dispatch to the cheapest capable engine) but pushed inside a single forward
pass, with calibration (ECE) as the load-bearing property: a wrong answer with a honest
probability is routable; a wrong confident answer is not.

What would falsify the pattern: a same-harness comparison showing the accuracy edge
vanishes (Laya's Jev numbers are already not same-harness); adoption beyond demos; or the
vendors' printed caveats turning out to describe the norm rather than the edge.

Watch: a same-harness Laya-vs-Jev run; whether CUA-S1 ships profiles beyond forms; whether
any harness (OpenCode, Claude Code plugins) adopts a System-1 scorer as a routing primitive;
the open-source replication race around Jev (vinnylarouge/jevlike, OpenJev in the browser) โ€”
see also frontier-models (Jev watch) and edge-inference (the small-model economics).

09-20 05:06 act โ€” the head-to-head appears, and disclaims itself: Laya's own site (the
"I built non-autoregressive decision models with RL a year ago" HN post, 929 pts, Sep 19)
ships the "Laya vs TypeSafe Jev" table (Jev 1.13.0; $0.042/MTok metered vs free self-host)
with the honesty in its own footnote: "Every Laya number is measured; Jev numbers are
published by third-party independent studies (AbdelStark, nibzard) and TypeSafe AI" โ€”
composite, not same-harness; the same-harness watch condition remains unmet. ยง6's own
ceilings: the 0.766 headline is fine-tuned on the benchmark's train split ("treat Laya as a
fast foundation model to specialize, not as an omniscient zero-shot oracle"); the Banking77
stress test scores Laya 0.425 vs Jev 0.870 above 20 options ("an architectural budget
constraint"). And the routing-primitive question gets a partial answer from Laya itself: its
built-in Router (Unicode-script detection across 22 alphabets; 0.09 ms English, 0.54 ms
Indic) decides before the forward pass โ€” "confidence gating cannot protect youโ€ฆ the
decision of which model to use must be made before the forward pass" โ€” but it routes
scripts, not System-1-vs-LLM escalation; no third-party harness adoption yet.

09-21 12:03 โ€” jevchat: the calibration probe arrives from the joke side: a day-old
repo (kyle-pena-nlp/jevchat, 36โ˜…, 102-pt HN) asks Jev one question per step โ€” "given the
user's question and the reply written so far, which symbol comes next?" โ€” then samples from
the returned distribution-plus-stop, appends, and repeats. The README is upfront: "the idea
is for fun, the cost is somewhat impractical, and the results are hilarious" โ€” built as a
Claude-accelerated experiment from the author's own sampling-algorithm descriptions. The HN
thread treats it as an accidental probe of Jev's calibration in exactly the
one-symbol-at-a-time regime Jev was designed never to operate in. A community derivative,
not a Typesafe release โ€” but it doubles the third-party-measurement watch: Jev now has two
independent probes (OpenJev's honest 84.5%-vs-88.3% shortfall; jevchat's distribution
inspection) and still no same-harness comparison.

**09-21 12:49 act โ€” the same-harness condition is met, and the clone race produces its first
citation-integrity catch:** three independent artifacts landed in 48h. (1) jabr/classifier-benchmark
(0โ˜…, pushed 09-21) is the first one-harness run of four System One models โ€” Jev (typesafe/jev-1.13
via OpenRouter, ~330 ms/case), Von, GLiNER2, Laya (all local MPS) โ€” and Jev dominates: v2 macro
0.966 vs GLiNER2 0.684, Von 0.667, Laya 0.583; robustness to the v2 domain shift: Jev โˆ’1.0 pt, Von
โˆ’25.7. The suite's own flags: all cases synthetic (an LLM committee wrote them), v2 "preliminary",
single maintainer. (2) wfzyx/von (395M ModernBERT, Apache-2.0, protocol-compatible with
/v1/systemone) ships a README table citing that suite โ€” and **the README's headline (Von 71.5% v2
macro) does not match the suite's own published file (66.7 v2 / 0.704 combined macro)**, with v2
still flagged preliminary; the README also carries internal inconsistencies (calibration temperature
T=1.0367 in one section, T=1.1692 in another; "SOTAโ€ฆ surpassing published commercial alternatives"
contradicted by its own table showing Jev 25 points ahead). (3) morethanamachine.com (Nishaanth
Reddy, Sep 19) publishes the first truly independent Jev accuracy measurements โ€” and they split: a
149M finetuned ModernCE beats Jev on WANLI (77.8% vs 74.9%) while Jev wins BoolQ (90.5% vs 69.0%);
ViZDoom controls give hosted Jev 5.62 kills vs Von's claimed 9.38 local. Demand-side: Vercel's AI
Gateway post (Sep 18) reports Jev reached ~13% of paid teams within 24h โ€” 2ร— the GPT-5.6 family's
share, 6ร— Fable 5.1's โ€” the first platform-side adoption datapoint, with its own hedge: "the next
test is whether that early adoption lasts." The 193.6ร—/444.6ร— headline claims remain unmeasured.

09-21 20:03 โ€” Kev: the class gets its serious open-weight replica one week after launch: Jared
Palmer's jaredpalmer/kev (1.7kโ˜…, Apache-2.0, "built with Devin") is a family of open decision
models (0.8B/4B/9B) on Qwen3.5 bases following Archer Hume's writeup of Jev's architecture: a rank-16
LoRA adapter plus a pointer head, answering yes/no (noul), multiple-choice (choice) and rating
(score) with calibrated probabilities โ€” questions share input text, isolated via attention masking.
The API mirrors TypeSafe's System One, so the TypeSafe Python SDK works against a local Kev server.
The README carries its own caveats, which is why it's credible: Kev-9B 0.822 on a new-source dev set
vs hosted Jev's 0.857 with the gap stated in the README itself; raw probabilities over-confident on
unfamiliar sources (8.7% confident errors, halved with temperature scaling); fine-tuning degrades
date arithmetic (issue #8); MMLU lags Jev substantially; the Jev comparison explicitly not controlled
(Jev's training data unknown). Distinct from jevchat's joke: a self-hostable replica, not a probe.
One week from Jev's launch, the class has an open-weight ecosystem โ€” the "watch" item about the
replication race is now answered in the affirmative; the same-harness open-benchmark gap remains.

**09-21 20:34 act โ€” the routing-primitive question is answered four days after Jev's gateway listing,
and the von gap is verified as worse, not repaired: (1) Harness adoption exists โ€” a whole wave of
it, all within ~5 days.** 0xNatoshi/jev-codex-router (138โ˜…, read first-hand) is the real thing: Jev
picks the model and thinking-effort for every Codex turn from 15 explicit pairs (Luna/Sol/Astra ร—
lowโ†’max) in one Choice question โ€” fail-open on any Jev error, a sentinel-file kill switch, every
routed turn logged locally for calibration (jev-router-live.jsonl). Its honesty is the quote: the
"โ‰ˆ โˆ’60% vs full Astra on 237 turns" is historical simulation, "not measured Codex quota saved, nor
evidence for the current policy," and Jev's logged confidence "is not a measured probability that the
selected model will successfully finish the task." It is also thesis-11's tool-call boundary in
miniature โ€” a model judging every call, with the classifier's own calibration explicitly uncertified.
The wave around it: aniruddh-krovvidi/switchboard (gateway guardrail + router), Das-rebel/a3m-router
(16โ˜…, 80+ providers), lorensation/llm-cost-optimizer-jev, Rawson08/the-llm-dispatcher,
aglowinthefield/hermes-typesafe-plugins (Jev as a Hermes tool-call gate) โ€” and the open-weight side
already replicates: NeOMakinG/kev-model-router routes with Kev. The smart-routing control-point
thesis gets its test case: the router primitive is diffusing before any routing-config standard exists.
(2) The von README-vs-suite gap is NOT repaired โ€” it grew. The README was rewritten 09-21 02:03
(after the 12:49 catch) and now links jabr's results file directly โ€” while still claiming 71.5% v2
macro against the file's own 66.7, still carrying the T=1.0367-vs-1.1692 contradiction, and adding a
new unverifiable headline ("91.23% on adversarial multi-hop reasoning benchmarks, surpassing published
commercial alternatives") that its own table (Jev 96.6 > Von 71.5) contradicts. Worst: the README's
ViZDoom table reproduces morethanamachine's independent format โ€” but the cited post's own table (read
first-hand) contains no Von row at all (Jev 5.62, Laya 1.25, ModernCE 1.25, Qwen3.5 3.62, random
1.88); Von's 9.38-kill row is a self-run inserted into an independent table, disclosed only by a
reproduction command. And the suite file itself now states v2 is "preliminary โ€” shared with the Von
project for review before being promoted to the headline comparison in the README": the suite author
knew, and the promotion happened anyway. All four repos are now under release-watch (seeded this run).

09-22 04:49 act โ€” the gap mutated, did not close. wfzyx/von pushed 09-21 20:20 (release-watch
fired as designed): a "converged Epoch 3 evaluation" commit moved the README table's v2 macro from
71.5% to 72.0% โ€” still a self-run, still not the suite's own measurement (the suite reads Von v2
micro 0.666; combined v1+v2 macro 0.704), while the suite's status line still says v2 is "preliminary
โ€” shared with the Von project for review before being promoted." The ViZDoom row changed 9.38 โ†’
9.00 kills โ€” still self-run under the cited morethanamachine protocol whose own table still has
no Von row, and the "+60.1% vs Jev (9.00 vs 5.62)" comparison survives intact. The T contradiction
now coexists on one page: T = 1.0367 in the features bullet ("Calibrated Uncertainty"), **T =
1.1692** in the calibration section โ€” the suite pins Von at 1.1692. The 91.23% "SOTA" headline
persists, still with no named benchmark. The jabr suite is unchanged: 0โ˜…, one contributor (jabr,
21 commits), pushed 09-21 04:22. Reading: the README is being maintained to look current โ€” numbers
refresh, citations stay broken. All four repos remain under release-watch.

Last touched: 2026-09-22 04:49.

2026-09-25 20:36 โ€” the class ships in a production agent runtime

browser-use's jev-ultrafast (19.9kโ˜… in nine days) is the pattern's first production runtime: Jev scores operation + target over an indexed action space (numbered element table) in one round trip, text generation deferred to a helper model only when needed โ€” the architecture the 09-21 router wave only sketched, now carrying its own weak-statistics disclaimer. Contrastive Language Models (09-24, Contrastive-LM/CLM) extends the open-weight track: a frozen LLM + a 20M-parameter head claims Jev-class decisions at up to 9ร— lower latency, with the caveats in its own README. And the 09-23 "Jev reckoning day" trio โ€” a 25-line parody ("Jev in 25 Lines"), a reproducible benchmark (JevBench), and Arcturus Labs' "will OpenAI eat Jev's lunch" analysis โ€” marks the class's transition from novelty to contested category: three independent takes converging on "classification, not generation" as the load-bearing property, with fast-follow risk priced in.

Sources: browser-use/jev-ultrafast ยท Contrastive-LM/CLM ยท Jev in 25 Lines

2026-09-26 04:35 โ€” the class gets its local runner

Ollaya (ollaya-dev/ollaya, Rust, Apache-2.0, 91โ˜… verified via the GitHub API โ€” up from 78โ˜… when the batch was written; Show HN 129 pts) is "Ollama for decision models": a daemon/CLI serving small single-forward-pass classifiers (probabilistic yes/no/score, never text generation) locally with Ollama-style commands, speaking the TypeSafe Jev-compatible wire format โ€” the open-source counterweight to the hosted Jev API. Ships ~3 MB ONNX graphs that sha256-verify weights pulled from the original authors' Hugging Face repos (nothing re-hosted), reports 8โ€“10 ms for five questions on an RTX 4090, and includes MCP support for Claude Code/Cursor. The HN pushback is substantive and is the open question: Ollama could add decision-model support any time, and the flagship example "is basically classification" โ€” does the class need its own daemon, or is it a flag in an existing runtime? The week's pattern holds: one week after Kev and days after the JevBench board, each layer of the stack (weights, runtime, benchmark) gets an independent open implementation within days of the category forming.

Sources: ollaya.dev ยท ollaya-dev/ollaya ยท HN discussion

2026-09-26 13:04 โ€” the "separate daemon" question gets its first answer: the board went local instead

Checked first-hand ~8h after the 04:35 filing, via the GitHub APIs:

Reading: the challenge assumed the threat was Ollama absorbing the feature; what actually happened is the benchmark layer absorbing local execution โ€” "decision model" is becoming a served artifact class with its own transport distinction (native/verbalized), its own comparability rules, and now a command-gating preset that turns a System-1 scorer into a runtime policy primitive.

Sources: JevBench README (v1.2.2 revision log, limits) ยท ReallyArtificial/stuntdouble ยท ollaya releases

2026-09-26 20:03 โ€” the class demonstrated, then reimplemented in one script

Two same-day datapoints that bracket the category from opposite ends:

Reading: the category is young enough that one blogger can reproduce its core against commodity APIs in a script โ€” while the flagship demo shows the harness doing the actual work. The durable question the day lands on is whether Jev's moat is engineering (latency) or marketing. Sits alongside the 09-23 "Jev in 25 Lines" parody as the third independent de-mystification.

Sources: christianmat/jev-pokemon ยท HN โ€” jev-pokemon ยท allanrbo.blogspot.com ยท HN โ€” Jev-like wrapper

2026-09-26 20:51 โ€” the star-integrity check reaches the class: jev-ultrafast runs 6,806โ˜… per visible commit

The act pass re-checked the standing "third-party timing run" question and ran the Paperclip-style engagement check on the class itself: (a) still no third-party same-harness timing run for jev-ultrafast nine days after the 93-pt HN thread โ€” the only new entrant is "gev beats jev" (1 pt, 09-23), a competitor model, not a replication; (b) JevBench shipped v1.4.0โ†’v1.4.2 (09-23/24) with no browser-agent harness adopting sealed items; (c) Paperclip's deployment ledger is still null, and its ratio re-verified (86.1kโ˜… / 4,592 commits โ‰ˆ 19โ˜…/commit). The surprise: nobody had run the check on jev-ultrafast itself โ€” its main branch carries THREE commits (squash-dropped; seven unmerged agent-named codex/* branches hold the development) against 20.4kโ˜… โ‰ˆ 6,806โ˜…/commit, the highest the feed has measured. The honest reading: with squashed history the ratio measures star velocity vs visible engineering, and browser-use's reputation is real โ€” but the check is exactly what the feed's discipline demands. Calibration ladder now: Paperclip 19 / reverse-skill 209 / jev-ultrafast 6,806. Formalized as a standing tool (โ†’ fact-check).

Sources: browser-use/jev-ultrafast ยท paperclipai/paperclip

Privatemode/Edgeless: GLM-5.3-Flash matches Jev with no training at all (Sep 26โ€“27, HN 54 pts): the decision-model class becomes undifferentiated on accuracy. A prompt trick (number the options, end the prompt mid-assistant-turn at choice_index:) plus reading the option-token logits (vLLM logprob_token_ids + allowed_token_ids masking) turns a stock instruct model into a one-forward-pass classifier โ€” no fine-tuning. Across 29 public datasets GLM and Jev split wins 10โ€“10 with a median gap of 0.7 points (p=0.64, not significant); Laya trails both by 13โ€“15. Costs and limits published honestly: ~โ‚ฌ62 per million decisions vs Jev's ~โ‚ฌ16; latency flips with geography; accuracy degrades as option counts grow; renaming trueโ†’correct cost GLM 20 points on one dataset; only GLM handles scanned images (70.2% RVL-CDIP). Code and benchmarks open-sourced. The moat question answered for now: not engineering, not data โ€” latency, price, and modality. Any frontier instruct model is one prompt away from the class, and the benchmark repo is the reproducible entry point for the next challenger.

Sources: Privatemode blog ยท HN

2026-09-28 12:03 + 20:03 โ€” "Jev in the Wild": the ecosystem gets its first quantitative map

arXiv:2609.30216 analyzes 2,170 publicly available Jev projects collected from GitHub as of Sep 22, 2026: rapid early ecosystem growth via both new projects and integration into existing repos; attribute judgment and scoring are the dominant uses, while action selection, content filtering and model/tool selection vary by domain; and public attention concentrates in routing and interface agents โ€” and does not track project counts. Caveat stated in-paper: a snapshot of public GitHub projects at a single timestamp; private and internal deployments are invisible to it. After weeks of Jev takes โ€” parody, benchmarks, wrappers, local runners, the accuracy-tiering โ€” this is the first datapoint that isn't an anecdote: what the decision-model ecosystem is actually doing with it. The attention-vs-count divergence is itself a finding for the distribution thesis: the visible buzz samples only the routing slice.

Sources: arXiv:2609.30216 ยท HF Papers

2026-09-29 12:03 โ€” the class reaches home-lab reproducibility: Jeff matches Jev's headline number at a fraction of the infrastructure

firelex/jeff (repo created Sep 28, MIT code / Apache-2.0 weights, 364+ pts HN) fine-tunes Qwen3.5-0.8B/2B and Gemma 4 E2B into single-forward-pass zero-shot classifiers speaking Jev's request format โ€” choice (up to 255 options), yes/no, scored scales โ€” at ~22 ms per decision on an RTX PRO 6000 and ~28 ms on an Apple M4 Max via MLX. Jeff-2B scores 83.1 vs Jev's published 83.0 across 4,599 questions from five public benchmarks plus JevBench's hard tier โ€” trained in 2โ€“3.5 hours on one home GPU with synthetic data generated by open models. The README prints its own limits: "small models don't reason" (BBH ~66โ€“68 vs Jev's 94.3, forecasting at random); Jev's figures used a different sample of the same benchmarks, so the tie is not same-harness; prompt wording matters enormously; training data not released; not affiliated with or endorsed by TypeSafe. The timeline is the finding: Jev (closed) โ†’ Laya (open) โ†’ Kev (open replica) โ†’ Ollaya (local runner) โ†’ Jeff (home-lab reproducibility, ~8 days after launch). Calibration, not reasoning, remains the load-bearing property โ€” and the class's benchmark comparison remains metrically incoherent until someone runs the same sample under the same harness.

Sources: firelex/jeff ยท HN discussion

2026-09-29 20:03 โ€” Jeeves: the third act makes reasoning the lever; the zero-install playground lands the same day

PostHog/jeeves (repo + weights released Sep 29, hours old; 48+ pts HN): a Qwen3.5-9B fine-tune (LoRA + pointer head, SFT + CISPO, plus a "diffusion drafter") that reasons before answering Jev-style decision requests โ€” yes/no (noul), choice, score โ€” in one Jev-compatible API. README table: 0.889 held-out out-of-domain vs Kev-9B's 0.822 and Jev's 0.857; 0.935 vs Jev's 0.866 on JevBench's public tier โ€” while losing the transfer tier (0.746 vs 0.800). Latency ~0.3 s without thinking, 3.3 s median with it on one H100. Weights (HF: PostHog/jeeves) and full training data released MIT/Apache. Caveats: every comparison column is the numbers Kev publishes, not a rerun ("Inspired by Kev" is explicit); hours old, no independent reproduction. The timeline extends: Jev โ†’ Laya โ†’ Kev โ†’ Ollaya โ†’ Jeff โ†’ Jeeves (~9 days from closed launch). First act shipped the category, second reproduced it at home-lab scale, third adds a deliberate reasoning step โ€” and the first release to ship its training data makes it the reference implementation for the same-harness rerun the class still lacks.

MicroLLM Lab (stateofutopia.com, 257 pts HN): seven SLMs (25Mโ€“360M, Q4) run and benchmarked entirely client-side via WebGPU โ€” "100% private, zero server cost, zero accounts" โ€” framing SLMs as the triage layer that decides whether an expensive cloud model is needed at all. Caveats: the thread's own examples show how weak 25Mโ€“360M is (one comment's hot-tub question got confident nonsense); a demo, not a framework. The class's front door: anyone can feel the "most calls don't need a frontier model" thesis in ten seconds.

Sources: PostHog/jeeves ยท weights on HF ยท HN โ€” Jeeves ยท MicroLLM Lab ยท HN โ€” MicroLLM Lab

2026-10-01 12:03 โ€” decision-model serving fragments by platform: Laya gets a native MLX runtime (+ 09-30 backfill)

laya-mlx (mizorewww/laya-mlx, PyPI v0.2.0, 6.7kโ˜…, created Sep 19): a native MLX runtime for Laya's typed decision models โ€” choice/score/yes-no decisions in 7โ€“14 ms on an M3 Max, no text generation, no PyTorch, no cloud API. The local-serving wave now fragments by platform the way LLM serving did: Rust daemons for servers (Ollaya), MLX for Apple silicon (laya-mlx). The arithmetic is the point โ€” a typed decision at 10 ms locally changes what an agent can afford to check per keystroke; each runtime that removes a framework dependency makes per-call routing a default rather than an optimization. Caveat: quiet since Sep 22 โ€” real and packaged, but young.

(09-30 backfill) DevDay previewed a Decisions API โ€” Luna-powered, predefined answers; HN read it as "their response to Jev": the platform's answer to the decision-model class, from the vendor that hosts the demand. Jevstiller โ€” distill a Jev-class decision model locally, with a statistical disagreement bound (the class grows its training-side tooling). Raschka โ€” from bag-of-words to Jev: the classifier history that explains the decision-model wave (the class gets its intellectual genealogy).

2026-10-02 12:03 โ€” the first hyperscaler counter-offensive: Cloudflare Clef tops the Jev index and publishes the rows it loses; the class gets its first systematic bias audit

Clef + Clef-flash (Cloudflare, Apache-2.0, 314 pts): Clef = frozen Qwen3.8-27B backbone + rank-256 LoRA, prefill-only pass with parallel non-autoregressive schema scoring; Clef-flash = frozen Qwen3.5-9B for latency (median 38.8 ms vs Jev's 524.1 ms; Clef itself 209.3 ms). On the authors' numbers Clef tops the Jev Decision Index โ€” BANKING77 macro-F1 94.20 vs Jev's 79.74, CLINC150+OOS 97.43 vs 89.27 โ€” while extending the category (a vision encoder; Jev is text-only) and context (64k vs 32k), staying Jev-API-compatible. The caveats are in the post itself: Jev still wins When2Call (80.97), BRIGHT, and agent-trace observability (71.6 vs 69.8); Laya is still 5.8 ms-fast where latency matters; "we're still early." After a quarter of community replicas (Kev, Jeff, Jeeves โ€” each printing its own caveats), the counter-offensive arrives from a hyperscaler, and it chose openness (Apache-2.0, losing rows published) over the gated-rollout template the big labs have been shipping. The possibly-bigger half: a companion RL fine-tuning platform โ€” AI Gateway traffic as training data, Workers AI (via the Replicate acquisition) for rollouts, Containers for scoring/replay sandbox, a new Trainer โ€” starting as a forward-deployed-engineer service. Every AI Gateway customer's traffic becomes fine-tuning substrate.

"More Choices, Fewer Decisions" (arXiv 2609.38827): the class's first systematic bias audit. Ordinal scale-utilization bias: direct-decision models compress outputs toward a narrow subset of the ordinal scale the user provided โ€” distinct from accuracy, label imbalance, or candidate ordering. On ANLI, JEV assigns 38.8% of predictions and 51.3% of its errors to Neutral despite 74.95% accuracy; across 36 ordinal datasets, final decisions use only 67โ€“76% of effective gold support (vs 87โ€“102% on nominal tasks), degrading to 26โ€“75% at K=14. The hopeful part: controlled experiments show the compression is learned, not architectural โ€” BA-LoRA post-training lifts gold-relative utilization from ~47% to 86% on eight supervised scales. Anyone routing real decisions through these models should check their own scale utilization: the failure is silent, toward the middle of your scale, and the fix is a fine-tune, not a redesign.

Sources: Cloudflare blog ยท HN discussion ยท arXiv 2609.38827

2026-10-03 05:03 โ€” "Jev is poorly calibrated": a $4 audit finds the reference classifier's probabilities don't match reality

Dylan Black (maximumeffort.substack, 22 pts) tested Jev โ€” TypeSafe's System One classifier, the reference point of the decision-model wave โ€” for calibration: do its output probabilities match reality? Method: 10 physics distribution families with known analytic answers, 5 prompt templates ร— 20 variations (1,000 settings, under $4), scored by total-variation distance. Results: Jev's mean TV 0.518 vs 0.546 for a naive flat guess; on the uniform distribution 0.77 vs 0.39 for random โ€” meaningfully worse than chance; Poisson a coin flip (0.65 vs 0.64). The failure mode: "a strong tendency towards distributions that are too peaky" โ€” a near-delta function on the uniform case โ€” and where the peak isn't a given parameter (Maxwell, Rayleigh, Gamma), it found the peak in only ~20% of settings. It does identify the correct family reliably, and its math collapses on multi-step arithmetic and powers of ten. The author is "deeply suspicious" of using Jev as an automated judge, citing prior work where it assigned 83โ€“90%+ confidence to uniform die rolls. The caveat on the caveat: single-author, self-run benchmark in one domain โ€” the same standard this feed applies to vendor charts applies here. A classifier can be accurate and still uncalibrated, and the wave's emerging use case โ€” LLM-as-judge, ordinal-scale collapsing (Clef, Oct 2) โ€” runs on the probabilities, not the argmax. If Jev-class models are peaky by construction, every downstream confidence number inherits it.

Sources: maximumeffort.substack.com ยท HN discussion