Edge / local inference engines (Aug 2026)

A cluster of projects unlocking huge models on tiny hardware. Shared technique: exploit MoE
sparsity โ€” keep the small shared core resident in RAM, stream routed expert weights from disk on
demand โ€” rather than quantizing the whole model.

The pattern

MoE models have a small active-per-token parameter count and a large mostly-idle expert set.
Streaming those experts from SSD/NVMe (with an LRU/LFU cache) turns multi-trillion-parameter models
into consumer-hardware workloads. "Zero quantization, zero distillation" is the common boast.

Projects

Memory-management comparison

Two distinct strategies are hiding under the shared "MoE sparsity" label. Worth keeping separate โ€”
they optimize for different constraints and fail differently.

A. Stream-and-cache (kimi-k3-in-c, TurboFieldfare, h3.c --ssd-streaming) โ€” keep the shared
core resident, stream routed experts from SSD/NVMe on demand, and cache the hot experts. Memory
footprint stays flat no matter how many experts exist; the cost is a cache miss on the first token
after a routing change.

B. Shrink the active set (Ling-3.0-tiny) โ€” make the active per-token footprint so small
(1.3B of 7.9B, KDA:MLA 3:1 hybrid attention) that the whole thing fits in RAM; no disk streaming
at all. Optimizes for latency and deterministic first-token time (<100ms) rather than total
parameter count.

The reusable insight: the engine choice is a trade between scale (A streams arbitrarily many
experts, but pays cache misses) and latency (B never misses, but is capped by what fits in RAM).
The cache policy (LRU vs LFU, per-layer vs global) is the tunable that separates the A-strategy
engines. Watch for the two strategies to merge โ€” a small-resident-core model that also streams
overflow experts on larger hardware.

On-device VLM (a third strategy, Aug 15)

This is neither stream-and-cache (A) nor shrink-the-active-set (B) โ€” it is the *small dense model +
official quantizations* path, complementary to the MoE-streaming engines above. On-device inference
now spans three strategies: stream huge MoEs from disk, shrink the active set, or ship a small model
with first-party quantization.

Fine-tuning with layer streaming (Aug 16)

The "stream the frozen base" trick now spans training, not just inference. Soup
(MakazhanAlpamys/Soup, Apache-2.0) lowers the hardware floor for local fine-tuning: a single YAML
drives SFT/DPO/KTO/ORPO and 20+ methods, and its layer streaming keeps the frozen base in system
RAM while streaming one decoder layer at a time into the GPU โ€” so an **8B model LoRA-finetunes on a
4GB laptop GPU (119.6 tok/s at 3.32GB peak VRAM on an RTX 3050). Results are verified bit-exact**
against a resident-GPU reference across nine architectures as a CI test. Same shape as strategy A
(stream-and-cache) applied to the training pass: the frozen parameters don't need to live in VRAM.
Beta: transformers + plain LoRA only (GRPO/PPO excluded โ€” generation re-reads every layer); migrates
Axolotl/LlamaFactory configs.

On-device training on Apple's Neural Engine (Aug 17 04:03)

A cluster of MIT projects reverse-engineer Apple's private ANE APIs (_ANEClient, _ANECompiler) to
run training โ€” not just inference โ€” on the Neural Engine, with no CoreML or Metal:

Signal: this extends thesis 3's "stream the frozen base" thread from inference to a genuinely new
on-device training substrate โ€” Apple's ANE was inference-only by design. Private APIs and ~5โ€“9%
utilization keep it research-grade for now.

"Will this run on my machine" becomes a tool (Aug 18)

As open models proliferate, the install problem has shifted from "how do I run this" to "does this
fit, and at what quantization" โ€” and two projects productize the answer:

Signal: the edge-inference story now has its selection and serving layers, not just the engines โ€”
llmfit answers "which model + quantization fits this box" and omlx answers "serve it as a persistent
server," both local-first.

Fit-to-measured-budget replaces preset compression โ€” as RAM stops being cheap (Aug 19)

Three independent projects converged on the same reframing within a fortnight: stop choosing a
compression preset, and solve an allocation problem against the bytes you actually measured.

The counterweight that makes this urgent โ€” memory stopped getting cheaper. TrendForce (Aug 17):
Germany's DDR5 retail price index climbed from 445% to 486% year-over-year in August โ€” a typical
kit at ~4.9ร— last year's price โ€” while Shenzhen's Huaqiangbei market saw **DDR5 24Gb +14.29%
week-over-week to $48, 16Gb to $40, and DDR4 8Gb 3200 +12.82% WoW to $22**. TrendForce forecasts
server DRAM contract prices up 13โ€“18% QoQ in 3Q26, calls the market undersupplied, and expects the
server DRAM shortage to run into 2027; Tom's Hardware's retail datapoint is **128 GB of DDR5 for
$3,399** (headline only โ€” its article body is paywalled). Cause: AI-datacenter and HBM demand pulling
fab capacity off commodity parts.

Signal โ€” the two halves of this file now pull against each other. **MoE sparsity + disk streaming
lowered the model's floor (thesis 3's original claim); DRAM pricing just raised the machine's
floor. So the optimization pressure has moved from "make the model smaller" to "spend the exact
bytes you have**" โ€” which is why fit-solvers (Shoehorn, llmfit) and OS-level device-memory QoS
(dmemcg) showed up in the same window. The cgroup work matters beyond gaming for the same reason:
it is the first mainline primitive for arbitrating VRAM between a local model and everything else on
the desktop.

Unsloth becomes a desktop app โ€” run and train collapse into one local tool (Aug 19)

unslothai/unsloth (Apache-2.0, 73,546 stars, pushed Aug 18) quietly changed category: the repo
description now reads "Local UI to run and train LLMs and diffusion models," and Unsloth Desktop
shipped for Windows/macOS/Linux across a fast release train (v0.1.70-beta โ†’ v0.1.800-beta, Aug 11โ€“14)
with no-code training, RAG, MCP, and remote Cloudflare access. The newest release runs **Qwen3.8-27B
locally in ~17 GB RAM** via Dynamic GGUFs plus NVFP4 quants, claims ~10% faster GGUF inference at
lower VRAM, and "Fast FP8 10ร— faster MiniMax-H3 inference (3 minutes vs 30)" with model splitting to
fit smaller GPUs; also landed AMD RDNA 3/4 + Strix Halo support, memory-based context sizing on Mac,
per-model llama-server arguments, and tool calling + web search for external providers.

The trigger is a stack of three model drops landing in one tool inside a fortnight โ€” Desktop's launch
(Aug 11โ€“13), Meta Muse Glimmer support (Aug 10), Qwen3.8 support (Aug 14). Signal: Unsloth was a
fine-tuning library you imported; it is now the local-first GUI for running and adapting a model
on the same hardware, with MCP wired in โ€” collapsing the gap between "try a model" and "adapt a model"
for people who never open a notebook. Together with Shoehorn and llmfit, the local stack now has
selection, fitting, serving, and adaptation as ordinary desktop software.

Ling-3.0 base checkpoints go MIT + a domain-token lesson (08-21 04:03)

Speculative decoding + sparse long-context (08-23 04:03)

Daedalus-150M โ€” designing the KV cache away for the CPU tier (08-24 12:03)

Daedalus-150M (arXiv 2608.20210) is a 150M-parameter LM built for CPU inference: only 6 of 18 blocks use full
attention, 12 use short convolutions "whose memory is two timesteps wide," so two-thirds of the network never
re-reads a growing cache. Trained from scratch on 59.9B tokens with 4-bit weights, it beats GPT-2 124M, Pythia-160M,
OPT-125M and MobileLLM-125M on a pre-registered five-task benchmark (47.31 vs a 42.20 bar) despite those seeing
3ร—โ€“1000ร— more data, and decodes 1.76ร— faster than a same-size all-attention control at 2K context. The point is the
ablation: same data, same size, only the architecture changes โ€” so the KV cache (the main memory cost of
long-context LLMs) is isolated as the lever. Where FreeToken streams experts against a live budget, Daedalus attacks
the other memory cost โ€” the cache itself โ€” by removing attention from most of the network. The edge arc extends: the
KV cache is not a given; it is a design choice the CPU/edge tier can mostly decline.

Second Thought โ€” reasoning in the idle window (08-25)

Second Thought (arXiv 2608.13667, SMU) exploits the "reasoning idle window" in ReAct agents โ€” the time spent
waiting on tool execution and observations โ€” by forking four auxiliary reasoning branches (verification, recall,
rehearsal, fallback) the instant each Thought phase ends, decoding them concurrently with the main loop and merging
when the observation arrives. Training-free. Across 3 benchmarks ร— 3 LLMs it cuts main-thread decoding by up to 43%
(~20% average) and, against a compute-matched control, hits higher Pass@1 at 1.3โ€“3.2ร— less sequential decoding. Not
an edge/quantization trick โ€” it is "think while you wait" at the agent-runtime layer, scaling reasoning without
user-perceived latency or retraining; directly relevant to any runtime that idles on tool I/O.

Apple's 2nm turn โ€” 512 GB / 1.2 TB/s on-device frontier-ish inference (08-26 04:03)

The local-inference ceiling just moved. Apple unveiled the M6 โ€” its first 2nm chip โ€” in a new Mac mini
(dual 16-core Neural Engine, up to 4ร— AI performance over the prior mini, $899) and the M5 Ultra
(quad-die UltraFusion, up to 36-core CPU / 80-core GPU) in a Mac Studio โ€” **512 GB unified memory and 1.2 TB/s
bandwidth**, enough for "hundreds of billions of parameter" on-device models, with LLM prompt processing up to
9.8ร— the M1 Ultra. The Mac Pro is discontinued, making Studio the top desktop. Ships Sep 22 on macOS 27. The
~4.3โ€“4.5ร— AI claims are Apple's own numbers, and the price jumps ($899 mini / $5,499 Studio) reflect the
DRAM-cost environment (the memory-economics note above). Thesis 3's local-inference turn now has a
consumer-adjacent machine that can hold frontier-ish weights resident โ€” the hardware half of bandwidth-adaptive
serving (FreeToken's 284B-on-a-desktop / 753B-on-one-workstation numbers) stops being theoretical.

llama.cpp v0.3.0 + Perplexity Portable Computer โ€” the local stack's runtime and a productized local-first agent (08-26 12:03)

QAH + CarWatch + Groq 3 LPX (08-26 20:19)

The Mask Is Not the Model + ALPHABET (08-27 04:15)

colibri + Baidu Unlimited-OCR โ€” the no-GPU MoE and the constant-KV decoder (08-28 12:15)

ODS โ€” the local-AI installer becomes its own category (09-01 12:22)

slotstream + Tiel-Coder โ€” expert streaming fragments, quant surgery matures (09-02)

Baseten's efficient frontier โ€” which techniques trade tradeoffs and which erase them (09-02)

The M4 Pro blueprint โ€” the concrete middle of the local-LLM market (09-02)

WebLLM โ€” the browser as the zero-install end (09-03)

2026-09-05 04:03

Signal-free eviction; compilation as an alternative to serving (09-05 20:03)

Minima: W4A4 on every linear layer, recurrent ones included (09-06 04:03)

Speculative decoding's printed counter-case (09-08)

vLLM's blog (AMD + Embedded LLM teams, Aug 23; HN trigger Sep 7, 118 pts) walks five drafting methods โ€” native MTP,
Gemma 4 MTP, EAGLE-3, DFlash, DSpark โ€” on Instinct MI300X/MI355X with ROCm. Peak measured: 2.83ร— output-token
throughput (Qwen3.5-122B-A10B, mean accepted length 5.01, acceptance 80.2%); Qwen3.6-35B-A3B with DFlash at 1.77โ€“2.06ร—;
optimal proposal length N = 3โ€“11 by model and dataset. The value is the printed counter-cases: **EAGLE-3's largest
measured MATH500 value remained below the no-spec baseline โ€” it can be slower.** The TL;DR hedges immediately
(results "depended on the model family, draft checkpoint, workload, and acceptance behavior," all numbers "from our
test environment") and insists configs be chosen "using representative workloads and end-to-end measurements."
Speculative decoding has become reflexive advice; the post's discipline is the model for presenting it. Standard
vendor-benchmark caveat applies: AMD measuring AMD.

Quantization with confidence intervals: 4-bit holds, 1-bit collapses (09-09)

Quesma's engineering post (Aug 26, trending Sep 8; read first-hand) is the rare quantization study with
Wilson 95% confidence intervals and an agentic benchmark: ~$3,000 of rented L40S/H100/H200 time
(Terminal-Bench 2.1 alone ~$2,308), llama.cpp, Unsloth GGUF quants of Qwen3.8 27B against a replicated BF16
baseline on GPQA Diamond / IFBench / Terminal-Bench 2.1 (89 tasks, 98k context). Results: Q4_K_M (17 GB)
โ€”"you won't notice a difference on these benchmarks" โ€” matched BF16 on Terminal-Bench while fitting a 24 GB
card; UD-Q2_K_XL (10.7 GB) โ€” "things break a bit" but still "the level of Opus 4.7 or Gemini 3.1 Pro,"
writing ~25% more tokens per solved task (IFBench unchanged down to 2-bit); UD-IQ1_S/M (6.2 GB) โ€”
"scores are around the random guessing level, with the smallest model being below that threshold," and at
xhigh reasoning effort it exhausts its token budget and returns empty answers โ€” longer reasoning hurts
the broken quant. This directly contradicts Unsloth's "retain around 72% top-1% accuracy" marketing:
"that missing ~28% is decisive." The caveats are printed, which is why they're citable: the exact tested
quant files were replaced upstream (Aug 19), Q8_0 was accidentally skipped on Terminal-Bench (interpolated),
KV-cache quantization untested (F16 throughout), and an Aug-16 llama.cpp build was required.

2026-09-09 12:03 โ€” Kimi K3 at 1 tok/s from four SSDs; gpu-lexer

2026-09-10 04:03 โ€” on-device inference becomes a priced SDK tier

2026-09-10 20:03 โ€” the streaming school gets its maintained engine; "will it run" becomes one command; memory stacks on the die

2026-09-14 04:03 โ€” local speech becomes a one-app category; the CUDA-on-Windows moat gets another friction-reducer

2026-09-16 04:03 โ€” ambient local AI ships as furniture; a memory-claim-only MoE trends

2026-09-16 12:03โ†’20:03 โ€” the clean-room GPU driver; local voice makes engines swappable

2026-09-17 12:03โ†’20:03 โ€” NVIDIA makes Rust a native CUDA language; ternary packing beats the 1.58-bit "floor"

2026-09-18 04:03 โ€” colibri re-trends with exact numbers still in print

2026-09-18 12:03โ†’20:03 โ€” ternary gets a second open challenger; quant culture keeps publishing its error bars; the browser replication arrives

Sources: Seoul Economic Daily ยท
HN discussion

2026-09-21 20:03 โ€” the disk-streaming school reaches continual learning: experts as files, one 8 GB GPU

volotat/mini-AGI (Alexey Borsky, MIT, Show HN 136 pts) makes training and inference the same
operation: a byte-level model (256 byte values + 9 structural markers, no tokenizer), PonderNet-style
adaptive halting applied up to 24 times per character, and a growing/pruning Mixture-of-Experts pool
where each expert is a file on disk paged onto the GPU as needed (~540M total params, 32
resident) โ€” the disk-streaming trick from agent-stack's MoE-serving school applied to the
continual-learning problem. The headline result is anti-forgetting: running the trunk at 0.1ร— the
experts' learning rate held measured forgetting to **+0.0067 nats after 524k characters โ€” 99.84%
retained, vs ~50% for other configurations**. Trains from scratch on a single 8 GB CUDA GPU
(reference rig: RTX 3070 Laptop).

The README does the honest-framing work: "as of now this is a small toy-level model," **weights not
published** ("a couple of weeks away"), outputs repetitive, and the nats/char benchmark carries ~0.03
run-to-run variance from nondeterministic CUDA expert dispatch. Treat it as an existence proof that
continual learning fits in modest hardware โ€” measured in nats, not vibes โ€” not a capable model.
Watch: the promised weights release (the claim is unfalsifiable until then), and whether the
expert-paging scheme survives contact with real workloads.

Sources: volotat/mini-AGI ยท
HN discussion

2026-09-22 12:03 โ€” the M5 Ultra review: the local-agent-fleet verdict from someone who actually lives on it

Federico Viticci (MacStories, 236-pt HN) reviews the M5 Ultra Mac Studio โ€” the first UltraFusion
quad-die design (two dual-die M5 Max chips), 80-core GPU, 819 GB/s โ†’ 1.2 TB/s, 256 GB unified
memory (512 GB variant late October). Local-AI numbers with Qwen3.8-Flash-Next 4-bit via oMLX:
prompt processing +150% vs M3 Ultra (~2,733 tok/s), ~108 vs 70 tok/s generation at 16K context,
60โ€“85 tok/s even at 256K, time-to-first-token at 256K halved to ~102s. **Concurrency is the quiet
win**: three parallel requests hit 81.5 tok/s combined (+23%) where the M3 Ultra gained only 4% โ€”
the property that actually matters for agent fleets.

The verdict matters more than the numbers: Viticci now runs his daily agent stack *entirely
on-device* (a 99-day agent research stack at zero API cost). The caveats are unusually clean: an
RTX 5090 still beats it on raw generation (~25% faster) for models that fit in 32 GB; setup is
"not something I would ever recommend" to casual users; the hardware costs more than years of
cloud subscriptions โ€” and the review states no price, the one spec that decides everything.
Same lane, same day: Dettmers' ecosystem drop claims Qwen 3.6 35B-A3B at ~450 tok/s on a Mac via
1.5-bit quantization and DeepSeek V4.1 (550B) on a 128 GB MacBook with automatic context
compression โ€” advocacy with concrete caveats (detail โ†’ frontier-models). The consumer
local-agent endpoint is now being priced in public by people who depend on it, not by vendors.

Sources: MacStories review ยท
HN discussion

2026-09-22 20:03 โ€” gzip as a language model, honestly reported

gzipt (pure-stdlib Python, by the nathan.rs author, 196-pt HN) primes DEFLATE's 32 KiB window with a corpus and scores continuations as len(compress(context + candidate)) โ€” shorter means more "predicted". Two tricks make it work at all: beam search over multi-byte spans (gzip emits integer byte counts, so single-byte steps tie and drown in quantization noise), and keeping only the last tail bytes in the scoring context, since DEFLATE favors cheap nearby matches and full history collapses into verbatim self-copying. The Shakespeare sample comes out recognizably play-formatted and garbled; the author's own verdict is "kind of?" โ€” citing DeepMind's "Language Modeling Is Compression" (arXiv 2309.10668), whose footnote already recorded that gzip-based generation "ended up performing poorly". Worth keeping twice over: a working, zero-trained-parameter demonstration of the compression=prediction equivalence, and a model of how to report a negative result (no benchmarks claimed, caveats in line). The beam-over-byte-spans construction is the actual novelty over the 2023 paper.

Sources: nathan.rs ยท arXiv:2309.10668 ยท HN discussion

2026-09-26 12:40 โ€” W4A4 becomes a library call, not a research project

NVIDIA Model-Optimizer 0.47.0 (NVIDIA/Model-Optimizer, Apache-2.0, 4,513โ˜…, +359/day trending; release Sep 23): a unified library spanning quantization (FP8/NVFP4), pruning, NAS, distillation, speculative decoding and sparsity, with export to TensorRT-LLM, vLLM and SGLang. The trending trigger is a fresh W4A4 tutorial (Sep 16): NVFP4 weights+activations with QAT on Qwen3.6-35B-A3B claiming 1.30ร— vLLM throughput over BF16 and 3.1ร— smaller checkpoints. Standard discount applies: NVIDIA's own tutorial numbers on Nemotron-adjacent models, not an independent benchmark โ€” but the Minima NVFP4 W4A4 line (noted 09-15) now has reproducible tooling an ordinary engineer can run without a research team. W4A4 (weights and activations at 4 bits) is the current post-training-quantization frontier.

Sources: NVIDIA/Model-Optimizer ยท Releases

Prompt-lookup drafting 42ร— faster in llama.cpp โ€” pure data structures, zero accuracy change (Sep 27, HN): n-gram speculation spent 165 ยตs per drafted token on a 541 MB corpus; four optimizations cut it to 3.98 ยตs on an M4 Pro โ€” kill per-step map copying (4.5โ€“25.6ร— on drafting alone), a segmented flat hash map, sorted vectors replacing inner maps (64% of 2-grams have a single follower, so hash maps were waste), and Lemire's immutable constmap for the static cache (6.3โ€“16ร— faster loads). The load-bearing honesty: acceptance rates are "almost identical to the original implementation" โ€” this is caching, not better speculation, and a single-machine benchmark. The honest version of a local-inference speedup claim: states explicitly it changed no model behavior, only made the same guesses cheaper.

Sources: jadidbourbaki.github.io ยท HN

2026-09-27 20:03 โ€” bandwidth-adaptive MoE serving on a gaming PC, trigger unverified

FreeToken (FlashML-org/FreeToken, 13,873โ˜…, v0.1.3 Sep 16, arXiv 2608.16157): datacenter-scale MoE serving on the desktop โ€” bandwidth-adaptive CPU-GPU co-execution of experts, LRU expert caching and elastic VRAM reallocation, targeting DeepSeek-V4-Flash, Qwen3.6-35B-A3B and GLM-5.2 in MXFP4/NVFP4/FP8/BF16 behind OpenAI/Anthropic-compatible APIs on RTX 30/40/50. Repo and paper both read; the caveats travel with it: "blistering interactive speeds" is the project's own framing with no independent benchmark verified, and earlier HN submissions scored only single digits โ€” the star spike lacks a clear external trigger. MoE sparsity plus adaptive expert placement remains the credible path to 290B-class models on consumer hardware (thesis 3's school); worth watching independently once benchmarks replicate.

Sources: FlashML-org/FreeToken ยท arXiv 2608.16157

2026-09-28 04:03 โ€” Ternary Bonsai 2 GGUF tops HF trending at 3.3M downloads; VoiceStudio re-trends as the day's fastest riser

PrismML's Ternary-Bonsai-2-27B-gguf is #1 on Hugging Face trending, 3.34M downloads (weights updated Sep 25) โ€” the 09-18 release now carrying demand numbers: Qwen3.8-27B quantized nearly whole (embeddings, attention/MLP, LM head) to ternary {โˆ’1,0,+1} at a claimed 1.72 bits/weight โ€” ~54 GB FP16 โ†’ ~6 GB, claimed "98.2% of FP16 intelligence retained" (84.78 vs 86.32 avg over 14 thinking-mode benchmarks), ~47 tok/s on an M5 Max; Apache-2.0, MLX companion. The catches are on the model card and they are structural: it requires Prism's custom llama.cpp fork โ€” stock llama.cpp silently loads it as Q2_0, "producing garbage" โ€” quality gaps concentrate in knowledge/reasoning (โˆ’5.7) and vision (โˆ’5.2), all benchmarks self-reported. The trending rank is not independent validation; the fork requirement is exactly the tooling gap that must close before the retention claim can be tested by people outside Prism.

VoiceStudio is the day's fastest GitHub riser (+3,060โ˜…/day, 39.7kโ˜…, pushed Sep 27) โ€” the 09-14 entry's local voice studio now with a demand spike: dense release cadence (v0.5.4โ†’v0.5.6 in three days) plus aggregator virality, while its Show HN flopped at 6 points (GitHub-side trend). The agent-relevant part of the design: a local API + MCP server, so agents drive voice pipelines as tooling โ€” local voice as agent infra, not just a desktop app. Caveats stand: "646 languages"/3-second cloning self-reported, consent-gated analytics.

Sources: prism-ml/Ternary-Bonsai-2-27B-gguf ยท PrismML-Eng/llama.cpp ยท debpalash/VoiceStudio ยท VoiceStudio releases

2026-09-28 05:15 โ€” Ternary Bonsai 2's fork requirement is closing upstream; the first independent measurement exists โ€” and measures the wrong thing for the claim

Checked first-hand this run (GitHub API + HF model card, act pass):

The upstreaming campaign is real and mid-flight. Prism's maintainers are landing the Hadamard-fold support in ggml-org/llama.cpp per-backend: merged โ€” ggml-cpu F16-input FWHT #27779 (09-18), Metal F16-input #29094 (09-20), Metal FWHT block>512 #29095 (09-25), CUDA F16-input #29096 (09-26), SYCL FWHT block>512 #29243 (09-27); open โ€” CUDA block>512 #29100, Vulkan #29101. The strategy is notable: no new GGML types โ€” Hadamard+sign-flip support rides official Q2_0; PQ2_0/PTQ1_0 stay fork-only ("increased maintenance work", khosravipasha in-thread). The community PR implementing both types (#29077) was closed at the maintainer's request โ€” "leave this for PrismML to submit themselves."

But stock llama.cpp still cannot run it today. The Q2_0 testing build (Ternary-Bonsai-2-27B-gguf-dev, 6,898 downloads) loads fine and per its own model card "outputs gibberish with no warning" โ€” the inverse activation transform exists only in the Prism fork. The watch item's fork-requirement clause: in progress, vendor-driven, not closed.

The first independent measurement exists, and its author draws exactly the right line. zhaoyilun/bonsai2-27b-mtp-repro measured MTP speculative-draft acceptance on the folded 27B: found and fixed a folded terminal-norm-gain bug (acceptance 35.6%โ†’40.5%; 0.8B 10.7%โ†’28.4% vs a 25.5% unfolded reference), and swept context depth to 191k tokens โ€” acceptance rises (65.8% @8k โ†’ 84.1% @191k), refuting the compounding-fold-error-with-context worry for speculation. But the same comment states: everything above measures draft/target agreement, not model accuracy โ€” long-context accuracy is "arithmetic, not measurement" (~0.032 nats/token paired KL at m=3 โ†’ ~32 nats by 1k tokens by the chain rule, never measured to 190k). So "98.2% of FP16 intelligence" still has no independent quality benchmark; the HN-reported long-context accuracy drop circulating in #29058 is second-hand paraphrase. The same thread holds a fully reverse-engineered format spec (QuentinDanblon, read out of the Prism fork and checked against the published GGUFs) โ€” the format is now public knowledge, so independent implementations are possible even before upstream lands.

Sources: ggml-org/llama.cpp #29058 ยท zhaoyilun/bonsai2-27b-mtp-repro ยท Ternary-Bonsai-2-27B-gguf-dev

2026-09-28 12:03 + 20:03 โ€” CoyoPedal: full-size neural amp modeling on a $10 microcontroller

CoyoPedal (dashersw/coyopedal, GPL-3.0, 100-pt Show HN): a Neural Amp Modeler guitar rig on the ~$10 Waveshare ESP32-S3-Touch-AMOLED board โ€” full-size NAM A2 captures (a 23-layer, eight-channel WaveNet) at 48 kHz in block floating point with hand-written Xtensa kernels, split across both cores in 64-frame blocks. Drives a class-compliant USB interface as USB host; the touchscreen UI is written in TSX compiled to native C++ โ€” no JavaScript engine on the device โ€” and a WASM build runs the identical DSP and model in the browser. Honest envelope: momentum slowed after the Show HN bump. A different flavor of edge inference than the LLM track โ€” real-time NN DSP on a microcontroller โ€” with a web-tooling-to-native (TSXโ†’C++) pipeline worth stealing beyond audio.

Sources: dashersw/coyopedal ยท Browser demo

2026-09-29 04:03 โ€” disaggregated quantization: prefill accuracy becomes a free variable

Sources: arXiv:2609.26333 ยท ISTA-DASLab GGUF

2026-09-29 05:06 โ€” act: the Bonsai fork gap gets its number โ€” stock llama.cpp PPL 1,258,507

PR #29600 ("Runtime support for Prism Bonsai 2 27B", opened 09-28 17:44Z by bri-prism โ€” Prism's own maintainer, so the upstreaming stays vendor-driven as read on 09-28) puts the fork requirement's cost in the PR body itself, measured with llama.cpp's own KL-divergence harness: under the Prism runtime the Q2_0 GGUF reaches PPL 10.2343 with max KLD 5.3e-5 and 99.975% same-top-p against the reference; on unpatched master the same file scores PPL 1,258,506.97 ยฑ 65,204 โ€” the model card's "silently loads as Q2_0, producing garbage" is now a number, not an adjective. Also new since 09-28: perf follow-ups #29602 (Metal FWHT) and #29605 (SYCL FWHT), both open; the previously-open CUDA #29100 and Vulkan #29101 remain unmerged. Not merged yet โ€” stock llama.cpp still cannot run Bonsai 2 today, and "98.2% of FP16 intelligence" still has no independent quality benchmark. (The PR's own AI-usage disclosure: Claude Code was used to develop and test it.)

Sources: ggml-org/llama.cpp #29600

2026-09-29 12:03 โ€” the hardware floor keeps dropping: a $60 ESP32-S3 cluster runs a 1.58-bit LLM over an SPI daisy-chain

Low-Zi-Hong/ESP32s3-LLM-Cluster (created Aug 6, pushed Sep 26, 90โ˜…, 53+ pts HN): a 0.4B-parameter LLM quantized to 1.58-bit ternary (BitNet-style) weights, sliced across seven ESP32-S3 nodes connected by SPI daisy-chain โ€” each node holds a slice of the weights, collectively performing inference on roughly $60 of microcontrollers. Caveats: a hobby build with no releases; a 0.4B model at BitNet precision is far below useful-model quality; the HN thread debates whether it counts as "real" distributed compute as much as it discusses results. Slow, but real โ€” the ternary floor keeps shrinking the hardware floor for edge LLMs, the same direction as Bonsai 2's 1.76 bits/weight above.

Sources: Low-Zi-Hong/ESP32s3-LLM-Cluster ยท HN discussion

2026-10-01 04:03 โ€” kernels tuned per-hardware on-device: Magnitude's self-optimizing inference engine

Magnitude (YC S25, magnitudedev/magnitude, Rust, Apache-2.0, 5.6kโ˜…, Launch HN 83 pts): an inference engine that tunes its own kernels on-device for your exact hardware before a model runs (~1 minute per download, per the founders) โ€” claiming "up to 2ร— faster than llama.cpp: 92% faster decode on Metal, 19% on CUDA" and "27% less memory per agent," with one-click connect for Pi, OpenCode, Hermes and Codex. The caveats come from the founders' own thread: the headline benchmark is "a simple prose-repetition taskโ€ฆ Moby Dick up to 64k contextโ€ฆ repeat the last section," the MLX comparison is "rough benchmarking," and rigorous numbers are "soon." Per-hardware kernel tuning is how you serve a small decision model cheaply at the edge โ€” and "up to 2ร—" measured on prose repetition is exactly the claim shape this feed discounts until the promised rigorous numbers land. Watch: the rigorous bench, independent Metal/CUDA timings, whether the ~1-minute tuning cost holds across model+hardware pairs.

2026-10-02 12:03 โ€” the DRAM squeeze becomes contractual: Micron's take-or-pay through 2030

Micron's Q4 call (Sep 30, 277 pts) put the supplier side under the RAM-aisle shock this feed covered Sep 30 โ€” and it is structural, by contract. CEO Sanjay Mehrotra: supply-demand will be "much tighter in calendar 2027 and 2028 than in 2026" โ€” "even with any new clean room space coming up in 2028, we see continuing tight supply conditions," with "a structural gap between DRAM supply and demand growth rates." The formalization: FY2026 net income $84B vs $8.5B the prior year, Q4 revenue $54.2B (+379% YoY), 90% datacenter gross margins, 26 multi-year strategic customer agreements representing over 35% of revenue through 2030, a majority with floor-and-ceiling price bands, HBM bit shipments expected to outgrow conventional DRAM through 2028, 1H-FY27 capex ~$25B. Take-or-pay with price floors means the consumer-market shortage is not a 2026 blip โ€” it is contractually guaranteed scarcity through 2030. For the local-inference thesis this is the cost floor hardening under everything here: fit-to-measured-budget and disk-streaming matter more exactly as RAM stops being cheap; spec machines and predict inference costs accordingly.

Sources: The Stack ยท HN discussion

2026-10-03 05:03 โ€” antirez ships ds4: the llama.cpp moment arrives as narrow hand-written C

antirez/ds4 (Salvatore Sanfilippo โ€” the Redis creator โ€” MIT, C, 22,878โ˜…): "a narrow C inference engine for high-memory Mac, CUDA and ROCm machines" that runs DeepSeek V4 / V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next (vision included) entirely on your own hardware. Deliberately "not a generic GGUF runner": asymmetric quantization compresses the routed experts to ~2-bit while keeping shared/critical paths at higher precision (a 284B-class model on 64 GB+ machines), and "KV cache as a disk citizen" persists long prefixes to SSD, resumable by prompt hash. Three interfaces โ€” CLI, an OpenAI/Anthropic-style server, and ds4-agent โ€” share one model state and cache. Stated numbers: M5 Max 128 GB at Q2 does 790.2 t/s prefill / 39.4 t/s generation at 2K context; DGX Spark 825.8/18.1. The dormancy note, verified: the repo was created in May and last pushed Sep 20; the project site went up Sep 17 โ€” today's HN post (25+ pts) surfaces a five-month-old working tool, not a launch. Why it matters: the runner layer for MoE-era frontier models is being won by narrow, per-model-family hand-tuned C โ€” from the author who shipped the last generation's infrastructure software. Watch whether "narrow on purpose" beats "runs everything" the way it did for Redis vs. generic KV stores.

Sources: dwarfstar.sh ยท antirez/ds4 ยท HN discussion