Edge / local inference engines (Aug 2026)
A cluster of projects unlocking huge models on tiny hardware. Shared technique: exploit MoE
sparsity โ keep the small shared core resident in RAM, stream routed expert weights from disk on
demand โ rather than quantizing the whole model.
The pattern
MoE models have a small active-per-token parameter count and a large mostly-idle expert set.
Streaming those experts from SSD/NVMe (with an LRU/LFU cache) turns multi-trillion-parameter models
into consumer-hardware workloads. "Zero quantization, zero distillation" is the common boast.
Projects
- kimi-k3-in-c โ
FareedKhan-dev/kimi-k3-in-c, Apache 2.0. 176KB C99 binary runs Moonshot Kimi K3 (2.78T params) on 8.24GB RAM. MXFP4-packed experts streamed from NVMe, 16/896 experts active, O_DIRECT trunk streaming, expert LRU cache. Byte-identical to the PyTorch reference. - TurboFieldfare โ
drumih/turbo-fieldfare, Apache 2.0. Swift+Metal engine for Gemma 4 26B-A4B on ~2GB RAM (Apple Silicon). ~1.35GB shared core resident, per-layer 16-slot LFU expert cache. - Ling-3.0-tiny โ
inclusionAI/Ling-3.0-tiny(Ant Group Bailing), MIT. 7.9B MoE (1.3B active), KDA:MLA 3:1 hybrid attention, 128 experts. ~90 tok/s on M4 Pro MacBook, <100ms first token. - Muse Glimmer โ
meta-models/Muse-Glimmer-30B(Meta), Apache 2.0. 30B distilled from Muse Spark 1.2, ~17GB 4-bit quantized, 233 tok/s on RTX 5090 via DFlash speculative decoding. - Needle 2 โ
cactus-compute/needle(Cactus Compute), Apache 2.0. 45M params โ 14MB C++ binary; no MLP layers (Walsh-Hadamard transforms), hashed n-gram tables. 500โ800 tok/s on Raspberry Pi 5. Deployed in the Pebble Index 01 smart ring. - h3.c โ
antirez/h3-metal, MIT. C/ObjC + Metal engine for MiniMax H3 omni-modal on Apple Silicon; mmap-from-safetensors loading,--ssd-streamingcuts DiT memory 36.5โ2.0 GiB.
Memory-management comparison
Two distinct strategies are hiding under the shared "MoE sparsity" label. Worth keeping separate โ
they optimize for different constraints and fail differently.
A. Stream-and-cache (kimi-k3-in-c, TurboFieldfare, h3.c --ssd-streaming) โ keep the shared
core resident, stream routed experts from SSD/NVMe on demand, and cache the hot experts. Memory
footprint stays flat no matter how many experts exist; the cost is a cache miss on the first token
after a routing change.
- kimi-k3-in-c โ largest scale (2.78T โ 8.24GB RAM). Expert LRU cache,
O_DIRECTtrunk streaming, 16/896 experts active. Byte-identical to the PyTorch reference is the notable claim. - TurboFieldfare โ tightest footprint (~2GB). Per-layer 16-slot LFU cache, ~1.35GB shared core resident. LFU (not LRU) because the active-expert set is small and hot per layer.
- h3.c โ the general mechanism exposed as a flag (
--ssd-streaming), applied to the DiT (diffusion) stack, not just the LLM โ proves the trick is modality-agnostic.
B. Shrink the active set (Ling-3.0-tiny) โ make the active per-token footprint so small
(1.3B of 7.9B, KDA:MLA 3:1 hybrid attention) that the whole thing fits in RAM; no disk streaming
at all. Optimizes for latency and deterministic first-token time (<100ms) rather than total
parameter count.
The reusable insight: the engine choice is a trade between scale (A streams arbitrarily many
experts, but pays cache misses) and latency (B never misses, but is capped by what fits in RAM).
The cache policy (LRU vs LFU, per-layer vs global) is the tunable that separates the A-strategy
engines. Watch for the two strategies to merge โ a small-resident-core model that also streams
overflow experts on larger hardware.
On-device VLM (a third strategy, Aug 15)
- LFM2.5-VL-3B โ
LiquidAI/LFM2.5-VL-3B, lfm1.0 license. A ~3.1B vision-language model (LFM2.5-2.6B backbone + SigLIP2 NaFlex encoder) built for the GUI-agent niche โ reading screens and grounding objects locally on phones/laptops that can't host a 27B model. 228 tok/s on Apple M5 Max, ~20 tok/s on a Galaxy S26 Ultra in under 3.3 GB; ScreenSpot-v2 80.7, RefCOCO P@1 87.9, ChartQA 81.3, 16 languages. Official GGUF/ONNX/MLX quantizations ship.
This is neither stream-and-cache (A) nor shrink-the-active-set (B) โ it is the *small dense model +
official quantizations* path, complementary to the MoE-streaming engines above. On-device inference
now spans three strategies: stream huge MoEs from disk, shrink the active set, or ship a small model
with first-party quantization.
Fine-tuning with layer streaming (Aug 16)
The "stream the frozen base" trick now spans training, not just inference. Soup
(MakazhanAlpamys/Soup, Apache-2.0) lowers the hardware floor for local fine-tuning: a single YAML
drives SFT/DPO/KTO/ORPO and 20+ methods, and its layer streaming keeps the frozen base in system
RAM while streaming one decoder layer at a time into the GPU โ so an **8B model LoRA-finetunes on a
4GB laptop GPU (119.6 tok/s at 3.32GB peak VRAM on an RTX 3050). Results are verified bit-exact**
against a resident-GPU reference across nine architectures as a CI test. Same shape as strategy A
(stream-and-cache) applied to the training pass: the frozen parameters don't need to live in VRAM.
Beta: transformers + plain LoRA only (GRPO/PPO excluded โ generation re-reads every layer); migrates
Axolotl/LlamaFactory configs.
On-device training on Apple's Neural Engine (Aug 17 04:03)
A cluster of MIT projects reverse-engineer Apple's private ANE APIs (_ANEClient, _ANECompiler) to
run training โ not just inference โ on the Neural Engine, with no CoreML or Metal:
- ANE (
maderix/ANE) โ the proof of concept: forward + backward on Stories110M, ~91โ115 ms/step. - Orion (
mechramc/Orion) โ a graph compiler with "Delta Compilation" (8.5ร faster weight updates) and stable 1,000-step training of a 110M transformer in ~22 min. - ANEForge (
sbryngelson/ANEForge) โ a pip-installable Python binding (~75 tok/s, 8โ16ร more energy-efficient than GPU on tested models).
Signal: this extends thesis 3's "stream the frozen base" thread from inference to a genuinely new
on-device training substrate โ Apple's ANE was inference-only by design. Private APIs and ~5โ9%
utilization keep it research-grade for now.
"Will this run on my machine" becomes a tool (Aug 18)
As open models proliferate, the install problem has shifted from "how do I run this" to "does this
fit, and at what quantization" โ and two projects productize the answer:
- llmfit โ
AlexsJones/llmfit, MIT, ~32k stars, Rust CLI. Detects RAM/CPU/GPU/VRAM/backend, then scores hundreds of models across memory-fit, estimated speed, quality, and context โ using a memory-bandwidth model with a ~80-GPU lookup table โ and picks the highest quantization that fits. It correctly sizes MoE models by active parameters (Mixtral 8x7B drops from ~23.9GB to ~6.6GB), andllmfit benchmeasures real tok/s that users contribute back via PR to replace estimates.llmfit recommend --jsonis built for scripts/agents, andllmfit planinverts the question to "what hardware do I need for this model?" โ hardware detection + quantization selection as a one-command, agent-scriptable answer. - omlx โ
jundot/omlx, Apache-2.0, ~19k stars, SwiftUI macOS app (originated from vllm-mlx). Runs LLMs/VLMs natively on Apple Silicon via MLX and exposes OpenAI/Anthropic-compatible APIs on localhost. The standout is a two-tier KV cache โ a hot RAM tier plus a cold SSD tier persisted as safetensors that survives restarts โ plus continuous batching, multi-model serving with LRU eviction, an 8GB-below-RAM memory enforcer, and MCP/structured-output support (LLMs, VLMs, OCR, embeddings, rerankers, optional distributed multi-Mac inference). Apple Silicon's unified memory is the best budget host for local models, and omlx turns it into a real (SSD-backed, batching) server โ another step toward the Mac-as-inference-node.
Signal: the edge-inference story now has its selection and serving layers, not just the engines โ
llmfit answers "which model + quantization fits this box" and omlx answers "serve it as a persistent
server," both local-first.
Fit-to-measured-budget replaces preset compression โ as RAM stops being cheap (Aug 19)
Three independent projects converged on the same reframing within a fortnight: stop choosing a
compression preset, and solve an allocation problem against the bytes you actually measured.
- Shoehorn (MIT, Rust, created Aug 13) โ inverts quantization selection. Instead of picking a preset that ignores the machine, it "starts from the memory you actually have, subtracts what inference itself needs, and solves a per-tensor mixed-precision assignment" against the remainder. Reported fits are extreme: "routinely using 99.99% of the budget, sometimes to the byte," with a worked example of 519.2 MiB of a 519.2 MiB budget โ 99.998% used, 13 KB slack for
unsloth/Qwen3-4B-GGUF. The quantizer is written from scratch in Rust (no llama.cpp code linked) and emits standard GGUF v3, with llama.cpp only as the inference backend, so nothing downstream changes;shoehorn uimeasures the machine, streams the fit, and reports the perplexity cost before you chat. Targets macOS Apple Silicon, Linux x86-64 (NVIDIA/AMD), Windows x86-64 (NVIDIA), profiles from 8 GB to 128 GB, contexts 4kโ32k. Very young โ 37 stars at time of check, so treat the 99.998% figure as an author demo, not an independent result. - Linux VRAM overcommit โ Valve contractor Natalie Vock shipped work stopping Linux from evicting a foreground game's VRAM to system RAM under GPU memory pressure. It builds on the
dmemcgroup controller (dmemcg, co-developed with Maarten Lankhorst/Intel + Maxime Ripard/Red Hat, already mainline) and adds six kernel patches plus two userspace helpers โdmemcg-boosterand a KDE Plasma "Foreground Booster" fork โ so the foreground app wins VRAM and background apps are evicted first. Covers AMDamdgpuand Intelxe; NVIDIA has no equivalent mechanism. Worked example: background apps left only 6.1 GB of an 8 GB card for a title needing 7.4 GB; the patches hand over 1 GB back. Available now via CachyOS (Linux 7.0rc7-2+) and thelinux-dmemcgAUR package. - llmfit (above) โ the same shape one level up: measure the box, then pick the highest quantization that fits, sizing MoE by active parameters.
The counterweight that makes this urgent โ memory stopped getting cheaper. TrendForce (Aug 17):
Germany's DDR5 retail price index climbed from 445% to 486% year-over-year in August โ a typical
kit at ~4.9ร last year's price โ while Shenzhen's Huaqiangbei market saw **DDR5 24Gb +14.29%
week-over-week to $48, 16Gb to $40, and DDR4 8Gb 3200 +12.82% WoW to $22**. TrendForce forecasts
server DRAM contract prices up 13โ18% QoQ in 3Q26, calls the market undersupplied, and expects the
server DRAM shortage to run into 2027; Tom's Hardware's retail datapoint is **128 GB of DDR5 for
$3,399** (headline only โ its article body is paywalled). Cause: AI-datacenter and HBM demand pulling
fab capacity off commodity parts.
Signal โ the two halves of this file now pull against each other. **MoE sparsity + disk streaming
lowered the model's floor (thesis 3's original claim); DRAM pricing just raised the machine's
floor. So the optimization pressure has moved from "make the model smaller" to "spend the exact
bytes you have**" โ which is why fit-solvers (Shoehorn, llmfit) and OS-level device-memory QoS
(dmemcg) showed up in the same window. The cgroup work matters beyond gaming for the same reason:
it is the first mainline primitive for arbitrating VRAM between a local model and everything else on
the desktop.
Unsloth becomes a desktop app โ run and train collapse into one local tool (Aug 19)
unslothai/unsloth (Apache-2.0, 73,546 stars, pushed Aug 18) quietly changed category: the repo
description now reads "Local UI to run and train LLMs and diffusion models," and Unsloth Desktop
shipped for Windows/macOS/Linux across a fast release train (v0.1.70-beta โ v0.1.800-beta, Aug 11โ14)
with no-code training, RAG, MCP, and remote Cloudflare access. The newest release runs **Qwen3.8-27B
locally in ~17 GB RAM** via Dynamic GGUFs plus NVFP4 quants, claims ~10% faster GGUF inference at
lower VRAM, and "Fast FP8 10ร faster MiniMax-H3 inference (3 minutes vs 30)" with model splitting to
fit smaller GPUs; also landed AMD RDNA 3/4 + Strix Halo support, memory-based context sizing on Mac,
per-model llama-server arguments, and tool calling + web search for external providers.
The trigger is a stack of three model drops landing in one tool inside a fortnight โ Desktop's launch
(Aug 11โ13), Meta Muse Glimmer support (Aug 10), Qwen3.8 support (Aug 14). Signal: Unsloth was a
fine-tuning library you imported; it is now the local-first GUI for running and adapting a model
on the same hardware, with MCP wired in โ collapsing the gap between "try a model" and "adapt a model"
for people who never open a notebook. Together with Shoehorn and llmfit, the local stack now has
selection, fitting, serving, and adaptation as ordinary desktop software.
Ling-3.0 base checkpoints go MIT + a domain-token lesson (08-21 04:03)
- Ling-3.0 base checkpoints โ the
Ling-3.0-tinyentry above gained a research-grade sibling: Ant Group/inclusionAI releasedLing-3.0-tiny-base(7.9B/1.3B active) andLing-3.0-flash-base(124B/5.1B active) plus six checkpoints spanning pre-training, mid-training and WSM-merged stages, all under MIT. Intermediate training stages are what researchers normally never see โ continued pre-training and MoE ablation on a frontier-adjacent model become possible. - RollTab โ a 125M decoder-only transformer that continues live MIDI piano on an iPhone (Core ML, INT8, ~108 notes/s on iPhone 15). The transferable lesson is tokenization: a single NOTE token carrying five categorical fields (event type, pitch, delta onset, duration, velocity) run once per note instead of once per field is what makes 125M feel real-time. Domain-specific tokenization beating brute scale โ the on-device mirror of "spend the exact bytes you have."
Speculative decoding + sparse long-context (08-23 04:03)
- Liquid AI DSpark โ self-contained speculative-decoding draft checkpoints (1.2B/2.6B/8B-A1B) that accelerate LFM2.5 with guaranteed-identical greedy output (draft tokens accepted only when they match the target distribution): up to 3.18ร H100 throughput (428โ1362 tok/s on MATH500), 2.87ร on M4 Max (136โ389 tok/s), 57% average latency cut on multi-tool function calling, day-one llama.cpp + SGLang. A pure ~3ร speedup with zero quality loss, spanning data-center to MacBook โ the "spend the exact bytes you have" optimization now has a guaranteed-lossless speculative variant (thesis 3).
- KeysAndValues (AWS, arXiv 2608.19920) โ a fine-tuning method for long-context sparse attention that works for any KV-cache policy on a single A100 40GB, letting the model co-adapt with the policy and often beating exact sequence-parallel attention; ships H2O kernels + an OSS library. Removes the sequence-parallel requirement that made long-context sparse fine-tuning impractical on modest hardware.
- Known repos, new facts.
jundot/omlx(~20.3k stars) added a DeepSeek-V4-Flash M2-Ultra kernel and cut ANE compilation memory 35.8GBโ4.7GB (0.6.3rc2);AlexsJones/llmfit(~33.6k) continues its "measure and share" tok/s PR loop. Both were covered 08-18 โ the dedup-window widening (see agent-stack) is what should have framed these as updates, not fresh discoveries. - FreeToken (arXiv 2608.16157, submitted 2026-08-17;
FlashML-org/FreeToken, Apache-2.0, 2,824โ , created 2026-07-20, pushed daily) โ "Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution" (Shuo Yang, Xiaoze Fan, Melissa Pan, Haocheng Xi, Zhe Wang, Shanlin Sun, Kurt Keutzer, Song Han, Matei Zaharia, Chenfeng Xu, Ion Stoica). The strongest instance yet of this file's thesis, and it generalizes it. Where the earlier work streamed experts against a fixed plan, FreeToken treats the whole personal machine โ GPU, CPU, RAM, PCIe, disk โ as "a unified, elastic inference platform" and, instead of a fixed offloading strategy, "continuously maps computation and model state onto the resources actually available," co-designing model layout/loading, expert residency, CPUโGPU execution, agentic state reuse and runtime memory management. Verified on the abstract page: 35B on an 8 GB laptop GPU, 284B on a gaming desktop, the 753B GLM-5.2 on a single workstation GPU, and 20+ MoE models. Read the speedup carefully: the 1.3โ2.1ร mean decode gain over llama.cpp / Ollama / KTransformers / MoE-Infinity is not on the abstract page โ it lives in the PDF, so cite it as a paper claim, not an abstract-verified figure. The sharpest line is the motivation, not the numbers. FreeToken justifies adaptivity by arguing that agent workloads "continuously change their execution pattern" โ prefill-heavy tool reads, decode-heavy reasoning, bursts of state reuse โ so a static offloading plan is wrong most of the time by construction. Local serving is now being designed against agentic variance rather than chat, which is why the optimization target moved again: 08-19 was fit-to-a-measured-budget (static), this is fit-to-the-budget-you-have-right-now (dynamic). It closes the arc that started with the DRAM price shock โ if you cannot buy the bytes, schedule them. - FlashPrefill V2 (arXiv 2608.19758;
qhfan/FlashPrefillv2, Apache-2.0) โ block-sparse prefill attention with a mean-correction term that suppresses approximation error at extreme sparsity, plus a PackGQA sparse operator with warp specialization and pingpong pipelining (FP8/BF16), paged KV cache, continuous batching and a drop-in SGLang backend. Reports up to 47.26ร over FlashAttention-2 (FP8) / 27.19ร (BF16) at 128K context on an H20. Freshness/credibility caveat, checked first-hand: the repo was created 2026-08-19 and had 8 stars when read โ a two-day-old artifact with a 47ร headline and no third-party replication. The right posture is the InferenceX one (see frontier-models): a kernel claim this large belongs on a standing, continuously-run harness before it is treated as a fact. Note also the axis: FreeToken optimizes the edge; FlashPrefill V2 optimizes datacenter long-context serving. Same year, opposite ends of the hardware curve, and only the first one is about the machine on your desk.
Daedalus-150M โ designing the KV cache away for the CPU tier (08-24 12:03)
Daedalus-150M (arXiv 2608.20210) is a 150M-parameter LM built for CPU inference: only 6 of 18 blocks use full
attention, 12 use short convolutions "whose memory is two timesteps wide," so two-thirds of the network never
re-reads a growing cache. Trained from scratch on 59.9B tokens with 4-bit weights, it beats GPT-2 124M, Pythia-160M,
OPT-125M and MobileLLM-125M on a pre-registered five-task benchmark (47.31 vs a 42.20 bar) despite those seeing
3รโ1000ร more data, and decodes 1.76ร faster than a same-size all-attention control at 2K context. The point is the
ablation: same data, same size, only the architecture changes โ so the KV cache (the main memory cost of
long-context LLMs) is isolated as the lever. Where FreeToken streams experts against a live budget, Daedalus attacks
the other memory cost โ the cache itself โ by removing attention from most of the network. The edge arc extends: the
KV cache is not a given; it is a design choice the CPU/edge tier can mostly decline.
Second Thought โ reasoning in the idle window (08-25)
Second Thought (arXiv 2608.13667, SMU) exploits the "reasoning idle window" in ReAct agents โ the time spent
waiting on tool execution and observations โ by forking four auxiliary reasoning branches (verification, recall,
rehearsal, fallback) the instant each Thought phase ends, decoding them concurrently with the main loop and merging
when the observation arrives. Training-free. Across 3 benchmarks ร 3 LLMs it cuts main-thread decoding by up to 43%
(~20% average) and, against a compute-matched control, hits higher Pass@1 at 1.3โ3.2ร less sequential decoding. Not
an edge/quantization trick โ it is "think while you wait" at the agent-runtime layer, scaling reasoning without
user-perceived latency or retraining; directly relevant to any runtime that idles on tool I/O.
Apple's 2nm turn โ 512 GB / 1.2 TB/s on-device frontier-ish inference (08-26 04:03)
The local-inference ceiling just moved. Apple unveiled the M6 โ its first 2nm chip โ in a new Mac mini
(dual 16-core Neural Engine, up to 4ร AI performance over the prior mini, $899) and the M5 Ultra
(quad-die UltraFusion, up to 36-core CPU / 80-core GPU) in a Mac Studio โ **512 GB unified memory and 1.2 TB/s
bandwidth**, enough for "hundreds of billions of parameter" on-device models, with LLM prompt processing up to
9.8ร the M1 Ultra. The Mac Pro is discontinued, making Studio the top desktop. Ships Sep 22 on macOS 27. The
~4.3โ4.5ร AI claims are Apple's own numbers, and the price jumps ($899 mini / $5,499 Studio) reflect the
DRAM-cost environment (the memory-economics note above). Thesis 3's local-inference turn now has a
consumer-adjacent machine that can hold frontier-ish weights resident โ the hardware half of bandwidth-adaptive
serving (FreeToken's 284B-on-a-desktop / 753B-on-one-workstation numbers) stops being theoretical.
llama.cpp v0.3.0 + Perplexity Portable Computer โ the local stack's runtime and a productized local-first agent (08-26 12:03)
- llama.cpp v0.3.0 (ggml-org) โ the reference local-inference runtime's first 0.x major bump in a long while. The
mtmdmultimodal library adds dots3-note vision and audio (a new DSA-ISWA KV cache type), WebP decoding via ffmpeg, a Pillow-accurate resize algorithm, and a fix for videos whosemoovatom sits at end-of-file; GLM-4.5-Air gains MTP, DeepSeek 4 gets a tensor-split mode, and the core bumps to ggml v0.22.0 (meta-backend tensor split, per-op Metal kernels with parallel compilation, a proper non-in-placeggml_clamp). Multimodal + video handling consolidate into the one binary most local-AI tooling builds on โ a first-class update signal for the whole local-inference ecosystem (thesis 3). - Perplexity Portable Computer โ "local-first with opt-in cloud" productized, on NVIDIA DGX Spark. A fully on-device version of Perplexity's Computer agent platform built in close cooperation with NVIDIA: local models (Qwen 3.8 27B or Perplexity's post-trained PPLX 27B), agent harness, tool router, connectors, and an OS-level sandbox all run locally, and local work consumes zero token credits โ escalation to 15+ cloud frontier models requires explicit approval and returns text-only advice with no local-file access. On its Local Knowledge Work Bench it scores 82.6% (85.4% with PPLX 27B) vs Pi 77.6 / Hermes 74.0, and uses ~70% fewer tokens than Pi on BrowseComp. The argument to track: local agents need a co-designed harness, not a general-purpose one โ the small-model agent debate reframed as a harness-design problem (thesis 12's lever at the edge), landing on the hardware ceiling above. Independence check (08-26 12:27): the benchmark is still vendor-run โ Perplexity says it plans to open-source Local Knowledge Work Bench but hasn't, and no third-party reproduction exists. The co-design mechanism has independent support: the harness-premium literature (thesis 12, arXiv:2605.30621) found weak models fail to load and adhere to general-purpose harnesses (skill-load 0.251, adherence 0.52โ0.13) โ exactly the "small models buckle under general-purpose harness assumptions" failure Portable Computer names. Perplexity's own breakdown credits ~5 of the ~12 pts over Pi to the harness stack (base Qwen also beats Pi ~5) + only 2.8 to PPLX post-training โ treat as a directional claim, not a spec, until the benchmark is open-sourced. A DIY replication path (Ollama + Qwen3.8-27B + OpenCode on a 24 GB GPU) exists but is not an independent benchmark.
QAH + CarWatch + Groq 3 LPX (08-26 20:19)
- QAH โ quantization-aware healing makes a 4-bit model beat its bf16 source (arXiv 2608.20953, Multiverse Computing). Replaces the degraded intermediate teacher in QAT/QAD: the 4-bit student is distilled directly from the original full-precision model via KL divergence. Applied to GPT-OSS 120B โ 60B โ MXFP4, the QAH student matched or beat its bfloat16 source on 7 of 9 benchmarks (AA-LCR 42.7 vs 35.3, AIME 2025 76.3 vs 70.7, Aider 40.9 vs 38.2) and edged the 120B teacher on LiveCodeBench โ at ~half the weights and compute per token. On GPT-OSS 9B it peaks ~7ร faster than QAT and stays within ~2 points of peak for 1,200 steps while QAT loses ~19. Ships open-weight as HyperNova-60B (Apache-2.0). Caveat: Multiverse's own measurements on its own pipeline (proprietary compression, GPT-OSS-only) โ "beats bf16" is a result to reproduce, not yet an independent fact. The fit-to-budget quantization turn (thesis 3) gains a "compression that heals" variant: if 4-bit + half the parameters can match full precision, the dominant serving cost of open models drops.
- CarWatch โ a Raspberry Pi 5 as a fully offline car agent (
ThinkOffApp/CarWatch, AGPL-3.0, 171โ , Show HN). Serves Qwen3.6-35B-A3B locally (~14.3 GB quant, ~3.5 tok/s) with RAG over the 745-page owner's manual, reads OBD-II via a Bluetooth ELM327, and issues make-safe cloud commands (lock doors, close windows) through Home Assistant. Hands-free voice is fully on-device โ continuous VAD โ whisper.cpp โ grounded answer. A ~$100 device running a 35B local model is a concrete "local AI" end state; the split between read-only OBD-II access and explicitly make-safe commands is a sane safety model for an on-device agent. - NVIDIA Groq 3 LPX โ a decode engine enters full production (~3,400 tok/s on Gemma 4 31B at 100K ctx). At Hot Chips 2026 NVIDIA announced the decode-phase LPU chip from the Groq acquisition (complementary to Vera Rubin) is in full production. Artificial Analysis measured ~3,400 output tok/s on Gemma 4 31B at 100K context (zhidx cites a 3,431 tok/s median at 100K, nearly flat vs 10K; 4,767 SPEED-Bench coding median). Each rack holds 256 LP30 accelerators (128 GB on-chip SRAM, 640 TB/s scale-up, liquid-cooled); Nebius is the first cloud via its Token Factory platform. The hardware bet that multi-turn agent workloads, not chat, are the binding inference constraint โ a "decode engine" complementary to the fit-to-budget turn. Caveat (verified 08-26 20:19): the headline number is Artificial Analysis-measured but on a private pre-release endpoint (Aug 21), not production serverless โ an independent evaluator, but not yet a production measurement; the 4ร/30ร claims are vendor projections.
The Mask Is Not the Model + ALPHABET (08-27 04:15)
- Causal leakage in shipped hybrid models โ "The Mask Is Not the Model" (arXiv 2608.22876). The field's default causal-correctness check โ inspecting attention masks โ is fundamentally insufficient; the paper formalizes prefix invariance and ships a one-page, two-forward-pass audit that scores each layer. Testing 8 released checkpoints via 192 injected-fault trials, it found real defects in two: Zamba2 and Nemotron-H leak information exactly at chunked-scan boundaries in their recurrent/scan component. The mask is correct, but inter-chunk aggregation leaks โ "causality is a graph-level property." Mask inspection "detected none, while our audit localized all 192/192 to the exact layer." Why it matters: causal leaks in shipped, widely-downloaded open models mean future-context contamination in pretrained weights โ and the lesson extends to every scan/aggregation architecture now shipping, including the new DeltaNet/QSA hybrids (Qwen3.8-Flash-Next) and Kimi K3's linear-attention core. The audit tooling is cheap (one page, two forward passes); the open question is whether it gets applied to the new hybrids before they ship (โ agenda).
- ALPHABET โ a 6,437-parameter linear-time sequence model approaches a Bayes oracle (arXiv 2608.24051). Compresses temporal history into stable complex "pole modes" via a direct bank (resynthesis into the feature trajectory), an independent cascaded bank, and an affine head that reads only modal energies and lag moments โ an "explicitly auditable prediction interface" at width D=64. On a Gaussian control task its learned descriptor approaches the Bayes oracle where raw autocovariances perform at chance; mean rank 3.97 across an 82-task registry; 5.02ร faster inference / 3.93ร faster training than nine baselines; each mode energy ties to a frequency-localized measurement of the second-order spectrum. The extreme end of the "tiny efficient models" trend (sits beside Daedalus's KV-cache-elimination) โ and an auditable internal state (modal energies, not black-box activations) is a genuine differentiator for control tasks where you need to know why the model decided.
- The audit tooling got productized the same week โ VIDRAFT AX-RAY (08-27 04:30, verified first-hand). The Mask-paper authors (VIDRAFT, a Korean AI company, CEO Kim Min-sik) shipped the diagnostic as AX-RAY / FINAL-Bench Diagnostics โ a public AI-safety catalog of 117 diagnostic items across three axes and eleven operational categories (prefix invariance / causal leak, chunked-scan & masking consistency, KV-cache path consistency, cross-implementation consistency), mapping items to legal/regulatory/ethical contexts and treating confirmed causal leakage as a blocking defect regardless of score. It is being positioned as the verification layer for South Korea's government "cybersecurity-specialized AI foundation model" project โ the audit now has a regulatory customer. Watch, answered-for-now (08-27 04:30): no published prefix-invariance audit for Qwen3.8-Flash-Next (Gated DeltaNet + QSA) or GLM-5.3-Flash (sparse + linear) โ by the labs or a third party โ as of this run. The root cause is now a code-level census item: in
transformers5.7.0,modeling_zamba2.py+modeling_nemotron_h.pypermute then reduce over the output chunk axis wheremodeling_mamba2.pyreduces over the input chunk axis โ and the defect fires only when fast kernels are absent (CPU/CI slow path). The audit is cheap enough that its absence from the new hybrid releases is now itself a signal.
colibri + Baidu Unlimited-OCR โ the no-GPU MoE and the constant-KV decoder (08-28 12:15)
- colibri (
JustVugg/colibri, Apache-2.0, pure C, 26.3kโ ) โ the strongest "no-GPU frontier" engine yet. Treats VRAM, RAM and NVMe as one memory hierarchy: the ~19,456 routed experts of a 744B MoE live on disk (~370 GB) and are streamed on demand through a per-layer LRU cache with learned hot-pins, batch-union reads,O_DIRECTand dual-SSD mirroring. Runs GLM-5.2 (744B), Kimi K3 (2.8T), Inkling (975B), DeepSeek-V4-Flash, Qwen3.6 and OLMoE โ "none of them needs a GPU"; speed is disk-bound and a GPU only helps. v1.8.0, active maintenance (77 open issues, 40 PRs). Expert streaming collapses the assumption that frontier MoE inference needs a datacenter โ the same pressure that makes 2.8T-parameter models claimable by a laptop (thesis 3, alongside kimi-k3-in-c / FreeToken). - Baidu Unlimited-OCR (
baidu/Unlimited-OCR, MIT, 24.7kโ ) โ one-shot long-horizon document parsing with a constant KV cache. Replaces all decoder attention layers of a DeepSeek-OCR-style pipeline with Reference Sliding Window Attention (R-SWA): a globally-visible reference segment of visual tokens plus a 128-token sliding decode window keeps the KV cache constant, so dozens of pages transcribe in a single forward pass instead of page-by-page loops that reset memory. The 3B-total / 500M-active MoE decoder compresses a 1024ร1024 PDF page to 256 visual tokens (16ร), with single-page ("gundam") and multi-page ("base") modes. Reaches SOTA on OmniDocBench v1.5/v1.6 single-page end-to-end parsing; authors argue R-SWA generalizes to ASR and translation. "Soft forgetting" is the actual fix for the KV-growth wall โ a general attention pattern, not a wrapper (thesis 3, beside Daedalus-150M's cache elimination).
ODS โ the local-AI installer becomes its own category (09-01 12:22)
Osmantic/ODS(Apache-2.0, 5.6kโ , v2.6.0 stable) โ acurl | bashinstaller (PowerShell block on Windows, Docker required) that assembles a full local stack: llama-server, Open WebUI, LiteLLM, Whisper, Kokoro TTS, the Hermes agent, n8n, Qdrant, SearXNG and ComfyUI. Auto-detects NVIDIA / AMD (incl. Strix Halo unified memory) / Intel Arc / Apple Silicon / CPU, picks a model tier to fit the VRAM/RAM envelope, and "bootstrap mode" serves a 1.5B model in under 2 minutes while the real model downloads in background and hot-swaps in. Every service is a drop-in extension (manifest + compose file) managed by anodsCLI; local-first by default, cloud/hybrid optional. The integration tax is the product โ local-AI installers are becoming a category at exactly the moment the new hardware wave (Strix Halo, Mac Studio clusters) gives people machines to point them at. Caveats: ~1.4k open PRs against 3.2k commits is an unusual maintenance shape; the "sovereign human right" framing is the project's own marketing; no third-party benchmarks of the assembled stack.
slotstream + Tiel-Coder โ expert streaming fragments, quant surgery matures (09-02)
- slotstream (
carloslfu/slotstream, Show HN, 82 pts) โ a fifth-plus parallel implementation of the expert-streaming thesis; the fragmentation is the finding. A single Swift/MLX binary running Qwen3.8-Flash-Next (125B MoE, 104GB at 4-bit) on Macs that can't hold it in RAM: a ~3.8GB resident dense trunk plus the 32GB n-gram table stay in unified memory, while the 68GB of routed experts (512 per layer, 10 active) are read on demand viapreadinto a fixed pool of cache slots shared across all 48 layers, auto-resized every 15s. Measured ~12 tok/s warm on a 48GB M5 Pro (~32GB peak). The right kind of claim: greedy decoding byte-identical between 4GB and 24GB caches, "enforced as a standing test" โ falsifiable, in CI. Stated limits: this one model only; the entire prompt prefills before the first token (~70s at 8k tokens); 32k context; no tools, images, or JSON-schema outputs (HTTP 400); non-48GB figures are estimates. The top comment lists at least five prior repos doing essentially the same thing (mlx-moe-offload, streamlx, mlx-moe, mlx-flash, deepseek-v4-flash-mlx) and asks for collaboration rather than another README โ the space is fragmenting exactly like the routing-DSL layer did (โ smart-routing): many engines, no shared implementation, and the differentiation is in cache policy and API compatibility. - Tiel-Coder-35B-A3B (peculiar-ragdoll, GGUF of MIT Ornith-1.5-35B-A3B, 87.8k downloads) โ template + imatrix surgery now plausibly rivals frontier-medium agentic coding at 22GB. A custom coding-weighted imatrix and a new "Sharp" chat template; the card claims 12/25 fixes on SWE-bench-Live, "the same as Opus 4.6 (medium)," at an 8.6-minute median per attempt, and the best multi-turn conversation of any local model measured (67.2 Claw-Eval vs the base's 65.3), vision inherited from the base's BF16 mmproj projector. The load-bearing caveat is the card's own: "Benchmarks are one run per problem on SWE-bench-Liveโฆ treat small differences as noise" (n=25) โ the honest headline, in the same class as the disclaimer-stripping cases in fact-check but on the virtuous side. Bonus quality-control story: a side note documents that Ornith's original MTP head shipped as random-init weights until a trained one was re-uploaded Aug 23, verified by kurtosis statistics โ check the checkpoints, not the card.
Baseten's efficient frontier โ which techniques trade tradeoffs and which erase them (09-02)
- Philip Kiely imports portfolio theory into inference engineering: every deployment sits on a latencyโthroughput efficient frontier, and techniques divide into those that move you along it (batch sizing, tensor/expert/attention-data parallelism) and those that push the frontier out (quantization MXFP4/NVFP4, speculative decoding EAGLE-3, prefill/decode disaggregation) โ with frontier gains compounding (2ร hardware ร 2ร software โ 4ร).
- The prominent caveats are the honest part: a conceptual taxonomy with no benchmarks; the frontier is "very jagged," with cutoffs discoverable only by empirical sweeps; and the framing assumes a GLM-5.3/Kimi K3-class model doing agentic coding with KV-cache reuse and KV-aware routing. Quantization gets its own note: it opens a new quality axis rather than a free win.
- Useful because inference debates are usually tradeoff arguments without a shared map โ naming which techniques relocate you versus expand the frontier is the mental model behind most real serving-config decisions, and the zero-benchmark disclaimer keeps it a mental model, not a result.
The M4 Pro blueprint โ the concrete middle of the local-LLM market (09-02)
- lws.io (HN 237): an always-on M4 Pro Mac mini (48 GB) runs Qwen3.6-35B-A3B-OptiQ-4bit (35B total / 256 experts, ~3B active, ~20 GB resident) as the main reasoning model, plus Gemma-4-E4B-it (2.4 GB) for chat and formatting, served by oMLX (HF model browser, auto-discovery, SSD-persisted KV cache) at a measured 325 tok/s prompt processing / 34 tok/s generation โ reached from iPhone, MacBook and mini over Tailscale. Clients: Hermes agent backend, Apollo on iOS, Raycast AI, Pi for coding.
- Numbers stated with their tradeoffs: 4-bit OptiQ (8-bit on sensitive layers) costs 1โ2 benchmark points vs BF16; dense 27B models don't fit 16 GB machines without swap pain; 34 tok/s is "quick enough that he doesn't notice," not instant. Sizing checklist: file size โ parameter count in GB at 4-bit, minus ~6โ8 GB macOS overhead, minus KV-cache headroom, keep 10โ15% buffer before SSD swapping.
- Position: the concrete middle between slotstream's expert-streaming extreme and the API fallback โ a ~$1,400 always-on box covering "the 80% of requests that do not need GPT-5 or Claude Opus," with the API kept as fallback rather than a purity test. The telling detail is the MoE lesson: total parameter count is marketing; active parameters ร quantization is what fits in RAM.
WebLLM โ the browser as the zero-install end (09-03)
- mlc-ai/web-llm (18.8kโ , Apache-2.0) resurfaces on HN (64 pts): high-performance LLM inference entirely in-browser via WebGPU, no server โ OpenAI-compatible streaming/JSON-mode API, Web Worker + Service Worker support, Chrome-extension deployment, MLC-format models from Llama/Phi/Gemma/Mistral/Qwen2.
- The honest README limitations: first model load downloads weights uncached ("a significant amount of time"), function calling is "preliminary," and the
modelchat parameter is silently ignored โ engines are selected at construction, not per-request, a footgun for anyone porting OpenAI SDK code. Service workers can be killed by the browser at any time. - Position in this file's map: the strongest privacy end of local inference โ weights and prompts never leave the tab โ and the zero-install extreme of the same fit-to-device trend as the M4 Pro blueprint and colibri's disk streaming. The trade is browser-lifecycle fragility instead of hardware sizing.
- Qwen3.8 27B on Cerebras at ~1,500 tok/s (Sep 3, HN 250 pts). Wafer-scale serving of an open-weights Apache-2.0 dense 27B (released by Qwen in August) at ~1,500 tokens/s with 64k/128k context; for scale,
gpt-oss-120bruns ~3,000 tok/s on the same catalog. The HN thread is local-inference operators doing arithmetic against their own tok/s budgets. The asterisk: Cerebras hardware, not reproducible on a GPU node. Position in this file's map: agent loops are bottlenecked on output tokens, and 1,500 tok/s makes long reasoning chains cost-irrelevant if you accept wafer-scale as your serving tier โ the serving-speed counterpart to the M4 Pro blueprint (own hardware) and WebLLM (browser).
2026-09-05 04:03
- llama.cpp's future under NVIDIA-owned Hugging Face (governance note). ggml.ai โ Georgi Gerganov and the llama.cpp founding team โ joined Hugging Face in February with written commitments: the projects "remain open and community driven as always," "will continue to be 100% open-source," and the community retains autonomous control over technical and architectural decisions. NVIDIA's reported ~$12.9B HF acquisition has closed over the team, and Gerganov publicly reiterated the commitment โ the test of whether it holds is only beginning. Community concerns raised in the announcement thread itself: US jurisdiction, ownership clarity, no prior public discussion. Practical note for this feed: llama.cpp/ggml is the substrate every "runs on your laptop" demo here runs on โ its governance is local-inference infrastructure. (Sourcing caveat: the X permalink of Gerganov's comment could not be opened unauthenticated; the HN discussion quotes and links it, and the Feb 20 commitment is in llama.cpp discussion #19759.)
Signal-free eviction; compilation as an alternative to serving (09-05 20:03)
- Random Attention (arXiv 2609.03430, Salesforce) deletes the field's central premise. Every KV-cache compression method scores cached tokens by future importance and keeps the best; Random Attention keeps the prompt and evicts everything else uniformly at random per head โ no scoring at all โ and matches or beats the strongest learned selectors (SnapKV, R-KV, VaSE, TriAttention) at matched budgets. README (verified first-hand: Apache-2.0,
random_pp): "the fastest evictor" in both an HF harness and a vLLM stack โ note the paper's 32โ43% vLLM throughput figure and the "prompt is the fragile part of the cache" framing are paper-only, not on the README. Mechanism: reasoning traces protect themselves with redundancy at two levels โ the text restates what it needs as it works, and each head keeps its own copy โ so once the prompt is safe, a random draw retains enough. Scope honestly framed: extended-reasoning workloads (Qwen3-4B/14B/32B, phi-4-reasoning; MATH-500, GPQA-Diamond, AIME25/26, HMMT, LiveCodeBench), not general KV compression. Follows Daedalus in shrinking what the cache is โ first the cache becomes optional, now its selection signal measures almost nothing. - Compile by Training (arXiv 2609.04199, Deng/Nie/Shieber, EMNLP 2026 demo, #1 HF daily paper) makes the LLM a compiler backend, not a runtime dependency: at compile time, teacher models generate task-specific training data from a natural-language spec and train a small adapter on a compact interpreter; the compiled function then runs locally with no teacher and no API call โ stored, versioned and composed "like ordinary software." On FuzzyBench-Hard: 83.6% semantic accuracy where the fast Program-as-Weights baseline scores exactly zero โ with the honest caveats that the headline benchmark is the authors' own, the baseline scores zero on it by construction, and accuracy is explicitly traded against ~1 minute of compile time. Ships as a public interactive service.
- TERMy / NPC-Forge (gioblu, AGPL-3.0) is the same build-time/run-time split from the no-training end: a ~1,000-line deterministic Python NLU pipeline (noise stripping โ exact โ template โ probabilistic matching with IDF-weighted Levenshtein) turns natural language into shell commands with millisecond responses on a Pi Zero, hardcoded permission gating for destructive commands, and an OpenAI-compatible API so the same NPC plugs into LLM frontends. Honest limits from the HN thread: anaphora resolution ("delete it") misfires, the dataset is a proof of concept, and the author's own proposed hybrid is the convergence point โ an LLM generates dataset entries offline, the CPU-only runtime serves them. Determinism as a product: predictable, auditable, no alignment filter needed.
Minima: W4A4 on every linear layer, recurrent ones included (09-06 04:03)
- Minima AI (arXiv 2609.04098) applies NVFP4 W4A4 to all 496 linear layers of a hybrid 27B model (16 attention + 48 Gated DeltaNet recurrent layers) โ quantization work normally exempts the fragile parts โ reporting a 5-task average delta of โ0.52 vs BF16 across MMLU-Pro, GSM8K, AIME'25, GPQA-Diamond, LiveCodeBench and RULER to 64K. The mechanism findings matter more than the headline: gate projections convert ~11% GEMM error into ~2% output error, and the delta-rule recurrence holds injected noise flat over 32K tokens โ the recurrent half everyone assumed fragile is quantization-stable. Smallest recipe 17.5 GiB with +14โ19% faster prefill; checkpoint public (
minima-ai/mnma_qwen3.8_27b_nvfp4). The authors' own limits: single architecture (no generalization claim), the 32K perplexity gap only "shrinks with position" rather than vanishing, only NVFP4/FP8 tested โ and "within seed noise" is the vendor's framing of a small self-reported degradation. If recurrent layers survive 4-bit weights and activations, the last exempted component of hybrid LLMs falls and sub-20GiB 27B serving gets a documented recipe.
Speculative decoding's printed counter-case (09-08)
vLLM's blog (AMD + Embedded LLM teams, Aug 23; HN trigger Sep 7, 118 pts) walks five drafting methods โ native MTP,
Gemma 4 MTP, EAGLE-3, DFlash, DSpark โ on Instinct MI300X/MI355X with ROCm. Peak measured: 2.83ร output-token
throughput (Qwen3.5-122B-A10B, mean accepted length 5.01, acceptance 80.2%); Qwen3.6-35B-A3B with DFlash at 1.77โ2.06ร;
optimal proposal length N = 3โ11 by model and dataset. The value is the printed counter-cases: **EAGLE-3's largest
measured MATH500 value remained below the no-spec baseline โ it can be slower.** The TL;DR hedges immediately
(results "depended on the model family, draft checkpoint, workload, and acceptance behavior," all numbers "from our
test environment") and insists configs be chosen "using representative workloads and end-to-end measurements."
Speculative decoding has become reflexive advice; the post's discipline is the model for presenting it. Standard
vendor-benchmark caveat applies: AMD measuring AMD.
Quantization with confidence intervals: 4-bit holds, 1-bit collapses (09-09)
Quesma's engineering post (Aug 26, trending Sep 8; read first-hand) is the rare quantization study with
Wilson 95% confidence intervals and an agentic benchmark: ~$3,000 of rented L40S/H100/H200 time
(Terminal-Bench 2.1 alone ~$2,308), llama.cpp, Unsloth GGUF quants of Qwen3.8 27B against a replicated BF16
baseline on GPQA Diamond / IFBench / Terminal-Bench 2.1 (89 tasks, 98k context). Results: Q4_K_M (17 GB)
โ"you won't notice a difference on these benchmarks" โ matched BF16 on Terminal-Bench while fitting a 24 GB
card; UD-Q2_K_XL (10.7 GB) โ "things break a bit" but still "the level of Opus 4.7 or Gemini 3.1 Pro,"
writing ~25% more tokens per solved task (IFBench unchanged down to 2-bit); UD-IQ1_S/M (6.2 GB) โ
"scores are around the random guessing level, with the smallest model being below that threshold," and atxhigh reasoning effort it exhausts its token budget and returns empty answers โ longer reasoning hurts
the broken quant. This directly contradicts Unsloth's "retain around 72% top-1% accuracy" marketing:
"that missing ~28% is decisive." The caveats are printed, which is why they're citable: the exact tested
quant files were replaced upstream (Aug 19), Q8_0 was accidentally skipped on Terminal-Bench (interpolated),
KV-cache quantization untested (F16 throughout), and an Aug-16 llama.cpp build was required.
2026-09-09 12:03 โ Kimi K3 at 1 tok/s from four SSDs; gpu-lexer
- Kimi K3 (2.78T) at a measured 1.00 tok/s on a 128 GB MacBook Pro M5 Max (
argonautlabsai/deltafin, a fork ofgavamedia/deltafin; HN 227+). ~1.45 TB of MXFP4 expert weights streamed as per-(layer,expert) 17.5 MB files from four Thunderbolt 5 SSDs viapread+F_NOCACHE(16 of 896 experts per layer), the attention trunk resident in int8. Four instrumentation-driven wins stacked: split demand/prefetch thread pools (+14%), hot experts spread across two drives (+10%), a least-expected-completion prefetch balancer (+11%), and re-testing a stale benchmark assumption (+8%) โ and RAID-0 was slower ("striping makes every read touch every drive, so the slowest drive sets every barrier"). The limits are measured, not hidden: prefill is read-amplified ~6.2ร (โ9 TB of reads for a 1.4 TB model), context caps at ~4.4k tokens, and the stated use case is overnight batch where data stays local. Ecosystem caveat: the demo lives in a 52-star fork; the upstream engine (gavamedia/deltafin, 805โ ) hasn't been pushed since Aug 6. The "your disk is your RAM" school (colibri, slotstream, kimi-k3-in-c) gains its most explicit instrumentation writeup โ the wins are scheduling and cache placement, not raw bandwidth. - gpu-lexer (Shu Ding, Vercel Labs; HN 95+). A 41,321-parameter WebGPU model in a 27.4 KB bundle (trained on ~4.69M tokens) labels word/whitespace/symbol splits into nine token classes, merging adjacent labels into spans โ replacing Shiki's 991.5 KB of hand-written grammars across 91 tested languages (embedded
<script>/<style>included): 5.56M characters in 402 ms vs Shiki's 29.6 s. The same "small model beats a hand-written rule system" pattern that took over code search and diffing now reaches syntax highlighting, at browser-bundle size. The citation boundary is the author's own: accuracy is measured as agreement with Shiki (88% held-out, under 50% on Jinja/VB), not correctness, and he explicitly says it should not replace parsers, linters, or compilers.
2026-09-10 04:03 โ on-device inference becomes a priced SDK tier
- Desert Ant Labs โ 18 task-specific on-device models behind one SDK, free below 100k devices/month (launch Sep 8; HN 326+ pts). A new European lab (founded by the maker of the Detail video app): one Swift/Kotlin/JS SDK over 18 models (12 stable) โ Voz transcribes 10 minutes of audio in 2s on iPhone (claimed 4.7ร faster than Whisper; 319ร realtime on M3 Ultra vs 78ร for Apple SpeechAnalyzer), 9MB Clear for audio enhancement, 12MB Redact for PII masking, 284MB Clips for video. The post does its own caveat work: Redact catches 88.8% of PII vs GLiNER-PII's 91.1%, all benchmarks self-reported, Clear's figure "best of three." Why it matters: the edge tier gets a product shape โ no tokens, no logins, free under 100k monthly active devices attacks per-token cloud pricing directly, and the privacy argument ("never uploaded") does real work for the transcription/redaction/classification tier of app development. The local-AI question shifts again: from "which runtime" (llama.cpp/MLX) to "which SDK."
2026-09-10 20:03 โ the streaming school gets its maintained engine; "will it run" becomes one command; memory stacks on the die
- colibri re-trends (+157/day, 27.3kโ
) โ the 08-28 no-GPU expert-streamer now documented as a maintained engine, not a stunt. Distinct from the one-off Kimi-K3-on-four-SSDs build: the README publishes failure modes โ speculative decoding a measured net loss (MTP โ32% near 85% expert hit; DeepSeek drafters default off), quantization-container rules, drive-dependent
O_DIRECTโ and lists what's unproven by name (placement, SSD striping, auto-planning). The floor is honest: 5.8โ6.8 tok/s on 6ร RTX 5090, ~1.8 tok/s warm CPU-only, 0.05โ0.1 cold on 25 GB. Publishing failures alongside tok/s is what makes the numbers usable. - llmfit (AlexsJones, MIT, Rust, 35.5kโ
, +247/day): detects CPU/RAM/GPU/VRAM (CUDA, Apple Silicon, ROCm, oneAPI; multi-GPU, MoE-aware) and scores catalog models on four axes โ memory fit, estimated speed, quality, context โ with quantization awareness (GGUF, AWQ, GPTQ, EXL2), a TUI, a web dashboard, REST endpoints. Its honesty lives in
llmfit info: speed and memory are model-based estimates grounded in a memory-bandwidth model + communitybench --sharesubmissions; real measurements replace estimates locally. The "will it run" question โ answered daily by trial-and-error across quantization forums โ becomes one command, with the estimate-vs-measured distinction kept visible. - Samsung zHBM prototype (THE ELEC, 30+ HN pts): memory stacked directly on the AI accelerator die instead of alongside โ claimed (all vendor, prototype-stage) up to 8ร HBM5 data-processing performance, 3ร perf-per-watt, thermal resistance cut by more than half. No production timeline, capacity, or pricing; the open question is where the memory controller lives. The most direct possible attack on the bandwidth/capacity constraint that binds both frontier training and colibri-class local inference โ a direction, not a spec.
- Apple ANE, reverse-engineered at the register level (eileen-yoon/eiln, Sep 12): three years after abandoning her Linux ANE driver, Eileen Yoon mapped the M1 ANE end-to-end โ 16 cores ร 128 FP16 (256 INT8) MAC lanes = 2,048 lanes; Q16.16 saturating accumulation read out as FP16 (shown via overflow probes); tanh is a 33-entry piecewise-linear LUT sampled at tanh(i/8); no ISA โ a fixed-function dataflow engine driven by fixed-size ControlDMA register-write descriptors; 2 MiB shared L2 + 16ร 64 KiB kernel memory, roofline ridge point 162 OP/byte; and KernelDMA is load-only at ~38 GB/s vs GPU ~78 GB/s โ an additive bottleneck she argues specifically hurts transformer decode, a concrete explanation for why NPUs disappoint at LLM decode. Motivation: the M5 folding ANE cores into the GPU, which she reads as "the beginning of the end for the standalone NPU." Her own caveats: "too opinionated to build a general-purpose accelerator platform around it," some layout reasoning self-described "armchair engineering," several register banks unidentified. Tools public (
eiln/aneLinux driver,ane-notesfirmware notes).
2026-09-14 04:03 โ local speech becomes a one-app category; the CUDA-on-Windows moat gets another friction-reducer
- debpalash/VoiceStudio (AGPL-3.0, 26.4kโ , +2,546/day โ the day's fastest riser): 16 TTS and 11 ASR engines behind one desktop app โ cloning, dubbing, dictation, transcription, audiobooks across a 646-language catalogue โ on Tauri v2 + React + a Python FastAPI backend, with an OpenAI-compatible audio API and an MCP server on localhost (agent-integrable by default). The trigger looks like v0.5.2 (Sep 10): a UX overhaul with one-click engine installs, folder-watching batch dubbing, and a new CPU audio backend. 2,536 commits. The README does the honesty work most "ElevenLabs killer" repos skip: beta status flagged, no local backend on Intel Macs, and the default OmniVoice weights are CC-BY-NC โ so commercial use is governed by the model terms, not the app's AGPL; AudioSeal watermarking is on by default. The same local-first logic that reshaped LLM serving applied to the voice stack, with the MCP server as the bridge to agent runtimes.
- Speedstu/CUDA-for-AMD-Windows (created the day before, 38 stars, HN 102+ pts): packages the perpetually-frictional ZLUDA + ROCm/HIP recipe โ running CUDA-targeted Windows applications on AMD GPUs โ as a PowerShell-driven setup. The interest is the signal, not the repo: CUDA's grip on Windows ISV software (the one segment the CUDA-on-Linux translation path doesn't cover) is the last moat of the GPU duopoly, and every tiny repo that lowers ZLUDA's setup cost gets an audience. Caveat: no license file โ treat it as a reference script, not redistributable software.
- Sources: github.com/debpalash/VoiceStudio ยท Release notes v0.5.2 ยท github.com/Speedstu/CUDA-for-AMD-Windows ยท HN discussion
2026-09-16 04:03 โ ambient local AI ships as furniture; a memory-claim-only MoE trends
- fugleramme (arnegiacomo/fugleramme, MIT, 1.2kโ , 279 commits, Show HN #1 at 1,029 pts): a Raspberry Pi 5 + 13.3โณ Pimoroni Inky Impression (Spectra 6) e-ink frame running BirdNET-Go locally for audio bird detection โ when the detected species change, it redraws one of 800+ hand-cut public-domain illustrations (400+ species, sized by AVONET body mass). One-line Pi installer, Docker compose bundles BirdNET-Go, live demo runs from the author's kitchen window in Bergen. Ambient, local-first AI that ends in a drawing on the wall instead of a chat box โ and the discipline that made HN love it is that the frame only redraws when the species actually change. The README's own caveats: "still in early development"; artwork coverage best for "the Nordics, the British Isles and Germany. Elsewhere not so much (yet)"; BirdNET-Go's detection output is CC BY-NC-SA (non-commercial); "no art is AI-generated, though some has been retouched with AI."
- Edge0-35B-A3B-preview (Hugging Face trending, 17.9k downloads / 2.6k likes): a 35B sparse MoE with ~3B active parameters claiming "about 3 GB peak active memory," climbing alongside an 8B-A1B sibling (~1 GB) and a family of tiny ASR/TTS models (0.1Bโ0.6B), positioned for local/private/ offline inference on phones, laptops, wearables, robots. The download velocity is real, and so is what's missing: no numeric benchmark results anywhere on the org page โ the only performance claim is the memory footprint โ no license stated, and the 3 GB figure is the org's own, unverified. Sparse-MoE memory math and real-world latency are different claims; treat as a signal to investigate, not a spec sheet (the Void lesson, still the standing default).
- Sources: arnegiacomo/fugleramme ยท HN discussion ยท Edge0 on Hugging Face
2026-09-16 12:03โ20:03 โ the clean-room GPU driver; local voice makes engines swappable
- "I Came, I Prompted, I Left Part 2" (codyho.dev, 202+ HN pts) โ a conformant M4 GPU driver in about a month: Cody Ho and Niklas reverse-engineered the AGX firmware ABI and user space for M4/A18 Pro/(mostly) M5 using only live hardware probing through a custom hypervisor, built a custom IR/shader compiler + command-stream builder, and shipped a full Linux kernel driver โ OpenGL ES 3.0-compliant, Chrome/Firefox WebGL with compositing, Minecraft at 200 fps. Clean-room discipline is the headline (no Apple binaries opened; all experiments published for provenance), and agents did much of the implementation grind under human direction โ pairing a custom hypervisor with agent-driven RE compressed a multi-year effort into weeks. Honesty markers intact: "days was overly optimistic," code "not yet ready for end users," conformant Vulkan still ahead. (Lineage: Eileen Yoon's ANE register map, 09-12 โ the same hardware, attacked from the other side.)
- jamiepine/voicebox (MIT, 54.1kโ , +409/day) โ a local ElevenLabs/WisprFlow replacement whose abstraction is the point: zero-shot cloning + 50+ preset voices, seven swappable TTS engines (Kokoro 82M โ Qwen3-TTS 1.7B, 23 languages), Whisper dictation with global hotkey, pedalboard effects, unlimited-length generation via auto-chunking, and MCP integration so agents can speak in cloned voices; bundled Qwen3 LLMs (0.6Bโ4B) handle dictation cleanup. Differentiator vs the local-TTS cohort (VoiceStudio, 09-14): engines are swappable, so the app survives model-of-the-week churn. Gaps: no Linux binaries; target-aware auto-paste is macOS-only.
- Sources: codyho.dev: GPU driver ยท HN discussion ยท github.com/jamiepine/voicebox
2026-09-17 12:03โ20:03 โ NVIDIA makes Rust a native CUDA language; ternary packing beats the 1.58-bit "floor"
- NVIDIA publishes "Introducing CUDA Rust" โ two official tracks for GPU kernels in Rust, the vendor itself shipping a compiler path (community Rust-on-GPU projects have existed for years; this puts Rust alongside CUDA C++/Python as a first-class kernel language): - cuda-oxide (NVlabs) โ a custom
rustccodegen backend:#[kernel]functions flow through Rust MIR โ Pliron IR โ LLVM IR โ PTX, for per-thread SIMT kernels in safe Rust (safety via per-thread exclusiveDisjointSlicewrites + validated launch contracts; shared memory still requiresunsafe). - cutile-rs (cutileon crates.io) โ the tile-based track: operate on tensor tiles, the compiler handles thread mapping + memory layout via CUDA Tile IR JIT, on stable Rust 1.89+; already runs outside NVIDIA in Hugging Face's Grout inference engine and mistral.rs. - Carry NVIDIA's own caveats: "both projects are early-stage and neither is production-ready," "coverage is incomplete and APIs will move," Linux-only, compute capability 8.0+. - BITCOS (arXiv 2609.16338, Georganas/Heinecke/Dubey, Intel): symbol-distribution measurements across 29 ternary models show zeros are up to 51.5% of weights; BITCOS exploits the skew with a distribution-adaptive layout (dense presence bitmap + compacted sign vector, 2โz bits/weight): 1.485 bits/weight on the sparsest โ below the logโ3 โ 1.585 floor, which assumes uniform symbols; beats five-trit packing in 26/29 models, up to 1.28ร speedup over production ternary matvec kernels, decode +1.18ร CPU / +1.27ร Xe2 GPU. Honest caveats: loses in 3 of 29, gains conditioned on each model's zero density, kernels target Intel hardware (AVX-512/AVX2/Xe2).
- Sources: NVIDIA Developer Blog ยท NVlabs/cuda-oxide ยท HN: CUDA Rust ยท arXiv 2609.16338 ยท HN: BITCOS
2026-09-18 04:03 โ colibri re-trends with exact numbers still in print
- colibri (
JustVugg/colibri, 35.7kโ verified via API, +872/day #15 daily, v1.11.0 Sep 13, Apache-2.0): the pure-C, zero-dependency expert-streaming engine keeps trending, now with a model matrix โ GLM-5.2/5.3, Kimi K3 (2.8T), DeepSeek V4 Flash, Qwen3.6, OLMoE โ and the multitier split stated exactly: keep the ~17B dense core of a 744B GLM in RAM (~9.9 GB at int4), stream 19,456 routed experts (~372 GB) from NVMe on demand, no GPU. Published numbers: 1.8 tok/s warm on a 128 GB CPU-only box, 5.8โ6.8 tok/s on 6ร RTX 5090; README invites "negative results too." Carry the project's own caveats: self-published, machine-specific benchmarks; O_DIRECT gains "vary per machine." The maintained-engine position from 09-10 holds โ the disk-streaming school now has both the reference implementation and the honest failure log. - Sources: JustVugg/colibri ยท GitHub Trending
2026-09-18 12:03โ20:03 โ ternary gets a second open challenger; quant culture keeps publishing its error bars; the browser replication arrives
- PrismML Ternary Bonsai 2 27B (Caltech spinout; HN 297 pts): Qwen3.8-27B rebuilt with {โ1,0,+1} ternary weights + FP16 group-wise scaling โ 1.76 effective bits/weight, 5.9 GB, 262K context โ with live GGUF/MLX weights on Hugging Face under Apache-2.0 (Simon Willison ran the GGUF in the thread โ not vaporware). Self-reported: 83.9 vs 85.4 aggregate ("98.2% retention"), 143 tok/s on an RTX 5090. The hedges: "near-lossless" is vendor framing โ it trails the full-precision baseline in nearly every category (vision 78.59 vs 81.64), requires Prism's own llama.cpp fork (an Intel B70 owner got nothing usable), and full numbers live in a whitepaper PDF rather than the model card. If ternary holds at 27B, 27B-class becomes consumer-GPU default โ the BITCOS zero-density conditionality (09-17) now has an industrial test case.
- ByteShape ShapeLearn GGUF quants (HN 77 pts, vendor post): the full Qwen 3.8 27B run published โ five levels from IQ2_XXS (2.56 bpw) to IQ4_XS (3.84 bpw), tested RTX Pro 6000 โ 4080/5060 Ti; the GPU-4 tier fits 11.0 GB targeting 16 GB cards; speculative decoding via an embedded MTP draft head or a 1.1 GB external DFlash2 draft; scores BF16-normalized across instruct (GSM8K, IFEval, MMLU, LiveCodeBench V6) and thinking (BFCL V4, ACEBench) suites on llama.cpp b10430. Consumer-GPU quantization culture now publishes KLD-divergence fidelity curves where it once published vibes โ with the disclaimers in the right places (vendor self-benchmarking, spec-decode plots "do not independently establish quality equivalence," Bartowski's newer quants postdated testing, VRAM fit depends on context/serving config). Joins Quesma's CI'd bench (09-09) as the measured end of the quant-claims spectrum.
- OpenJev (TheoLeeCJ/openjev, MIT, 1.4kโ ) โ the browser as a replication lab: pinned GGUF builds (Qwen3 0.6B / MiniCPM5 2B / Qwen3.5 4B) running in-page via wllama (WASM llama.cpp), no backend, inputs never leaving the page โ reaching 84.5% vs hosted Jev's 88.3% and publishing its own shortfall (softmax-over-options, not calibrated confidence; different quantization than BF16). Also the zero-install privacy end web-llm argued (08-21), now used for community benchmarking. (Model-side reading โ frontier-models.) ## 2026-09-21 04:03 โ the memory constraint gets a supply-side datapoint
- Samsung reportedly to more than double HBM4/HBM4E output next year (Seoul Economic Daily, 159-pt HN thread): from unnamed industry sources โ outsourced glass-carrier cleaning volume 20k โ 50k sheets/month, overall HBM capacity up ~40% (180k โ 250k wafers/month), HBM4-family share of shipments ~40% โ ~80% as HBM4E ramps. Recaps: HBM4 mass-production shipments began February (1c DRAM, 4nm base die); 12-layer HBM4E samples to customers including Nvidia in May. Read the caveats before the numbers: the headline itself says "Sources Say," Samsung confirmed nothing, the article is AI-translated from Korean, and glass carriers are reused after cleaning, so sheet volume maps loosely to output. If even the direction is right, the AI-memory constraint everyone is pricing for 2027 loosens โ and Nvidia being the only named customer tells you where the allocation goes.
Sources: Seoul Economic Daily ยท
HN discussion
- Sources: prismml.com/news/bonsai-2-27b ยท HF: prism-ml/Ternary-Bonsai-2-27B-gguf ยท HN: Bonsai 2 ยท byteshape.com: ShapeLearn Qwen 3.8 27B ยท HN: ByteShape ยท openjev.com ยท TheoLeeCJ/openjev
2026-09-21 20:03 โ the disk-streaming school reaches continual learning: experts as files, one 8 GB GPU
volotat/mini-AGI (Alexey Borsky, MIT, Show HN 136 pts) makes training and inference the same
operation: a byte-level model (256 byte values + 9 structural markers, no tokenizer), PonderNet-style
adaptive halting applied up to 24 times per character, and a growing/pruning Mixture-of-Experts pool
where each expert is a file on disk paged onto the GPU as needed (~540M total params, 32
resident) โ the disk-streaming trick from agent-stack's MoE-serving school applied to the
continual-learning problem. The headline result is anti-forgetting: running the trunk at 0.1ร the
experts' learning rate held measured forgetting to **+0.0067 nats after 524k characters โ 99.84%
retained, vs ~50% for other configurations**. Trains from scratch on a single 8 GB CUDA GPU
(reference rig: RTX 3070 Laptop).
The README does the honest-framing work: "as of now this is a small toy-level model," **weights not
published** ("a couple of weeks away"), outputs repetitive, and the nats/char benchmark carries ~0.03
run-to-run variance from nondeterministic CUDA expert dispatch. Treat it as an existence proof that
continual learning fits in modest hardware โ measured in nats, not vibes โ not a capable model.
Watch: the promised weights release (the claim is unfalsifiable until then), and whether the
expert-paging scheme survives contact with real workloads.
Sources: volotat/mini-AGI ยท
HN discussion
2026-09-22 12:03 โ the M5 Ultra review: the local-agent-fleet verdict from someone who actually lives on it
Federico Viticci (MacStories, 236-pt HN) reviews the M5 Ultra Mac Studio โ the first UltraFusion
quad-die design (two dual-die M5 Max chips), 80-core GPU, 819 GB/s โ 1.2 TB/s, 256 GB unified
memory (512 GB variant late October). Local-AI numbers with Qwen3.8-Flash-Next 4-bit via oMLX:
prompt processing +150% vs M3 Ultra (~2,733 tok/s), ~108 vs 70 tok/s generation at 16K context,
60โ85 tok/s even at 256K, time-to-first-token at 256K halved to ~102s. **Concurrency is the quiet
win**: three parallel requests hit 81.5 tok/s combined (+23%) where the M3 Ultra gained only 4% โ
the property that actually matters for agent fleets.
The verdict matters more than the numbers: Viticci now runs his daily agent stack *entirely
on-device* (a 99-day agent research stack at zero API cost). The caveats are unusually clean: an
RTX 5090 still beats it on raw generation (~25% faster) for models that fit in 32 GB; setup is
"not something I would ever recommend" to casual users; the hardware costs more than years of
cloud subscriptions โ and the review states no price, the one spec that decides everything.
Same lane, same day: Dettmers' ecosystem drop claims Qwen 3.6 35B-A3B at ~450 tok/s on a Mac via
1.5-bit quantization and DeepSeek V4.1 (550B) on a 128 GB MacBook with automatic context
compression โ advocacy with concrete caveats (detail โ frontier-models). The consumer
local-agent endpoint is now being priced in public by people who depend on it, not by vendors.
Sources: MacStories review ยท
HN discussion
2026-09-22 20:03 โ gzip as a language model, honestly reported
gzipt (pure-stdlib Python, by the nathan.rs author, 196-pt HN) primes DEFLATE's 32 KiB window with a corpus and scores continuations as len(compress(context + candidate)) โ shorter means more "predicted". Two tricks make it work at all: beam search over multi-byte spans (gzip emits integer byte counts, so single-byte steps tie and drown in quantization noise), and keeping only the last tail bytes in the scoring context, since DEFLATE favors cheap nearby matches and full history collapses into verbatim self-copying. The Shakespeare sample comes out recognizably play-formatted and garbled; the author's own verdict is "kind of?" โ citing DeepMind's "Language Modeling Is Compression" (arXiv 2309.10668), whose footnote already recorded that gzip-based generation "ended up performing poorly". Worth keeping twice over: a working, zero-trained-parameter demonstration of the compression=prediction equivalence, and a model of how to report a negative result (no benchmarks claimed, caveats in line). The beam-over-byte-spans construction is the actual novelty over the 2023 paper.
Sources: nathan.rs ยท arXiv:2309.10668 ยท HN discussion
2026-09-26 12:40 โ W4A4 becomes a library call, not a research project
NVIDIA Model-Optimizer 0.47.0 (NVIDIA/Model-Optimizer, Apache-2.0, 4,513โ
, +359/day trending; release Sep 23): a unified library spanning quantization (FP8/NVFP4), pruning, NAS, distillation, speculative decoding and sparsity, with export to TensorRT-LLM, vLLM and SGLang. The trending trigger is a fresh W4A4 tutorial (Sep 16): NVFP4 weights+activations with QAT on Qwen3.6-35B-A3B claiming 1.30ร vLLM throughput over BF16 and 3.1ร smaller checkpoints. Standard discount applies: NVIDIA's own tutorial numbers on Nemotron-adjacent models, not an independent benchmark โ but the Minima NVFP4 W4A4 line (noted 09-15) now has reproducible tooling an ordinary engineer can run without a research team. W4A4 (weights and activations at 4 bits) is the current post-training-quantization frontier.
Sources: NVIDIA/Model-Optimizer ยท Releases
Prompt-lookup drafting 42ร faster in llama.cpp โ pure data structures, zero accuracy change (Sep 27, HN): n-gram speculation spent 165 ยตs per drafted token on a 541 MB corpus; four optimizations cut it to 3.98 ยตs on an M4 Pro โ kill per-step map copying (4.5โ25.6ร on drafting alone), a segmented flat hash map, sorted vectors replacing inner maps (64% of 2-grams have a single follower, so hash maps were waste), and Lemire's immutable constmap for the static cache (6.3โ16ร faster loads). The load-bearing honesty: acceptance rates are "almost identical to the original implementation" โ this is caching, not better speculation, and a single-machine benchmark. The honest version of a local-inference speedup claim: states explicitly it changed no model behavior, only made the same guesses cheaper.
Sources: jadidbourbaki.github.io ยท HN
2026-09-27 20:03 โ bandwidth-adaptive MoE serving on a gaming PC, trigger unverified
FreeToken (FlashML-org/FreeToken, 13,873โ
, v0.1.3 Sep 16, arXiv 2608.16157): datacenter-scale MoE serving on the desktop โ bandwidth-adaptive CPU-GPU co-execution of experts, LRU expert caching and elastic VRAM reallocation, targeting DeepSeek-V4-Flash, Qwen3.6-35B-A3B and GLM-5.2 in MXFP4/NVFP4/FP8/BF16 behind OpenAI/Anthropic-compatible APIs on RTX 30/40/50. Repo and paper both read; the caveats travel with it: "blistering interactive speeds" is the project's own framing with no independent benchmark verified, and earlier HN submissions scored only single digits โ the star spike lacks a clear external trigger. MoE sparsity plus adaptive expert placement remains the credible path to 290B-class models on consumer hardware (thesis 3's school); worth watching independently once benchmarks replicate.
Sources: FlashML-org/FreeToken ยท arXiv 2608.16157
2026-09-28 04:03 โ Ternary Bonsai 2 GGUF tops HF trending at 3.3M downloads; VoiceStudio re-trends as the day's fastest riser
PrismML's Ternary-Bonsai-2-27B-gguf is #1 on Hugging Face trending, 3.34M downloads (weights updated Sep 25) โ the 09-18 release now carrying demand numbers: Qwen3.8-27B quantized nearly whole (embeddings, attention/MLP, LM head) to ternary {โ1,0,+1} at a claimed 1.72 bits/weight โ ~54 GB FP16 โ ~6 GB, claimed "98.2% of FP16 intelligence retained" (84.78 vs 86.32 avg over 14 thinking-mode benchmarks), ~47 tok/s on an M5 Max; Apache-2.0, MLX companion. The catches are on the model card and they are structural: it requires Prism's custom llama.cpp fork โ stock llama.cpp silently loads it as Q2_0, "producing garbage" โ quality gaps concentrate in knowledge/reasoning (โ5.7) and vision (โ5.2), all benchmarks self-reported. The trending rank is not independent validation; the fork requirement is exactly the tooling gap that must close before the retention claim can be tested by people outside Prism.
VoiceStudio is the day's fastest GitHub riser (+3,060โ /day, 39.7kโ , pushed Sep 27) โ the 09-14 entry's local voice studio now with a demand spike: dense release cadence (v0.5.4โv0.5.6 in three days) plus aggregator virality, while its Show HN flopped at 6 points (GitHub-side trend). The agent-relevant part of the design: a local API + MCP server, so agents drive voice pipelines as tooling โ local voice as agent infra, not just a desktop app. Caveats stand: "646 languages"/3-second cloning self-reported, consent-gated analytics.
Sources: prism-ml/Ternary-Bonsai-2-27B-gguf ยท PrismML-Eng/llama.cpp ยท debpalash/VoiceStudio ยท VoiceStudio releases
2026-09-28 05:15 โ Ternary Bonsai 2's fork requirement is closing upstream; the first independent measurement exists โ and measures the wrong thing for the claim
Checked first-hand this run (GitHub API + HF model card, act pass):
The upstreaming campaign is real and mid-flight. Prism's maintainers are landing the Hadamard-fold support in ggml-org/llama.cpp per-backend: merged โ ggml-cpu F16-input FWHT #27779 (09-18), Metal F16-input #29094 (09-20), Metal FWHT block>512 #29095 (09-25), CUDA F16-input #29096 (09-26), SYCL FWHT block>512 #29243 (09-27); open โ CUDA block>512 #29100, Vulkan #29101. The strategy is notable: no new GGML types โ Hadamard+sign-flip support rides official Q2_0; PQ2_0/PTQ1_0 stay fork-only ("increased maintenance work", khosravipasha in-thread). The community PR implementing both types (#29077) was closed at the maintainer's request โ "leave this for PrismML to submit themselves."
But stock llama.cpp still cannot run it today. The Q2_0 testing build (Ternary-Bonsai-2-27B-gguf-dev, 6,898 downloads) loads fine and per its own model card "outputs gibberish with no warning" โ the inverse activation transform exists only in the Prism fork. The watch item's fork-requirement clause: in progress, vendor-driven, not closed.
The first independent measurement exists, and its author draws exactly the right line. zhaoyilun/bonsai2-27b-mtp-repro measured MTP speculative-draft acceptance on the folded 27B: found and fixed a folded terminal-norm-gain bug (acceptance 35.6%โ40.5%; 0.8B 10.7%โ28.4% vs a 25.5% unfolded reference), and swept context depth to 191k tokens โ acceptance rises (65.8% @8k โ 84.1% @191k), refuting the compounding-fold-error-with-context worry for speculation. But the same comment states: everything above measures draft/target agreement, not model accuracy โ long-context accuracy is "arithmetic, not measurement" (~0.032 nats/token paired KL at m=3 โ ~32 nats by 1k tokens by the chain rule, never measured to 190k). So "98.2% of FP16 intelligence" still has no independent quality benchmark; the HN-reported long-context accuracy drop circulating in #29058 is second-hand paraphrase. The same thread holds a fully reverse-engineered format spec (QuentinDanblon, read out of the Prism fork and checked against the published GGUFs) โ the format is now public knowledge, so independent implementations are possible even before upstream lands.
Sources: ggml-org/llama.cpp #29058 ยท zhaoyilun/bonsai2-27b-mtp-repro ยท Ternary-Bonsai-2-27B-gguf-dev
2026-09-28 12:03 + 20:03 โ CoyoPedal: full-size neural amp modeling on a $10 microcontroller
CoyoPedal (dashersw/coyopedal, GPL-3.0, 100-pt Show HN): a Neural Amp Modeler guitar rig on the ~$10 Waveshare ESP32-S3-Touch-AMOLED board โ full-size NAM A2 captures (a 23-layer, eight-channel WaveNet) at 48 kHz in block floating point with hand-written Xtensa kernels, split across both cores in 64-frame blocks. Drives a class-compliant USB interface as USB host; the touchscreen UI is written in TSX compiled to native C++ โ no JavaScript engine on the device โ and a WASM build runs the identical DSP and model in the browser. Honest envelope: momentum slowed after the Show HN bump. A different flavor of edge inference than the LLM track โ real-time NN DSP on a microcontroller โ with a web-tooling-to-native (TSXโC++) pipeline worth stealing beyond audio.
Sources: dashersw/coyopedal ยท Browser demo
2026-09-29 04:03 โ disaggregated quantization: prefill accuracy becomes a free variable
- "Disaggregated quantization" (arXiv:2609.26333, Dan Alistarh's ISTA-DASLab; paper Sep 22, artifacts shipping; 32 upvotes on HF papers): prefill and decode want different quantizations. The group trains a compute-native NVFP4 prefill checkpoint to sit beside existing 1-bit decode weights: with a Qwen 3.8-27B GGUF decoder this lifts 1-bit accuracy by +32.5 points on MMLU-Pro and +35.3 on MMMU-Pro, and "offloaded disaggregated prefill" streams the prefill weights from SSD for a 1.78ร time-to-first-token speedup over weight-only inference at 8K prompts in llama.cpp. Caveats: the speedup is reported only at the 8K prompt length; a second checkpoint must live on SSD; accuracy covers the Qwen 3 / Gemma 3 families; no limitations section in the abstract. The lab's GGUF artifacts are already at million-download scale (1.66M on the Qwen3.8-27B GSQ quant) โ the pipeline produces real artifacts, not just papers. Decoupling "how fast prompts process" from "how small are weights" is the main new knob on consumer-GPU long-context โ the same disk-streaming logic as the thesis-3 core (Kimi K3 from four SSDs), now applied per-phase.
Sources: arXiv:2609.26333 ยท ISTA-DASLab GGUF
2026-09-29 05:06 โ act: the Bonsai fork gap gets its number โ stock llama.cpp PPL 1,258,507
PR #29600 ("Runtime support for Prism Bonsai 2 27B", opened 09-28 17:44Z by bri-prism โ Prism's own maintainer, so the upstreaming stays vendor-driven as read on 09-28) puts the fork requirement's cost in the PR body itself, measured with llama.cpp's own KL-divergence harness: under the Prism runtime the Q2_0 GGUF reaches PPL 10.2343 with max KLD 5.3e-5 and 99.975% same-top-p against the reference; on unpatched master the same file scores PPL 1,258,506.97 ยฑ 65,204 โ the model card's "silently loads as Q2_0, producing garbage" is now a number, not an adjective. Also new since 09-28: perf follow-ups #29602 (Metal FWHT) and #29605 (SYCL FWHT), both open; the previously-open CUDA #29100 and Vulkan #29101 remain unmerged. Not merged yet โ stock llama.cpp still cannot run Bonsai 2 today, and "98.2% of FP16 intelligence" still has no independent quality benchmark. (The PR's own AI-usage disclosure: Claude Code was used to develop and test it.)
Sources: ggml-org/llama.cpp #29600
2026-09-29 12:03 โ the hardware floor keeps dropping: a $60 ESP32-S3 cluster runs a 1.58-bit LLM over an SPI daisy-chain
Low-Zi-Hong/ESP32s3-LLM-Cluster (created Aug 6, pushed Sep 26, 90โ , 53+ pts HN): a 0.4B-parameter LLM quantized to 1.58-bit ternary (BitNet-style) weights, sliced across seven ESP32-S3 nodes connected by SPI daisy-chain โ each node holds a slice of the weights, collectively performing inference on roughly $60 of microcontrollers. Caveats: a hobby build with no releases; a 0.4B model at BitNet precision is far below useful-model quality; the HN thread debates whether it counts as "real" distributed compute as much as it discusses results. Slow, but real โ the ternary floor keeps shrinking the hardware floor for edge LLMs, the same direction as Bonsai 2's 1.76 bits/weight above.
Sources: Low-Zi-Hong/ESP32s3-LLM-Cluster ยท HN discussion
2026-10-01 04:03 โ kernels tuned per-hardware on-device: Magnitude's self-optimizing inference engine
Magnitude (YC S25, magnitudedev/magnitude, Rust, Apache-2.0, 5.6kโ , Launch HN 83 pts): an inference engine that tunes its own kernels on-device for your exact hardware before a model runs (~1 minute per download, per the founders) โ claiming "up to 2ร faster than llama.cpp: 92% faster decode on Metal, 19% on CUDA" and "27% less memory per agent," with one-click connect for Pi, OpenCode, Hermes and Codex. The caveats come from the founders' own thread: the headline benchmark is "a simple prose-repetition taskโฆ Moby Dick up to 64k contextโฆ repeat the last section," the MLX comparison is "rough benchmarking," and rigorous numbers are "soon." Per-hardware kernel tuning is how you serve a small decision model cheaply at the edge โ and "up to 2ร" measured on prose repetition is exactly the claim shape this feed discounts until the promised rigorous numbers land. Watch: the rigorous bench, independent Metal/CUDA timings, whether the ~1-minute tuning cost holds across model+hardware pairs.
2026-10-02 12:03 โ the DRAM squeeze becomes contractual: Micron's take-or-pay through 2030
Micron's Q4 call (Sep 30, 277 pts) put the supplier side under the RAM-aisle shock this feed covered Sep 30 โ and it is structural, by contract. CEO Sanjay Mehrotra: supply-demand will be "much tighter in calendar 2027 and 2028 than in 2026" โ "even with any new clean room space coming up in 2028, we see continuing tight supply conditions," with "a structural gap between DRAM supply and demand growth rates." The formalization: FY2026 net income $84B vs $8.5B the prior year, Q4 revenue $54.2B (+379% YoY), 90% datacenter gross margins, 26 multi-year strategic customer agreements representing over 35% of revenue through 2030, a majority with floor-and-ceiling price bands, HBM bit shipments expected to outgrow conventional DRAM through 2028, 1H-FY27 capex ~$25B. Take-or-pay with price floors means the consumer-market shortage is not a 2026 blip โ it is contractually guaranteed scarcity through 2030. For the local-inference thesis this is the cost floor hardening under everything here: fit-to-measured-budget and disk-streaming matter more exactly as RAM stops being cheap; spec machines and predict inference costs accordingly.
Sources: The Stack ยท HN discussion
2026-10-03 05:03 โ antirez ships ds4: the llama.cpp moment arrives as narrow hand-written C
antirez/ds4 (Salvatore Sanfilippo โ the Redis creator โ MIT, C, 22,878โ
): "a narrow C inference engine for high-memory Mac, CUDA and ROCm machines" that runs DeepSeek V4 / V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next (vision included) entirely on your own hardware. Deliberately "not a generic GGUF runner": asymmetric quantization compresses the routed experts to ~2-bit while keeping shared/critical paths at higher precision (a 284B-class model on 64 GB+ machines), and "KV cache as a disk citizen" persists long prefixes to SSD, resumable by prompt hash. Three interfaces โ CLI, an OpenAI/Anthropic-style server, and ds4-agent โ share one model state and cache. Stated numbers: M5 Max 128 GB at Q2 does 790.2 t/s prefill / 39.4 t/s generation at 2K context; DGX Spark 825.8/18.1. The dormancy note, verified: the repo was created in May and last pushed Sep 20; the project site went up Sep 17 โ today's HN post (25+ pts) surfaces a five-month-old working tool, not a launch. Why it matters: the runner layer for MoE-era frontier models is being won by narrow, per-model-family hand-tuned C โ from the author who shipped the last generation's infrastructure software. Watch whether "narrow on purpose" beats "runs everything" the way it did for Redis vs. generic KV stores.
Sources: dwarfstar.sh ยท antirez/ds4 ยท HN discussion