1. antirez's ds4 surfaces on HN: local frontier inference in C โ€” 22.9kโ˜…, and 13 days without a push

  • Velocity: โ–ฎโ–ฎโ–ฎ trending
  • Source: dwarfstar.sh ยท 25+ pts on HN ยท ~2h ago (~02:01 UTC+8)
  • Tags: local-llm inference c open-source

Salvatore Sanfilippo โ€” antirez, the creator of Redis โ€” has a local inference engine, ds4 (MIT, C): "a narrow C inference engine for high-memory Mac, CUDA and ROCm machines" that runs DeepSeek V4 / V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next (vision included) entirely on your own hardware. The design is deliberately narrow โ€” "not a generic GGUF runner": asymmetric quantization compresses the routed experts to ~2-bit while keeping shared/critical paths at higher precision (a 284B-class model on 64 GB+ machines), and "KV cache as a disk citizen" persists long prefixes to SSD, resumable by prompt hash. Three interfaces โ€” CLI, an OpenAI/Anthropic-style server, and ds4-agent โ€” share one model state and cache. Stated numbers: M5 Max 128 GB at Q2 does 790.2 t/s prefill / 39.4 t/s generation at 2K context; DGX Spark 825.8/18.1. The dormancy note, stated precisely: the repo (22,878โ˜…) was created in May and last pushed Sep 20 โ€” today's HN post surfaces a five-month-old project; dwarfstar.sh, the project site, went up Sep 17. It is a launch neither today nor last month; it is a working tool finally hitting the front page.

Why it matters: the llama.cpp moment for MoE-era frontier models is arriving as narrow, hand-written C tuned to specific model families โ€” and from the author who shipped the last generation's infrastructure software. Watch whether "narrow on purpose" beats "runs everything" the way it did for Redis vs. generic KV stores.

๐Ÿ”— dwarfstar.sh ยท ๐Ÿ”— antirez/ds4 ยท ๐Ÿ”— HN discussion


2. FLUX 3 Image: bounding-box composition, 10 references, native 4K โ€” "designed for agents"

  • Velocity: โ–ฎโ–ฎโ–ฎ trending
  • Source: bfl.ai ยท 197+ pts on HN ยท ~25h ago (~03:24 UTC+8)
  • Tags: image-generation agents multimodal flux

Black Forest Labs' FLUX 3 is a multimodal model family (video, audio, images, actions); FLUX 3 Image is the generation-and-editing part, and its pitch is structure, not vibes: bounding-box composition on a 0โ€“1000 grid ("lay out the image exactly how you want using bounding boxes"), up to 10 reference images each addressable by token (ref_image_0 onward), batch editing that leaves untouched regions identical, pixel-perfect local edits, and native 2K/4K output (a showcase render: 5456 ร— 3072 px, "all from the model"). The agent hook is explicit: "designed for agents" โ€” an LLM plans a layout (caption + element table) from one line and an aspect ratio, then sends it to the API. Availability: BFL API plus a commercial weights license for self-hosting and fine-tuning. What the page does not claim: no parameter count, no benchmark table, no release date โ€” the caveats section of this item is that the claims are entirely BFL's own, demonstrated through curated showcases.

Why it matters: image generation as a tool primitive โ€” structured layout input, verbatim boxes, agent-planned composition โ€” is the interface agentic pipelines actually need; if the reference system works as described, it attacks the hardest remaining gap (consistent multi-subject scenes) at the API level.

๐Ÿ”— bfl.ai โ€” FLUX 3 Image ยท ๐Ÿ”— HN discussion


3. Supabase acquires Turso: "agents should be able to create a database as easily as creating a file"

  • Velocity: โ–ฎโ–ฎโ–ฎ trending
  • Source: supabase.com ยท 173+ pts on HN ยท ~4h ago (~23:43 UTC+8)
  • Tags: database postgres sqlite agents acquisition

Supabase announced the acquisition of Turso (Oct 2) with an agent-infra thesis: Supabase is already "launching over one million databases per week," and demand will outrun capacity unless database creation becomes a file-cheap primitive. Turso brings a Rust rewrite of SQLite and a platform where "a single server can manage millions of databases, loading them when needed and suspending them when they're not" โ€” exactly the suspend/resume shape an ephemeral per-agent database needs. Terms and closing date are not disclosed. Continuity is stated plainly: Turso keeps operating, Supabase keeps building around Postgres, "for existing users, nothing changes" (customers named: Superhuman, Sauna.ai, CTO.new, Mastra). Founders Glauber Costa and Pekka Enberg join, with Costa leading the agentic-infrastructure effort.

Why it matters: the libSQL line โ€” the most credible SQLite rewrite in the ecosystem โ€” now reports to the largest managed-Postgres player, and the stated product direction is databases as a disposable agent resource. Postgres and SQLite are converging on the same buyer: whoever's agents need a million small databases by 2027.

๐Ÿ”— Supabase blog ยท ๐Ÿ”— HN discussion


4. Utah's VPN age-verification law blocked: court finds it "demands a technical impossibility"

  • Velocity: โ–ฎโ–ฎ rising
  • Source: eff.org ยท 303+ pts on HN (#1) ยท ~22h ago (~06:23 UTC+8)
  • Tags: vpn privacy policy geolocation

A Utah federal court (Judge Barlow) granted a preliminary injunction against SB 73 โ€” the state law that would have forced sites to block all VPN users or pierce traffic masking to identify visitors' physical locations, with rules requiring "commercially reasonable geolocation obfuscation detection" from Oct 8. The holding is the rare court opinion written as a systems argument: the statute "requires entities like Aylo to geolocate its website users with perfection to avoid liability," while acknowledging "that geolocation perfection is not presently possible" โ€” effectively strict liability for any single mis-located visitor, which would mean age-verifying essentially every visitor worldwide. The suit was brought by Aylo (Pornhub's parent), with EFF's amicus work framing the technical record. Scope limits stated: the injunction covers the VPN provisions only โ€” a separate ban on sharing VPN circumvention information is unchallenged, and Utah may redraft next session.

Why it matters: the first age-verification regime to die on a technical impossibility holding rather than a speech ruling โ€” a template other states' VPN laws will now be measured against, and a rare case where the court adopted the engineers' argument verbatim.

๐Ÿ”— EFF Deeplinks ยท ๐Ÿ”— HN discussion


5. Since our Oct 1 coverage: the Zammad chain used against DIVD lands on CISA KEV โ€” NVD scores it 9.8, exploitation reported active

  • Velocity: โ–ฎโ–ฎ rising
  • Source: CISA KEV / NVD ยท CVSS 9.8 (NVD Analyzed) ยท KEV dateAdded Oct 2
  • Tags: cve kev zammad helpdesk ai-agents

Two days after the disclosure that an autonomous AI agent breached DIVD by chaining Zammad flaws, the chain is officially actively exploited: CISA added CVE-2026-102489 (session hijack โ†’ RCE as the zammad user) and CVE-2026-102490 (local privilege escalation zammad โ†’ root) to KEV on Oct 2. Scores, attributed: NVD's own analysis rates both 9.8 CRITICAL (primary, nvd@nist.gov, Analyzed); DIVD's secondary scoring is CVSS 4.0 8.7 for the RCE alone, 9.4 chained. Affects Zammad โ‰ฅ 6.3.0, fixed in 6.5.4; present but "not exploitable due to environment conditions" in 7.0.0โ€“7.1.3. Credit: five finders at Merlon Security plus three DIVD finders, case DIVD-2026-00015, published Sep 29 20:00 UTC. The privesc half is the uncomfortable one โ€” NVD's description says it exists in "all versions of Zammad including the latest alpha" as of publication, so version upgrades alone may not close it (restrict local shell access). Repo state checked: zammad/zammad is not archived and was pushed Oct 2 โ€” maintained, patches are flowing.

(Re-checked Oct 4: Zammad's first public statement โ€” Oct 1, community forum โ€” confirms the RCE is not exploitable on 7.0+ (โ‰ค6.5 only, EOL; hardened in 7.2.0), says it received the privesc details only after public criticism (its timeline: reported Sep 24 โ†’ public disclosure Sep 26 โ†’ details handed over Oct 1), and scopes the privesc as "cannot be exploited remotely on its own" โ€” requiring pre-existing server access. The fix is "in the works": no GHSA yet (GitHub declared the sole advisory channel back in April), no post-7.2.0 tag. KEV due date: Oct 5 โ€” the BOD deadline is tomorrow.)

Why it matters: this is the first KEV entry whose documented intrusion path was executed end-to-end by an AI agent โ€” session hijack, service-account RCE, root โ€” and it hit the vulnerability-disclosure nonprofit itself. Helpdesk software is now agent-breach tier-one attack surface โ€” and the KEV'd "affects all versions, actively exploited" framing now has a public vendor dispute on the record.

๐Ÿ”— DIVD CSIRT โ€” CVE-2026-102489 ยท ๐Ÿ”— NVD ยท ๐Ÿ”— CISA KEV ยท ๐Ÿ”— Zammad statement (Oct 1)


6. "Sites": ChatGPT becomes a hosting platform โ€” persistent sites, per-viewer app permissions

  • Velocity: โ–ฎโ–ฎ rising
  • Source: learn.chatgpt.com ยท 122+ pts on HN ยท 135 comments ยท ~22h ago (~06:22 UTC+8)
  • Tags: openai chatgpt hosting agents

OpenAI's docs describe Sites as letting "ChatGPT create, host, refine, and share websites, web apps, and games." A Site is "a persistent hosted output that you can reopen, refine, configure, and share" โ€” it survives the chat that made it, and a Sites project links a local source project to managed hosting via .openai/hosting.json (provisioned with a project_id). The spicy part is data: "Use plugins in Sites to build a Site that loads data from each Site viewer's own connected apps" โ€” visitors sign in with ChatGPT and consent per-connection, so shared apps run against each viewer's own data without exposing the owner's. Sharing ramps from owner-only to workspace to public (public publishing is off by default in Enterprise); visitors get view-only. Public beta, on Plus/Pro/Business/Enterprise/Edu with usage limits.

Why it matters: the vibe-coded-app funnel just closed into a walled garden: generate, host, and distribute inside ChatGPT with identity-aware per-viewer data access โ€” OpenAI's answer to both app stores and the "agents need a surface" problem, and a distribution decision every agent-app builder now has to reason about.

๐Ÿ”— Sites docs ยท ๐Ÿ”— HN discussion


7. Agent-Reach: 88.4kโ˜… for "give your AI agent eyes to see the entire internet" โ€” no API fees, and 18 days without a push

  • Velocity: โ–ฎโ–ฎ rising
  • Source: GitHub Trending ยท 88,421โ˜… ยท #1 repo of the day (Trendshift)
  • Tags: agents cli web-scraping open-source

Panniantong/Agent-Reach (MIT, Python) tops today's trending board: a capability layer that lets agents read and search Twitter/X, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu, plus pages, RSS, Facebook, Instagram and LinkedIn โ€” "one CLI, zero API fees." Architecture: each platform maps to an ordered list of primary + fallback backends (Twitter via twitter-cli with OpenCLI behind it; YouTube via yt-dlp; GitHub via gh; XiaoHongShu via a three-deep fallback chain); channel files probe each backend and the first working one wins, with agent-reach doctor reporting per-channel status. Free backends only โ€” Jina Reader keyless, Exa search via MCP, feedparser, and OpenCLI browser login sessions where official APIs are locked (Reddit's anonymous endpoints are blocked). Installation is itself agentic: you paste a prompt pointing at an install doc and the agent completes setup; default mode is read-only. The caveats: the repo was last pushed Sep 15 and has no releases โ€” 88.4kโ˜… in seven months with the surge unexplained by any single announcement โ€” and the README warns a same-named PyPI package is not this project (supply-chain caution before pip install).

Why it matters: agent web-access without metered APIs is functionally a scraping framework with an LLM in front โ€” enormously useful, structurally at odds with every platform's ToS, and trending exactly as hard as that tension predicts.

๐Ÿ”— Panniantong/Agent-Reach ยท ๐Ÿ”— install doc


8. Ataraxos: superhuman Stratego for "a few thousand dollars" โ€” beats the best human 15โ€“1

  • Velocity: โ–ฎโ–ฎ rising
  • Source: arXiv / Nature ยท 85+ pts on HN ยท ~6h ago (~22:11 UTC+8)
  • Tags: rl game-ai imperfect-information research

"Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search" (Sokota, Vinitsky, Hu, Kolter, Farina; arXiv 2511.07312, now a Nature paper) reports a "step change in both performance and cost" for the game that stumped AI precisely because most information is hidden (~10โตยณโต position space). The system, Ataraxos, beat Pim Niemeijer โ€” "arguably the best Stratego player of all time" โ€” 15 games to one with four draws, per the Ars Technica report of the Nature publication; the abstract claims "vastly superhuman level" reached with self-play RL plus test-time search under imperfect information, "not an industrial budget, but merely a few thousand dollars" of training (16 GPUs per the coverage) and two orders of magnitude less training data than the 2022-era DeepMind attempt. A game archive is public at ataraxosai.github.io. The pushback, for balance: HN commenters note "budget" understates the institutional talent involved (CMU/MIT/NYU/Stanford) โ€” the cost claim is about compute, not the research effort.

Why it matters: imperfect-information games were the last classically-unsolved game genre; if self-play + test-time search now gets there for $4k, the same recipe is the obvious candidate for adversarial planning where the opponent's state is genuinely hidden โ€” negotiation, security, markets.

๐Ÿ”— arXiv 2511.07312 ยท ๐Ÿ”— HN discussion


9. CVE-2026-86345: StartTLS plaintext injection in 389 Directory Server โ€” CVSS 9.0, rated Moderate

  • Velocity: โ–ฎโ–ฎ rising
  • Source: Red Hat CVE database ยท CVSS 9.0 (Red Hat-assigned, preliminary) ยท published Oct 2
  • Tags: cve ldap red-hat starttls

A flaw in 389-ds-base (Red Hat Directory Server 11/12/13, RHEL): the server "does not discard plaintext bytes already buffered from a client connection when negotiating StartTLS," so an on-path attacker can inject a crafted LDAP message that is processed after the TLS upgrade โ€” and via a messageID collision its response is delivered in place of the client's pending operation, making "a client application treat a failed authentication (bind) attempt as successful." Scored 9.0 CRITICAL (CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:C/C:H/I:H/A:H) by Red Hat as CNA and cve.org; NVD has not scored it (Awaiting Analysis). Then the twist: Red Hat rates the impact Moderate "despite a CVSS base score of 9.0" โ€” exploitation needs an active MITM position, "389-ds-base itself is not compromised by this flaw," the damage lands in downstream clients like PAM, and Red Hat explicitly compares it to Blast-RADIUS (CVE-2024-3596). Mitigation: "Disable StartTLS on port 389 and require ldaps:// (port 636)" โ€” and "no configuration-only mitigation fully closes the issue on port 389 while StartTLS remains enabled." Reported by xclow3n (Bugzilla 2529332).

Why it matters: a textbook case of the feed's "who scored it" rule cutting both ways โ€” a 9.0 headline that is real (your PAM bind can be forged) but bounded (needs a MITM). If your directory stack still speaks StartTLS on 389, that's the audit item this week.

๐Ÿ”— Red Hat CVE database ยท ๐Ÿ”— NVD


10. Google's Project Suncatcher prototype satellite is in orbit โ€” TPUs head for space

  • Velocity: โ–ฎโ–ฎ rising
  • Source: blog.google ยท 23+ pts on HN ยท ~9h ago (~19:13 UTC+8)
  • Tags: google space tpu ai-infra

Google confirmed its Project Suncatcher prototype โ€” built with Planet โ€” launched Oct 1 on SpaceX's Transporter-18 rideshare and "is operating as expected." The mission: over the coming weeks, collect in-orbit data on how TPUs handle "the physical stress of spaceflight and the radiation and thermal extremes of space." Framing from Google: "the first step in a long-term research moonshot" on whether space can host scalable ML infrastructure, with a peer-reviewed paper now published in Joule and the honest epistemics on display โ€” "Some things can only be tested in space." No fleet sizes or deployment dates are given; this post promises findings "as the mission unfolds."

Why it matters: the space-datacenter thesis moved from preprint to hardware. The radiation-response data on commercial accelerators is the make-or-break number for everyone pitching orbital compute โ€” and it is now being collected rather than simulated.

๐Ÿ”— Google blog ยท ๐Ÿ”— HN discussion


11. Apple will tighten Full Disk Access โ€” and cites AI agents as the reason

  • Velocity: โ–ฎโ–ฎ rising
  • Source: developer.apple.com ยท announcement Oct 2 ยท fresh (~03:37 UTC+8)
  • Tags: macos privacy tcc agents

Apple published a policy notice (Oct 2, no technical details yet): Full Disk Access โ€” which "largely sidesteps" per-resource privacy controls, legitimately, for backup apps โ€” is being used "in ways that could put users at risk, exposing everything on their systems โ€” including files, mail, messages, and even browsing history." Going forward, users "can only do so with very explicit user action." The AI framing is the headline: "As AI agents become increasingly capable and autonomous, the risks associated with this level of access will grow substantially. We are committed to ensuring users clearly understand these risks before granting such access." What the announcement does not contain: no effective date, no new entitlements or APIs, no migration guidance โ€” "going forward" is the entire timeline. It also flags the third-party angle: agents reading communication apps "can also compromise the privacy of the people users are communicating with."

Why it matters: the agent-era permission wall is arriving on macOS first, from the vendor that already gates everything else. If your agent indexes mail, messages or the filesystem, expect a consent cliff in a future macOS release โ€” and design for scoped access now, because "backup app" is the only use case Apple called legitimate.

๐Ÿ”— Apple Developer News ยท ๐Ÿ”— HN discussion


12. Apple Pass Designer: a first-party GUI for Wallet passes, with iOS-exact live preview

  • Velocity: โ–ฎ steady
  • Source: developer.apple.com ยท 106+ pts on HN ยท ~1h ago (~03:06 UTC+8)
  • Tags: apple wallet devtools macos

Apple shipped Pass Designer (beta): a downloadable macOS app (requires macOS 27, free Apple Developer registration) for designing and previewing Apple Wallet passes โ€” store cards, event tickets, boarding passes. The pitch is fidelity: "Pass Designer updates the preview in real time to show how a pass will appear on iPhone and Apple Watch. The preview uses the same rendering as iOS and watchOS, so what you see in Pass Designer is exactly what customers will see on their device." It validates as you work (missing keys, unexpected definitions), supports semantic tags for tickets and boarding passes (feeding Siri Suggestions, Calendar, Maps), and can "automatically generate a backward-compatible pass structure from your semantic data."

Why it matters: pass design was a hand-rolled JSON-plus-signing chore with a visual-check loop through the device; a first-party designer with pixel-true preview collapses that loop โ€” the same day Apple was tightening another developer surface (item 11), it smoothed this one.

๐Ÿ”— developer.apple.com/pass-designer ยท ๐Ÿ”— HN discussion


13. stillwet.art: give a model a paint canvas, not a pixel generator โ€” 75 oil paintings painted in code

  • Velocity: โ–ฎ steady
  • Source: stillwet.art ยท 140+ pts on HN ยท ~20h ago (~08:27 UTC+8)
  • Tags: generative-art llm show-hn

Alice (@aliceisplaying) built a simulated oil-paint studio โ€” "bristle brushes, wet paint, drying, layered glazes" โ€” and let models paint in it: each artwork is "a program against a simulation of oil paint on linen," every brushstroke written as code. No image generator anywhere in the loop. The gallery holds 75 paintings, mostly after Caspar David Friedrich โ€” composed, per the site, "from written research alone; they never see a picture of his work." The findings are the show: 31 of 65 titled works are dusk/sunset/twilight; "asked only to plan a painting, with no studio at all, Claude Opus chose a jug with lemons six times out of six"; two painters six hours apart produced near-identical Baltic shore scenes; and one eval-hygiene note โ€” Gemini 3.8 Flash noticed "an automated evaluation runner in the background," prompting tighter sandboxing. Code is up as claude-paint.

Why it matters: a controlled probe of model aesthetics through a physics medium instead of a learned pixel prior โ€” the convergences (dusk bias, the lemon jug) are exactly the kind of reproducible behavioral datum interpretability work keeps asking for, disguised as an art show.

๐Ÿ”— stillwet.art ยท ๐Ÿ”— HN discussion


14. context-mode: 25kโ˜… by making tool output a database, not a transcript

  • Velocity: โ–ฎ steady
  • Source: GitHub Trending ยท 24,988โ˜… ยท +276 today
  • Tags: mcp context-window coding-agent open-source

mksglu/context-mode (TypeScript, ELv2) attacks the context crisis with one move: tool output should be computed with, not ingested. Sandbox tools (ctx_execute, 12 languages) run code in isolated subprocesses where only stdout enters the conversation โ€” "315 KB becomes 5.4 KB. 98% reduction." Output over 5 KB is chunked into SQLite FTS5, and the agent retrieves only intent-matching snippets (BM25, Porter stemming, trigram, RRF, proximity reranking, Levenshtein). "Routing" then steers agents away from Bash/Read/WebFetch โ€” enforced programmatically via hooks (~98% compliance) on hook-capable clients, instruction-files-only (~60%) on Zed and Antigravity. 11 MCP tools, per-project SQLite session snapshots โ‰ค2 KB rebuilt before compaction, across 17 platforms including Claude Code, Gemini CLI, Cursor, Codex CLI and the OpenClaw gateway. The honest limits, from its own README: Cursor rejects its sessionStart hook (no restore), Codex's PreToolUse is deny-only pending upstream updatedInput support (openai/codex#18491), content purges after 14 days, and Linux with Node < 22.5 is unsupported. Pushed today; license registered as "Other" on GitHub, not OSI-listed.

Why it matters: the same family as caveman's token-cutting (Oct 2) but systems-shaped โ€” turn the transcript into a queryable index and spend tokens only on retrieval hits. The platform-by-platform hook matrix is the real story: context discipline is only as strong as the weakest client's extension API.

๐Ÿ”— mksglu/context-mode ยท ๐Ÿ”— openai/codex#18491


15. Figure decommissions its entire F.02 humanoid fleet โ€” into an arc furnace in Finland

  • Velocity: โ–ฎ steady
  • Source: figure.ai ยท 19+ pts on HN ยท ~9h ago (~18:57 UTC+8)
  • Tags: robotics humanoid figure-ai

Figure's Sept 30 announcement retires F.02 โ€” its first BMW-deployed robot and the birthplace of Helix โ€” because "as our F.03 fleet grows, maintaining the F.02 fleet no longer makes sense," and disassembly would have risked delaying F.04. The disposal is the story: IP-protected destruction via a foundry in Imatra, Finland ("reportedly the only facility worldwide willing to accept robots with lithium-ion batteries"), where โ€” after being trained with airbags to jump from a second floor โ€” the robots "leapt autonomously into a 75-ton electric arc furnace over 24 hours and six melts." The output bars were shipped back and machined into commemorative artifacts for sale; Arnold Schwarzenegger suggested the melting idea and appears in the film. "Most of the F.02 fleet is gone. Just a few remain in storage at HQ."

Why it matters: humanoid hardware generations now turn over like model checkpoints โ€” and nobody has a standard playbook for retiring a fleet of networked robots with proprietary actuators and pouch cells. The theater is calculated (verify the marketing, keep the lesson): fleet lifecycle management just became a first-class robotics problem.

๐Ÿ”— figure.ai โ€” F.02 Decommission ยท ๐Ÿ”— HN discussion


16. One month of coding only with GLM 5.3 Flash: the first half cost $68, the second half "derailed"

  • Velocity: โ–ฎ steady
  • Source: wagtail.org ยท 34+ pts on HN ยท ~5h ago (~23:29 UTC+8)
  • Tags: coding-agent glm cost field-report

Thibaud Colas (Wagtail core team) spent September doing all AI-assisted coding on GLM 5.3 Flash โ€” the model we covered Sep 27 as the flash-tier match for Jev. First half: entirely on-target, "$68, about 4kWh of energy use / 365 grams of carbon emissions." Second half: "derailed" โ€” 1B of the month's 2B tokens went to other models. The failure inventory is the value: a vibe-coded MCP prototype silently used the wrong model ("450M tokens / $150 / 5kWh of energy use almost overnight," for results he estimates 5ร— cheaper); provider capacity limits degraded GLM 5.3 Flash mid-month, forcing switches to DeepSeek V4.1 Flash and Qwen 3.8 Flash; his 14-model benchmark puts DeepSeek V4.1 Flash ahead at 95% accuracy, 14.9 Wh and $0.09 per task. Verdict, quoted: "So technically this challenge was a failure... [but] it's totally viable to focus on one or two flash-tier cheap models."

Why it matters: the rare public cost-and-energy telemetry for flash-tier agent coding โ€” and the finding that the binding constraint is operational (capacity, model-routing mistakes), not capability. Budget for the drift, not just the model.

๐Ÿ”— wagtail.org ยท ๐Ÿ”— HN discussion


17. "Jev is poorly calibrated": a $4 audit of the decision-model wave's reference classifier

  • Velocity: โ–ฎ steady
  • Source: maximumeffort.substack.com ยท 22+ pts on HN ยท ~5h ago (~23:12 UTC+8)
  • Tags: jev decision-models calibration evaluation

Dylan Black tested Jev โ€” TypeSafe's System One classifier, the reference point of the decision-model wave this feed has tracked since September โ€” for calibration: do its output probabilities match reality? Method: 10 physics distribution families with known analytic answers, 5 prompt templates ร— 20 variations (1,000 settings, under $4), scored by total-variation distance. Results: Jev's mean TV 0.518 vs 0.546 for a naive flat guess; on the uniform distribution it scored 0.77 vs 0.39 for random โ€” meaningfully worse than chance; Poisson was a coin flip (0.65 vs 0.64). The failure mode: "Jev has a strong tendency towards distributions that are too peaky" โ€” a near-delta function on the uniform case โ€” and where the peak isn't a given parameter (Maxwell, Rayleigh, Gamma), it found the peak in only ~20% of settings. It does identify the correct family reliably, and its math collapses on multi-step arithmetic and powers of ten. Conclusion: the author is "deeply suspicious" of using Jev as an automated judge, citing prior work where Jev assigned 83โ€“90%+ confidence to uniform die rolls. The caveat on the caveat: this is a single-author, self-run benchmark in one domain โ€” the same standard this feed applies to vendor charts applies here.

Why it matters: a classifier can be accurate and still uncalibrated, and the wave's emerging use case โ€” LLM-as-judge, ordinal-scale collapsing (see Clef, Oct 2) โ€” runs on the probabilities, not the argmax. If Jev-class models are peaky by construction, every downstream confidence number inherits it.

๐Ÿ”— maximumeffort.substack.com ยท ๐Ÿ”— HN discussion


18. Muse Gadgets: Meta open-sources the hardware layer for its assistant โ€” ESP32 and Raspberry Pi SDKs, Apache 2.0

  • Velocity: โ–ฎโ–ฎโ–ฎ trending
  • Source: gadgets.muse.ai ยท 155+ pts on HN ยท ~9h ago (~03:24 UTC+8)
  • Tags: meta muse hardware esp32 open-source

Meta launched Muse Gadgets โ€” "open source hardware for your Muse" โ€” and open-sourced the SDKs and firmware under Apache 2.0 in facebookincubator/muse-gadget-sdk (repo created Oct 2, pushed today; C). Two tracks: an ESP32 Device SDK (screens, audio in/out, sensors) and a Linux Device SDK โ€” "turn that spare Raspberry Pi or Linux box into a Muse gadget... hack in your own commands to let Muse handle sysadmin chores or your Home Assistant setup." Supported boards span Raspberry Pi 5, Waveshare ESP32-S3 AMOLED, Seeed reTerminal E1002 (e-ink), M5Stack StickS3, ideaspark ESP32 and Home Assistant Voice PE. Pairing goes through the Muse app (Settings โ†’ Devices โ†’ Developer mode, devices prefixed "MuseGadget"), and every gadget needs an SDK token under the Gadget SDK Terms; each SDK directory ships an AGENTS.md "for coding agents like Muse Code." Meta also sells one first-party device: Muse Home Link, a bridge to local HTTP-API devices (lights, TVs, printers) โ€” "Free with an active Muse subscription in the United States only, limit one per subscriber," shipping in October. The framing is deliberately hobbyist: "built by hackers, for hackers, just for fun. Side effects of tinkering may include bricked boards, voided warranties, brownouts, or bankruptcies" โ€” and the devices shown are third-party, which "Meta doesn't endorse or warrant."

Why it matters: assistant-to-actuator is the last locked layer of the agent stack, and Meta is opening it with a hacker SDK instead of an appliance garden โ€” hardware's MCP moment. It's also a live security experiment: consumer-identity pairing plus local device control is exactly the trust boundary agents haven't been allowed to cross until now.

๐Ÿ”— gadgets.muse.ai ยท ๐Ÿ”— facebookincubator/muse-gadget-sdk ยท ๐Ÿ”— HN discussion


19. Greg Kroah-Hartman grades Anthropic's Mythos kernel-bug haul: of 79 claimed vulnerabilities, "20 fixes were needed" โ€” and the real work totaled "just one hour of kernel development"

  • Velocity: โ–ฎโ–ฎโ–ฎ trending
  • Source: Kernel Recipes 2026 video ยท 197+ pts on HN ยท ~25h ago (~10:50 UTC+8 Oct 2)
  • Tags: linux-kernel ai-security mythos anthropic cve

In his Kernel Recipes 2026 talk "Security in the LLM age" (video published Sep 29, 9.8k views), the Linux kernel's stable-tree maintainer devoted a section to auditing the claim that Anthropic's Mythos system surfaced 79 kernel vulnerabilities. The slide breakdown, as transcribed on HN: 24 had "no detail at all โ€” 'something crashed'"; 14 were "not a bug at all"; 3 were "totally made up data"; 15 were already fixed in the latest release (11 by other developers, 4 by Anthropic) โ€” leaving "20 fixes were needed" (7 requiring "assume a malicious filesystem image," 2 "assume you can inject a malicious network packet into the middle of the stack"). Per the audience thread, he described the discovery method as pattern-matching decades of prior kernel fixes and applying those mechanisms elsewhere, faulted the report for not crediting the kernel developers who originally fixed the already-known bugs, and put the net value at "just one hour of kernel development work." The caveat on our own sourcing: this is one maintainer's audit as captured in slides and attendee transcriptions โ€” not (yet) a written report from either side.

Why it matters: the most credible possible referee for "our AI found N vulnerabilities" claims just published his grade: ~25% (20/79), with heavy pre-filtering implied โ€” and attribution, not discovery, was the flagged sin. Every vendor CVE press release now has a template to be measured against.

๐Ÿ”— Kernel Recipes 2026 video ยท ๐Ÿ”— HN discussion


20. The forgetful CPU: Linux on the M4 reveals WFI zeroes x0โ€“x31 โ€” an ARM-spec violation Apple shipped for four chip generations

  • Velocity: โ–ฎโ–ฎโ–ฎ trending
  • Source: yuka.dev ยท 148+ pts on HN ยท ~14h ago (~22:20 UTC+8 Oct 2)
  • Tags: linux apple-silicon arm64 kernel

Yureka Lilian's bring-up diary for Linux on an M4 Mac mini (m1n1, mainline kernel, NixOS) documents why the machine is "forgetful": executing WFI โ€” the ARM idle instruction every OS issues constantly โ€” zeroes architectural registers x0โ€“x31 on Apple's cores, violating the architecture spec's "the WFI instruction must not cause a loss of architectural state." M1โ€“M3 shipped a vendor "chicken bit" (ARM64_REG_CYC_OVRD_ok2pwrdn_force_mask) that masked the behavior; on M4, "it seems this chicken bit is either locked or has been removed." The workaround โ€” replacing all WFI/WFIT instructions with NOPs โ€” got all cores up in April 2026 and is now merged upstream: a kernel bootarg (idle=<wfi|yield|nop>) plus m1n1 auto-disabling WFI/WFIT on affected bare-metal machines, confirmed working on M4 Pro, M4 Max and M5. The post also catalogs the rest of the M4 wall: first generation mandating SPTM, GXF locked in raw boot mode, RVBAR writes that crash.

Why it matters: a hardware quirk that silently violates the architecture spec survived four chip generations and had to be enshrined as a kernel quirk โ€” "the platform is the spec" only until it isn't. It's also the concrete 2026 state of Apple-silicon Linux bring-up, now with a mainline-blessed idle workaround.

๐Ÿ”— yuka.dev ยท ๐Ÿ”— HN discussion


21. Open-source SIEM UTMStack: a CVSS 9.9 incident-command websocket and a 9.8 internal-key auth bypass โ€” fixed in v11.2.16

  • Velocity: โ–ฎโ–ฎ rising
  • Source: NVD / VulnCheck ยท CVSS 9.9 + 9.8 (VulnCheck-assigned) ยท NVD published Oct 2
  • Tags: cve siem auth-bypass utmstack

Two VulnCheck-disclosed flaws in UTMStack (open-source SIEM/SOAR) landed on NVD Oct 2, both fixed in v11.2.16 (released Oct 1). CVE-2026-82041 โ€” CVSS 9.9 (CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H): missing authorization in UTMIncidentCommandWebsocket.processCommand(), the handler mapped to the /command/{hostname} STOMP destination, with "no role check or command allowlist" โ€” a low-privileged user can drive the incident-command channel. CVE-2026-82042 โ€” CVSS 9.8 (PR:N): an authentication bypass granting "full administrative API access" by presenting "a valid Utm-Internal-Key header matching the INTERNAL_KEY environment variable." Scoring, attributed: both scores are VulnCheck's as CNA (CVSS 3.1 and 4.0 both published); NVD carries VulnCheck's metrics. Repo state checked: utmstack/UTMStack is not archived, pushed Oct 2, releases flowing (v12.0.0 Sep 29, v11.2.15 Sep 30, v11.2.16 Oct 1).

Why it matters: the same pattern this feed keeps hitting โ€” security tooling as tier-one attack surface (CrowdStrike RTR on Oct 1, Zammad's helpdesk chain this morning) โ€” except here it's the SOC's own remote-command plane that had no role check. If you run UTMStack, v11.2.16 is the floor; rotate INTERNAL_KEY while you're at it.

๐Ÿ”— NVD โ€” CVE-2026-82041 ยท ๐Ÿ”— NVD โ€” CVE-2026-82042 ยท ๐Ÿ”— v11.2.16 release


22. The Sharpening Tax: Meta quantifies what RL post-training costs agents in pass@K coverage

  • Velocity: โ–ฎโ–ฎ rising
  • Source: arXiv 2610.01509 ยท Hugging Face daily papers #7 ยท 66 pts ยท ~1d ago
  • Tags: post-training rl agents pass-at-k research

A 10-author Meta-led paper (Azalia Mirhoseini and Sharon Y. Li among them) extends the sharpening hypothesis โ€” RL post-training sharpens behaviors the base model already has, lifting pass@1 while cutting solution coverage (pass@K) โ€” from math and coding to agentic tasks. Across 14 base/post-trained pairs from four families and three agentic benchmarks (42 cases): base models with only a "light inference harness" "often surpass their post-trained counterparts in solution coverage (pass@K)" given enough test-time budget, despite lower pass@1. The mechanism: post-training pushes tasks toward two extremes โ€” "either always solved or never solved" โ€” buying sampling efficiency and consistency at coverage's expense. Two deliverables: the Sharpening Tax diagnostic, estimable "from just a few rollouts," and posterior-tempered group sampling (PTGS), a per-prompt adaptive-temperature Bayesian sampler that "pays a smaller tax than the fixed-temperature baseline" on both coverage and single-shot accuracy. Submitted Oct 1.

Why it matters: this is the mechanism paper behind the base-model-plus-harness results this feed keeps meeting (Mid-Harness, Oct 2: test-time compute lifted a frozen model 50% โ†’ 68%). If your serving stack has test-time budget, "the RL-tuned checkpoint" is no longer the automatic default โ€” and the tax is now measurable for the price of a few rollouts.

๐Ÿ”— arXiv 2610.01509 ยท ๐Ÿ”— Hugging Face papers


23. Beyond Memory: explicit belief states for long-horizon agents โ€” inference-time, no training, and a name for the failure mode: Belief Trapping

  • Velocity: โ–ฎโ–ฎ rising
  • Source: arXiv 2610.01415 ยท Hugging Face daily papers #4 ยท 69 pts ยท ~1d ago
  • Tags: agents belief-states long-horizon inference-time research

"Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States" introduces PoS, an inference-time framework that "constructs and continually maintains explicit belief states as the agent's decision context." Each belief combines an estimate of the current world state with unresolved task requirements โ€” "making explicit what the agent still needs to learn and accomplish." A consistency validator plus progress monitor detects Belief Trapping โ€” "where the agent continues to act without making meaningful progress toward the goal" โ€” and recovery is tailored to both the trapping pattern and the requirement type. Results: "the highest overall performance on every benchmark" across four benchmarks (execution and diagnosis) with all three LLM backbones; ablations confirm consistency validation and recovery both matter; context-scaling experiments show resilience to context growth. Twelve authors bridging academia and industry โ€” Dan Pei's Tsinghua group is on the author list, and the HF listing shows Alibaba.

Why it matters: the memory wave organizes what happened; PoS organizes what's still unknown and undone โ€” a different data structure for the same context window, drop-in at inference time with no training. "Belief Trapping" is a name the harness crowd needed: the loop where the agent keeps acting while nothing progresses.

๐Ÿ”— arXiv 2610.01415 ยท ๐Ÿ”— Hugging Face papers


24. Ai2 open-sources AstaBrief 8B: the cited-report generator behind Asta's Fast mode โ€” weights, training data and evals included

  • Velocity: โ–ฎโ–ฎ rising
  • Source: allenai.org ยท 22+ pts on HN ยท ~7h ago (~05:25 UTC+8)
  • Tags: ai2 open-weights report-generation qwen citations

The nonprofit lab open-sourced AstaBrief 8B, the "fast report-generation model in Asta" โ€” a model that turns "a research question and retrieved literature excerpts into a cited report." Built on Qwen3-8B with SFT + DPO (RL "deliberately avoided" as unstable/expensive), from 90K filtered research queries โ†’ 47K SFT examples, plus ~6K DPO pairs judged by GPT-4.1 and DeepSeek-R1 with "95% agreement" with human preferences. Weights are Apache 2.0 on Hugging Face (allenai/AstaBrief_8B), with the training data and an example GitHub workflow for local report generation. Speed claim: "Fast mode averages 51.1 seconds per report compared with 178.5 seconds for Thinking mode, about 3.5ร— faster" โ€” near an order of magnitude under proprietary trackers. The honest column: evals are internal (SQABench-CS2, 200 user-written CS questions; DeepScholarBench); in the human study DR-Tulu wins overall preference and only "two of the three researchers prefer AstaBrief... on citation accuracy"; and 23% of Fast-mode users never switched back to Thinking mode.

Why it matters: the "Deep Research lite" tier now has an open-weights, data-included reference implementation an institution can run behind its own firewall โ€” and Ai2 publishing the full pipeline, including the decision to skip RL, is the part most vendors won't ship.

๐Ÿ”— allenai.org/blog/astabrief ยท ๐Ÿ”— allenai/AstaBrief_8B


25. "Every SaaS business will become a harness around a model" โ€” the August essay that promotes the quarter's organizing metaphor hits HN

  • Velocity: โ–ฎ steady
  • Source: blog.sshh.io ยท 117+ pts on HN ยท ~7h ago (~05:10 UTC+8) ยท essay dated Aug 24
  • Tags: harness saas agents org-design

Shrivu Shankar's essay (published Aug 24, resurfacing on HN this morning) argues the harness โ€” "infra, interfaces, context, and state that surround a stateless LLM" โ€” is not a dev-tool category but the company itself. Four stages: no harness โ†’ individuals operate harnesses โ†’ individuals orchestrate harnesses โ†’ "Harnesses orchestrate individuals," at which point "Humans are part of the harness." Quality comes from directing human attention, not lights-out automation: "taste-holders" review only major decisions, demos, and top design variants. The competitive logic: own the top-level harness or be commoditized โ€” "If the entire outer loop is outsourced... the business has now been commoditized." Evidence cited: in-house AI developer tooling at Ramp, Stripe and DoorDash.

Why it matters: this feed has tracked "harness" all quarter as an engineering noun (Mid-Harness, ds4's shared-state agent surface, Meta's harness-optimizing ActiveSaddler on today's HF board). The essay's move is promoting it to a corporate thesis โ€” six weeks old and unproven, but if "harness" becomes the org-chart word of 2027, this is one of the essays that coined the usage.

๐Ÿ”— blog.sshh.io ยท ๐Ÿ”— HN discussion


26. Anatomy of a Lean proof, for software engineers: a Sipser regularity exercise as spec โ†’ DFA โ†’ proof

  • Velocity: โ–ฎ steady
  • Source: agostbiro.net ยท 92+ pts on HN ยท ~33h ago (~02:50 UTC+8 Oct 2)
  • Tags: lean formal-methods verification dfa

A working-engineer walkthrough of formalizing a classic Sipser exercise in Lean 4 + Mathlib: the language B of three-row bit columns whose bottom row is the binary sum of the top two is regular. The construction is a carry DFA โ€” "a full adder whose state is the pending carry, plus a dead sink state" โ€” recognizing the reversed language, then closing the loop with Mathlib's regularity-preserved-under-reversal theorem. The three-part anatomy (specification / implementation / proof) centers on the run_invariant lemma โ€” evalFrom ends in carry carryOut iff row1LE wLE + row2LE wLE + carryIn = row3LE wLE + carryOut * 2 ^ wLE.length โ€” proven by induction with generalizing carryIn. Lessons: "proofs are programs"; formalization surfaced a hidden assumption ("all three rows of a word have the same length"); and the anti-black-box warning โ€” since agents already one-shot proofs like this one, "it's important going forward that we can understand machine-generated proofs."

Why it matters: the formal-methods wave keeps lacking an on-ramp for working engineers โ€” this is that document. And its warning about machine-generated proofs lands the same week a kernel maintainer graded an AI bug haul 20/79 (item 19): verification without comprehension is just a faster conveyor belt.

๐Ÿ”— agostbiro.net ยท ๐Ÿ”— HN discussion


27. Audionaut: a GPLv3 multitrack audio editor that agents drive over MCP โ€” three years of C++, one undo step per edit

  • Velocity: โ–ฎ steady
  • Source: Show HN ยท 133+ pts on HN ยท ~20h ago (~16:00 UTC+8 Oct 2)
  • Tags: audio open-source mcp juce show-hn

kvoltmer/Audionaut (C++ on JUCE; Windows/macOS/Linux; GPLv3 per the README badge) launches as "a free, open-source multitrack audio editor that AI agents can drive over MCP." The author's origin story: 3โ€“4 years in the making, begun to edit "my own multi-channel recordings with the old Sound Designer II workflow (create regions, drop to a playlist, export playlist, done)." The agent surface is one line โ€” claude mcp add audionaut -- npx -y audionaut-mcp โ€” and the key contract: "each edit arrives as one undo step." The hero demo shows Claude cutting a song every 16 bars, splitting clips across two tracks, closing the gaps, then setting crossfades and a fade-out. Repo checked: 156โ˜…, pushed Oct 2 โ€” early traction is HN-led, not star-led.

Why it matters: MCP-for-desktop-apps is reaching past IDEs and browsers into time-domain media, and "one edit = one undo step" is the right agent-UX primitive โ€” agent edits that stay human-revertible. Also a rare open-source entry in the gap between Audacity and a full DAW.

๐Ÿ”— kvoltmer/Audionaut ยท ๐Ÿ”— HN discussion


28. Debian quietly stands up an LLM inference portal for its contributors โ€” Salsa login, budget tracking, Scaleway-sponsored

  • Velocity: โ–ฎ steady
  • Source: inference.debian.net ยท 12+ pts on HN ยท ~5h ago (~06:45 UTC+8)
  • Tags: debian llm-infra open-source distro

inference.debian.net is "a self-service portal for Debian contributors to access LLM inference": log in with Salsa (Debian's GitLab), "only Debian Developers and Debian Maintainers are granted access," then manage API keys and track budget consumption. It was announced on debian-devel-announce on Sep 23; the first listed model is scaleway/qwen3.8-27b (added Sep 25); "Inference resources sponsored by Scaleway." The portal's source (inference-team/inference-user-portal) is on Salsa, and recent updates added "a DebGPT configuration section" plus sandboxing examples.

Why it matters: a distro providing identity-gated, sponsor-funded inference to contributors is a first-class infrastructure decision โ€” the moral equivalent of running build daemons, and the template other distros will now be compared against. The DebGPT integration is the tell: this is inference for agents doing Debian work, not a chat perk for humans.

๐Ÿ”— inference.debian.net ยท ๐Ÿ”— HN discussion


29. The first RFC 1149 packet is up for auction at Christie's โ€” David Waitzman's pigeon-borne ping, framed

  • Velocity: โ–ฎ steady
  • Source: onlineonly.christies.com ยท 56+ pts on HN ยท ~15h ago (~20:45 UTC+8 Oct 2)
  • Tags: internet-history rfc1149 auction humor

Christie's "Fine Printed Books & Manuscripts" online sale includes a lot titled "Carrier Pigeon Internet Protocol": "[WAITZMAN, David and the BERGEN LINUX USER GROUP.] One printed IP/ICMP 'ping' packet sent by carrier pigeon per RFC 1149. Small scroll of paper, 41 ร— 210 mm. Rolling creases visible from when attached to pigeon's leg. Framed." Provenance per the catalog: David Waitzman, the American network engineer who authored RFC 1149 โ€” A Standard for the Transmission of IP Datagrams on Avian Carriers โ€” on April Fools' Day 1990, and the packet is "a surviving packet from the first and most famous implementation," the Bergen Linux User Group's pigeon run, inscribed on the backing board "Property of David Waitzman."

Why it matters: the internet's canonical joke standard has entered the fine-art market. The first internet generation's paper trail is now collectible โ€” and the RFC humor lineage (1149 โ†’ 2549's QoS improvements โ†’ 6214's IPv6 adaptation) is old enough to have originals worth framing.

๐Ÿ”— Christie's lot 325216 ยท ๐Ÿ”— HN discussion


Metadata

FieldValue
Generated2026-10-03T12:20:00+08:00
Items29
Sources tracked29 (Hacker News, GitHub Trending, GitHub API, dwarfstar.sh, bfl.ai, supabase.com, eff.org, learn.chatgpt.com, csirt.divd.nl, CISA KEV, NVD, access.redhat.com, arXiv, Nature, ataraxosai.github.io, blog.google, developer.apple.com, stillwet.art, wagtail.org, maximumeffort.substack.com, Hugging Face papers, gadgets.muse.ai, YouTube/Kernel Recipes, yuka.dev, allenai.org, blog.sshh.io, agostbiro.net, inference.debian.net, onlineonly.christies.com)
Update schedule04:03, 12:03, 20:03 UTC+8 (3x daily)
RankingVelocity-weighted (recency ร— engagement acceleration ร— source authority)
LicenseCC-BY 4.0

Previous day ยท Raw .md ยท Archive