Smart routing โ€” "route before compute" (Aug 2026)

A cross-cutting pattern that appeared in three independent projects in a single batch: a
classification/routing layer that inspects each unit of work and sends it to the *cheapest
engine that can do it well* โ€” instead of running everything through the most expensive engine.

The pattern

Classify first, dispatch second. Each request/page/inference gets a cheap "which engine?" decision,
then goes to the smallest capable model/parser. The saving comes from not sending work to the
heavy path: the bulk of units are handled by a cheap engine, and only the genuinely hard tail
reaches the expensive one.

Four instances (same shape, different domain)

  1. Model routing โ€” NeMo Switchyard (NVIDIA-NeMo/Switchyard, Apache 2.0, Rust). Translates between OpenAI Chat / Anthropic Messages / OpenAI Responses and routes each request across a pool of models (vLLM, NIM, Ollama, any OpenAI-compatible endpoint). Built-in routers (verified from the repo's routing table): llm_classifier (content decides weak vs strong tier), stage_router (conversation signals route most turns without an extra model call), escalation (llm_classifier mode="escalation" โ€” weak tier first, a judge decides whether to escalate), random (fixed A/B split), plus passthrough (single target, no routing decision). LangChain cut cost 74% by routing only 7% of calls to a frontier model โ€” at a 6% accuracy tradeoff (145 multi-turn Deep Agents tasks); the internal benchmark claims frontier-level accuracy at ~1/3 the cost of Claude Opus 4.8 alone. (The repo confirms the mechanics โ€” Apache 2.0, ~755 stars, pre-alpha; the 74%/7% + Opus figures come from NVIDIA's blog, which launched Switchyard alongside the 30B-MoE Nemotron 3.5 Lightning.)
  1. Document routing โ€” Firecrawl pdf-inspector (firecrawl/pdf-inspector, MIT, Rust). Reads a PDF's internal structure (font encodings, text operators, image coverage) without rendering and classifies each page TextBased/Scanned/ImageBased/Mixed in ~10โ€“50ms. Text pages get native extraction; only the rest go to OCR. Skipping OCR on the ~54% text-based PDFs is how Firecrawl made its hosted parser 3.5โ€“5ร— faster. Ships Python (PyO3) / Node (napi-rs) / WASM bindings plus pdf2md / detect-pdf CLIs; 0.875 on opendataloader-bench.
  1. Inference escalation โ€” Needle 2 (cactus-compute/needle, MIT). 45M-param / 14MB model that solves problems as function calls and returns structured JSON with a calibrated confidence score; low-confidence results escalate to a bigger model. Runs the whole session locally (~28MB RAM), so the expensive path is only taken on the tail.
  1. Search sub-agent โ€” Toast 1 (mixedbread). A specialized search agent that decomposes a query into sub-queries, gathers evidence, inspects sources, and curates context before a generalist frontier model answers โ€” claiming frontier-class quality at up to 10ร— lower cost and 12ร— faster. On Databricks' OfficeQA Pro V2, GPT-5.6 Sol + Toast 1 hit 70% at ~$1.15/task vs Claude Fable 5 at 60% for ~$4/task; on Harvey's Legal Agentic Benchmark it cut token usage from 80.6M โ†’ 23M while preserving quality. The classify-then-cheap-specialist shape applied to retrieval: the search/decomposition work is offloaded to a specialized model so the frontier model only does final synthesis.

Why this matters

Four different domains โ€” LLM serving, document parsing, on-device agents, search/retrieval โ€” but the same
optimization: **the expensive engine (frontier LLM / GPU OCR / cloud inference) should only ever
see the tail of the distribution.** As multi-model and multi-parser workloads proliferate, "which
engine serves which unit" becomes its own layer โ€” a new control point that the router owner
controls.

A fifth instance โ€” voice-stack routing (Aug 18)

Speko (YC S26, SpekoAI/gateway, MIT, Go) is "OpenRouter for Voice AI" โ€” the same
classify-then-cheap-specialist shape applied to a stack instead of a single engine. Send criteria
(accuracy/latency/cost, language, region) and it benchmarks 50+ providers / 140+ models across the
STT, LLM, and TTS layers, picks the winner, and returns provider + model + scores in response
headers. The MIT gateway runs as a local sidecar (BYOK, no call-home); hosted routing costs 5% over
provider rates; public boards publish WER/latency/cost-per-minute at benchmarks.speko.ai.

Signal: voice stacks rot because nobody re-benchmarks after launch โ€” continuous independent evals plus
a drop-in gateway turn "which STT/TTS for Spanish medical calls" into an answered, routeable
question. It is the first routing instance where the routed unit is a multi-layer pipeline
(STTโ†’LLMโ†’TTS) rather than a single model call โ€” the classify-first pattern scaling from "which model"
to "which stack."

Router lock-in map (verified 2026-08-13)

"Where does lock-in form?" โ€” comparing the four routing approaches against what a router controls
(policy, signal, catalog):

  1. Hosted aggregator โ€” OpenRouter (SaaS, ~$10B valuation, ~1.5 quadrillion tokens/yr). Default routing is inverse-square price-weighted (with a 30s outage window) plus an "Auto Exacto" step that tiers providers by tool-call quality; a per-request provider object overrides it (order, sort, only, max_price, allow_fallbacks). Pass-through token pricing ("no markup"), with the margin on ~5.5% credit fees + ~5% BYOK. Lock-in = one key, one bill, and a model catalog + routing policy you don't own. Its "Fusion" multi-model fan-out (up to 8 models + a judge) is a proprietary value-add independent testing measured at ~4ร— a solo frontier call.
  2. Vendor router โ€” NeMo Switchyard (NVIDIA, Apache 2.0). Routes on top of the inference stack (NIM, vLLM); NVIDIA frames it as "orchestration software on top of the chips." Lock-in = routing coupled to NVIDIA's accelerator/NIM stack.
  3. Self-hosted OSS gateway โ€” LiteLLM (MIT, ~40K stars). Router = load balancing across model_group deployments, fallback chains, retries, budgets, rate limits, virtual keys. No vendor lock-in โ€” the "lock" shifts to your own config being the control point (Postgres + Redis state).
  4. Confidence-gated escalation โ€” Needle 2 (MIT). The escalate-or-not decision is a calibrated confidence score embedded in the model's output. Lock-in = the escalation policy is owned by the confidence model; if proprietary, the "when to pay for the frontier" decision is unauditable.

Where lock-in forms โ€” three vectors, all of which are the router decision itself:
(a) ownership of the policy (you in LiteLLM; the vendor in OpenRouter/Switchyard),
(b) ownership of the signal (Switchyard's classifier, OpenRouter's Auto Exacto tiers, Needle's
confidence), (c) ownership of the catalog + billing (OpenRouter's 70+ providers + one bill;
NVIDIA's NIM catalog). There is no shared routing-config standard yet โ€” each has its own DSL
(LiteLLM YAML, OpenRouter provider object, Switchyard router types). That fragmentation is the
lock-in surface: an "MCP for routing" would commoditize it, and nobody has shipped one.

The standard is emerging (Aug 15 20:31)

"Who ships a shared routing-config DSL?" now has two concrete answers โ€” neither yet the winner:

  1. BitRouter (bitrouter/bitrouter, Apache 2.0, ~220 stars, 821 commits, local-first Rust proxy). The first router to make three primitives routable under one gateway, not just model calls: - Models โ€” cross-protocol translation (OpenAI Chat/Responses, Anthropic Messages, Gemini), multi-account failover, streaming. - Capabilities โ€” an MCP gateway (proxy MCP servers so agents discover/call tools across hosts) plus an AgentSkills gateway (tracks/exposes SKILL.md skills per the agentskills.io standard); both fold into one ToolEntry type surfaced at GET /v1/tools. - Agents โ€” an ACP (Agent Client Protocol) gateway making sub-agents first-class routable primitives (local stdio today; remote with ACP v2). Policy is declarative: bitrouter.yaml declares providers/presets, and a git-owned policy-lock.yaml is "the only live route authority" โ€” tier targets, canonical routes, capability guardrails, decision certificates โ€” produced by a self-improving act โ†’ observe โ†’ evaluate โ†’ learn loop. Runs under the harness (Claude Code, Codex, OpenCode, Pi-Agent) via a base-URL env swap. Its one validated objective is cost: gpt-5.5 on Terminal-Bench 2.1 โˆ’32.8% cost at โˆ’1.1pp accuracy (76.1% vs 77.3%), self-described as a mechanism study, not a leaderboard entry.
  2. Semantic Router DSL (arXiv 2603.27299 โ€” "From Inference Routing to Agent Orchestration: Declarative Policy Compilation with Cross-Layer Verification"; Chen, Liu, He, Liu). A non-Turing-complete declarative routing-policy language: one source file compiles into verified decision nodes for LangGraph/OpenClaw, Kubernetes artifacts (NetworkPolicy, Sandbox CRD, ConfigMap), YANG/NETCONF payloads, and protocol-boundary gates (MCP, A2A). Because it emits only policy-decision logic (no sequencing/loops/side effects), the compiler guarantees exhaustive routing, conflict-free branching, dead-branch detection, and audit traces structurally coupled to the decision logic. This is a position paper โ€” architectural claims, no measured results yet.
  1. MCP-native routing (the protocol itself โ€” 2026-07-28 stateless rewrite). The Model Context Protocol's July 28 2026 "stateless core" rewrite is, de facto, the MCP-native routing extension this question kept predicting โ€” but it arrived as the protocol, not a third-party DSL. It drops the initialize/initialized handshake, Mcp-Session-Id, and sticky sessions (a remote server "can now run behind a plain round-robin load balancer"), moves protocol metadata into per-request _meta, adds server/discover for connection-free capability discovery, and โ€” the routing part โ€” adds two mandatory routing headers, Mcp-Method and Mcp-Name, so gateways / WAFs / rate limiters route, throttle, and meter agent traffic without opening the JSON-RPC body (tool params can also be copied into headers for fine-grained routing; results carry ttlMs/cacheScope; Multi Round-Trip Requests put server-initiated state in the payload, not an open SSE stream). It is not a routing policy DSL โ€” but it makes routing a protocol-native, commodity transport concern, which is exactly what would commoditize the BitRouter/DSL lock-in. Two IETF drafts extend the same idea to cross-protocol routing headers (draft-hood-agtp-composition: Authority-Scope + Budget-Limit; draft-gaikwad-agent-proxy-modes: proxy gateway routing layers).

The shape has shifted again: the question is no longer "which standalone DSL wins" โ€” it is "does a
routing-policy DSL survive once the transport (MCP's stateless core + Mcp-Method/Mcp-Name
headers, AGTP for cross-protocol) makes basic routing a commodity?" The likely end-state is a
two-layer split, not one winner: MCP/AGTP own the how-to-route-a-request transport layer, while
the policy (what tier gets what call, and who may change it) stays a git-owned artifact (BitRouter's
policy-lock.yaml) or a verified-compiled research DSL (Semantic Router). The lock-in surface moved
from absence of a standard โ†’ choice of standard โ†’ transport vs policy.

Watch for

A sixth instance โ€” A2A agent-network routing (Aug 19 20:03)

Sprix SAGE Router (wang2122/sprix-sage-router, MIT, Python, 362 stars, v0.2 research preview) is a
decision layer that sits between A2A protocol discovery and task execution, choosing mid-run
whether the incumbent agent should continue alone (SELF), recruit collaborators while keeping
ownership (COLLABORATE), or transfer full ownership (HANDOFF). It composes task-DAG roles,
schedules dependencies, and updates trust from execution evidence under permission/budget/deadline
constraints, using a learned outcome model plus beam-search team composition. The README's 2,500-task
simulation (0.634 vs 0.507 incumbent-only quality) is flagged synthetic.

Signal: as A2A (now a Linux Foundation protocol) matures, the open problem shifts from "can agents
talk" to "when should they collaborate vs hand off" โ€” the missing middle layer between discovery and
execution. This is "route before compute" applied one level above the model router: the routed unit
isn't a model call but a whole sub-task's ownership. The learned, evidence-based SELF/COLLABORATE/
HANDOFF decision is the A2A-era answer to the "who owns the router decision" lock-in question โ€” with
the same caveat as every learned router (the outcome model is a black box and the eval is synthetic).

A seventh instance โ€” the router's ownership becomes a supply-chain question (08-21 04:03)

OpenRouter is joining Stripe (announced Aug 19; the sale has not closed). The multi-provider
router a large share of agent stacks call instead of talking to vendors now has a parent company โ€”
and the post is explicit about what that means: "same mission, same name, same product, same
roadmap," "if you build on OpenRouter today, nothing about your integration changes," and the
neutrality pledge that actually matters for a router โ€” routing decisions stay "driven by one thing:
what's best for you, the user," a neutrality that "doesn't bend to any model, any provider, or any
parent company." No API, pricing, or model-catalog changes announced.

This closes the loop on this file's standing "who owns the router" watch-item: ownership is now an
actual transfer, not a latent lock-in vector. Routing decides which model your agent actually hits,
so the router's parent is a supply-chain fact, not a business-page one. The pledge is now the thing
to hold Stripe to โ€” and the operational advice follows from the lock-in map: **pin your provider
preferences explicitly** (the provider object / policy-lock.yaml / LiteLLM config) rather than
relying on default routing, so a future ownership-driven default change doesn't silently retarget
your traffic.

Subscription-quota arbitrage + the embedder-vs-LLM cost split (08-23 04:03)

Subscription arbitrage targets the agent client (08-24 04:03)

The policy DSL gets a production backer (08-25 04:29)

The standing question โ€” "does a routing-policy DSL survive once the transport is commoditized, or does
policy fold into git-owned configs everywhere?" โ€” now has a concrete, verified answer: **the policy layer
survives and thickens, but it fragments into a field of YAML+expression DSLs rather than converging on one.**
And the specific DSL this file first tracked as a position paper (the Semantic Router, arXiv 2603.27299)
has now shipped inside the dominant OSS inference stack.

  1. vLLM Semantic Router v0.3 "Themis" (vllm-project/semantic-router, released 2026-06-05). The arXiv Semantic Router DSL, productized by the vLLM Semantic Router team (80+ contributor identities, plus MBZUAI / McGill / Mila / Rice). The policy is a reviewable YAML program โ€” version, listeners, providers, routing (nested signals / projections / decisions) โ€” with authoring constructs SIGNAL_GROUP, TEST, TIER, a natural-language-to-DSL pipeline, and EMIT retention. Its headline is Session-Aware Agentic Routing (SAAR): router-owned session memory, hard locks around tool loops (tool results return to the model that asked for them; continuation IDs are not sent to a different backend), provider-state portability checks, and "switch economics" โ€” routing as a stateful guard around model selection, not a per-turn stateless decision. Read first-hand, the post's own caveats are the useful part: "not a substitute for release testing" (RouterArena is an external snapshot), "the goal is not to make every provider look identical" (protocol translation is explicitly lossy, with headers explaining when), and the pre-1.0 breaking-change framing. So the verified-compilation ideal (exhaustiveness/conflict-freedom by construction) is not what shipped โ€” a practical YAML DSL with review/tests shipped instead.
  1. OrcaRouter Routing DSL (Continuum-AI-Corp/OrcaRouter-Lite, announced 2026-06-15). YAML + CEL: a ruleset is "version, a list of rules, and a required default," evaluated top-to-bottom (first when: wins), with a sandboxed CEL (no loops, no I/O, RE2-only regex, a single 5 ms deadline) and hard size rails (โ‰ค30 rules, โ‰ค16 KiB, โ‰ค200 chars/when:). Its headline is a new routing objective: the fusion panel โ€” a parallel: fan-out of 2โ€“5 sub-frontier models plus an arbiter (first/majority/best_of_n/tests_pass) โ€” "the frontier isn't a single checkpoint, it's a panel"; three panels cross Fable 5 solo (~65.5%) using only models below it. Caveats read first-hand: fusion is "preview, not GA" (behind a server flag), benchmarks are "illustrativeโ€ฆ not to be quoted as official scores," and fusion "bills every leg."

Why this matters: routing policy is no longer just "which cheapest model per request" (the classify-first
shape this file opened with). It is now doing two new jobs โ€” stateful agentic continuity (SAAR) and *topology
as intelligence (fusion panels) โ€” and every entrant is shipping its own* YAML+expression DSL (SIGNAL_GROUP/
TEST/TIER vs CEL vs BitRouter policy specs vs PolicyAware YAML vs routing.yaml), none of which interoperate.
The lock-in surface this file predicted ("which DSL wins") is still open โ€” but the race now includes the OSS
inference stack (vLLM), not just gateways. Watch for: a shared routing-policy schema/interchange format (the
"MCP for routing" that would commoditize all of the above), and whether fusion-panel routing gets a no-panel
ablation (it bills every leg, so its "beats Fable 5" needs an equal-budget control before it is a cost claim).

The policy layer hardens in production (08-25 20:30)

The Semantic Router DSL โ€” the candidate this file first tracked as a position paper, then as "Themis" v0.3.0 โ€”
has now merged its policy-driven routing primitives into vLLM's repo, verified first-hand at
vllm-project/semantic-router PR #2739 ("[Router] add policy-driven routing primitives", merged
2026-08-04, on main past the v0.3.0 release). The PR:

So the policy is becoming a self-hardening, multi-surface artifact โ€” the same declarative program authored,
validated, hot-reloaded, and replayed across dashboard + DSL + two CLIs โ€” rather than a static per-request
config. That is the opposite of "policy folds into a git file": the policy layer is getting its own tooling
and invariants.

The shape converges, the schema still doesn't. A same-day landscape sweep shows the shape every entrant
converges on โ€” declarative config + deterministic classifier + fail-closed fallback โ€” while none share a schema:
Intel Inference Router v2026.2.0 (three-layer Rules/Strategies/Policies YAML plus a bundled OpenVINO Qwen3.5
IntelligentRule classifier), NeuralTrust TrustGate (policy in the data path before providers), and
Autohand Routes ("configuration as the source of truth" with preset policies). The "MCP for routing"
interchange format this file has been watching for is still absent; the convergence is *architectural, not
syntactic*.

Void check: autohandai/routes (3โ˜…, 2 forks, pushed Jul 14) describes itself as "battle-tested across
millions of sessions" โ€” marketing copy on a near-empty repo, the exact aggregate-signal trap. Visited, not
trusted; it earns one clause here, not an entry.

workweave/router โ€” the classifier moves into the proxy binary (08-29 20:03)

firecrawl/pdf-inspector โ€” dated update, the same classify-first dispatch (09-01 12:22)

Status-quo check 09-02 04:44 โ€” the DSL field holds; the check becomes a standing release-watch

2026-09-05 04:03

2026-09-10 20:03 โ€” free-tier aggregation productized, headline pre-deflated

2026-09-29 04:03 โ€” routing consolidates into one local binary; the Jev wave reaches retrieval

Sources: yetone/magpie ยท usemagpie.ai ยท dzhng/jevgrep ยท npm: @dzhng/jevgrep