jev-router is a new open source proxy that sits between Claude Code and Anthropic, using TypeSafe AI's Jev model to decide which Claude model (Haiku, Sonnet, or Opus, plus an optional Fable tier) should handle each message. Jev is described as a 'System One' model that answers typed yes/no, choice, or scale questions in 70-500ms for a fraction of a cent, rather than generating text like an LLM. The router asks Jev three calibrated questions per new message, applies confidence thresholds per tier, only ever raises a session's model tier (never lowers it, to preserve Anthropic's prompt cache), and exposes a local dashboard showing routing decisions, costs, and savings. It requires Node.js 22+, a Jev API key, and installs as a background service via npx.
Table of contents
What Jev isWhy Claude Code needs a routerHow jev-router worksWatch every decisionTry jev-routerQuestions this post answers
What is Jev from TypeSafe AI and how is it different from an LLM like Claude or GPT?
Jev is a 'System One' model that makes fast, cheap decisions rather than generating text. It answers one of three question types per call, a Noul (true/false probability), a Choice (probability across provided options), or a Score (placement on a scale), in 70-500 milliseconds, and costs $0.042 per million input tokens with free output tokens. It never answers outside the options you give it, unlike an LLM producing free-form text. Developers weighing when a fast classifier beats a full LLM call can track model routing patterns like this on daily.dev.
How does jev-router decide whether to send a Claude Code message to Haiku, Sonnet, or Opus?
For each new message, jev-router asks Jev a Choice between mechanical (Haiku 4.5), routine (Sonnet 5), complex or deep (Opus 5.5, or Fable 5.1 for deep if enabled), plus two Noul checks for production risk and prompt injection attempts. Each tier has its own confidence bar (85% fast, 60% balanced, 30% frontier); if Jev's top pick misses its bar, the router escalates to a stronger model, and within a session the tier only ever moves up, never down. Teams optimizing Claude Code spend without losing prompt-cache benefits can follow routing approaches like this via daily.dev.
Why does switching Claude Code between models mid-session hurt quality and cost more?
Switching models breaks Anthropic's prompt cache since the cache belongs to one model, forcing the next request to start cold, and handing off a half-finished task to a stronger model doesn't fully recover lost quality. An AWS study of 500 SWE-bench Verified tasks found that handing a Haiku 4.5 transcript to Opus 4.7 recovered only 47% of the quality gap, costing $1.61 per task versus $0.72 for starting directly on Opus. Anyone deciding between model hand-offs versus starting strong can dig into cost-quality tradeoffs like this on daily.dev.
1 Comment
Share this post