TypeSafe's Jev, released September 15, 2026, is described as a 'System One' model that returns probabilities (0-1) for narrow yes/no, choice, or scored questions instead of generating text, priced at $42 per billion input tokens with free output tokens and response times of 70-500ms. The author built two apps: one that gates pull requests by asking six questions and converting probabilities into pass/review/block verdicts, and another that estimates AWS workload costs by having Jev classify workload characteristics while deterministic code handles pricing. Calling TypeSafe directly was roughly twice as fast as going through OpenRouter (318ms vs 683ms). The piece stresses that schema validity isn't correctness - Jev can return a confidently wrong but well-formed probability - so any consequential action needs a threshold and human or stronger-model fallback rather than trusting the number outright.
Table of contents
What a request actually looks likeThree question typesApp one: gating a pull requestApp two: AWS workloadSchema validityWhat other people are doing with itWould I use itLinksQuestions this post answers
What is Jev by TypeSafe and how is it different from a normal chat model?
Jev is a 'System One' model released by TypeSafe that returns probabilities instead of conversational text. You send it a state object plus a set of narrow questions (yes/no, multiple choice, or ordered rubric) and it responds with a number between 0 and 1 for each, in 70 to 500 milliseconds, rather than generating an explanation or plan. Developers weighing probability-output models against chat-based LLMs for classification can track real build experiences on daily.dev.
How much does Jev cost to use and how fast is it compared to routing through OpenRouter?
Jev costs $42 per billion input tokens with output tokens free, since it returns probabilities rather than generated text. Calling TypeSafe directly took about 318 milliseconds versus 683 milliseconds through OpenRouter for the same request, roughly half the latency, though OpenRouter access was easier before getting off TypeSafe's waitlist. Anyone comparing API providers for latency and cost can follow practical benchmarks like this on daily.dev.
Why isn't a guaranteed structured output format like Jev's enough to trust an AI classification result?
Schema validity is not correctness: a model constrained to answer only within defined questions and criteria will never return a malformed response, but it can still return a confidently wrong probability that validates perfectly. The practical fix is putting a threshold in front of any consequential action, routing low-confidence answers to a person or stronger model rather than acting on them directly. Teams building automated gating or review pipelines can find similar reliability lessons on daily.dev.
Share this post