Jev, a specialized decision model from TypeSafe AI, makes a practical case for putting a fast decision model ahead of a larger AI model. In a test that sorted 100 inbound emails, Jev finished in 23.00 seconds for $0.00196. Claude Fable 5.1 took 218.05 seconds and cost $0.14811, while assigning the same category as Jev to 89 of the 100 messages. The result does not establish which model was right on every email. It does show why teams might separate quick classification from the more expensive work of composing an answer.

Three bars of different heights with a stopwatch beside the tallest and coins beside the shortest

▲ Differences in time and cost

That separation is the point of model routing: choosing which model should handle a request before sending it along. In Claude Code, Jev can make that initial choice while Claude remains available for tasks that need deeper reasoning or a written response. The routing decision itself has a cost, so the relevant question is whether it saves more than it adds across the full workflow.

A model built for bounded choices

TypeSafe AI designed Jev for three kinds of output: yes or no, one selection from a supplied list, or a rating on a scale with two to 10 steps. Each answer includes a confidence percentage. That makes Jev suited to questions such as which category an email belongs in or which model should receive a programming prompt—not to drafting the email or writing the program.

A workflow can use confidence to decide when to pass a question to Claude instead. One skill-routing configuration escalates decisions below 60% confidence. That cutoff is a workflow choice, not a guarantee that every decision above it is correct. Jev should not be assigned open-ended writing, summarization, mathematical calculations, calendar-date comparisons, or judgments about the quality of another model’s answer. It also should not autonomously execute financial transfers.

What the measured runs cost

The 100-email test used four categories: hot lead, warm lead, cold lead, and not a lead. Its measured times and charges were:

Model Time for 100 emails Total cost
Jev 23.00 seconds $0.00196
Claude Haiku 4.5 74.65 seconds $0.01959
Claude Fable 5.1 218.05 seconds $0.14811

Many blank message bubbles passing through a funnel into a few color-coded groups

▲ Email classification into categories

The 89 matching classifications between Jev and Claude Fable 5.1 measure agreement, not accuracy against a verified answer key. That distinction matters if a category will drive an important action. A team can check a sample against its own criteria and send uncertain or consequential cases to a stronger model or a person.

Other runs show where the same division of labor applies—and where it needs care:

  • Programming prompts: Jev routed 12 requests among Claude Haiku 4.5, Claude Sonnet 5, Claude Opus 5, and Claude Fable 5.1. Ten avoided the top-tier model. The routed set cost $0.506, compared with $1.380 if Claude Fable 5.1 had handled every request. The model-tier choices were judged optimal roughly 80% of the time; in two cases, Jev chose Haiku when Opus was preferable.
  • Routine checks: Across 150 yes-or-no decisions spanning five types of task, Jev’s median latency was 219 milliseconds per decision and the total charge was $0.0019. The tasks included spam flags, invoice checks, community-rule checks, refund eligibility, and customer-churn signals.
  • Comment sorting: Jev made 2,184 decisions on 312 YouTube comments, assessing seven questions per comment. The run took 72.5 seconds and cost $0.0106. A Claude Fable 5.1 cost of $3.48 for that work was a projection, not a measured charge. Jev identified 155 comments needing replies, 28 with content ideas, 97 showing buyer intent, and five as spam.

In the comment test, Jev’s categories agreed with the larger models 90% to 96% of the time, while spam determinations agreed 98% to 100% of the time. As with the emails, agreement is not a substitute for an independently verified answer key.

How to use the savings without overusing Jev

Jev’s listed OpenRouter price in September 2026 was $0.042 per million input tokens and $0.00 per million output tokens. A token is a unit used to measure model input or output. A typical routing decision was estimated at about $0.00002. Those prices make repeated small choices inexpensive, but the programming test shows that routing still has to choose a capable model when the task demands one.

Claude Code integration uses the jev-decisions package. The setup described for it involves obtaining an API key, checking credentials and latency with node scripts/selftest.mjs, connecting the hook through settings.json, and restarting the environment. The /jev on and /jev off commands control routing; /jev status reports classifications, latency, confidence, and cumulative cost. Those figures give teams a way to inspect what the router is doing rather than judging it solely by a lower bill.

Data handling sets another boundary. Standard accounts do not have the stated SOC 2, ISO, or GDPR compliance assurances, and zero data retention is limited to enterprise accounts. Sensitive customer names, credentials, and proprietary financial records should not go through a standard account. Jev’s useful scope also depends on the input it can read: in a separate file-search test, it found matches using descriptive text, folder names, and file paths, not image pixels. Poorly labeled images would therefore be a weak fit.

The practical takeaway

Jev appears most useful as a narrow first step: classify, select, or score; check the confidence; then send drafting and harder reasoning to a model suited to them. Start with a repeatable, low-risk decision set, measure the complete workflow cost, and review classifications against your own answer key. Keep a fallback for uncertain decisions and exclude sensitive data from standard accounts. The 100-email result shows the potential savings, while the routing mistakes and unverified ground truth show why oversight remains part of the design.