← all articles (German)Deutsche Fassung

JEV hands-on: not an LLM killer, but a fast specialist

JEV sorts business emails about as accurately as GPT-6 Luna, but answers four times faster and costs about 47% less. A hands-on test with 450 emails, open code and numbers you can check.

4× faster than GPT-6 Luna
47% cheaper than GPT-6 Luna
65% all four decisions correct
97% same decisions in two runs

My take up front: JEV will not replace large language models. For classification tasks, meaning clearly defined decisions, it is a serious alternative, though. I expect fast, cheap specialist models to become the standard for this kind of work, whether they come from TypeSafe, from OpenAI or from the open-weight world. OpenAI has already announced a direct counterpart, the Decisions API.

Two weeks of hype

On September 15, 2026, the start-up TypeSafe AI introduced JEV: a model that does not write text. You ask it questions with fixed answer options, and it picks one.

Interest was high. According to Vercel, more than twice as many paying teams started using JEV through its AI Gateway, a service developers use to call models from different providers, in the first 24 hours than for any earlier model launch. TypeSafe promises answers 40 to 200 times faster than frontier LLMs, at a fraction of the cost.

I could not find independent measurements at launch. So I measured it myself, on a task everyone knows: which email do I need to handle right away, and which one can wait? Because JEV is built for classification, I did not test it against the big frontier models but against an obvious alternative for this job: small, fast language models.

What JEV does differently

Instead of generating text, JEV answers questions you define in advance. TypeSafe calls this a "System One model": fast and intuitive, like the fast thinking in Kahneman's terms.

You send a state, for example an email, plus typed questions. For instance: "How urgent is this? immediately / soon / later" or "Is this phishing? yes/no". You get the answers back with probabilities, all questions in a single call.

TypeSafe advertises that JEV does not hallucinate. What that means is that the answer always fits the given schema. A wrong choice is still possible. In my test there was not a single invalid answer, but the same was true for Luna and Haiku, which also had to follow a fixed answer format. So I measured separately whether JEV answers consistently and how often it is right.

The test: 450 emails, four decisions

Each model makes four decisions per email in one call: urgency (immediately / soon / later), category (six classes), reply needed (yes/no) and suspected phishing (yes/no).

  1. 300 real emails from my private Gmail inbox. Before they went to any API, I pseudonymized them locally: detected names of people, private addresses and phone numbers were replaced, links shortened to their domain. Organization names and sender domains were kept because they matter for classification. I checked the result on a sample.
  2. 150 synthetic business emails with a much broader spread of classes than my private inbox, including hidden deadlines, ads with fake urgency and 24 phishing attempts. They were written in turns by GPT-6 Sol, Claude Opus 5.5 and Gemini 3.8 Flash. For privacy reasons I did not use real work emails. Most of the emails are in German.
  3. The reference labels ("ground truth") come from the same three models by majority vote. I decided the four cases without a majority myself. So the test measures agreement with this reference, not with a human-checked answer key.
  4. The candidates: JEV 1.13.0 against two fast, cheap LLMs, GPT-6 Luna (without reasoning) and Claude Haiku 4.5. JEV also ran with questions phrased in German instead of English; the main results refer to the English questions. Every email ran twice, with each model's requests sent one after another so that latency stays measurable.

According to my invoices, the whole test cost about 10 euros, almost all of it for the large reference models. JEV itself cost just under 10 US cents for 1,811 requests, which matches TypeSafe's billing to the cent.

The results: similar accuracy, much faster

All numbers in this section refer to the 150 synthetic business emails. Costs are calculated from the reported tokens and list prices. JEV reaches a similar accuracy to GPT-6 Luna, but answers four times faster and costs about 47% less.

JEV, GPT-6 Luna and Claude Haiku 4.5 at a glance

Across both runs, JEV got all four decisions right for 65% of the emails, Luna for 64%. In the first run, that was 97 versus 96 out of 150 emails. Statistically, this difference cannot be told apart from chance (McNemar test, p = 1.0). That does not prove the two are equally good, only that the test finds no difference. Claude Haiku reached 48% and is clearly behind, statistically too.

In practice, it also matters which mistakes happen. Of the emails that needed immediate attention, JEV caught 87%, Luna 92% and Haiku 97%. In return, Haiku also flagged noticeably more emails as "immediately" by mistake. Of the 24 phishing attempts, JEV and Luna each missed four per run on average.

JEV's median latency was 257 milliseconds, Luna's just over a second (1,031 ms) and Haiku's slightly more (1,091 ms). Per 1,000 emails, JEV cost just under 5 US cents, Luna just over 9 cents and Haiku 2.25 dollars. For a million emails a month, that would be about 49, 92 and 2,250 dollars. With JEV you only pay for input; output tokens are free.

TypeSafe itself promises answers 40 to 200 times faster. For this, the vendor compares JEV's 70 to 500 milliseconds with the documented response times of frontier LLMs, 3 to 329 seconds. In the demo, GPT-5.6 Terra runs with standard reasoning. Against small, fast models, JEV was four times faster in my test. Against the three large reference models, the factor was 8 to 15: GPT-6 Sol with medium reasoning effort took 2.1 seconds at the median, Claude Opus 5.5 2.4 and Gemini 3.8 Flash 3.9 seconds, the latter two with default settings.

JEV also leads on consistency for the business emails. With identical input, it gave the same four decisions twice in 97% of cases, Luna in 93%, Haiku in 87%. On my private inbox, Luna (99.0%) and JEV (98.7%) were practically tied.

TypeSafe names English as JEV's primary language, where it is most accurate. Most emails in my test were German, and JEV still came close to Luna. With questions phrased in German, JEV reached 61.7% instead of 65.3%; this difference is not statistically significant (p = 0.58). Whether JEV would do better on English emails I cannot say: only 19 of the 150 emails were in English.

What the numbers don't tell you

On my real private inbox the picture is different: Claude Haiku led with 88.5%, Luna reached 81.5% and JEV 79.8%. However, 93% of these emails were newsletters, ads or notifications. For urgency, the simple rule "always later" was already right 96% of the time. Luna (97.5%) and Haiku (96.5%) were slightly above it, JEV (95.0%) slightly below. Only four emails needed immediate attention, and there was no phishing at all. So my private inbox says little about how well a model catches rare but important cases.

  • Luna and Haiku ran through Langdock with EU hosting, JEV directly at TypeSafe in the US. The latencies therefore include different network paths and infrastructure. I did not measure how much that contributes.
  • Haiku 4.5 is a year older than Luna. I also had to enforce the answer format for Claude via a tool call, which costs extra tokens.
  • The instructions to the LLMs mention "my private inbox", even for the business emails. And categories such as "personal" and "appointment" partly overlap. Both can only be fixed with a new run.
  • The reference labels are not perfect. On urgency, the three reference models fully agreed on only 91 of 150 emails. On top of that, the same models wrote the synthetic emails and labeled them. I could not find an advantage for a model's own family, though. Luna, for example, did best on the emails written by Opus, not on those written by Sol.
  • With 150 emails the estimates remain imprecise, and smaller differences can go unnoticed. The 95% confidence interval for JEV in the first run ranges from 57% to 72%, and possible errors in the reference labels are not even included.
  • The factor of 8 to 15 against the reference models is only a rough comparison. They ran with two parallel requests and not under the same conditions as the actual benchmark.
  • Whether JEV really reaches frontier level, as TypeSafe writes, my test cannot answer fairly. After all, the frontier models defined the reference themselves.
  • Pseudonymization is not anonymization. Content, organization names and domains can still allow conclusions. Anyone who wants to send sensitive emails to an API has to check that separately.

What this means for the big providers

JEV is no replacement for language models. It does not answer emails, summarize anything or write code. It chooses between options you defined beforehand.

For exactly this kind of task, however, the big providers are under pressure on price and speed. A week after JEV's launch, OpenAI released GPT-6 Sol and GPT-6 Luna at roughly half the price of their GPT-5.6 predecessors. Luna costs 0.10 dollars per million input tokens and 0.50 dollars per million output tokens. Whether the timing has anything to do with JEV, I don't know. Open-weight models such as DeepSeek or Xiaomi's MiMo are pushing prices down as well. In my test, Luna was close to JEV in accuracy and about 1.9 times as expensive.

The more direct response came on September 29: at DevDay, OpenAI introduced the Decisions API. It uses Luna for decisions between predefined answers, for example for classification, routing or choosing an agent's next step. At its core that is JEV's approach, now also offered by OpenAI. So far it is only available as a limited preview, and OpenAI has not announced a price. I have not tested it yet. A comparison with JEV would be the obvious next test.

Why speed matters so much here: a quarter of a second instead of a second sounds like little. But an agent that makes ten such decisions in a row would, at these latencies, wait 2.5 instead of 10 seconds for the model.

On the same day, OpenAI also introduced an Ultrafast mode, generally available for its top model GPT-6 Astra and in preview for GPT-5.6 Sol. According to OpenAI, it generates tokens up to six times faster in the API, but also costs six times as much: 60 dollars per million input tokens and 300 dollars per million output tokens. There is no EU processing for it. I did not measure Ultrafast. With Luna's token usage from my test, it would come to about 55 dollars per 1,000 emails; how many tokens Astra itself would need, I don't know. For tasks like mine, this is not an alternative.

There is an open alternative too: three days after JEV, Convai Innovations released Laya, a model with a similar decision interface that runs locally. For sensitive data such as work emails, that is a big advantage. On the developers' benchmark, a specially fine-tuned variant reached 76.6%, JEV 72.7%. The developers took the JEV figures from third-party measurements, though. Without adaptation, the base model reached only 36.2%, below the simple rule of always picking the most frequent answer (46.1%). So it does not work without your own training. I have not tested Laya myself yet.

Where JEV makes sense today

JEV is worth a test wherever many similar decisions have to be made quickly, for example with:

  • large volumes of tickets, emails, forms or content that need pre-sorting
  • tasks with fixed categories where the possible answers are known in advance
  • decisions someone is waiting for, such as routing in agents or moderation

Whether it is good enough for production depends on how many mistakes you can afford. Especially for phishing or moderation, speed alone is not an argument.

It gets interesting in combination. JEV returns a confidence value with every answer. In a simulation after the fact, using the data of the first run, I forwarded all emails below a confidence threshold to Luna. Accuracy rose from 65% to 72%, more than either model achieved on its own. A third of the emails were forwarded, and the cost was 0.08 dollars per 1,000.

Hybrid curve: share of forwarded emails versus accuracy

One caveat: I chose the threshold on the same data I measured on. The gain may therefore be overestimated. Whether it holds up on new emails would need a separate test.

Conclusion

For clearly defined decisions on the business emails, JEV reached a similar accuracy to GPT-6 Luna in my test, with four times lower latency, about 47% lower cost and more consistent answers. The 40 to 200 times speed advantage in TypeSafe's marketing applies to large models, not small ones. JEV does not replace large language models, because it does not produce free text. But the fact that OpenAI followed so quickly with the Decisions API shows how seriously the big providers take this kind of model.

Code, synthetic data, individual measurements and all metrics are openly available on GitHub. If you want to repeat the test with your own emails, you will also find the local pseudonymization there.

What do you think: where would you use a decision model like JEV, and where do you stick with an LLM?

This is a translation of the German original, including its update of October 2, 2026: OpenAI's pricing corrected (new models rather than a price cut), Decisions API added, figures on the private inbox, costs and question language corrected, methodology described in more detail.

Sources