I Benchmarked 11 Open-Source JEV Alternatives
After testing Jev against Tev, I ran 11 open-source JEV alternatives (imajev, Decider, JevK5, Kev, Laya and more) through the same 400 decisions. Here's how they compare on accuracy, cost and speed.

A few days ago I compared Jev with Tev, Together AI's $17 fine-tune. That left a bigger question open. Tev is only one of the open-source JEV alternatives people have put on Hugging Face since TypeSafe launched Jev, and I wanted to know how the rest do.
So I searched for "JEV alternatives". Every result was a list, and every list ranked the models by GitHub stars. None of them said how often a model picks the right answer.
I ended up building a benchmark for it. DecideBench runs 11 open decision models, Jev itself and 7 general-purpose LLMs through the same 400 decisions, and measures accuracy, cost per decision and latency for each. The results are also on a Hugging Face leaderboard.
The short version: one open-source JEV alternative, imajev-4b, gets within 3 points of Jev and costs less per decision. The most-starred one scored 59%.
How I tested them
The setup is the same one I used for JEV vs TEV. A decision model gets an input, a question and 3–6 options, and returns one option key. The 400 tasks come in pairs, and the two halves of each pair differ by one small edit that flips the right answer. It might be a negation, a return requested one day past the window, or an SQL statement pointed at the primary database instead of a replica. A model that matches on topic words gets one half of the pair wrong, so I report pair accuracy (both halves right) alongside plain accuracy.
The 400 items are the same ones. The difference is that this time every model that can take worked examples also saw one solved example per option. Those examples are why Tev scores 92.8% here, up from 90.0% without them. Jev scores about a point higher than in my last post even without examples. Three of the models (Laya, CLM and Julia-1) are encoders with nowhere to put examples, so they ran without them.
I self-hosted the open models on an NVIDIA DGX Spark and priced them by GPU time, at the median on-demand rate for an NVIDIA L4. For the hosted ones I used the tokens each API actually billed.
The results

| Model | Accuracy | Pair accuracy | Cost per 1M decisions | Latency p50 |
|---|---|---|---|---|
| Jev (closed, via AI Space) | 98.0% | 96.0% | $32 | 639 ms |
| imajev-4b | 95.0% | 90.5% | $23 | 401 ms |
| Tev, hosted on Together | 92.8% | 86.0% | $50 | 197 ms |
| Decider-4B | 89.0% | 79.0% | $26 | 468 ms |
| JevK5 v0.3 | 88.8% | 78.5% | $27 | 486 ms |
| Kev-4B | 79.0% | 66.5% | $22 | 356 ms |
| Kev-9B | 72.8% | 53.5% | $38 | 641 ms |
| Decider-2B | 63.5% | 41.5% | $11 | 204 ms |
| Laya typed-decisions 421M | 59.0% | 35.0% | $8 | 153 ms |
| CLM-v0.1-8B | 41.0% | 11.0% | $11 | 171 ms |
| Julia-1 144M | 35.0% | 9.0% | $4 | 64 ms |
| Random guessing | 22.2% | 4.5% | – | – |
Self-hosted latency was measured on the GPU box itself, with no network in between, so it isn't directly comparable to the hosted numbers. The DecideBench README has the full table, including self-hosted Tev and all 7 general LLMs.
What stood out
imajev-4b is the one to beat
imajev-4b scored 95.0% and got both halves right on 90.5% of pairs. That's ahead of Tev on accuracy and cost, and it's the only open decision model that got 90% or more in every one of the 8 task families.
It's still 3 points behind Jev. On 400 decisions that's 12 more wrong answers, and whether that's acceptable depends on what acts on the answer. One quirk I ran into: imajev caps the question field at 2,000 characters, so I had to move the worked examples into the input. If your prompts carry long policies, plan for that.
The most-starred model scored 59%
Laya has over 19,000 GitHub stars, far more than any other open-source JEV project. It's small and fast, at 421M parameters and 153 ms. On these tasks it scored 59%, with 35% of pairs fully right.
It's not a perfectly fair fight, because Laya ran without worked examples. But Jev without examples still scored 98.2%, and Tev 90.0%. The three encoder models share one failure: they often give both halves of a pair the same answer. They pick up the topic and miss the one clause that changes the decision. If your options are genuinely different topics, like billing vs shipping vs login, an encoder might be enough. Not if it hangs on a date.
Easy tasks hide the differences
Almost every model does well at routing and intent. Every model above 85% overall scored 96–100% on those two families. The models separate on approving an agent's action, checking a claim against evidence and ruling on a written policy, which are the decisions you'd least want an agent to get wrong.
Kev is the clearest example. Kev-4B scored 90–100% on sentiment, intent and triage, but 40% on claim checking and 48% on the returns policy. Kev-9B did worse than Kev-4B overall and cost almost twice as much to run. If you're trying a JEV alternative on your own data, start with your policy-heavy decisions, because that's where the differences show up.
Jev is still the cheapest way past 95%
Jev scored 98.0% for $32 per million decisions. The general LLMs that beat it cost 6–14× more. DeepSeek-V4-Flash reached 99.8% for $197, and GLM-5.3-Flash 99.2% for $192. imajev-4b saves about $9 per million decisions compared with Jev, which isn't much.
So price isn't the reason to pick an open source JEV alternative today. Control is: your data stays on your hardware, the model can't change under you, and you can fine-tune it.

Which JEV alternative I'd use
If I needed open weights, I'd start with imajev-4b. If I didn't want to run a GPU at all and needed a fast answer inside a user-facing request, I'd use Tev on Together, which answered in 197 ms at the median. For anything an agent acts on without a human checking, I'd still pay for Jev.
Whichever one you pick, the biggest lever is the same as in my last post. Write a clear one-line description for every option. In the JEV vs TEV runs, removing those descriptions cost Tev 6.8 points and Jev 4.5, while rewording the question moved both by less than a point.
What I didn't test
Some projects on the other lists aren't in here yet: SemIf (formerly OpenJev), NanoJev, jevlike, Nimble, Von and Rizzo Flow. The Jeff models and GLiNER2.5-Decide are already set up in the harness and are next.
A few caveats:
- Claude wrote the test items. I spot-checked the labels and had two other models check the worked examples, but a careful reader could still dispute a few. Jev and Tev miss the same 6 items, and 3 of those are arguably mislabelled.
- 400 decisions is small. Trust the overall numbers more than the per-family ones.
- I ran 4 requests at a time. A server batching more requests would bring the self-hosted costs down.
Try it yourself
The 400 items, the worked examples, every model's raw answers and the scoring code are on GitHub, and the dataset is on Hugging Face under CC BY 4.0. If you maintain a JEV alternative I haven't covered, open an issue and I'll run it.