
> Jev decision models, explained▋
Tiny models that answer yes-or-no questions with a probability on every option — the Jev decision wave, read from the release record, not the bench.
mrsaynothing· 7 October 2026· 6 min read
curl -fsSL https://ollama.com/install.sh | shThe short version: when a pipeline needs a decision — which queue, allow or deny, which label — a Jev decision model answers in one pass with a probability attached. Six checkpoints shipped local this week: five permissive, one non-commercial. Use them for decisions, never for reasons — and bench the numbers before you trust them.
The honesty line: this wave is days old, so I've read the record, not run it. Everything here comes from the Ollama announcement, the llama.cpp 0.6.0 release notes and the published checkpoint cards. The bench with real latencies is the follow-up — those numbers are the ones nobody can scrape from us.
- decision model — a tiny AI that answers a question with a pick and a confidence number, not a sentence
- probability — the confidence number: 0.87 means "87 in 100 times I'd say the same"
- routing — sending each incoming request to the right queue or worker, like a postie choosing a van
- calibration — whether the confidence number tells the truth; a calibrated 0.9 is right 9 times in 10
- GGUF — the file format local models ship in; the name in every download list
1 · The two shapes of answer
Chat models answer everything with prose. That is the wrong shape for half the things pipelines need. A router wants which queue. A gate wants allow or deny, how sure. A classifier wants one of six labels.
Ask a chat model for a decision and you get a sentence that means the right thing — usually — in different words every time, plus a parsing step that fails quietly.
You ask one person to sort the post, and they write you an essay about each letter. The other hands you a sorted tray with a sticky note: "87% sure this is the bills pile." Which one do you trust at 6am?
2 · What the Jev wave actually is
"Jev" is the name of the API contract: a model that takes a question, a fixed list of options, and returns a probability for each. TypeSafe defined it; the local tools adopted it within a week of each other.
Ollama shipped local support on September 29 — v0.35 gained its own /v1/systemone, with Nimble and Tev1 pullable at launch. llama.cpp followed in 0.6.0 (October 5): a /v1/systemone endpoint serving five open GGUFs — Julia-1, Laya, Kev-4B, lev, OpenJev — with Nimble landing days later.
These models are tiny — 144M to 27B parameters, where chat models start at 30 times that. Small is the point: a decision is one pass, and a model small enough to feel instant changes which jobs can afford to ask.
One public accuracy comparison exists — Bespoke Labs' 13 datasets, 3,880 human-labeled decisions: Nimble 75.7%, Tev1-4B 73.3%, Tev1-0.8B 63.5%, with hosted Jev 1.13 at 76.0%. Vendor-adjacent numbers, so treat them as a starting point — the bench that matters runs on YOUR labels.
picture it: a chat model is a senior consultant you book by the hour. A decision model is the rubber stamp on the front desk — it only says which tray, and it never gets tired.
3 · The models on the shelf
The table below is the llama.cpp shelf — the six checkpoints its 0.6.0 endpoint serves (Ollama's library carries its own set: Nimble, Tev1, Laya, Clef). Filter by license — OpenJev is the one NC entry to check before a commercial pipeline touches it. The llama.cpp runtime being MIT does not change what the weights allow.
| Checkpoint | Size | License | Notes |
|---|---|---|---|
| Julia-1 | 144M | Apache-2.0 | 2–20 options, 8,192-token combined limit |
| Laya | 421M | Apache-2.0 | English routing; don't assume multilingual |
| lev | 4B | Apache-2.0 | Text decisions |
| Kev-4B | 4B | Apache-2.0 | Reference-date preprocessing not bundled |
| Nimble | 9B | Apache-2.0 | LoRA on Qwen3.5-9B; text-only locally |
| OpenJev | 27B | CC BY-NC 4.0 | Vision input via its projector; non-commercial |
A seventh shape, Clef, shipped in the same 0.6.0 release (Apache-2.0, text and vision), and Cloudflare's Clef line pulls from Ollama too. Same idea, more doors.
The size ladder reads simply: the 144M–421M pair runs on idle CPU and handles clean routing; the 4B pair earns its keep when labels get subtle; the 27B adds eyes. Run nvidia-smi --query-gpu=memory.total --format=csv,noheader — at these sizes, almost any card qualifies.
picture it: the size is the salary. The 144M does the job for pocket change; you pay the 27B only when the job needs to look at pictures.
4 · Where it fits — and where it doesn't
It fits when the output is a decision: request routing (ticket → billing / sales / abuse), gates (allow-or-deny with a confidence score), and document triage at volume — classify with the small model, spend big-model tokens only on what it flags.
It doesn't fit when you need reasons. A label with a probability has no reasons attached. The moment the pipeline needs "why", you're back to a chat model — or to reading the documents yourself.
It doesn't fit when you need truth by consensus. Agreement between models is not verification. A unanimous wrong answer arrives in the most confident voice of all — exactly the failure mode our ledgers exist to catch.
The quiet feature: the probability is the answer
A chat model that says "this is probably billing" hides its own calibration — you get one phrasing, once. A decision model hands you the number, and 0.51 and 0.99 are different decisions: below your threshold, escalate to a human instead of acting. Whether that number tells the truth on YOUR labels is not in the release notes. That is what the bench is for — and why it's the next post.
verify: ollama run qwen3:0.6b "Reply YES or NO: is 7 prime?" — a one-token answer with no essay attached means the shape works on your box.
next on this wave: llama.cpp vs Ollama · Ollama not using your GPU — and everything local-LLM lives in the local LLM hub.
faq
— mrsaynothing
$ Related entries
llama.cpp vs Ollama — which one should you run?
The record read, not run: Ollama pins its llama.cpp engine at b11351 while upstream ships b11443. One is an appliance, one is the engine room.
2026-10-06 · 5 min read

llama.cpp vs Ollama: Which Should You Run in 2026?
llama.cpp vs Ollama: Ollama wraps llama.cpp for convenience, raw llama.cpp wins on speed and control. Benchmarks, GPU offload flags, and when each wins.
2026-09-10 · 7 min read

How to Run GGUF Models Locally: Ollama, llama.cpp & vLLM
How to run GGUF models locally: one-line Ollama pulls, llama.cpp straight off a Hugging Face URL, and how to pick the right quant for your VRAM.
2026-09-07 · 7 min read
