mrsaynothing.dev
What decision models can't do: six honest limits

> What decision models can't do: six honest limits▋

Two of the six new local decision models are non-commercial, none of them explain themselves, and the confidence number is a claim until you bench it. The wave, read against the grain.

mrsaynothing· 8 October 2026· 3 min read

The short version: use them for decisions, never for reasons. Two of the six are non-commercial. The printed limits — option counts, token ceilings, language gaps — are real and fail silently. And the confidence number is a claim until somebody benches it on real labels.

Who this is for: anyone about to pipe a fresh API into production because the release notes were exciting. This is the other half of the conversation — the first half lives here.

the request innocent-looking decision model stamps: 0.99 sure tray A the wrong one tray B where it belonged the stamp is not the proof
fig 1 — the confident tray: high probability sorts it, nothing checks it.

The six limits

  1. There are no reasons in the box. A decision model returns a label and a number. Why that label — the evidence, the counter-argument, the doubt — is not in the output, at any price. The moment your pipeline needs a why, you are back to a chat model or a human.
  2. Agreement is not verification. Ask three models and get three matching answers, and you have one opinion in triplicate. A unanimous wrong answer arrives in the confident tray — sorted, stamped, and wrong at 0.99. Consensus needs an independent reference to mean anything, and none of these models brings one.
  3. The license walls are real. Julia-1, Laya, lev and Kev-4B are Apache-2.0. Nimble and OpenJev are CC BY-NC 4.0 — non-commercial, full stop. The llama.cpp runtime is MIT, and that means nothing about the weights: the checkpoint card outranks the repo's license badge every time.
  4. The modality map has holes. Laya is English-only — don't route multilingual traffic through it. Nimble is text-only locally. OpenJev's vision input needs its own projector. Kev-4B ships without the reference-date preprocessing it expects. Each gap is printed; none of them is visible from the API's side.
  5. The limits are printed on the box. Julia-1 takes 2–20 options and an 8,192-token combined question — go past either and the failure is silent, not an error. Feeding a decision model a 12k-token document is not stretching it; it is using a different product than the one that shipped.
  6. Serving is not proof. A new endpoint returning numbers is a plumbing fact, not an accuracy fact. The published per-million-token prices and the serving support say nothing about whether the probability is honest on your labels. Only a bench on real traffic answers that — and no one has published one yet. That gap is the opportunity, and it's ours to take.

What holds despite all six

The scope gets smaller, not the value. One-pass routing, gating and triage — where the output is a decision and a human owns the consequences — is still the cleanest fit local AI has offered in a year. Small, fast, free of per-token bills, and honest about being a stamp. The failure mode to design for is the confident tray: set a threshold, escalate below it, and never let the stamp be the last word on anything that matters.

Remember this: a probability is a claim. Stamped ≠ sorted right. Bench on your labels before the tray is the truth.

The bench is next: Julia-1 against Laya on real routing labels, calibration per size, latency per hardware. If the confidence numbers hold, the wave earns it. If they don't — this post already said so.

$ Questions people ask

Can decision models explain their answers?

No. A decision model returns a label with a probability — the reasons are not in the output. The moment your pipeline needs a why, you are back to a chat model or a human.

Are all the local decision models open licensed?

No. Julia-1, Laya, lev and Kev-4B are Apache-2.0, but Nimble and OpenJev are CC BY-NC 4.0 — non-commercial. The llama.cpp runtime being MIT does not change what the model weights allow.

Does a high probability mean the answer is right?

Not until it's measured. The endpoint returning 0.99 is a claim, not evidence — published API prices and serving support say nothing about accuracy on your labels.

$ Share this post

$ Get the next argument by email

One email per post. Agree or tear it apart.

self-hosted · no third parties · one-click unsubscribe

One email, weekly digest. Zero pixels.

what is this?

← Back to blog