// the find
jaredpalmer/kev
Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own
Kev is a family of small, self-hostable decision models (0.8B to 27B parameters, built on Qwen bases) that answer yes/no, multiple-choice, and rating questions about a document and return calibrated probabilities instead of just a label. It's aimed at teams currently calling a hosted model like Jev for ticket routing, triage, or similar structured classification, who want to run the same workload locally or on their own cloud account and fine-tune it on their own labels.
Calibration is treated as a first-class concern rather than a footnote: every checkpoint ships a fitted temperature, and the benchmark harness reports Brier score, calibration error, and the rate of confident wrong answers alongside raw accuracy, so you can check whether the probabilities mean anything before setting an automation threshold on them. The eval suites are frozen with checksummed manifests, and a CI script re-derives the README's own numbers from committed run reports, which is a more serious anti-benchmark-rot setup than most repos bother with. Question isolation is solved properly at the model level (position IDs and attention masking, or separate rows where the architecture's recurrent layers ignore masks) rather than by convention, which avoids a real failure mode in multi-question batching. The fine-tuning guidance is backed by concrete before/after numbers rather than asserted — starting from the released checkpoint instead of the base model took one user's result from 0.33 to 0.84 on their own eval set.
The headline comparison to the hosted Jev has a real asterisk: on the metric that decides whether you can trust the model unsupervised, automating at a 5% error budget, Kev-4B/9B/27B sit at 0.52-0.69 against Jev's 0.70, so 'drop-in replacement' undersells how much of the gap is in exactly the place that matters. Kev-27B, the size closest to Jev on accuracy, ships as 51GB of full fine-tuned weights rather than an adapter, so it needs an 80GB GPU — the model that's actually competitive is also the one that's expensive to self-host. Knowledge-heavy questions depend entirely on the base model and lag visibly (MMLU-Pro 0.675 vs Jev's 0.840), meaning this is really a tool for classification over text you hand it, not a general Jev substitute, and that boundary isn't obvious until you're well into the benchmarks section. Date arithmetic is a known gap patched with an environment flag that appends day-count hints to the input rather than being fixed in the model, which is the kind of workaround that quietly breaks if your criteria involve deadlines and nobody remembers to set it.