How to train your own Jev-compatible decision model? Here's how Jeff did it

With Jeff you can run decision-making models compatible with TypeSafe’s Jev in your own code and fine-tune them with your own data in about half an hour on a single GPU. They are open weights models from 0.8B parameters, the smallest takes up 1.7 GB, and its author (firelex on GitHub) trained them entirely on local hardware. The code is MIT and the weights are Apache-2.0.

Before you start: Jeff only works with text in English. If your data is in Spanish or Portuguese, we’ll tell you what alternative to use below.

If you don’t know Jev yet, the idea is simple: you describe a situation, list the options in natural language, and the model returns a calibrated probability for each one in a single pass. It doesn’t generate text, so there’s nothing to parse. We explain how it works in Jev de TypeSafe: how AI makes decisions without generating text. Jev’s weights are not public. Jeff’s are, and so is the recipe they were trained with.

What is Jeff and what’s its relationship with Jev?

Jeff is a family of fine-tunes of Qwen3.5 and Gemma 4 trained for zero-shot classification and served behind an API compatible with Jev’s. It’s an independent project: the author clarifies that they are not affiliated with TypeSafe nor do they have their endorsement.

It was born as a fork of AutoJev, an open recipe by Denis Yarats that fine-tunes a 27B Qwen model to return decisions Jev-style. Jeff kept the core design of AutoJev:

  • a single model pass per decision
  • a trained reading of the response
  • a temperature tuned for calibration

From there, it brought the idea to small language models. There are three published on Hugging Face:

Model Parameters Weights (16 bits)
Jeff-Qwen3.5-0.8B 0.8B 1.7 GB
Jeff-Qwen3.5-2B 2B 4.2 GB
Jeff-Gemma4-E2B 2B effective (4.6B stored) 9.3 GB

Zero-shot means your categories don’t have to appear in the training data: support queues, user intents, moderation labels, voice commands, or even video game moves. You describe the options and Jeff chooses.

Is Jeff as good as Jev?

In classification, yes; in reasoning, it clearly falls behind, and the author themselves says so. The following figures are published by the project and compare Jeff with Jev’s published numbers, which were measured on a different sample from the same benchmarks.

Where Jeff is ahead:

  • On Financial PhraseBank, Jeff’s three models hover around 96%, versus the 77% published for Jev.
  • On RAGTruth they score between 86% and 89%, versus 77%.
  • On the overall score of five benchmarks, Jeff-2B reaches 83.1%, versus 83.0% for Jev.

Where it falls behind:

  • On BBH, Jeff scores between 64% and 68%, versus 94%.
  • On the hard level of JevBench, between 48% and 53%, versus 73%.

The author’s conclusion is that small models don’t reason. With 0.8B–2B parameters you get fast, calibrated choices among options you describe, not multi-step reasoning.

There are two external data points. A user who opened an issue in the repository reconstructed the author’s benchmark panel and reproduced their scores within a 0.6 point margin, so Jeff’s figures hold up. However, a Hacker News commenter who compared Jeff with Jev on their own classification tasks reported 70% versus 94%. It’s an isolated case, but it shows that zero-shot accuracy depends heavily on the task. That’s why the fine-tuning section below matters.

Jeff also leaves a curious result: bigger is not better. The 2B model beats the 0.8B on benchmarks, but performs worse on the author’s video game tests, which they describe as more risk-averse. To choose options quickly, the author recommends the 0.8B.

How fast is Jeff?

These are the median times per decision published by the author, measured on 200 questions of about 200 input tokens each:

Model RTX PRO 6000 Apple M4 Max (MLX) CPU (32 threads)
Jeff-Qwen3.5-0.8B 22 ms 28 ms 463 ms
Jeff-Qwen3.5-2B 24 ms 60 ms 708 ms
Jeff-Gemma4-E2B 29 ms — (MLX only runs Qwen) 1.0 s

For reference, the Doom runs published by Jev took between 114 and 212 ms per call, including network. The two measurements were not done on the same hardware. These are the author’s figures as of September 29, 2026.

How do you fine-tune Jeff with your own data?

The realistic path is to fine-tune a Jeff checkpoint with your own labeled examples, not retrain from scratch.

The author’s example is a voice navigation app. A fine-tuning with about 11,000 app-specific examples, in a single epoch, took around half an hour on a GPU. It boosted accuracy on held-out data from 31.7% to 95.8%, with about 40 ms per decision on an M4 Max. The README indicates it’s done with AutoJev’s autojev-train, using the --initial-checkpoint option to start from a Jeff checkpoint.

The complete recipe is also public and worth reading even if you don’t run it:

  • Training method: fine-tuning of all weights, one epoch, batches of 256, cross-entropy over option letters, and finally a single temperature adjusted for calibration. Checkpoints are chosen with a held-out development set, never with the benchmarks panel.
  • Data: 271,000 questions. Most are public datasets converted into decisions (inference, question-and-answer, sentiment, safety, fact-checking). Around 31,000 are synthetic questions written and reviewed by an open model, Qwen3.8-Flash-Next. No results from a closed model entered training; only one was used to review sample quality.
  • Leak filter: each training question is checked against the benchmarks panel and JevBench. This removed 50 near-duplicates.
  • Hardware: one RTX PRO 6000 (96 GB) for training, two DGX Spark with the master model, and a MacBook for testing. The 0.8B model trains in about two hours and the 2B model in about three and a half hours.

It’s local hardware, but high-end; someone pointed it out with irony on Hacker News. The recipe works without cloud GPU, but a full retraining requires a serious workstation. The part that most teams will actually use is fine-tuning with their own examples.

Two more details. The training data is not published: only the code, the weights, and the list of sources with their licenses in docs/data-sources.md, and some of those sources are share-alike (CC BY-SA). The author also acknowledges that at least half of each data family follows the format conventions of the benchmarks panel, which helps the score.

How do you install and use Jeff?

Jeff runs as a local server that speaks the /v1/systemone endpoint of Jev. According to the README:

uv sync
uv run hf download mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8b

# NVIDIA GPU or CPU (PyTorch)
JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve

# Apple silicon (MLX, much faster on a Mac; Qwen models only)
uv sync --extra mac
JEFF_BACKEND=mlx JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve

A request looks like this:

curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
  "model": "jeff-latest",
  "state": "Refund request: the customer says the parcel arrived crushed and wants their money back.",
  "questions": {
    "route": {"type": "choice", "instructions": "Which team should handle this?",
              "criteria": {"1": "Refunds and payments", "2": "Damaged or lost parcels", "3": "Account and login problems"}},
    "angry": {"type": "noul", "instructions": "Is the customer angry?"}
  }
}'

Each response includes a probability for each option, the selected option, and a confidence level. There are three types of questions:

  • choice picks one option from a list.
  • noul answers yes or no, as a probability.
  • score returns a point on a scale you describe.

You can send multiple independent questions in the same request.

How many options can Jeff handle?

In practice, 26. The README admits up to 255 options in a choice question and the server accepts them, but as of September 29, 2026, there is an open issue reporting that Qwen3.5 models never choose an option beyond number 26 (those encoded as AA or later).

The user’s test is clear. With 30 cities and Zurich at position 28, the 0.8B model chose Athens and gave Zurich a probability of 0.0004. With Zurich at position 3, it chose it with 0.998. On real intent datasets like BANKING77, CLINC150, and MASSIVE, the models got between 69% and 80% correct when the answer was among the first 26 options, and 0% when it was beyond. The user suspects that training never included questions with more than 26 options, though he presents it as a hypothesis. Jeff-Gemma4-E2B was not tested.

Until the author fixes it, if your list is longer, filter down to a short list of candidates first.

How to get more out of it?

The usage tips in the README work as a small style guide:

  • Reason in code and decide with Jeff. It’s a classifier, not a planner. Describe what each option leads to; if you ask it to predict the future, it won’t do better than random.
  • Wording matters a lot. In Frogger, giving the goal option the same words as the rest of the forward moves took an episode from 15 crossings to 23.
  • Use short keys and descriptive text. Write {"1": "Engagement letter"}, not long identifiers.

Does Jeff work for text classification in Spanish?

No. The model card indicates that Jeff only works with English text and that its calibration is tuned with the author’s development data. If your inputs are support tickets, emails, or voice commands in Spanish, Jeff is not the right tool today.

For text that isn’t in English, the closest local option is Ollaya. It runs open decision models behind the same Jev-compatible API, and its laya router sends text in other languages to a multilingual model.

Who does Jeff make sense for?

Jeff fits anyone who wants a decision step within their own code (route tickets, detect intents, moderate) that runs locally in tens of milliseconds, with a model they can retrain when zero-shot accuracy doesn’t cut it. It doesn’t replace Jev for judgments that require reasoning, it’s not multilingual, and today it won’t work for classifying among dozens of intents without filtering first.

What makes it worth an afternoon is that the entire pipeline is open: weights, training code, leak filter, and calibration step. You can take a 1.7 GB model, point it at your own labeled data, and measure the result yourself.