An easy-to-miss but worth-examining direction emerged in the open-source community: narrowing a general-purpose LLM assistant into a "one forward pass, one decision" classifier. Following TypeSafe's Jev, the independent project Jeff fine-tuned Qwen3.5 and Gemma 4 into three small decision models, released under Apache 2.0. A single 0.8B model makes a decision in about 22 ms, running on consumer GPUs and the Apple M4 Max.

What it actually does

The idea: you describe a situation in plain words and list candidate options; the model outputs a calibrated probability for each option from a single forward pass. No generated text, no parsing — the answer is "which one", not "why". The README is blunt about positioning: it is a classifier, not a planner. Reason in code, decide with the model.

The request format is fully compatible with TypeSafe's Jev, though the project explicitly states it has no affiliation with TypeSafe; the training code builds on the MIT-licensed open recipe AutoJev.

The numbers: small models beat the big one (partly)

On five public benchmarks totaling 4,599 questions, Jeff-Qwen3.5-2B scored 83.1 overall, matching Jev's published 83.0; the 0.8B reached 79.1. The breakdown is informative: on Financial PhraseBank all three Jeff variants scored above 96, beating Jev's 77.0; on RAGTruth, Jeff-2B's 88.9 tied AutoJev-27B for the best. But on reasoning-heavy BBH, JudgeBench and JevBench hard, the small models lag clearly — Jev scored 73.3 on JevBench hard while Jeff-2B got 53.3.

The pattern matches intuition: for classification and grounding, fine-tuned small models are enough; for multi-step reasoning, parameter count is a hard constraint.

All local hardware, no cloud GPUs

The whole project ran on one RTX PRO 6000 workstation GPU: the 0.8B trained in about 2 hours, the 2B in about 3.5. Synthetic training data was written by the open model Qwen3.8-Flash-Next on two DGX Sparks; a closed model was only used to spot-check sample quality, and no closed-model output is in the training data. Training details are open too: full-weight fine-tuning, one epoch, batch 256, cross-entropy over option letters, plus a fitted temperature for calibration.

Half an hour of fine-tuning: 31.7% to 95.8%

The most practical data point: when zero-shot is not enough, a short fine-tune on your own examples pays off enormously. The official voice-navigation fine-tune used about 11k in-app examples and half an hour on one GPU, moving held-out accuracy from 31.7% to 95.8%, at about 40 ms per decision. Another extreme case: fine-tuning the 0.8B on 600,000 Lichess positions labelled by Stockfish for 3.5 hours lifted puzzle-solving from 15.5% to 55.8% on 1,000 held-out puzzles, and one GPU keeps up with about 600 human blitz games at once. The repo also ships zero-shot game tests — Doom, Frogger, Pac-Man — where the 0.8B matched the hand-coded rule bot in Doom.

Limits, and the "so what"

The project is honest about boundaries: at most 26 options per question, English and text only, no multi-step reasoning at 0.8B–2B, and the 2B plays games worse than the 0.8B. Note that benchmark scores don't predict gameplay: the lower-scoring 0.8B is the sweet spot.

The industry takeaway: massive amounts of "routing, intent recognition, moderation, command extraction" steps inside agent frameworks don't need a hundred-billion-parameter model. One workstation GPU plus a few hours of fine-tuning yields a 22-ms, calibratable, Apache-2.0 commercial-friendly decision unit. As inference cost becomes the bottleneck of agent scaling, the division of labor — big models plan, small models decide — is turning from papers into a reproducible engineering path. Next time you write an if-else chain inside an agent, ask yourself: could this branch decision go to a 0.8B model?

(Project: firelex/jeff, weights under mstrasser on Hugging Face; code MIT, weights Apache 2.0)