"Which team should handle this ticket?" "Is this review positive or negative?" Most of what software asks an LLM is not an essay but a typed decision: options fixed, answer confined to a set, ideally with a confidence attached. Privatemode (Edgeless Systems) pushed this scenario to its limit last week: GLM-5.3-Flash returns a full decision in a single forward pass and a single token. On Hacker News the post has gathered 133 points and 58 comments.

Three moves, zero fine-tuning

The recipe has three steps. First, pack state, question, and numbered options into the prompt as JSON, and ask the model to answer with an option index. Second, end the prompt inside the answer: prefill choice_index:, so the model's next token must be an index. Third, instead of reading the token the model emits, read the probabilities it assigned to all option indexes at that position, normalize, and take the max.

The core insight: since we already know the shape of the JSON, having the LLM predict the whole object is waste — what we want is its typed judgement under constraint. No fine-tuning, no distillation; the model runs exactly as it ships. The comparison targets, Jev (TypeSafe) and Laya (Convai), are purpose-trained "System One" decision models. Privatemode's bet: a general LLM with the right sampling posture can reach the same tier without retraining.

Results across 29 datasets

The benchmark covers 29 public labeled datasets, 2 to 151 options, spanning intent routing, sentiment, topic classification, moderation, entailment, QA, legal text, and scanned documents, in English and German. Across the 28 text datasets, GLM-5.3-Flash and Jev each win 10; 8 are within one percentage point; the median gap is 0.7 points in Jev's favor, not statistically significant (p=0.64). The authors also report that even at temperature 0, identical runs drift on up to 3.5% of answers, caused by batching and floating-point arithmetic — smaller differences are treated as noise.

The number of options matters more than the choice of system. On TREC, going from 6 to 42 options, Jev drops from 92.1% to 85.6%, GLM-5.3-Flash from 91.2% to 79.6%, Laya from 88.4% to 51.2%.

Cost and latency are reported honestly. A million decisions cost about EUR 62 with GLM-5.3-Flash versus EUR 16 with Jev — the dedicated model is four times cheaper; measured from Germany, latency was 180 ms versus 264 ms. The limits are self-disclosed too: on CLINC150 with 151 options, GLM-5.3-Flash splits the query into two requests and reaches 87.5% accuracy (Jev 78.4%), but takes 719 ms (Jev 249 ms); the ceiling is 191 option indexes, because GLM-5.3-Flash's tokenizer splits larger numbers into multiple tokens. Images are the exclusive capability: on RVL-CDIP, 1,600 scanned business documents in 16 classes, it scores 70.2%, while Jev and Laya are text-only and cannot answer at all.

Label noise and the price of reasoning

Two findings deserve their own section. First, some of the remaining errors are in the labels: about 17% of banking77 examples admit two defensible answers (e.g. get_physical_card vs order_physical_card), so the theoretical ceiling there is about 85%, not 100%. Second, reasoning is more accurate but priced per token: letting the same GLM-5.3-Flash reason before answering lifts accuracy in every band, 89.9% vs 85.5% (two options) up to 82.0% vs 79.2% (21-80 options), at the cost of hundreds of tokens per decision and about EUR 350 per million — against EUR 62 for the single-token setup. The expensive part is not the model; it is the act of thinking itself.

So what

The real audience for this method is anyone locked out by dedicated decision models. Jev is a closed API with a fixed model; here is an MIT-licensed Python library that runs against any vLLM backend, fetches token IDs from the server, and switches models without code changes. Reproducibility is rare in its class: the benchmark methodology, frozen dataset specs, test harness, and aggregation scripts are all open, and every number in the post can be recomputed. Which specialized niche will "general model + inference engineering" flatten next? Drop your pick in the comments.