Fitting a 27B reasoning model onto a single consumer GPU used to mean distilling a smaller model and losing capability. OrcaRouter offers another path: quantize the weights to an average of 3.21 bits. On September 28, the routing vendor released OrcaSAQ2 27B on Hugging Face under Apache-2.0 — its first open-weights model, built on Alibaba's Qwen3.8-27B.
The compression numbers
The original BF16 checkpoint is 54 GB; the quantized build is 12.3 GB — a 77.2% footprint reduction, roughly 4.4x smaller. On fidelity, WikiText-2 perplexity moves from 5.6468 to 5.6482, just +0.02%; Top-1 token agreement is 93.2%; mean KL divergence is 0.031. All figures come from the official model card's self-reported measurements, produced by running these exact OrcaSAQ2 weights against the BF16 reference through the same evaluation path.
Architecturally, OrcaSAQ2 keeps the base model's hybrid attention design: 48 Gated DeltaNet layers plus 16 full-attention layers out of 64, hidden size 5120, vocabulary 248,320, with 262K context, thinking mode, tool calling, and the MTP speculative-decoding head all intact. The cost: the vision tower is dropped, making this a text-only model.
What it means for deployment
A 12.3 GB checkpoint means a 16 GB GPU holds the weights with about 3.7 GB left for KV cache; the official practical starting point is around 32K interactive context, against an architectural ceiling of 262K. Serving runs on vLLM. Measured under a 15.7 GiB memory cap: 65.3 tok/s single-stream with MTP off, 90.1 tok/s with MTP on — a 38% throughput gain — plus 332 tok/s at 8 concurrent streams and 333 tok/s at 16. The card also notes MTP consumes extra compute and KV capacity, so batched workloads should benchmark both configurations before choosing.
Agentic benchmarks are likewise vendor-reported: SWE-bench Verified 70.0% and Terminal-Bench 2.1 58.4%. The model card explicitly frames these as public reference points rather than strict apples-to-apples comparisons — the 70.0 comes from a 12.06 GB 27B checkpoint, edging past Qwen3-Coder-480B-A35B's 69.6.
Why long-horizon agents are the point
The most informative part of the model card is not the compression ratio but its evaluation philosophy: perplexity asks how similar the next-token distribution is, while long-horizon evaluation asks whether the model can still finish the job after many decisions. Quantization error invisible in single-turn QA compounds across the plan → act → observe → recover loop — one wrong tool call changes the environment state, and every later step builds on polluted state.
That is the core risk of low-bit quantization in the agent era: short benchmarks hide degradation; long trajectories expose it. By publishing BF16 fidelity metrics alongside downstream numbers, OrcaSAQ2 at least gets the methodology right.
Two caveats
First, the quantization method itself is proprietary. SAQ (sensitivity-aware quantization) is a name; calibration strategy, precision allocation, and packing techniques are all undisclosed, so anyone wanting to reproduce or build on it receives a compressed artifact, not a method. Second, every fidelity and benchmark number is vendor-reported — the card itself says public scores use different agent stacks and should not be read as a strict model-only ranking, and llm-releases.com flags the entry as vendor-reported too. Apache-2.0 opens the weights, not the method.
So what
If quantization is truly near-lossless, the deployment economics of agents get rewritten: one 16 GB GPU running a 27B long-horizon agent versus renting a cloud frontier API becomes a very different cost equation. But remember the line — the point is not 3 bits; the point is what survives at 3 bits. Benchmark on your own long trajectories; that beats any leaderboard.
Full details and every number: the official model card at https://huggingface.co/orcarouter/OrcaSAQ-2-27B