Unsloth turns Qwen3.5-0.8B into a fast decision engine
Unsloth has added decision-model fine-tuning to its platform, enabling developers to convert small language models into fast, structured classifiers that bypass slow text generation.

Unsloth has integrated decision-model fine-tuning into its desktop application and open-source training library. Instead of generating autoregressive text, the new system trains small large language models to output structured, calibrated probabilities for specific labels, yes-or-no questions, or ordered scales. By utilizing a small classification head based on Cloudflare's Clef design alongside low-rank adaptation (LoRA), the system bypasses the traditional decoding loop. This eliminates output-token latency and the risk of malformed JSON.
The training process is highly efficient, requiring minimal hardware and time. For instance, fine-tuning the Qwen3.5-0.8B model on a single Nvidia L4 GPU took just 10 minutes, consumed 4GB of peak VRAM, and boosted its aggregate accuracy across three benchmarks from 38.3% to 74.3%. On a separate 3,000-row test set, the same 0.8B model achieved 78% accuracy. Other supported models include Llama-3.2-3B, which reached 76.8% accuracy in 21 minutes using 8GB of VRAM, and Gemma-2-9B, which hit 79.5% accuracy in 49 minutes with 20GB of VRAM. The Laya-8B model, further fine-tuned from a pretrained decision model, reached 80.2% accuracy in 44 minutes using 18GB of VRAM. Additionally, Qwen3.5-4B achieved 78.4% accuracy in 28 minutes using 10GB of VRAM, while a shortened 10-minute run reached 76% accuracy.
To achieve these results, Unsloth used rank-64 LoRA adapters for a single epoch. The training mixture combined typed decisions with 12 datasets, including AG News, ARC, BANKING77, BoolQ, CLINC150, CommonsenseQA, MMLU, MNLI, prompt-injection examples, SNLI, SST-5, and WANLI. The starter configuration for developers employs 4-bit base weights, rank-16 adapters, and a learning rate of 2e-4.
For practitioners, this approach allows a single forward pass to evaluate multiple related queries simultaneously, such as identifying a support ticket's intent, department, and refund status. The system's built-in Decision API, compatible with Laya and Jev endpoints, returns structured categories and probability scores. Developers can also use a built-in calibrate method to measure expected calibration error, helping ensure that confidence scores align with actual accuracy before deploying models locally.
This is our own summary of reporting by AlphaSignal


