Models

Unsloth turns Qwen3.5-0.8B into a fast decision engine

Unsloth has added decision-model fine-tuning to its platform, enabling developers to convert small language models into fast, structured classifiers that bypass slow text generation.

AlphaSignal19 hrs agoModels
Image: AlphaSignal

Unsloth has integrated decision-model fine-tuning into its desktop application and open-source training library. Instead of generating autoregressive text, the new system trains small large language models to output structured, calibrated probabilities for specific labels, yes-or-no questions, or ordered scales. By utilizing a small classification head based on Cloudflare's Clef design alongside low-rank adaptation (LoRA), the system bypasses the traditional decoding loop. This eliminates output-token latency and the risk of malformed JSON.

The training process is highly efficient, requiring minimal hardware and time. For instance, fine-tuning the Qwen3.5-0.8B model on a single Nvidia L4 GPU took just 10 minutes, consumed 4GB of peak VRAM, and boosted its aggregate accuracy across three benchmarks from 38.3% to 74.3%. On a separate 3,000-row test set, the same 0.8B model achieved 78% accuracy. Other supported models include Llama-3.2-3B, which reached 76.8% accuracy in 21 minutes using 8GB of VRAM, and Gemma-2-9B, which hit 79.5% accuracy in 49 minutes with 20GB of VRAM. The Laya-8B model, further fine-tuned from a pretrained decision model, reached 80.2% accuracy in 44 minutes using 18GB of VRAM. Additionally, Qwen3.5-4B achieved 78.4% accuracy in 28 minutes using 10GB of VRAM, while a shortened 10-minute run reached 76% accuracy.

To achieve these results, Unsloth used rank-64 LoRA adapters for a single epoch. The training mixture combined typed decisions with 12 datasets, including AG News, ARC, BANKING77, BoolQ, CLINC150, CommonsenseQA, MMLU, MNLI, prompt-injection examples, SNLI, SST-5, and WANLI. The starter configuration for developers employs 4-bit base weights, rank-16 adapters, and a learning rate of 2e-4.

For practitioners, this approach allows a single forward pass to evaluate multiple related queries simultaneously, such as identifying a support ticket's intent, department, and refund status. The system's built-in Decision API, compatible with Laya and Jev endpoints, returns structured categories and probability scores. Developers can also use a built-in calibrate method to measure expected calibration error, helping ensure that confidence scores align with actual accuracy before deploying models locally.

This is our own summary of reporting by AlphaSignal

More in Models