2026-10-07 · 27 min · fine-tuning · lora · calibration · small-models · qwen
Why read this
Notabletop 60%Reads Unsloth's decision trainer end to end: the Clef head, the soft-label loss, the data mixture, and which run each headline number came from.
- Original analysis
- Runs on a consumer GPU
- Widely used
LLM architectureMixed licencesPractitioner tool
How this was scored
- Is it new?
- 1 of 3: An incremental tweak
- Can I trust it?
- 2 of 3: Measures key facts from files, code or configs
- Can I run it?
- 2 of 3: Open code or weights with real limits
- Will I understand it?
- 2 of 3: Mechanism from first principles with figures
- Can I act on it?
- 2 of 3: A concrete recipe, numbers or comparison
- Will it last?
- 1 of 3: Relevant for months
- Does it affect many?
- 2 of 3: A widely used model, tool or lab release
- Only here?
- 2 of 3: A teardown or measurement few others did
Score 62 of 100, ranked 214 of 454 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
This site has written about decision models more than twenty times, and every one of those pieces looked at them from the outside. Jev's receipts, what RLCD might be, how llama.cpp serves the readout, the open reproductions, Julia-1's mask tokens. Inference, serving, benchmarks. The training side was always the part I had to reason about, because nobody who trained one published the trainer in a form I could read line by line.
Unsloth just did. Its post says you can now "train your own Decision model like Jev locally", and the code sits in the main unslothai/unsloth repository: FastDecisionModel, DecisionTrainer, a dataset builder, and a vendored copy of Cloudflare's head. So I read all of it, at commit 08f87d0, along with the guide and the notebook.
The surprise came early. The "before" number in every row of Unsloth's table is not Qwen3.5 answering decision questions badly. It is a randomly initialised head. The notebook says so in one sentence, and once you see it the post's numbers rearrange themselves.
- license
- Apache-2.0 / AGPL-3.0
- branch
- main
- tests
- 4019 files
- source
- 131.4 MB
- commit date
- 2026-10-07
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-07 at 08f87d0 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history
Two sentences, two different measurements
The post makes two claims that do not agree: "We increased Qwen3.5 0.8B's aggregate accuracy from 20.7% to 74.3% across 3 decision benchmarks", and later "increasing downstream accuracy from 30–37% to 78%". The guide's table resolves both.

Take the 0.8B row: 36, 7 and 19 before; 73, 74 and 76 after. The plain mean of the first three is 20.67, and of the last three 74.33. So "aggregate accuracy" is an unweighted average of three benchmark scores. The guide says the test set was 3,000 rows: 2,000 from typed-decisions, 500 each from BANKING77 and CLINC150. Weighted by what was actually scored, the same cells give 28.3% before and 73.7% after. The direction is identical; the starting point moves eight points depending on how you average, and the unweighted version flatters the jump because BANKING77's 7% counts as much as 2,000 typed decisions.
The second sentence is lifted from the guide, which reads "boosting downstream accuracy from 30–37% (which is roughly chance level) to as high as 78%". The 30–37% brackets the typed-decisions starting points, 36% and 33%, and the 78% is the best single number on the card: the 2B's typed-decisions score, or equally the 0.8B's "Holdout acc". Those two 78s are different things. The holdout column is accuracy on the rows split_holdout set aside from the training mixture, at most 400 decisions (HOLDOUT_MAX = 400, unsloth/models/decision.py:37), drawn from the same distribution the model trained on. "As high as" is doing a lot of work in that sentence.
Then the starting points themselves. A typed-decisions "uniform" guess, the same probability on every option, scores 0.308 on the dataset's own card, so 33 to 36% is indeed about chance. But why should a fresh model be at chance at all? Qwen3.5-0.8B can read a customer email. The answer is in the notebook, right before the evaluation call:
Let's check accuracy before training. The head is new, so it's about the same as guessing (about 30% accuracy in our runs).
And in the script that produced the table, scripts/train_decision_from_lm.py:135, the baseline is computed straight after from_pretrained attaches a newly seeded head and before a single step:
# scripts/train_decision_from_lm.py:129-135
items, report = FastDecisionModel.build_dataset(rows, processor, model)
train, holdout = FastDecisionModel.split_holdout(items, args.seed)
eval_items = {}
for name, eval_rows in evals.items():
eval_items[name], eval_report = FastDecisionModel.build_dataset(eval_rows, processor, model)
base = {name: _record_metrics(model, processor, its) for name, its in eval_items.items()}So 20.7% is not "Qwen3.5 0.8B's aggregate accuracy". It is the accuracy of 27 million random weights sitting on top of Qwen. The jump to 74.3% measures how much a classifier learns from training, which is a fine thing to measure, but it does not say how much better the fine-tuned model is than the base model prompted well. Nothing in the repository measures that. The image's "Before" panel, a paragraph of chat text that routes the ticket to billing correctly, is the base model doing exactly that, and the number in the table does not describe it.
One oddity is worth a sentence. An untrained 0.8B head gets 19% on CLINC150, where a guess among 151 intents would get under 1%, while the 2B starts at 1% on both intent sets. I will come back to why a random Clef head is not quite random.
What FastDecisionModel builds
Point FastDecisionModel.from_pretrained at an ordinary language model and it checks whether the folder is a Laya or Clef checkpoint. If it is neither, it quietly picks the Clef path:
# unsloth/models/decision.py:1826-1830
if kwargs.get("decision_head") is None and _is_plain_lm(
model_name, subfolder, token, revision, local_files_only
):
kwargs["decision_head"] = "clef"load_lm_as_decision_model (unsloth/models/decision_from_lm.py:144) then loads the backbone through FastModel, optionally in 4-bit, and builds a head whose size depends only on the backbone's hidden width:
# unsloth/models/decision_from_lm.py:29-42
def default_head_config(hidden_size: int, width: Optional[int] = None) -> dict:
# Clef's head: width 1024, 2 routing + 4 decoder layers, 16 heads, ff 4096; small backbones get width 512.
if width is None:
width = 1024 if hidden_size >= 3072 else 512
width = int(width)
heads = max(1, width // 64)
return dict(
hidden_size = int(hidden_size),
width = width,
routing_layers = 2,
layers = 4,
heads = heads,
feedforward = 4 * width,
)Qwen3.5-0.8B has a hidden size of 1,024 (its config.json), so it gets the small head: 512 wide, 8 attention heads, a 2,048 feed-forward. The head's weights are seeded from random_state and moved to float32. That float32 matters later, when we count memory.
The prompt is the other half of the design. Every question about a state, and every option of every question, is written into one prompt. This is Cloudflare's encoder, vendored byte for byte into unsloth/_vendor/clef/joint_schema_model.py, with its SHA-256 pinned in clef_manifest.json and checked by a test:
# unsloth/_vendor/clef/joint_schema_model.py:110-138 (abridged)
schema_ids = _tokens(tokenizer, "\n\nSCHEMA FIELDS:\n")
for question_index, (question_id, question) in enumerate(record["questions"].items()):
schema_ids.extend(_tokens(tokenizer,
f"\nFIELD {question_index + 1}\nID: {question_id}\nTYPE: {question['type']}\nINSTRUCTION: "))
question_start = len(schema_ids)
schema_ids.extend(_tokens(tokenizer, render(instructions)))
question_end = len(schema_ids)
schema_ids.extend(_tokens(tokenizer, "\nALLOWED OPTIONS:\n"))
for option_index, (option_id, description) in enumerate(question_options(question)):
schema_ids.extend(_tokens(tokenizer, f"OPTION {option_index + 1}: "))
option_start = len(schema_ids)
semantics = {"option_id": option_id}
if description is not None:
semantics["description"] = description
schema_ids.extend(_tokens(tokenizer, render(semantics)))
option_spans.append((option_start, len(schema_ids)))The state goes first, after a fixed system prompt ("Decide every field jointly"), and the assistant turn closes with an empty <think> block and JOINT SCHEMA DECISIONS:. The model never generates anything after that colon. What the encoder keeps is a list of spans: where each instruction sits, and where each option's little JSON object sits. Everything downstream reads those spans.
I re-implemented the text branch of this encoder and ran it with the Qwen3.5-0.8B tokenizer on the first customer-service row of typed-decisions' train split. The prompt is 970 tokens. The state is 233 of them; the schema is 683. For a five-question record with twenty options, the questions cost nearly three times what the input does. That ratio is the price of deciding everything in one pass, and it is also why the widget's step 2 looks mostly orange and green.
- actionchoicegold answer_directly at 0.74 · What should the assistant do next with this conversation?
- categorychoicegold account at 0.97 · What is this customer conversation primarily about?
- churn_riskscoregold 3 at 0.94 · How likely is this customer to stop doing business with us because of this interaction?
- needs_humannoulgold true at 0.65 · This conversation requires a human agent rather than automated handling.
- urgencyscoregold 3 at 0.40 · How time-sensitive is this conversation?
Gold is not one label. It is the mean of three teacher samples, so needs_human is 0.65 true and urgency spreads 0.28 / 0.32 / 0.40 over its top three levels.
encode_record and the Qwen3.5-0.8B tokenizer, on typed-decisions train row tr_customer_service_000000; targets are that row's gold. The predictions in step 4 are a slider, not a model: soft cross-entropy bottoms out at the target's own entropy when the prediction equals the gold, while the hard-label loss keeps falling all the way to one-hot.What a Clef head is
The site first met this head as a 128-million-parameter cross-attention transformer bolted onto a 27B Qwen. Unsloth's version is the same class, and its own JointSchemaHead only subclasses it to replace the forward with a batched one that pools spans in chunks; the parameters and state dict are identical (unsloth/models/clef.py:468-470). So it is worth reading the reference once, slowly.
# unsloth/_vendor/clef/joint_schema_model.py:281-337 (abridged)
class JointSchemaHead(torch.nn.Module):
def __init__(self, hidden_size, width, routing_layers, layers, heads, feedforward, dropout=0.0):
self.hidden_norm = torch.nn.LayerNorm(hidden_size)
self.memory_projection = torch.nn.Linear(hidden_size, width, bias=False)
self.question_projection = torch.nn.Linear(hidden_size, width, bias=False)
self.option_question_projection = torch.nn.Linear(hidden_size, width, bias=False)
self.global_projection = torch.nn.Linear(hidden_size, width, bias=False)
self.option_context_projection = torch.nn.Linear(hidden_size, width, bias=False)
self.option_lexical_projection = torch.nn.Linear(hidden_size, width, bias=False)
self.type_embedding = torch.nn.Embedding(3, width)
self.evidence_layers = torch.nn.ModuleList([EvidenceRoutingLayer(...) for _ in range(routing_layers)])
self.layers = torch.nn.ModuleList([torch.nn.TransformerDecoderLayer(...) for _ in range(layers)])
self.residual_scorer = torch.nn.Sequential(
torch.nn.Linear(width * 4, width), torch.nn.GELU(), torch.nn.Dropout(dropout), torch.nn.Linear(width, 1))
self.prior_logit_scale = torch.nn.Parameter(torch.zeros(()))
self.joint_logit_scale = torch.nn.Parameter(torch.zeros(()))
self.residual_gate = torch.nn.Parameter(torch.zeros(()))Read it in the order data moves. The backbone's final hidden states, all 970 of them in our example, are layer-normed and projected down to the head's width; that is the memory. Each question becomes one vector, the mean hidden state over its instruction span. Each option becomes a query built from three things added together: the mean hidden state over its span, the mean of the output-embedding rows for its tokens, and its question's vector. Two routing layers let every option query cross-attend to the whole memory, so an option can go looking for evidence anywhere in the state.
Then the questions become "fields". A field is the question vector plus an attention-weighted summary of its own options, plus the very last token's state, plus a learned row for its type (noul, choice or score). Four standard decoder layers run over the fields: they self-attend, which is where one question's answer can inform another's, and cross-attend to the memory. This is the "joint" in joint schema head. It is also a different answer from the ones the AgentJev piece and what decision models cannot do laid out, where options either share a sequence or never meet. Here options share the backbone's causal sequence, and questions additionally talk to each other in the head.
The score for each option is where it gets interesting:
# unsloth/_vendor/clef/joint_schema_model.py:433-457 (abridged)
anchor = functional.normalize(question_vectors[len(record_logits)] + global_vector, dim=-1)
lexical_anchor = functional.normalize(lexical, dim=-1)
prior_scale = self.prior_logit_scale.clamp(max=math.log(100.0)).exp()
prior = prior_scale * torch.matmul(lexical_anchor, anchor)
options = self.option_norm(routed)
repeated_field = field.unsqueeze(0).expand_as(options)
cosine = functional.cosine_similarity(repeated_field, options, dim=-1)
features = torch.cat([repeated_field, options, repeated_field * options,
torch.abs(repeated_field - options)], dim=-1)
residual = self.residual_scorer(features).squeeze(-1)
joint_scale = self.joint_logit_scale.clamp(max=math.log(100.0)).exp()
joint = joint_scale * cosine + residual
record_logits.append(prior + torch.sigmoid(self.residual_gate) * joint)There are two terms. The joint term is the learned one: a cosine between the field and the routed option, plus a small MLP over the usual matching features (both vectors, their product, their absolute difference). The prior term has almost no learned parameters. It is the cosine between an option's averaged embedding rows and the prompt's own hidden state at the question and the final token. Qwen3.5 ties its input and output embeddings, so this is close to asking the language model, logit-lens style, how much the prompt's last state "points at" the option's words.
All three scalars start at zero. So at initialisation the prior's scale is e⁰ = 1, the gate is σ(0) = 0.5, and the head's score is the embedding cosine plus half of a random joint term. That is my reading of the 0.8B's 19% on CLINC150 before training: the prior is a zero-shot signal built into the architecture, strong enough to beat a 151-way guess, diluted by random noise. Why the 2B starts at 1% instead I cannot tell from the code; a different random draw against a wider hidden state is the obvious suspect, and nothing in the repository tests it.
If the "mean of the option's own tokens" choice sounds familiar, it is the same lesson the RLCD piece quoted from kotoba-lang/typed-decisions: a fresh marker token "does not learn", the option's own words do. Clef pools option spans, both through the backbone and straight from the embedding table, and so does every other head here that trains well.
How big is it? I wrote the parameter count out from the layer shapes. With Cloudflare's own configuration (hidden 5,120, width 1,024) the formula gives 128,056,324, and with clef-flash's (hidden 4,096) 121,762,820. Both match the safetensors headers this site counted in week three to the parameter, so the formula is right. At Unsloth's default for Qwen3.5-0.8B, hidden 1,024 and width 512, the head is 27,324,932 parameters. On the 2B, 30,472,708. For a sense of scale, the whole 0.8B checkpoint is 873,438,784 parameters, more than a quarter of which is the 248,320-row embedding table.
What it is trained on, and against what
The data format is the System One request body plus a gold answer, one row per state, which is why the guide can say "the same dataset format as Laya and Clef". Here is the row the widget uses, trimmed:
{
"state": "{\"account\": {\"tier\": \"enterprise\", \"tenure_months\": 60, ...}, \"thread\": [...]}",
"questions": {
"action": {"type": "choice",
"instructions": "What should the assistant do next with this conversation?",
"criteria": {"answer_directly": "The assistant can resolve this itself ...",
"escalate_to_human": "Hand off to a human agent ...", "...": "..."}},
"churn_risk": {"type": "score", "instructions": "How likely is this customer to stop ...",
"criteria": ["No sign of dissatisfaction.", "...", "Imminent: threatening to cancel, dispute or leave."]},
"needs_human": {"type": "noul", "instructions": "This conversation requires a human agent ..."}
},
"gold": {
"action": {"label": "answer_directly",
"probabilities": {"answer_directly": 0.743333, "escalate_to_human": 0.2, "...": 0.0}},
"needs_human": {"label": "true", "noul": 0.65, "probabilities": {"false": 0.35, "true": 0.65}}
}
}The gold is soft. typed-decisions' card explains why: each case was labelled three times by "a teacher of roughly 4B-class capability" at temperature 0.7, and the gold is the mean. When build_dataset meets a probabilities dict it keeps the whole distribution as the target (_target_for, decision.py:373-398), so a 0.65 "true" stays 0.65.
The loss is short enough to quote whole:
# unsloth/models/decision.py:401-403
def _soft_cross_entropy(logits, target, mask):
logits = logits.float().masked_fill(~mask, -1e4)
return -(target * torch.log_softmax(logits, -1)).sum(-1).mean()Cross-entropy against the teacher's distribution, padded options masked out. Step 4 of the widget shows what that rewards: the loss bottoms out at the target's own entropy when the prediction equals the gold, 0.754 nats for action and 1.106 for the genuinely ambiguous urgency, and pushing more mass onto the label makes it worse, not better. The model is being taught to reproduce the teacher's uncertainty, not to be confident.
The function directly below it is more revealing about intent than about behaviour:
# unsloth/models/decision.py:406-416 (abridged)
def _decision_loss(logits, target, mask, ordinal = None, label_smoothing = 0.0,
brier_weight = 0.0, ordinal_weight = 0.0):
# Cloudflare's Clef recipe: label-smoothed cross entropy, a Brier term for calibration, and
# partial credit on score questions (expected distance from the gold level).
if not (label_smoothing or brier_weight or ordinal_weight):
return _soft_cross_entropy(logits, target, mask)Unsloth implemented Cloudflare's full recipe, label smoothing plus a Brier term plus an ordinal penalty for score questions, and then shipped it switched off. DecisionTrainer reads its defaults from CLEF_RECIPE, and CLEF_RECIPE = {} (decision.py:42). Every term defaults to zero, as does permute_fields, Cloudflare's trick of shuffling question order per record. The CE-plus-Brier objective that two independent reproductions converged on in the RLCD piece is one argument away (brier_weight = 1.0), but it is not what produced the table.
Calibration is done afterwards, by temperature. FastDecisionModel.calibrate fits one head-wide temperature with L-BFGS on the held-out rows and then one per question type, and it reports accuracy and ECE with each half of the rows scored at temperatures fitted on the other half. The comment above it is the most candid line in the repository:
# unsloth/models/decision.py:1668-1670
# Calibrated against being right (the gold label), not the soft gold distribution: Clef's
# confidence is read as the chance the answer is correct, and soft gold targets left a tuned
# model underconfident (confidence 0.61 at accuracy 0.78 on typed-decisions).That is the tension the widget's three loss boxes show. Training on soft targets teaches the teacher's spread, the spread is wider than the model's real error rate, and a model that is right 78% of the time says 0.61. So the trainer learns the spread and the calibrator sharpens it back toward "how often am I right". The fitted temperature is folded into the head when you save. It is a sensible, cheap fix, and it means "calibrated probabilities" here means a post-hoc temperature per question type, not anything learned.
The typed-decisions rows are only part of the run behind the table. The script mixes in 12,000 rows from public datasets, split evenly across sources:
# scripts/train_decision_from_lm.py:26-40
DEFAULT_SOURCES = (
"banking77", "clinc150", "mnli", "snli", "wanli", "boolq", "ag_news",
"sst5", "mmlu", "commonsense_qa", "arc", "prompt_injections", "xlam",
)That is thirteen sources; the guide lists twelve, without xlam. Salesforce's function-calling set is gated on the Hub, and the script skips any source it cannot load, so a run without accepted terms ends up with exactly the guide's twelve. I think that is what happened; the guide does not say.
Each converted row goes through augment_row (decision_datasets.py:568): option ids renamed to letters or random codes 30% of the time, instructions paraphrased or dropped, an extra yes/no question derived from a choice half the time ("Is the answer to … this: …?"), fields shuffled, and choice questions with more than 24 options cut to the gold option plus 23 random others. Then it is decontaminated against the evaluation states by 13-word shingles (Decontaminator, decision_datasets.py:645).
Now put those facts next to the benchmark names. BANKING77 and CLINC150 are in DEFAULT_SOURCES. The model trains on their train splits and is scored on their test splits. Decontamination keeps test sentences out of training, and that is all it does. So the 74% on BANKING77 is supervised intent classification on a model that has seen roughly a thousand labelled BANKING77 messages (12,000 rows over twelve sources), never more than 24 intents at a time, and is then tested on all 77 at once. That last mismatch is my best guess for why the bigger 2B lands at 58% there while beating the 0.8B on typed-decisions; it is a guess, and no ablation in the repository separates it.
Three recipes behind one post
The post says "LoRA (r=64) for one epoch". The guide's code, a few screens down, uses r = 16 and num_train_epochs = 2, and its settings section says "We used 2 epochs on typed-decisions". One of the replies asked which recipe reproduces the table. They are not in conflict once you see there are three:
| Table in the post | Guide code and notebook | Studio, "Decision model" on a plain LLM | |
|---|---|---|---|
| Source | scripts/train_decision_from_lm.py | the guide, Qwen3_5_(4B)-Decision | llm_decision_defaults.yaml |
| Model | Qwen3.5-0.8B and 2B | Qwen3.5-4B, 4-bit | any |
| Data | 12 sources + typed-decisions | typed-decisions only | your upload |
| LoRA | r=64, alpha 64 | r=16, alpha 16 | r=64, alpha 64 |
| Epochs, LR | 1, 1e-4 (head 3e-4) | 2, 2e-4 (head 1e-4) | 1, 1e-4 |
| Batch | 8 × 2 accumulation | 8 × 4 accumulation | 8 × 2 accumulation |
| Max length | 4,096 | 2,048 | 4,096 |
The library's own default sits with the table: FastDecisionModel.get_peft_model defaults to r = 64, lora_alpha = 64 (decision.py:1910-1914). The training call that produced the table is this one:
# scripts/train_decision_from_lm.py:178-182, 137-165 (abridged)
model = FastDecisionModel.get_peft_model(
model, r = args.lora_r, lora_alpha = args.lora_r, random_state = args.seed
)
t = DecisionTrainer(
model = model, tokenizer = processor, train_dataset = train,
head_learning_rate = args.head_lr, # 3e-4: the fresh head learns 3x faster
args = TrainingArguments(
per_device_train_batch_size = 8, gradient_accumulation_steps = 2,
learning_rate = 1e-4, lr_scheduler_type = "cosine",
num_train_epochs = 1, weight_decay = 0.01, bf16 = True, seed = 3407,
),
)
out = t.train()
calibration = FastDecisionModel.calibrate(model, processor, holdout) if holdout else {}DecisionTrainer.create_optimizer (decision.py:1534) splits parameters into backbone and head groups and gives the head its own learning rate. The LoRA adapters go on every attention, gated-delta-net and MLP projection of the language layers; the vision tower and the output embeddings the head reads stay frozen. On the 0.8B, r = 64 over those modules is 40,894,464 trainable parameters, by my count from the weight shapes. Add the head and about 68 million parameters learn, under 8% of the checkpoint.
Why is one epoch at r = 64 enough? Unsloth publishes no ablation, so this is my reading rather than theirs. The task is narrow: given text and a schema, pick among options. Most of what the model needs is already in the backbone, and the prior term proves the head can read it on day zero. The training signal is a teacher's soft labels with a ceiling: typed-decisions' card measures teacher self-agreement at 0.735, so past about 73% you are fitting the teacher's noise. One epoch over roughly 13,000 records is enough to reach that ceiling, and the notebook makes the same point from the other side: 60 steps at r = 16 on Qwen3.5-4B, about 1.7 epochs of typed-decisions alone, "reached 75% to 78% test accuracy in our runs". Rank does not look like the bottleneck. A rank of 16 gets the same number on the dataset that matters most, with a quarter of the adapter.

Where the 4 GB goes
The guide's second table puts the 0.8B at 4GB, the 2B at 8GB and Llama 3.2 3B at 4.1GB, with Gemma 4 E4B at 14.4GB. It does not say whether the 0.8B was loaded in 4-bit, what the batch was, or how memory was measured. The script records torch.cuda.max_memory_allocated() as peak_gb, but its --load-in-4bit flag defaults to off, and I cannot tell which way the table's run went. So I counted what I could.
- ■ frozen backbone, 4-bit linears0.97 GB
- ■ LoRA r=64: 40.9M params + grads + Adam0.65 GB
- ■ Clef head: 27.3M params + grads + Adam0.44 GB
- ■ layer checkpoints: in system RAM0.00 GB
- counted total2.07 GB
last_hidden_state and a handful of embedding rows for the option tokens, so this tensor never exists.JointSchemaHead at Unsloth's default width. A lower bound: it omits the activations of the layer being recomputed, the CUDA context and allocator slack, and assumes the vision tower loads in 16-bit. Unsloth does not say which of these settings produced its 4 GB figure.The counted part is small. At r = 64, a 1,024-token batch of 8 and a 4-bit backbone, the 0.8B needs about 0.97 GB of weights (the 4-bit linears are about a quarter of a gigabyte; the 16-bit embedding table and vision tower are most of it), 0.65 GB for the LoRA adapters with their gradients and Adam state, and 0.44 GB for the head, which trains in float32 and so costs sixteen bytes a parameter. That is 2.07 GB before any activations. Unsloth's gradient checkpointing offloads the per-layer checkpoints to system RAM (clef.py:372-374 has to work around it for the head), which keeps another 0.40 GB off the card. With a 16-bit backbone the count is 2.80 GB. Either way, 4 GB with the recompute working set and the CUDA context on top is believable.
The interesting line is the dashed one. A language-model fine-tune scores every position against the vocabulary, and Qwen3.5's vocabulary is 248,320 entries. For the same batch, float32 logits would be 8.1 GB; at the 4,096-token maximum, 32.5 GB. Unsloth's ordinary LM path already avoids materialising that tensor with chunked cross-entropy, so this is not a fair comparison with Unsloth's own SFT. It is the right comparison with the naive way of turning an LLM into a classifier, which is to train it to emit the answer token. The decision head reads last_hidden_state and a few dozen embedding rows for the option tokens, and the projection to 248,320 words simply never runs. The guide puts it in one sentence: "It never writes text, so it doesn't spend memory on the LLM's word predictions."
The thing that does grow is the prompt. The schema is written out in full for every record, so a record with many options is long whatever its state, and max_seq_length cuts the end of the state, never the questions. Batches are length-grouped (_LengthGroupedBatches, decision.py:1308), so one long record does not pad seven short ones.
Against Jev and the open reproductions
The fairest place to put these numbers is typed-decisions' own leaderboard, because that card is unusually strict about what is comparable.

Its "2,000" test rows are 2,000 decisions, five questions on each of 400 cases. Jev 1.13.0 scores 0.727 there zero-shot, never having seen these workflows. Unsloth's 0.8B scores 73% after training on the 1,200-case train split, so it goes in the card's other table, "Fitted or fine-tuned on train", next to Laya typed-decisions at 0.766, OpenDecider-nano at 0.796 and od1-typed-decisions at 0.797. Among models that trained on these four workflows, a 0.8B at 73% is mid-pack, ahead of the ModernBERT-base specialist (0.646) and behind the tuned 4B and 400M entries. The 2B's 78% sits above the 0.735 ceiling, where the card says plainly that a model "is learning the teacher's quirks". Julia-1 is the useful cross-check: its CPU run reproduced at 1,451 of 2,000, which is 72.6%, and Julia's provenance file names these four workflows among its training sets.
BANKING77 and CLINC150 have no common leaderboard to sit on. Cloudflare's card reports macro-F1 of 94.2 and 97.4 for the full 27B Clef on its own Decision Index suite, against 79.7 and 89.3 for Jev; those are a different metric on a different sample, so they bound the picture rather than rank anything. The more pointed comparison is the one the site made for Julia: Maxime Rivest's learning curve, where a 17M encoder fine-tuned on 9,493 BANKING77 labels reaches 91.5%. Unsloth's 0.8B gets 74% with about a tenth of that data, never shown more than 24 of the 77 intents in one example. One reply under the post put it better than I can: "We're going back to 2021 and teaching people how to train classifiers". That is accurate, and it is not an insult. A classifier that reads a typed schema at request time is useful. It just should be measured as one.
Unsloth's own screenshots are honest about what a short run gives you.


Neither image is the 93% of the announcement card. Both are what a few minutes of training looks like, and I would rather the guide showed them than not.
The first reply under the post is the most useful thing in the thread. Leo Borcherding fine-tuned Laya, the encoder-based decision model that Unsloth also trains (through the other branch of FastDecisionModel, not the Clef head), to decide whether a message needs a calculator, on a 4 GB AMD RX 6500 XT: 40 steps at r = 64, 25 seconds, "50% → 93% on decisions it never saw".

His chart has the lesson the headline does not. Ten categories go to 100%. CSV and script requests, which the base model got right every time, fall to 59%, and weather and news stay at zero before and after. Forty steps taught the new boundary and moved an old one. Unsloth ships the knob for exactly this, an opt-in KL penalty against the starting model with adapters disabled (kl_weight, decision.py:1344), and it defaults to zero.
What I would use it for
I came in wanting to know how such a model is trained, and the answer is less exotic than the category's marketing. You take any LLM, write the input and the whole question schema into one prompt, pool hidden states over the instruction and option spans, run a 27-million-parameter cross-attention head over them, and train adapters and head together with cross-entropy against soft labels. Then you fit a temperature. Nothing in that recipe is Jev's secret, because none of it needs to be.
What Unsloth adds is the plumbing that makes it a one-liner: the dataset builder that turns twelve public datasets into typed questions with augmentation and decontamination, the batched head that does not hold a float32 copy of every hidden state, offloaded checkpoints, the Decision API endpoint that serves the result as jev-latest. That is real work, and it is the part I would use.
Three things I would do differently from the defaults. Turn on brier_weight, since a decision model's selling point is its probabilities and the trainer already has the term. Report the base model prompted for the same task, not a random head, as the "before". And run a held-out evaluation on something you did not train on, because every number on the card is either in-distribution or has its train split in the mixture.
A licensing note, since it is easy to miss: Unsloth's package is Apache-2.0, but the decision-model modules (decision.py, clef.py, decision_from_lm.py, decision_datasets.py) carry SPDX-License-Identifier: AGPL-3.0-only headers. The vendored Clef head stays Apache-2.0, following Cloudflare. If you embed the trainer in a product rather than just running it, read that header first.
- architecture
- Qwen3_5ForConditionalGeneration
- task
- image-text-to-text
- library
- transformers
- license
- apache-2.0
- safetensors
- 13 shards
- largest file
- 4.99 GB
- files
- 25
- downloads
- 18
- likes
- 458
repo last modified 2026-10-01
How I checked
I shallow-cloned unslothai/unsloth at 08f87d0 and unslothai/notebooks at 59d5a1b and read the decision path: unsloth/models/decision.py, clef.py, decision_from_lm.py, decision_datasets.py, the vendored _vendor/clef/joint_schema_model.py and its manifest, scripts/train_decision_from_lm.py, and Studio's decision defaults. I ran none of it and trained nothing. The guide was read from its Markdown export on 2026-10-07, and the post's thread and replies through the fxtwitter mirror.
Model dimensions come from config.json and from the safetensors headers of unsloth/Qwen3.5-0.8B and unsloth/Qwen3.5-2B, read by HTTP range request. Head parameter counts are my formula over the head's layer shapes, checked against Cloudflare's two published heads. The LoRA count sums in and out features over the modules in decision.py's fallback target regex; Unsloth's fast path may also include two tiny gate projections per linear-attention layer, which would add about 2.4M at r = 64. The token counts in the first widget come from my own re-implementation of encode_record's text branch with the Qwen3.5 tokenizer. The memory widget is arithmetic, not a measurement, and leaves out working memory. typed-decisions figures are from its dataset card; the Clef, Julia-1 and BANKING77 learning-curve numbers are from this site's earlier articles, linked where they appear. What I could not check: which settings produced the 4 GB figure, whether the table's run skipped xlam, and why the 2B starts lower and ends lower than the 0.8B on the intent sets.