Qwen3.5-2B-DSpark

A DSpark draft model for speculative decoding with Qwen/Qwen3.5-2B as the verifier, trained with speculators. The drafter proposes 8 tokens at a time and the verifier checks them in one forward pass, so output is identical to running the verifier alone — a lossless speedup. At temperature 0, mean acceptance length is 3.89 tokens committed per verification step, up to 5.40 on math_reasoning. Compared with the verifier alone, throughput is 2.00× higher on average, up to 2.88× on math_reasoning and 2.62× on HumanEval, and it is faster than the verifier's native MTP drafter on eight of the nine benchmark subsets.

Training code: rasyosef/train-dspark-draft-models.

Trained on 100,000 samples.

Usage

vLLM

vLLM loads the verifier automatically from the config — don't pass it separately.

vllm serve rasyosef/Qwen3.5-2B-DSpark --port 8000 \
  --gpu-memory-utilization 0.8 \
  --override-generation-config '{"temperature": 0}'

Then query the OpenAI-compatible endpoint at http://localhost:8000/v1.

llama.cpp

llama-server -hf unsloth/Qwen3.5-2B-GGUF:BF16 \
  -hfd rasyosef/Qwen3.5-2B-DSpark-GGUF:BF16 \
  --spec-type draft-dspark --spec-draft-n-max 8 --spec-draft-n-min 0 \
  -fa on -ngl 99

The block size is read from the sidecar metadata (n-max is clamped to it). Speculative decoding is exact: the target verifies every proposed token, so greedy output equals the target alone; per-response timings report draft_n / draft_n_accepted.

Details

5 Qwen3 layers (hidden size 2048, intermediate size 6144, 16 attention heads over 8 KV heads, head dim 128, all layers sliding-window attention with a 2048-token window), ~0.5B params, bfloat16. Block size 8, draft vocabulary reduced to 50,000 from the verifier's 248,320 (selected by token frequency over the training data), aux hidden-state layers 1/6/11/16/21, confidence head with Markov (vanilla, rank 256).

Trained for 4 epochs at lr 4e-4 (AdamW, cosine schedule with 4% warmup) on 100,000 filtered Open PerfectBlend prompts with responses regenerated by the verifier itself, split 96/4 into train and validation, with a {"ce": 0.1, "tv": 0.9} loss. Samples prepared at 2048 tokens; training sequence length 8192, up to 1024 anchors per sample. Verifier hidden states were pulled on demand from a running vLLM server during training and deleted after use rather than staged to disk up front.

Evaluation

Drafter results are at temperature 0 (greedy decoding), from evaluate.py throughput across the nine RedHatAI/speculator_benchmarks subsets. acceptance_length is mean tokens committed per verification step, including the bonus token — floor 1.0, ceiling 9.0 at block size 8.

Acceptance

subset acceptance_length pos_0 pos_1 pos_2 pos_3 pos_4 pos_5 pos_6 pos_7
math_reasoning 5.405 87.4% 75.8% 64.9% 56.1% 48.5% 41.5% 35.9% 30.3%
HumanEval 4.766 84.1% 69.6% 57.5% 47.6% 39.3% 32.4% 25.8% 20.3%
qa 4.074 73.2% 54.1% 42.0% 35.5% 30.2% 27.3% 24.2% 20.9%
rag 3.986 80.1% 61.9% 47.6% 37.4% 29.1% 20.3% 13.0% 9.2%
tool_call 3.619 73.5% 53.8% 40.8% 30.0% 23.0% 17.4% 13.3% 10.0%
question 3.299 69.4% 46.4% 32.1% 24.4% 19.7% 15.5% 12.4% 10.0%
writing 3.107 69.3% 44.9% 30.0% 21.4% 16.4% 12.3% 9.4% 7.0%
summarization 2.548 65.9% 38.7% 22.3% 13.3% 7.8% 3.8% 2.1% 0.9%
translation 2.527 68.2% 41.0% 23.5% 11.0% 4.7% 2.3% 1.3% 0.7%

Weighted across all subsets: 3.892 over 120,525 verification steps. math_reasoning and HumanEval are highest, and both still accept more than one draft in five at pos_7, as does qa. summarization and translation are lowest and fall below 1% by pos_7.

Throughput (tokens/s)

Output throughput with a single concurrent request on one A100. Baseline is Qwen3.5-2B with no speculative decoding. MTP is the verifier's native MTP drafter, run at 4 and 8 draft tokens. The fastest configuration in each row is bolded. The last column gives DSpark's throughput divided by baseline throughput for each subset.

subset baseline (no drafter) MTP (4 draft tokens) MTP (8 draft tokens) DSpark (this model) DSpark speedup
math_reasoning 233.4 410.1 357.0 671.8 2.88×
HumanEval 233.5 412.8 376.4 612.2 2.62×
qa 228.7 386.7 311.4 472.9 2.07×
rag 221.5 401.2 373.3 499.5 2.26×
tool_call 224.5 360.5 376.2 560.4 2.50×
question 231.6 345.4 229.3 359.7 1.55×
writing 233.0 345.1 229.4 365.3 1.57×
summarization 219.5 293.8 220.5 316.6 1.44×
translation 168.4 196.6 152.9 188.1 1.12×
average speedup 1.00× 1.57× 1.31× 2.00× 2.00×

DSpark has the highest average speedup at 2.00× (unweighted mean of per-subset speedups), ahead of MTP at 4 draft tokens (1.57×) and at 8 (1.31×), and it is the fastest configuration on eight of the nine subsets. The gains are largest on math_reasoning (2.88×), HumanEval (2.62×) and tool_call (2.50×), and smallest on summarization (1.44×) and translation (1.12×). Translation is the one subset where MTP at 4 draft tokens is ahead (196.6 against 188.1); it is also the smallest subset, with fewer than 900 verification steps in each run.

MTP at 8 draft tokens is slower than MTP at 4 on every subset except tool_call, and it only matches or falls below the baseline on question, writing, summarization and translation.

Acceptance comparison with MTP (4 and 8 draft tokens)

Acceptance length for DSpark and for the verifier's native MTP drafter on the same subsets at temperature 0. The highest value in each row is bolded and named in the last column. At 4 draft tokens the ceiling on acceptance_length is 5.0 rather than 9.0.

subset MTP (4 draft tokens) MTP (8 draft tokens) DSpark (this model) best
math_reasoning 3.832 4.743 5.405 DSpark
HumanEval 3.785 4.862 4.766 MTP (8 draft tokens)
qa 3.408 4.227 4.074 MTP (8 draft tokens)
rag 3.812 4.844 3.986 MTP (8 draft tokens)
tool_call 3.309 4.041 3.619 MTP (8 draft tokens)
question 3.062 3.379 3.299 MTP (8 draft tokens)
writing 3.054 3.281 3.107 MTP (8 draft tokens)
summarization 2.695 3.018 2.548 MTP (8 draft tokens)
translation 2.873 3.115 2.527 MTP (8 draft tokens)
weighted 3.398 4.106 3.892 MTP (8 draft tokens)

MTP at 8 draft tokens has the highest acceptance length on eight of the nine subsets and on the weighted average. Its lead over DSpark is widest on rag (4.844 against 3.986) and translation (3.115 against 2.527), and narrowest on question (3.379 against 3.299). DSpark is highest only on math_reasoning (5.405 against 4.743). MTP at 4 draft tokens is below MTP at 8 on every subset. DSpark is above it on seven subsets and below it on summarization and translation.

DSpark is faster than MTP at 8 draft tokens even on the eight subsets where MTP accepts more tokens per step, so each of its draft-and-verify steps takes less time.

Limitations

Works only with Qwen3.5-2B and is not usable as a standalone model. The verifier is vision-language, but the drafter was trained on text-only prompts and was evaluated on text-only benchmarks, so acceptance on image inputs is untested. All results are for greedy decoding with a single concurrent request; acceptance with sampling is typically lower, and throughput under load is not reported. The verifier's native MTP drafter at 8 draft tokens has a higher acceptance length on eight of the nine subsets, though lower throughput on all nine. MTP at 4 draft tokens is faster than DSpark on translation, where DSpark's speedup over the verifier alone is smallest (1.12×). Acceptance at the last draft position is under 1% on summarization and translation and about 10% or less on rag, tool_call, question and writing, so a block size of 8 is partly wasted there. Real-world speedup depends on your traffic mix, and because verification is lossless, the verifier's own behavior and biases carry through unchanged.

License

Apache 2.0, inherited from the verifier. The speculators training code is Apache-2.0.

Downloads last month
159
Safetensors
Model size
0.5B params
Tensor type
I64
·
BF16
·
BOOL
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for rasyosef/Qwen3.5-2B-DSpark

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(456)
this model
Quantizations
1 model

Collection including rasyosef/Qwen3.5-2B-DSpark