Instructions to use rasyosef/Qwen3.5-2B-DSpark with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use rasyosef/Qwen3.5-2B-DSpark with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("rasyosef/Qwen3.5-2B-DSpark", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Qwen3.5-2B-DSpark
A DSpark draft model for speculative decoding with Qwen/Qwen3.5-2B as the verifier, trained with speculators. The drafter proposes 8 tokens at a time and the verifier checks them in one forward pass, so output is identical to running the verifier alone — a lossless speedup. At temperature 0, mean acceptance length is 3.89 tokens committed per verification step, up to 5.40 on math_reasoning. Compared with the verifier alone, throughput is 2.00× higher on average, up to 2.88× on math_reasoning and 2.62× on HumanEval, and it is faster than the verifier's native MTP drafter on eight of the nine benchmark subsets.
Training code: rasyosef/train-dspark-draft-models.
Trained on 100,000 samples.
Usage
vLLM
vLLM loads the verifier automatically from the config — don't pass it separately.
vllm serve rasyosef/Qwen3.5-2B-DSpark --port 8000 \
--gpu-memory-utilization 0.8 \
--override-generation-config '{"temperature": 0}'
Then query the OpenAI-compatible endpoint at http://localhost:8000/v1.
llama.cpp
llama-server -hf unsloth/Qwen3.5-2B-GGUF:BF16 \
-hfd rasyosef/Qwen3.5-2B-DSpark-GGUF:BF16 \
--spec-type draft-dspark --spec-draft-n-max 8 --spec-draft-n-min 0 \
-fa on -ngl 99
The block size is read from the sidecar metadata (n-max is clamped to it). Speculative decoding is exact: the target verifies every proposed token, so greedy output equals the target alone; per-response timings report draft_n / draft_n_accepted.
Details
5 Qwen3 layers (hidden size 2048, intermediate size 6144, 16 attention heads over 8 KV heads, head dim 128, all layers sliding-window attention with a 2048-token window), ~0.5B params, bfloat16. Block size 8, draft vocabulary reduced to 50,000 from the verifier's 248,320 (selected by token frequency over the training data), aux hidden-state layers 1/6/11/16/21, confidence head with Markov (vanilla, rank 256).
Trained for 4 epochs at lr 4e-4 (AdamW, cosine schedule with 4% warmup) on 100,000 filtered Open PerfectBlend prompts with responses regenerated by the verifier itself, split 96/4 into train and validation, with a {"ce": 0.1, "tv": 0.9} loss. Samples prepared at 2048 tokens; training sequence length 8192, up to 1024 anchors per sample. Verifier hidden states were pulled on demand from a running vLLM server during training and deleted after use rather than staged to disk up front.
Evaluation
Drafter results are at temperature 0 (greedy decoding), from evaluate.py throughput across the nine RedHatAI/speculator_benchmarks subsets. acceptance_length is mean tokens committed per verification step, including the bonus token — floor 1.0, ceiling 9.0 at block size 8.
Acceptance
| subset | acceptance_length | pos_0 | pos_1 | pos_2 | pos_3 | pos_4 | pos_5 | pos_6 | pos_7 |
|---|---|---|---|---|---|---|---|---|---|
| math_reasoning | 5.405 | 87.4% | 75.8% | 64.9% | 56.1% | 48.5% | 41.5% | 35.9% | 30.3% |
| HumanEval | 4.766 | 84.1% | 69.6% | 57.5% | 47.6% | 39.3% | 32.4% | 25.8% | 20.3% |
| qa | 4.074 | 73.2% | 54.1% | 42.0% | 35.5% | 30.2% | 27.3% | 24.2% | 20.9% |
| rag | 3.986 | 80.1% | 61.9% | 47.6% | 37.4% | 29.1% | 20.3% | 13.0% | 9.2% |
| tool_call | 3.619 | 73.5% | 53.8% | 40.8% | 30.0% | 23.0% | 17.4% | 13.3% | 10.0% |
| question | 3.299 | 69.4% | 46.4% | 32.1% | 24.4% | 19.7% | 15.5% | 12.4% | 10.0% |
| writing | 3.107 | 69.3% | 44.9% | 30.0% | 21.4% | 16.4% | 12.3% | 9.4% | 7.0% |
| summarization | 2.548 | 65.9% | 38.7% | 22.3% | 13.3% | 7.8% | 3.8% | 2.1% | 0.9% |
| translation | 2.527 | 68.2% | 41.0% | 23.5% | 11.0% | 4.7% | 2.3% | 1.3% | 0.7% |
Weighted across all subsets: 3.892 over 120,525 verification steps. math_reasoning and HumanEval are highest, and both still accept more than one draft in five at pos_7, as does qa. summarization and translation are lowest and fall below 1% by pos_7.
Throughput (tokens/s)
Output throughput with a single concurrent request on one A100. Baseline is Qwen3.5-2B with no speculative decoding. MTP is the verifier's native MTP drafter, run at 4 and 8 draft tokens. The fastest configuration in each row is bolded. The last column gives DSpark's throughput divided by baseline throughput for each subset.
| subset | baseline (no drafter) | MTP (4 draft tokens) | MTP (8 draft tokens) | DSpark (this model) | DSpark speedup |
|---|---|---|---|---|---|
| math_reasoning | 233.4 | 410.1 | 357.0 | 671.8 | 2.88× |
| HumanEval | 233.5 | 412.8 | 376.4 | 612.2 | 2.62× |
| qa | 228.7 | 386.7 | 311.4 | 472.9 | 2.07× |
| rag | 221.5 | 401.2 | 373.3 | 499.5 | 2.26× |
| tool_call | 224.5 | 360.5 | 376.2 | 560.4 | 2.50× |
| question | 231.6 | 345.4 | 229.3 | 359.7 | 1.55× |
| writing | 233.0 | 345.1 | 229.4 | 365.3 | 1.57× |
| summarization | 219.5 | 293.8 | 220.5 | 316.6 | 1.44× |
| translation | 168.4 | 196.6 | 152.9 | 188.1 | 1.12× |
| average speedup | 1.00× | 1.57× | 1.31× | 2.00× | 2.00× |
DSpark has the highest average speedup at 2.00× (unweighted mean of per-subset speedups), ahead of MTP at 4 draft tokens (1.57×) and at 8 (1.31×), and it is the fastest configuration on eight of the nine subsets. The gains are largest on math_reasoning (2.88×), HumanEval (2.62×) and tool_call (2.50×), and smallest on summarization (1.44×) and translation (1.12×). Translation is the one subset where MTP at 4 draft tokens is ahead (196.6 against 188.1); it is also the smallest subset, with fewer than 900 verification steps in each run.
MTP at 8 draft tokens is slower than MTP at 4 on every subset except tool_call, and it only matches or falls below the baseline on question, writing, summarization and translation.
Acceptance comparison with MTP (4 and 8 draft tokens)
Acceptance length for DSpark and for the verifier's native MTP drafter on the same subsets at temperature 0. The highest value in each row is bolded and named in the last column. At 4 draft tokens the ceiling on acceptance_length is 5.0 rather than 9.0.
| subset | MTP (4 draft tokens) | MTP (8 draft tokens) | DSpark (this model) | best |
|---|---|---|---|---|
| math_reasoning | 3.832 | 4.743 | 5.405 | DSpark |
| HumanEval | 3.785 | 4.862 | 4.766 | MTP (8 draft tokens) |
| qa | 3.408 | 4.227 | 4.074 | MTP (8 draft tokens) |
| rag | 3.812 | 4.844 | 3.986 | MTP (8 draft tokens) |
| tool_call | 3.309 | 4.041 | 3.619 | MTP (8 draft tokens) |
| question | 3.062 | 3.379 | 3.299 | MTP (8 draft tokens) |
| writing | 3.054 | 3.281 | 3.107 | MTP (8 draft tokens) |
| summarization | 2.695 | 3.018 | 2.548 | MTP (8 draft tokens) |
| translation | 2.873 | 3.115 | 2.527 | MTP (8 draft tokens) |
| weighted | 3.398 | 4.106 | 3.892 | MTP (8 draft tokens) |
MTP at 8 draft tokens has the highest acceptance length on eight of the nine subsets and on the weighted average. Its lead over DSpark is widest on rag (4.844 against 3.986) and translation (3.115 against 2.527), and narrowest on question (3.379 against 3.299). DSpark is highest only on math_reasoning (5.405 against 4.743). MTP at 4 draft tokens is below MTP at 8 on every subset. DSpark is above it on seven subsets and below it on summarization and translation.
DSpark is faster than MTP at 8 draft tokens even on the eight subsets where MTP accepts more tokens per step, so each of its draft-and-verify steps takes less time.
Limitations
Works only with Qwen3.5-2B and is not usable as a standalone model. The verifier is vision-language, but the drafter was trained on text-only prompts and was evaluated on text-only benchmarks, so acceptance on image inputs is untested. All results are for greedy decoding with a single concurrent request; acceptance with sampling is typically lower, and throughput under load is not reported. The verifier's native MTP drafter at 8 draft tokens has a higher acceptance length on eight of the nine subsets, though lower throughput on all nine. MTP at 4 draft tokens is faster than DSpark on translation, where DSpark's speedup over the verifier alone is smallest (1.12×). Acceptance at the last draft position is under 1% on summarization and translation and about 10% or less on rag, tool_call, question and writing, so a block size of 8 is partly wasted there. Real-world speedup depends on your traffic mix, and because verification is lossless, the verifier's own behavior and biases carry through unchanged.
License
Apache 2.0, inherited from the verifier. The speculators training code is Apache-2.0.
- Downloads last month
- 159