Qwen3-1.7B-DFlash2

A DFlash2 draft model for speculative decoding with Qwen/Qwen3-1.7B as the verifier, trained with speculators. The drafter proposes 7 tokens at a time (block size 8) and the verifier checks them in one forward pass, so output is identical to running the verifier alone β€” a lossless speedup. Mean acceptance length is 3.75 tokens committed per verification step, up to 4.91 on math_reasoning. That gives a throughput speedup over the verifier alone of 3.35Γ— on math_reasoning and 2.86Γ— on HumanEval, and 2.48Γ— averaged across nine task types.

Trained on 100,000 samples.

Usage

vLLM loads the verifier automatically from the config β€” don't pass it separately.

vllm serve rasyosef/Qwen3-1.7B-DFlash2 --port 8000 --gpu-memory-utilization 0.8

Then query the OpenAI-compatible endpoint at http://localhost:8000/v1.

Details

4 Qwen3 layers (hidden size 2048, intermediate size 6144, 16 attention heads over 8 KV heads, head dim 128), about 0.3B parameters. The first three layers use sliding-window attention with a 2048-token window and the last layer uses full attention. Block size 8, target hidden-state layers 2/10/18/26.

Trained for 3 epochs at lr 3e-4 (cosine schedule with 4% warmup) on 100,000 Open PerfectBlend prompts regenerated by the verifier itself, split 96/4 into train and validation. Prompts prepared at 2048 tokens; training sequence length 8192, up to 1024 anchors per sample. Verifier hidden states were pulled on demand from a running vLLM server during training and deleted after use rather than staged to disk up front.

Evaluation

evaluate.py throughput across the nine RedHatAI/speculator_benchmarks subsets. acceptance_length is mean tokens committed per verification step, including the bonus token β€” floor 1.0, ceiling 8.0 with 7 drafted tokens per step.

subset acceptance_length pos_0 pos_1 pos_2 pos_3 pos_4 pos_5 pos_6
math_reasoning 4.912 86.5% 74.0% 63.1% 53.7% 45.2% 37.8% 30.9%
HumanEval 4.201 83.1% 67.0% 53.2% 41.5% 32.2% 24.5% 18.6%
translation 3.558 78.3% 59.0% 43.1% 30.3% 20.9% 14.5% 9.8%
question 3.487 75.2% 54.6% 39.7% 29.5% 21.8% 16.1% 11.8%
writing 3.472 74.9% 54.3% 39.8% 29.2% 21.5% 16.0% 11.5%
rag 3.406 75.5% 54.6% 39.4% 28.2% 19.9% 13.7% 9.2%
tool_call 3.263 74.0% 52.5% 36.8% 25.6% 17.7% 11.8% 7.9%
qa 3.187 71.7% 50.1% 34.5% 24.4% 17.3% 12.2% 8.5%
summarization 2.645 66.9% 42.0% 25.1% 14.5% 8.4% 4.7% 2.7%

Weighted across all subsets: 3.753 over 355,511 verification steps.

Acceptance is highest where the verifier's next token is most predictable β€” math and code. math_reasoning leads HumanEval by about 0.7 tokens and holds a margin at every position in the block; math_reasoning's pos_6 (30.9%) is well above HumanEval's (18.6%), and its pos_4 (45.2%) is above summarization's pos_1 (42.0%). translation, question, writing, and rag cluster between 3.41 and 3.56. tool_call (3.26) and qa (3.19) sit just below that group, and summarization is lowest at 2.64.

Throughput (tokens/s)

Mean output throughput (tokens/s) across the same nine subsets, measured on a single A100 at concurrency 1. Baseline is Qwen3-1.7B with no speculative decoding, measured on every subset; without a drafter, throughput barely depends on the task (218.7–223.3 tokens/s). DSpark is rasyosef/Qwen3-1.7B-DSpark, a drafter for the same verifier measured under the same conditions. Speedup is this model's mean throughput relative to that subset's baseline.

subset baseline (no drafter) DSpark DFlash2 (this model) speedup
math_reasoning 220.1 725.8 738.1 3.35Γ—
HumanEval 219.7 586.5 627.8 2.86Γ—
translation 223.3 555.9 578.0 2.59Γ—
rag 218.7 506.0 526.2 2.41Γ—
qa 222.5 477.3 518.5 2.33Γ—
question 221.3 476.0 511.3 2.31Γ—
writing 220.9 483.8 509.3 2.31Γ—
tool_call 219.9 471.7 502.3 2.28Γ—
summarization 219.4 384.5 411.8 1.88Γ—
average speedup 1.00Γ— 2.35Γ— β€” 2.48Γ—

DFlash2 averages a 2.48Γ— speedup (unweighted mean of per-subset speedups). The biggest gains are on math_reasoning at 3.35Γ— (220.1 β†’ 738.1 tokens/s) and HumanEval at 2.86Γ—. translation is next at 2.59Γ— and rag at 2.41Γ—; qa, question, writing, and tool_call sit in a tight band between 2.28Γ— and 2.33Γ—. summarization is lowest at 1.88Γ—.

DFlash2 is faster than DSpark on every subset, 2.48Γ— against 2.35Γ— on average. The margin is smallest on math_reasoning (738.1 vs 725.8 tokens/s, about 2%) and largest on qa (518.5 vs 477.3, about 9%); the other seven subsets gain 4–7%.

Means and medians agree to within about 4% on every subset. On most subsets the median is the higher of the two (by 2–4% on HumanEval, translation, rag, question, and summarization), so a typical request does slightly better than the mean suggests. math_reasoning is the exception: its mean runs 3.5% above its median, so a typical request there sees closer to 3.24Γ—.

Limitations

Works only with Qwen3-1.7B and is not usable as a standalone model. Acceptance falls off steeply past the first few positions on prose-like traffic, and on summarization in particular the later positions of the block are mostly wasted β€” the gains concentrate in math and code. Tool-call traffic sees only prose-level acceptance with this drafter. Throughput was measured at concurrency 1 on one A100; real-world speedup depends on your traffic mix, hardware, and load, and because verification is lossless, the verifier's own behavior and biases carry through unchanged.

License

Apache-2.0, matching the verifier. The speculators training code is also Apache-2.0.

Downloads last month
24
Safetensors
Model size
0.3B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for rasyosef/Qwen3-1.7B-DFlash2

Finetuned
Qwen/Qwen3-1.7B
Finetuned
(1257)
this model

Collection including rasyosef/Qwen3-1.7B-DFlash2