Instructions to use rasyosef/Qwen3-1.7B-DFlash2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use rasyosef/Qwen3-1.7B-DFlash2 with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("rasyosef/Qwen3-1.7B-DFlash2", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Qwen3-1.7B-DFlash2
A DFlash2 draft model for speculative decoding with Qwen/Qwen3-1.7B as the verifier, trained with speculators. The drafter proposes 7 tokens at a time (block size 8) and the verifier checks them in one forward pass, so output is identical to running the verifier alone β a lossless speedup. Mean acceptance length is 3.75 tokens committed per verification step, up to 4.91 on math_reasoning. That gives a throughput speedup over the verifier alone of 3.35Γ on math_reasoning and 2.86Γ on HumanEval, and 2.48Γ averaged across nine task types.
Trained on 100,000 samples.
Usage
vLLM loads the verifier automatically from the config β don't pass it separately.
vllm serve rasyosef/Qwen3-1.7B-DFlash2 --port 8000 --gpu-memory-utilization 0.8
Then query the OpenAI-compatible endpoint at http://localhost:8000/v1.
Details
4 Qwen3 layers (hidden size 2048, intermediate size 6144, 16 attention heads over 8 KV heads, head dim 128), about 0.3B parameters. The first three layers use sliding-window attention with a 2048-token window and the last layer uses full attention. Block size 8, target hidden-state layers 2/10/18/26.
Trained for 3 epochs at lr 3e-4 (cosine schedule with 4% warmup) on 100,000 Open PerfectBlend prompts regenerated by the verifier itself, split 96/4 into train and validation. Prompts prepared at 2048 tokens; training sequence length 8192, up to 1024 anchors per sample. Verifier hidden states were pulled on demand from a running vLLM server during training and deleted after use rather than staged to disk up front.
Evaluation
evaluate.py throughput across the nine RedHatAI/speculator_benchmarks subsets. acceptance_length is mean tokens committed per verification step, including the bonus token β floor 1.0, ceiling 8.0 with 7 drafted tokens per step.
| subset | acceptance_length | pos_0 | pos_1 | pos_2 | pos_3 | pos_4 | pos_5 | pos_6 |
|---|---|---|---|---|---|---|---|---|
| math_reasoning | 4.912 | 86.5% | 74.0% | 63.1% | 53.7% | 45.2% | 37.8% | 30.9% |
| HumanEval | 4.201 | 83.1% | 67.0% | 53.2% | 41.5% | 32.2% | 24.5% | 18.6% |
| translation | 3.558 | 78.3% | 59.0% | 43.1% | 30.3% | 20.9% | 14.5% | 9.8% |
| question | 3.487 | 75.2% | 54.6% | 39.7% | 29.5% | 21.8% | 16.1% | 11.8% |
| writing | 3.472 | 74.9% | 54.3% | 39.8% | 29.2% | 21.5% | 16.0% | 11.5% |
| rag | 3.406 | 75.5% | 54.6% | 39.4% | 28.2% | 19.9% | 13.7% | 9.2% |
| tool_call | 3.263 | 74.0% | 52.5% | 36.8% | 25.6% | 17.7% | 11.8% | 7.9% |
| qa | 3.187 | 71.7% | 50.1% | 34.5% | 24.4% | 17.3% | 12.2% | 8.5% |
| summarization | 2.645 | 66.9% | 42.0% | 25.1% | 14.5% | 8.4% | 4.7% | 2.7% |
Weighted across all subsets: 3.753 over 355,511 verification steps.
Acceptance is highest where the verifier's next token is most predictable β math and code. math_reasoning leads HumanEval by about 0.7 tokens and holds a margin at every position in the block; math_reasoning's pos_6 (30.9%) is well above HumanEval's (18.6%), and its pos_4 (45.2%) is above summarization's pos_1 (42.0%). translation, question, writing, and rag cluster between 3.41 and 3.56. tool_call (3.26) and qa (3.19) sit just below that group, and summarization is lowest at 2.64.
Throughput (tokens/s)
Mean output throughput (tokens/s) across the same nine subsets, measured on a single A100 at concurrency 1. Baseline is Qwen3-1.7B with no speculative decoding, measured on every subset; without a drafter, throughput barely depends on the task (218.7β223.3 tokens/s). DSpark is rasyosef/Qwen3-1.7B-DSpark, a drafter for the same verifier measured under the same conditions. Speedup is this model's mean throughput relative to that subset's baseline.
| subset | baseline (no drafter) | DSpark | DFlash2 (this model) | speedup |
|---|---|---|---|---|
| math_reasoning | 220.1 | 725.8 | 738.1 | 3.35Γ |
| HumanEval | 219.7 | 586.5 | 627.8 | 2.86Γ |
| translation | 223.3 | 555.9 | 578.0 | 2.59Γ |
| rag | 218.7 | 506.0 | 526.2 | 2.41Γ |
| qa | 222.5 | 477.3 | 518.5 | 2.33Γ |
| question | 221.3 | 476.0 | 511.3 | 2.31Γ |
| writing | 220.9 | 483.8 | 509.3 | 2.31Γ |
| tool_call | 219.9 | 471.7 | 502.3 | 2.28Γ |
| summarization | 219.4 | 384.5 | 411.8 | 1.88Γ |
| average speedup | 1.00Γ | 2.35Γ | β | 2.48Γ |
DFlash2 averages a 2.48Γ speedup (unweighted mean of per-subset speedups). The biggest gains are on math_reasoning at 3.35Γ (220.1 β 738.1 tokens/s) and HumanEval at 2.86Γ. translation is next at 2.59Γ and rag at 2.41Γ; qa, question, writing, and tool_call sit in a tight band between 2.28Γ and 2.33Γ. summarization is lowest at 1.88Γ.
DFlash2 is faster than DSpark on every subset, 2.48Γ against 2.35Γ on average. The margin is smallest on math_reasoning (738.1 vs 725.8 tokens/s, about 2%) and largest on qa (518.5 vs 477.3, about 9%); the other seven subsets gain 4β7%.
Means and medians agree to within about 4% on every subset. On most subsets the median is the higher of the two (by 2β4% on HumanEval, translation, rag, question, and summarization), so a typical request does slightly better than the mean suggests. math_reasoning is the exception: its mean runs 3.5% above its median, so a typical request there sees closer to 3.24Γ.
Limitations
Works only with Qwen3-1.7B and is not usable as a standalone model. Acceptance falls off steeply past the first few positions on prose-like traffic, and on summarization in particular the later positions of the block are mostly wasted β the gains concentrate in math and code. Tool-call traffic sees only prose-level acceptance with this drafter. Throughput was measured at concurrency 1 on one A100; real-world speedup depends on your traffic mix, hardware, and load, and because verification is lossless, the verifier's own behavior and biases carry through unchanged.
License
Apache-2.0, matching the verifier. The speculators training code is also Apache-2.0.
- Downloads last month
- 24