iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Speculative decoding speeds up autoregressive text generation by having a faster drafter propose several tokens and asking the target model to verify them in fewer sequential decoding steps. EAGLE-3 drafts autoregressively using learned predictions and target-model features; DFlash drafts a block in one diffusion-model forward pass; xPress adds a causal refinement step to DFlash drafts. Which method is faster depends on the drafter’s overhead, how much the target accepts, and the serving workload—not on a paper’s maximum speedup alone.
What is speculative decoding?
A standard autoregressive model generates one token at a time: each new token depends on the preceding context, so generation requires a sequence of target-model steps. Speculative decoding introduces a faster proposer, or drafter, that predicts multiple candidate tokens. The target model then verifies those candidates in parallel. If verification accepts a useful run, the system can advance farther per target-model decoding iteration.
The trade-off is between the cost of drafting and verification and the amount of useful draft text accepted. A draft that is quick to produce but often rejected may not improve end-to-end speed. Conversely, a more expensive drafter can be worthwhile if it reliably reduces the target model’s sequential work.
“Lossless” or distribution-preserving describes the verification method under its assumptions: it can preserve the target model’s output distribution. It does not mean every run emits the same sampled text, that drafting has no cost, or that every workload runs faster. See the DFlash paper and the vLLM overview of parallel drafting, dated July 28, 2026.
#1 Best Overall
How do EAGLE-3, DFlash, and xPress differ?
| Method | How it drafts | Reported result and scope | Practical consideration |
|---|---|---|---|
| EAGLE-3 | Learned autoregressive token prediction with target-model features fused using training-time test. | The paper reports up to 6.5× speedup in its experiments; it is a maximum, not a general production expectation. EAGLE-3 paper | Draft tokens are generated autoregressively, so evaluate the sequential drafting work as part of total cost. |
| DFlash | A lightweight block-diffusion drafter produces a draft block in one forward pass, conditioned on target-model context features. | The authors report over 6× lossless acceleration across tested models and tasks, and up to 2.5× higher speedup than EAGLE-3 in their experiments. DFlash paper | Check implementation settings, including whether sample_from_anchor matches the model configuration. |
| xPress | A lightweight causal refiner restores dependencies between positions in a block-diffusion draft. | Against the original DFlash drafter on Qwen3-8B across seven math, code, and chat benchmarks, the authors report about 30% average acceptance-length gain (up to 56%) and about 1.3× average end-to-end decoding throughput (up to 1.7×). xPress paper | These figures describe the specified comparison and test suite, not a direct ranking against EAGLE-3 or a guarantee for another workload. |
What does EAGLE-3 change?
EAGLE-3 is a learned, autoregressive drafting method. It predicts candidate tokens and uses fused features from multiple layers of the target model; the paper describes using training-time test in this feature-fusion approach. This makes its proposal mechanism different from simply running a smaller standalone language model beside the target.
The paper’s up-to-6.5× result is the highest reported experimental speedup, not a multiplier to apply to an arbitrary deployment. The result depends on the tested models, tasks, and setup. The official EAGLE repository covers EAGLE-1, EAGLE-2, and EAGLE-3 and lists checkpoints. Before planning around a checkpoint, verify that it supports the intended target model and serving configuration.
Rank #2
What is DFlash’s block-diffusion approach?
DFlash uses a lightweight diffusion drafter to propose a block of tokens in one forward pass. It conditions that draft on context features extracted from the target model. This replaces a token-by-token autoregressive draft path with block drafting; the target still verifies the proposal.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →The DFlash paper reports over 6× lossless acceleration across its tested models and tasks, and a comparative maximum of 2.5× higher speedup than EAGLE-3 in its experiments. These are paper results, not a matched universal comparison or a forecast for a particular GPU, model, or concurrency level.
Rank #3
For vLLM, the Speculators DFlash guide describes the implementation and calls out the need to set sample_from_anchor to match the model configuration. Follow the current guide for version-specific setup rather than assuming a setting or checkpoint transfers unchanged between releases.
What does xPress add to DFlash?
Block diffusion can draft multiple positions together, but the resulting positions may not retain the dependencies that a causal sequence captures. xPress adds a lightweight causal refinement step to restore those dependencies and improve acceptance in the reported experiments.
Its reported gains are against the original DFlash drafter on Qwen3-8B across seven math, code, and chat benchmarks. The acceptance-length and throughput figures in the comparison table should be read together: higher acceptance alone does not guarantee faster user-visible generation. The xPress README describes a paper harness and a vLLM V1 integration; that documents a project path, not compatibility with every model or vLLM release.
Free tools Windows power users keep installed
One-click scans. No signup required.
How should you benchmark speculative decoding in vLLM?
There is no apples-to-apples universal ranking in the headline figures above: the methods’ results come from different models, task suites, and comparisons. Benchmark the exact deployment you care about with matched conditions. A practical comparison should hold constant:
- Target-model checkpoint, prompts, context lengths, and output-length distribution.
- Decoding and sampling settings, including temperature or other sampling controls where applicable.
- GPU or other accelerator, precision, available memory, batch size, and request concurrency.
- Serving framework and version, drafter/checkpoint versions, and warm-up procedure.
- Workload mix, including short answers and long structured generation.
Compare each speculative method with the same target model running without speculation under those conditions. Repeat runs consistently and record the software versions and configuration so a result can be reproduced. Use representative prompts rather than relying on a single favorable example.
Measure end-to-end performance, not just acceptance
Collect end-to-end tokens per second and latency, including time to first token when it matters to the application. Also record acceptance rate or accepted length, drafter overhead, verifier cost, memory use, and output-quality or distribution checks. Acceptance is a useful diagnostic: it helps explain why a method performs as it does. It is not a substitute for throughput and latency, which capture the cost of the entire serving path.
Keep output lengths comparable when interpreting throughput. A method can accept many tokens yet lose its advantage to drafting overhead, and a workload dominated by short responses may behave differently from long generation. The studies establish method-specific experimental results and implementation paths; they do not establish what a local benchmark will show.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsDoes speculative decoding preserve output quality?
With a correct verification procedure and its required assumptions and configuration, speculative decoding can preserve the target model’s output distribution. That is a distribution-level property, not a claim that speculative and non-speculative runs must produce identical sampled responses. It also does not establish that every implementation or configuration is correct. Check the method’s documentation for the target model and sampling mode, and include output or distribution checks in your deployment evaluation.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

