vLLM Speculative Decoding in 2026: P-EAGLE vs DFlash vs DSpark for Faster, Cheaper LLM Serving
P-EAGLE, DFlash, and DSpark bring parallel drafting to vLLM in 2026, generating a whole block of draft tokens in one pass for up to 1.69x higher throughput than EAGLE-3. How they work and how to serve them.