Sources
DFlash 2
Published figures cited for comparison (Inco AI, H200, SGLang, block size 8, temperature 1.0
with xhigh reasoning): acceptance length 4.80 mean for Qwen3.8-27B against MTP 4.28 and a
community DSpark drafter 3.62; throughput 2.7-3.4x autoregressive at concurrency 1.
Models
Engines
Evaluation data
Prior art referenced by DFlash 2
- Canon Layers; Dynamic Short Convolutions; Convolution for Large Language Models —
cited in the announcement as the basis for the two-tap dynamic depthwise convolution.
- Modal, Speculation Is All You Need — cited on speculative decoding for low-latency serving.