DeepSeek DSpark speeds inference 60-85%
Archive item — written before sources were shown.
DSpark cuts DeepSeek-V4-Flash latency 60-85% using semi-autoregressive generation and confidence-scheduled verification, deployed under live production traffic.
DeepSeek published DSpark (arxiv:2607.05147) on July 6, a speculative decoding approach combining a semi-autoregressive backbone with confidence-scheduled verification. In production on DeepSeek-V4-Flash under live user traffic, DSpark delivers 60 to 85% per-user generation speedup over the prior MTP-1 baseline; DeepSeek-V4-Pro sees 57 to 78% speedup. The improvement comes from the confidence scheduler dynamically adapting verification length per request based on estimated survival probabilities and engine throughput profiles.
- 01DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generationarxiv.org · primary
- 02DeepSeek open-sources DSpark, a new framework to speed up LLM inference by up to 85%venturebeat.com · reporting
