THE BEST NEWS OF THE WEB
Lien affiliéNike Sac de Sport Rigide SoccerVoir l'offre Amazon(s'ouvre dans un nouvel onglet)
← Toutes les actus
AI

🇬🇧 DFlash Speculative Decoding Drafts Whole Token Blocks in Parallel for Up to 15x Higher Throughput on NVIDIA Blackwell

UC San Diego's DFlash replaces autoregressive drafting with a lightweight block diffusion model for speculative decoding. It drafts whole token blocks in a single forward pass and conditions on target hidden features through KV injection. The paper reports up to 6.08x lossless speedup on Qwen3-8B, while NVIDIA reports up to 15x throughput on Blackwell at fixed interactivity. D…

1 vote / personne · anonyme

1 source