SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification
Quick Answer
SonicSampler introduces a unified suite of tile-aware Triton kernels that optimize LLM sampling, achieving up to 16x speedup over existing methods while supporting dynamic sampling behaviors.
Quick Take
This innovative approach integrates the entire sampling pipeline into a single batched kernel, enhancing CUDA Graph execution efficiency for diverse workloads.
Key Points
- Achieves up to 10x speedup with a novel hierarchical two-stage top-k algorithm.
- Supports dynamic sampling behaviors like frequency penalties and speculative verification.
- Fully compatible with CUDA Graph for efficient execution.
- Optimizes the complete sampling pipeline into a single batched kernel.
- Demonstrates significant performance improvements across heterogeneous workloads.
DeepSignal Analysis
What happened
SonicSampler is a new suite of tile-aware Triton kernels designed to optimize the sampling process in large language model (LLM) inference. It integrates the entire sampling pipeline into a single batched kernel, achieving significant speedups in performance and supporting dynamic sampling behaviors.
Key evidence
- SonicSampler achieves up to 16x speedup over existing methods while maintaining flexible batched execution for diverse workloads.
- The approach includes a hierarchical two-stage top-k algorithm that provides up to 10x speedup compared to competitive baselines.
- The unified kernels support various dynamic sampling behaviors, such as grammar-constrained decoding and temperature scaling, within a single CUDA Graph-compatible kernel.
Why it matters
The development of SonicSampler addresses limitations in current LLM sampling implementations, which often require multiple kernel launches and lack support for dynamic behaviors. By optimizing the entire sampling pipeline, it enhances efficiency and flexibility, potentially improving the performance of applications relying on LLMs.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Sampling in LLM inference comprises a combinatorial set of logit processing, token selection, and verification operations for speculative decoding. However, existing implementations either accelerate only subsets of this pipeline, rely on multiple kernel launches, or assume homogeneous sampling behavior across a batch, limiting support for dynamic serving workloads and preventing efficient CUDA Graph execution. We present $\textbf{SonicSampler}$, a unified suite of tile-aware Triton kernels that vertically fuses the complete sampling pipeline into a fixed, workload-aware execution model. Our kernels support dynamic per-request sampling behaviors, including grammar-constrained decoding, repetition, frequency and presence penalties, logit bias, temperature scaling, top-$k$ / top-$p$ / min-$p$ filtering, and speculative verification - within a single batched kernel while remaining fully CUDA Graph-compatible. Central to our approach is a novel hierarchical two-stage top-$k$ algorithm that achieves up to $\textbf{10x speedup}$ over competitive baselines and exploits the low-entropy structure of LLM outputs to enable efficient selection over large vocabularies. Across heterogeneous speculative decoding workloads, SonicSampler achieves up to $\textbf{16x speedup}$ over state-of-the-art baselines while preserving flexible batched execution.
| Comments: | 26 pages, 12 figures |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.20475 [cs.AI] |
| (or arXiv:2607.20475v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.20475 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Pragaash Ponnusamy [view email]
[v1]
Sun, 24 May 2026 23:01:50 UTC (3,581 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.