Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX
Quick Answer
This study benchmarks confidential GPU inference on NVIDIA H100 under Intel TDX, revealing a 21.8% and 27.8% increase in time to first token for Mistral-7B and Qwen3-30B-A3B, respectively, while global token throughput decreases by 17.7% and 21.1%.
Quick Take
The findings indicate that while confidential execution maintains usable throughput, it incurs a performance penalty that must be considered in capacity planning.
Key Points
- Confidential mode increases average time to first token by 21.8% for Mistral-7B.
- Qwen3-30B-A3B shows a 27.8% increase in time to first token under confidential mode.
- Global token throughput drops by 17.7% for Mistral-7B and 21.1% for Qwen3-30B-A3B.
- Throughput gaps in closed-loop concurrency experiments range from 11.5% to 20.2%.
- Larger models reach saturation knee earlier in confidential execution mode.
DeepSignal Analysis
What happened
A benchmark study evaluated the performance of confidential GPU inference on the NVIDIA H100 using Intel TDX. The study found that enabling confidential execution increased the time to first token for two language models while decreasing global token throughput.
Key evidence
- The study reported a 21.8% increase in time to first token for the Mistral-7B model and a 27.8% increase for the Qwen3-30B-A3B model under confidential execution.
- Global token throughput decreased by 17.7% for Mistral-7B and 21.1% for Qwen3-30B-A3B when using confidential mode.
- In closed-loop concurrency tests, throughput gaps remained between 11.5% and 20.2%, with larger models reaching saturation earlier in confidential mode.
Why it matters
The findings highlight the trade-offs involved in implementing confidential computing for AI workloads. While it can maintain usable throughput, the performance penalties observed necessitate careful capacity planning, especially for larger models that exhibit earlier saturation.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Confidential computing is becoming a practical deployment requirement for AI inference workloads that process sensitive inputs or protect proprietary model assets. However, the performance cost of enabling confidential execution for GPU-accelerated large language model serving remains workload dependent and operationally important. This paper presents a benchmark study comparing standard non-confidential execution with confidential computing mode on a single NVIDIA H100 80GB GPU hosted in an Intel TDX confidential instance. The evaluation uses two representative language models, Mistral-7B v0.1 and Qwen3-30B-A3B, and measures time to first token, end-to-end request latency, per-request token generation throughput, global token throughput, and closed-loop request throughput under increasing concurrency. In fixed request-rate experiments, confidential mode increases average TTFT by 21.8% for Mistral-7B and 27.8% for Qwen3-30B-A3B, while global token throughput drops by 17.7% and 21.1%, respectively. In closed-loop concurrency experiments, throughput gaps remain in the 11.5-20.2% range, but the larger model reaches its saturation knee earlier under confidential mode. The results suggest that confidential GPU inference can retain usable throughput under load, but capacity planning must account for both the steady throughput penalty and the earlier saturation behavior observed for larger models.
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.19353 [cs.AI] |
| (or arXiv:2607.19353v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.19353 arXiv-issued DOI via DataCite |
Submission history
From: Wei Wang [view email]
[v1]
Wed, 20 May 2026 22:52:12 UTC (614 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.