NVIDIA documents TensorRT LLM adaptations for confidential Blackwell (B200) inference
NVIDIA published guidance and measured results showing how TensorRT LLM adapts to confidential computing on Blackwell B200 hardware, including CC-aware memory selection, async token readback, using the GPU %globaltimer for autotuning, and communication-path detection when NVLS multicast is unavailable. In their tests, CC-enabled runs retained about 96.1–98.2% of token throughput versus CC-off, with per-token latency (TPOT) within roughly 1.2–4.3% of baseline.