NEWS · MODELS · #1311
NVIDIA publishes Dynamo-Triton HSTU recsys-examples for production inference
NVIDIA's recsys-examples repository provides an end-to-end HSTU generative recommender inference workflow that combines Dynamo-Triton, PyTorch AOTI, FlexKV-backed KV caching, and NV Embedding Cache. On an RTX PRO 6000 Blackwell GPU, the workflow's KV-cache-backed AOTI deployment achieved up to 5.93x lower latency (batch size 8, 100% GPU KV-cache hit) versus the same AOTI configuration without caching, and the guide shows export, validation (Python and native C++), and serving steps.
KEY POINTS
- NVIDIA's recsys-examples repository provides an end-to-end HSTU generative recommender inference workflow that combines Dynamo-Triton, PyTorch AOTI, FlexKV-backed KV caching, and NV Embedding Cache.
- On an RTX PRO 6000 Blackwell GPU, the workflow's KV-cache-backed AOTI deployment achieved up to 5.93x lower latency (batch size 8, 100% GPU KV-cache hit) versus the same AOTI configuration without caching, and the guide shows export, validation (Python and native C++), and serving steps.
- This matters because the workflow and GPU-backed KV caching show a practical, high-throughput path to serve sequence-based generative recommenders with substantially reduced latency in production.
WHY IT MATTERS
This matters because the workflow and GPU-backed KV caching show a practical, high-throughput path to serve sequence-based generative recommenders with substantially reduced latency in production.