Co-designing AI models with speculative decoding to speed LLM inference
NVIDIA Developer published the third post in a series on AI model co-design that examines using speculative decoding to accelerate large language model (LLM) inference while aiming to preserve accuracy. The article discusses trade-offs and techniques for faster inference in the context of co-design work between models and systems.