Tech Meridian ← ENTITY INDEX
RU

COMPANY · ENTITY #715

Amazon SageMaker HyperPod

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

COMPANIES · 1 SOURCE · AWS Machine Learning

Amazon SageMaker HyperPod adds model caching to reduce inference cold starts

Amazon SageMaker HyperPod now supports model caching for inference: model weights and container images can be pre-loaded onto cluster nodes' local NVMe storage so pods read from local disk instead of downloading over the network. AWS says this cuts cold starts from tens of minutes to seconds; the announcement explains how it works and how to enable it.

7.0

MODELS · 1 SOURCE · AWS Machine Learning

Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

AWS Machine Learning provides a walkthrough for deploying Qwen3.8-2.4T-A95B, a 2.4‑trillion‑parameter open‑weight model, on Amazon SageMaker HyperPod using vLLM. The guide covers cluster provisioning, NVFP4 quantization, and creating an OpenAI‑compatible endpoint with built‑in reasoning, tool calling, and native MTP speculative decoding.

7.0