Tech Meridian ← ENTITY INDEX
RU

COMPANY · ENTITY #203

Amazon SageMaker

Related event timeline, sources and context from the news index.

EVENT TIMELINE

4

COMPANIES · 1 SOURCE · AWS Machine Learning

Amazon SageMaker launches HyperPod Inference Gateway for GPU-aware LLM routing

Amazon announced the SageMaker HyperPod Inference Gateway, a Kubernetes-native EKS addon that routes OpenAI-compatible inference requests using real-time GPU signals (KV cache, queue depth, LoRA residency, etc.) to reduce first-token latency and GPU waste without application changes. The two-tier system offers per-cluster intelligent routing and fleet-wide coordination, deployable via a single InferenceGatewayConfig resource and emitting Prometheus/CloudWatch metrics.

7.0

MODELS · 1 SOURCE · AWS Machine Learning

Walkthrough: Customize Qwen3-8B on SageMaker serverless to build an AI product-tagging system

AWS Machine Learning published a walkthrough that demonstrates customizing Qwen3-8B using supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR) via Amazon SageMaker serverless model customization, then deploying it for asynchronous inference to create a cost-efficient product tagging system.

5.0

COMPANIES · 1 SOURCE · AWS Machine Learning

Amazon SageMaker AI adds instance preference lists for training and processing jobs

Amazon SageMaker AI now supports instance preference lists for training and processing jobs. You can specify an ordered list of up to five instance types, and SageMaker AI will automatically launch the first type with available capacity, removing the need for manual retry loops and capacity‑watching scripts.

6.0

COMPANIES · 1 SOURCE · AWS Machine Learning

Amazon SageMaker HyperPod adds model caching to reduce inference cold starts

Amazon SageMaker HyperPod now supports model caching for inference: model weights and container images can be pre-loaded onto cluster nodes' local NVMe storage so pods read from local disk instead of downloading over the network. AWS says this cuts cold starts from tens of minutes to seconds; the announcement explains how it works and how to enable it.

7.0