Tech Meridian ← ENTITY INDEX
RU

COMPANY · ENTITY #837

OpenAI-compatible endpoint

Related event timeline, sources and context from the news index.

EVENT TIMELINE

3

COMPANIES · 1 SOURCE · AWS Machine Learning

Amazon SageMaker launches HyperPod Inference Gateway for GPU-aware LLM routing

Amazon announced the SageMaker HyperPod Inference Gateway, a Kubernetes-native EKS addon that routes OpenAI-compatible inference requests using real-time GPU signals (KV cache, queue depth, LoRA residency, etc.) to reduce first-token latency and GPU waste without application changes. The two-tier system offers per-cluster intelligent routing and fleet-wide coordination, deployable via a single InferenceGatewayConfig resource and emitting Prometheus/CloudWatch metrics.

7.0

CODING · 1 SOURCE · Cohere

Cohere publishes megakernel serving engine for North Mini Code with faster H100 decoding

Cohere describes a megakernel-based serving engine for its North Mini Code 30B model that runs BF16 on a single NVIDIA H100 and claims 1.25×–1.41× end-to-end speedup over vLLM, with a reported 292 tok/s (62% of SoL) at batch size 1 — about 1.58× faster than vLLM. The system supports production features (continuous batching, paged attention, ragged sequences), an OpenAI-compatible endpoint with tool calling, is implemented as a single CUDA file, and the code is available on GitHub.

7.0

MODELS · 1 SOURCE · AWS Machine Learning

Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

AWS Machine Learning provides a walkthrough for deploying Qwen3.8-2.4T-A95B, a 2.4‑trillion‑parameter open‑weight model, on Amazon SageMaker HyperPod using vLLM. The guide covers cluster provisioning, NVFP4 quantization, and creating an OpenAI‑compatible endpoint with built‑in reasoning, tool calling, and native MTP speculative decoding.

7.0