Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #5707

H100 GPUs

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

MODELS · 1 SOURCE · AWS Machine Learning

Amazon SageMaker HyperPod and Qumulo enable multi‑Region model training without copying data

Amazon and Qumulo validated an architecture pairing Amazon SageMaker HyperPod with Cloud Native Qumulo (CNQ) and Qumulo Cloud Data Fabric (CDF) so HyperPod clusters can train from a single dataset stored in a different AWS Region without replicating data. In a cross‑Region test (hub in us-east-2, spoke in us-west-2) running a 1.02B‑parameter LLaMA v3 job on two ml.p5.48xlarge instances per cluster (16 H100 GPUs total), the remote spoke converged to hub throughput (≈115–117 samples/sec) after a short NeuralCache warmup, achieving near‑full GPU utilization thereafter.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

RBS-Attention: radius-adaptive dual-branch sparse prefill for long-context LLMs

arXiv preprint arXiv:2609.20971v1 introduces RBS-Attention, a training-free sparse-prefill method that combines a centroid base branch with a radius-based rescue branch to avoid 'mean dilution' when selecting sparse blocks for long-context self-attention. On H100 GPUs the method reports up to 20.65× standalone prefill-attention speedup, 11.92× vLLM prefill-attention speedup and 5.97× end-to-end time-to-first-token speedup at 128K on Qwen3-30B-A3B-Instruct-2507-FP8, while retaining near-dense accuracy (88.65 vs 89.52 RULER on dense Qwen3-32B); evaluations include LongBench-v2, InfiniteBench and Video-MME.

7.0