RELEASE · CODING · #903
NVIDIA open-sources NVCRE, a Kubernetes controller to validate GPU cluster readiness
NVIDIA Cluster Readiness Engine (NVCRE) is an open-source Kubernetes controller that runs real distributed workloads across topology-aware node groups to validate GPU cluster readiness before production AI jobs start. It exposes Certification/Workflow/Job CRDs, adaptive fault isolation to pinpoint failing nodes, a WorkloadRun API for multi-node GPU setup, integrates with NVIDIA AI Cluster Runtime and NVSentinel (part of NVIDIA DSX OS), and is installable via nvcrectl; contributors can extend tests and workload adapters on GitHub.
KEY POINTS
- NVIDIA Cluster Readiness Engine (NVCRE) is an open-source Kubernetes controller that runs real distributed workloads across topology-aware node groups to validate GPU cluster readiness before production AI jobs start.
- It exposes Certification/Workflow/Job CRDs, adaptive fault isolation to pinpoint failing nodes, a WorkloadRun API for multi-node GPU setup, integrates with NVIDIA AI Cluster Runtime and NVSentinel (part of NVIDIA DSX OS), and is installable via nvcrectl; contributors can extend tests and workload adapters on GitHub.
- Running real distributed workloads to validate clusters helps catch slow GPUs, network degradations, or configuration issues before costly AI training jobs fail and reduces time operators spend bisecting failures.
WHY IT MATTERS
Running real distributed workloads to validate clusters helps catch slow GPUs, network degradations, or configuration issues before costly AI training jobs fail and reduces time operators spend bisecting failures.