Not established from the available sources.
NVIDIA · MODEL RELEASE TRACKER
NVIDIA Nemotron-3 Nano 30B
NVIDIA Nemotron-3 Nano 30B is presented as a generative AI model that can be deployed to Amazon SageMaker AI endpoints; it is used in an automated concurrency-sweep workflow to measure endpoint throughput and latency.CURRENT SNAPSHOT2/5 DIMENSIONS WITH DATA
The dimensions that change the decision.
Not established from the available sources.
Not established from the available sources.
Benchmarking (concurrency sweeps)The model is used in an automated concurrency-sweep workflow that measures throughput (tokens per second) and latency to identify endpoint saturation and right-size capacity.
DeployabilityCan be deployed to Amazon SageMaker AI endpoints using the native vLLM container.
VERIFIABLE FACTS
Every value stays attached to a source and date.
CAPABILITIES · Intended useDEVELOPER CLAIM
Presented for use with generative AI endpoints (used as a generative model in the concurrency-sweep example).
AVAILABILITY · DeployabilityDEVELOPER CLAIM
Can be deployed to Amazon SageMaker AI endpoints using the native vLLM container.
BENCHMARKS · Benchmarking (concurrency sweeps)DEVELOPER CLAIM
The model is used in an automated concurrency-sweep workflow that measures throughput (tokens per second) and latency to identify endpoint saturation and right-size capacity.
WHAT CHANGED