DirectorySecurity AdvisoriesPricing
Sign in
Directory
nemo-rl logo

nemo-rl

packaged by Chainguard

Last changed
Request a free trial

Contact our team to test out this image for free. Please also indicate any other images you would like to evaluate.

Tags
Overview
Comparison
Provenance
Specifications
SBOM
Vulnerabilities
Advisories

Chainguard Container for nemo-rl

NVIDIA NeMo RL is a scalable post-training library for reinforcement learning and supervised fine-tuning of large language models (GRPO, DPO, SFT, distillation and reward modeling), built on Ray, vLLM, PyTorch and Megatron.

Chainguard Containers are regularly-updated, secure-by-default container images.

Download this Container Image

For those with access, this container image is available on cgr.dev:

docker pull cgr.dev/ORGANIZATION/nemo-rl:latest

Be sure to replace the ORGANIZATION placeholder with the name used for your organization's private repository within the Chainguard Registry.

Compatibility Notes

The nemo-rl image is designed as a drop-in replacement for NVIDIA's nvcr.io/nvidia/nemo-rl container, version 0.7.0. It ships NeMo RL unmodified, with the dependency versions from upstream's uv.lock apart from security fixes. The following match NVIDIA's image, so existing scripts, recipes, and Kubernetes manifests work unchanged:

  • Layout. The source tree and work directory is /opt/nemo-rl. The driver venv is /opt/nemo_rl_venv, first on PATH, so bare python and ray are the driver's. The worker venvs are under /opt/ray_venvs. These paths link to the Chainguard tree under /opt/chainguard/nemo-rl.
  • Environment. NRL_CONTAINER=1 is set, and /opt/nemo_rl_container_fingerprint matches the shipped source tree, so NeMo RL's container fingerprint check passes. The other variables are listed under Configuration.
  • User and entrypoint. The container runs as root. Its entrypoint checks the NVIDIA driver and enables CUDA forward compatibility on older drivers, then runs its arguments, or bash when there are none.

The image differs from NVIDIA's in the following ways. Each difference is deliberate:

  • Launch recipes with python, not uv run. NeMo RL's documentation launches recipes with uv run examples/run_grpo.py .... uv run re-syncs the venv against uv.lock: that needs network access (it fails offline in NVIDIA's image too), would undo Chainguard's security fixes, and would try to rebuild the extensions built from source. This image ships no uv or uv.lock, so uv run fails with bash: uv: command not found. Run python examples/run_grpo.py ... instead.
  • Security fixes move some dependencies past upstream's pins. Most notably transformers 5.10.0 (upstream pins 5.5.0 to 5.8.1, depending on the backend), ray 2.56.0 (2.55.1), mlflow 3.16.1 (3.14.0), and wandb 0.30.0 (0.28.0), along with cryptography, aiohttp, hydra-core, and the OpenTelemetry packages. Each was tested against NeMo RL's example recipes. vLLM stays at 0.20.0, the version NeMo RL 0.7.0 pins.
  • No development tools in the driver venv. NVIDIA's driver venv also carries NeMo RL's dev, docs, and test dependency groups (pytest, ruff, pre-commit, sphinx, and their dependencies), which nothing at runtime imports. To run NeMo RL's own unit tests, add pytest to a derived image.
  • TORCH_CUDA_ARCH_LIST is 8.0 9.0 10.0 12.0, where NVIDIA's image sets 9.0 10.0. NVIDIA's value compiles no runtime kernels for A100, L4, or RTX PRO 6000 GPUs; this one is a superset.
  • The entrypoint requires a driver. Without a GPU and the NVIDIA Container Toolkit, the entrypoint stops with FATAL: libcuda.so.1 not found in the library path. instead of printing a warning and continuing. To run without a GPU, override it with --entrypoint /bin/bash.
  • LD_LIBRARY_PATH does not list the forward-compatibility libraries. The entrypoint enables them only when the host driver needs them. CUDA_VERSION and NVIDIA's other version-tracking variables are not set.
  • Not included: TensorRT, DALI, the EFA and HPC-X networking stack, apptainer, and the Nsight profilers. NeMo RL uses none of them. Add Nsight with Custom Assembly; see Profiling.
  • x86_64 only.

Prerequisites

  • NVIDIA GPUs of the Ampere (sm_80), Ada (sm_89), Hopper (sm_90), or Blackwell (sm_100, sm_120) generations. The image has been validated on A100, L4, H100, and RTX PRO 6000 GPUs. NeMo RL's default recipes are sized for GPUs with 40 GB or more memory; on smaller GPUs, lower policy.train_micro_batch_size.
  • An NVIDIA driver that supports CUDA 13.0 (R580 or later). On data center GPUs with an older driver, the entrypoint enables CUDA forward compatibility.
  • The NVIDIA Container Toolkit on the host, or the NVIDIA GPU Operator on Kubernetes.

Getting Started

Start the container with access to the GPUs:

docker run -it --rm \
  --gpus all \
  --shm-size=16g \
  --ulimit memlock=-1 \
  --ulimit stack=67108864 \
  -e HF_TOKEN \
  -v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
  -v "$PWD/results:/opt/nemo-rl/results" \
  -v "$PWD/logs:/opt/nemo-rl/logs" \
  cgr.dev/ORGANIZATION/nemo-rl:latest
  • --shm-size gives Ray's object store room in /dev/shm; Docker's 64 MB default makes Ray spill objects to disk. The --ulimit flags match NVIDIA's recommendations for NCCL and PyTorch.
  • The first mount caches downloaded models and datasets across runs. The other two keep checkpoints and logs, which the recipes write to results/ and logs/ under the work directory.

The shell starts in the work directory, /opt/nemo-rl; pwd prints /opt/chainguard/nemo-rl/src, the directory it links to. Check that the GPUs are visible:

python -c 'import nemo_rl, torch; print(nemo_rl.__version__, torch.cuda.device_count())'

On a host with two GPUs, this prints the NeMo RL version and the GPU count:

0.7.0 2

Run a short GRPO job on one GPU with the default math recipe, which trains Qwen/Qwen2.5-1.5B with DTensor v2 and generates with vLLM:

python examples/run_grpo.py --config examples/configs/grpo_math_1B.yaml \
  grpo.max_num_steps=10

SFT and DPO run the same way with examples/run_sft.py and examples/run_dpo.py. Their default configurations use gated Llama models, so either set HF_TOKEN for an account with access, or override policy.model_name with an ungated model such as Qwen/Qwen2.5-0.5B-Instruct.

To run a command without an interactive shell, pass it after the image name:

docker run --rm --gpus all --shm-size=16g cgr.dev/ORGANIZATION/nemo-rl:latest \
  python examples/run_grpo.py --config examples/configs/grpo_math_1B.yaml grpo.max_num_steps=10

Run your own NeMo RL code

To run a modified NeMo RL, mount your checkout of the same release over /opt/nemo-rl:

docker run -it --rm --gpus all --shm-size=16g \
  -v "$PWD/RL:/opt/nemo-rl" \
  cgr.dev/ORGANIZATION/nemo-rl:latest

The driver and workers then import nemo_rl from your checkout and keep using the image's venvs. NeMo RL logs a container fingerprint mismatch warning, because your checkout carries its own uv.lock; the warning is expected. A checkout mounted anywhere else still imports the nemo_rl shipped in the image.

Run on Kubernetes or across nodes

NeMo RL runs multi-node jobs on a Ray cluster. On Kubernetes, follow NeMo RL's KubeRay guide with cgr.dev/ORGANIZATION/nemo-rl:latest as the image, and note the following:

  • Replace uv run with python in its commands. Cloning NeMo RL onto the shared volume is only needed to run modified code; otherwise run the recipes from /opt/nemo-rl.
  • Mount a memory-backed volume (emptyDir with medium: Memory) at /dev/shm in every Ray pod, as the guide shows.
  • KubeRay starts Ray by overriding the container command, so the entrypoint does not run. On nodes whose driver is older than R580, add /usr/local/cuda/compat to LD_LIBRARY_PATH in the pod spec to enable forward compatibility.

Profiling

NeMo RL profiles its Ray workers with Nsight Systems (see NeMo RL's profiling guide). Add the nvidia-nsight-systems-cli package with Custom Assembly, then select workers with NRL_NSYS_WORKER_PATTERNS and the steps to profile with NRL_NSYS_PROFILE_STEP_RANGE:

NRL_NSYS_WORKER_PATTERNS='*policy*' NRL_NSYS_PROFILE_STEP_RANGE=2:3 \
  python examples/run_grpo.py --config examples/configs/grpo_math_1B.yaml \
  grpo.max_num_steps=3 checkpointing.checkpoint_dir=results/grpo-profile

Each profiled worker writes a report such as dtensor_policy_worker_v2_2:3_<pid>.nsys-rep to /tmp/ray/session_*/logs/nsight/ in the container. Add -v "$PWD/ray:/tmp/ray" to the docker run command to keep the reports after the container exits, and open them in the Nsight Systems desktop application.

The nvidia-nsight-systems-cli-collectx package adds the InfiniBand and NIC counters behind nsys --nic-metrics and --ib-switch-metrics-devices, and nvidia-nsight-compute-13.0 adds Nsight Compute.

Configuration

NeMo RL is configured through its YAML recipes, with command-line overrides for individual keys. For example, this runs the default GRPO recipe across both GPUs of a two-GPU host, generating with vLLM on one GPU and training on the other, and saves a checkpoint every 5 steps in its own directory under the mounted results/:

python examples/run_grpo.py --config examples/configs/grpo_math_1B.yaml \
  cluster.gpus_per_node=2 \
  policy.generation.colocated.enabled=false \
  policy.generation.colocated.resources.gpus_per_node=1 \
  grpo.max_num_steps=10 \
  checkpointing.checkpoint_dir=results/grpo-2gpu \
  checkpointing.save_period=5

See the NeMo RL documentation for the recipe keys.

The image sets the same environment as NVIDIA's image. Change these only to match a different host setup:

  • PATH puts the driver venv (/opt/nemo_rl_venv/bin) first, then /usr/local/cuda/bin for nvcc.
  • CUDA_HOME=/usr/local/cuda and CPLUS_INCLUDE_PATH=/usr/local/cuda/include/cccl point runtime kernel compilation (Triton, TorchInductor, and FlashInfer) at CUDA 13.0's nvcc, the CUDA runtime, cuRAND, and libcu++ headers. /usr/local/cuda does not carry the other CUDA math libraries (cuBLAS, cuFFT, and the rest); PyTorch loads its own copies from /opt/chainguard/nemo-rl/shared/nvidia.
  • NRL_CONTAINER=1, NEMO_RL_VENV_DIR=/opt/ray_venvs, and NEMO_GYM_VENV_DIR=/opt/gym_venvs tell NeMo RL it runs in a container and where its venvs are. /opt/gym_venvs is an empty, writable directory where NeMo Gym creates venvs at runtime.
  • TORCH_CUDA_ARCH_LIST="8.0 9.0 10.0 12.0" lists the GPU architectures for kernels compiled at runtime. Megatron refuses to start without it, so it is set even though the image does not include Megatron by default; Custom Assembly cannot set environment variables.
  • RAY_USAGE_STATS_ENABLED=0 turns off Ray's usage reporting, and RAY_ENABLE_UV_RUN_RUNTIME_ENV=0 stops Ray from propagating uv run to workers.
  • NVIDIA_VISIBLE_DEVICES=all and NVIDIA_DRIVER_CAPABILITIES=compute,utility,video tell the NVIDIA Container Toolkit to expose every GPU and the driver libraries.

Adding Backends

Add the following packages with Custom Assembly to enable more of NeMo RL. Each backend package activates itself when installed; no configuration is needed.

PackageEnables

nemo-rl-cuda-13.0-mcore

Megatron training (Megatron-Bridge and Megatron-LM), used by the *_megatron.yaml recipes

nemo-rl-cuda-13.0-sglang

SGLang generation and its router, used by the *_sglang.yaml recipes

nemo-rl-cuda-13.0-gym

The NeMo Gym environment actors. As upstream, NeMo Gym's environment servers build their own venvs in /opt/gym_venvs at runtime, which needs network access.

nemo-rl-cuda-13.0-flashinfer-jit-cache

Precompiled FlashInfer kernels for the vLLM and Megatron backends, so the first run skips JIT compilation. Large (about 9 GB), and only affects start-up time.

nemo-rl-cuda-13.0-sglang-flashinfer-jit-cache

The same for the SGLang backend (about 8 GB)

nemo-rl-cuda-13.0-mooncake

The Mooncake transfer engine behind the non-default mooncake_cpu data plane. Its bundled Go etcd client has open CVEs, which is why it is not included by default.

A recipe that needs a backend the image does not include fails when its Ray actor starts. NeMo RL tries to create the missing venv with uv, and the run stops with a ray.exceptions.RayTaskError from _env_builder() that ends with:

FileNotFoundError: [Errno 2] No such file or directory: 'uv'

Add the backend's package rather than installing uv.

Documentation and Resources

What are Chainguard Containers?

Chainguard's free tier of Starter container images are built with Wolfi, our minimal Linux undistro.

All other Chainguard Containers are built with Chainguard OS, Chainguard's minimal Linux operating system designed to produce container images that meet the requirements of a more secure software supply chain.

The main features of Chainguard Containers include:

For cases where you need container images with shells and package managers to build or debug, most Chainguard Containers come paired with a development, or -dev, variant.

In all other cases, including Chainguard Containers tagged as :latest or with a specific version number, the container images include only an open-source application and its runtime dependencies. These minimal container images typically do not contain a shell or package manager.

Although the -dev container image variants have similar security features as their more minimal versions, they include additional software that is typically not necessary in production environments. We recommend using multi-stage builds to copy artifacts from the -dev variant into a more minimal production image.

Need additional packages?

To improve security, Chainguard Containers include only essential dependencies. Need more packages? Chainguard customers can use Custom Assembly to add packages, either through the Console, chainctl, or API.

To use Custom Assembly in the Chainguard Console: navigate to the image you'd like to customize in your Organization's list of images, and click on the Customize image button at the top of the page.

Learn More

Refer to our Chainguard Containers documentation on Chainguard Academy. Chainguard also offers VMs and Libraries — contact us for access.

Trademarks

This software listing is packaged by Chainguard. The trademarks set forth in this offering are owned by their respective companies, and use of them does not imply any affiliation, sponsorship, or endorsement by such companies.

Licenses

Chainguard's container images contain software packages that are direct or transitive dependencies. The following licenses were found in the "latest" tag of this image:

  • ( Artistic-1.0-Perl

  • (MIT

  • Apache-2.0

  • Artistic-1.0-Perl

  • BSD-1-Clause

  • BSD-2-Clause

  • BSD-3-Clause

For a complete list of licenses, please refer to this Image's SBOM.

Software license agreement

Compliance

Chainguard Containers are SLSA Level 3 compliant with detailed metadata and documentation about how it was built. We generate build provenance and a Software Bill of Materials (SBOM) for each release, with complete visibility into the software supply chain.

SLSA compliance at Chainguard

This image helps reduce time and effort in establishing PCI DSS 4.0 compliance with low-to-no CVEs.

PCI DSS at Chainguard

Category
AI

The trusted source for open source

Talk to an expert
PrivacyTerms

Product

Chainguard ContainersChainguard LibrariesChainguard VMsChainguard OS PackagesChainguard ActionsChainguard Agent SkillsIntegrationsPricing
© 2026 Chainguard, Inc. All Rights Reserved.
Chainguard® and the Chainguard logo are registered trademarks of Chainguard, Inc. in the United States and/or other countries.
The other respective trademarks mentioned on this page are owned by the respective companies and use of them does not imply any affiliation or endorsement.