Quantization-Aware RL (QARL)#
Quantization-Aware RL (QARL) integrates NVIDIA Model Optimizer (ModelOpt) into the NeMo RL training loop, enabling quantization-aware training and generation for both GRPO and on-policy distillation workflows. QARL automatically quantizes a standard model at initialization, maintains quantizer state (amax values) throughout training, and transfers quantized state to vLLM during weight refit. By default, vLLM generation uses fake-quantized modules. For NVFP4 W4A16 rollout experiments, NeMo RL can instead stream packed real-quant ModelOpt NVFP4 weights into vLLM.
Overview#
In a standard NeMo RL loop, model weights are trained in full precision and refitted into vLLM for generation. QARL applies quantization-aware modules so that both the policy forward pass and rollout generation exercise quantized weights and, depending on the recipe, quantized activations. The policy backward pass remains in full precision, using the straight-through estimator to propagate gradients through the quantization nodes.
There are two vLLM rollout modes:
Fake-quant rollout: vLLM receives folded full-precision weights and runs fake-quantized layers. This is the default when
policy.generation.quant_cfgis set.Real-quant rollout: vLLM is initialized with ModelOpt NVFP4 kernels and receives packed NVFP4 weights plus scale tensors during every refit. Enable this with
policy.generation.real_quant: true.
See Verified Configurations for the workflow + recipe combinations that have been empirically validated, and Supported Quantization Formats for the full set of available formats. W4A4 (NVFP4_DEFAULT_CFG) converges for on-policy distillation but has been observed to have convergence issues on GRPO; W4A16 (NVFP4 weights, native-dtype activations) works for GRPO.
Verified Configurations#
The following workflow + quantization recipe combinations have been validated end-to-end (Megatron training + NVFP4-quantized vLLM generation + held-out validation):
Workflow |
Quantization |
Recipe |
Status |
Example Config |
|---|---|---|---|---|
QA-Distillation |
W4A4 |
|
✅ Converges |
|
QA-GRPO |
W4A16 |
|
✅ Converges |
|
QA-GRPO |
W4A4 |
|
⚠️ Known convergence issue |
|
QA-Distillation |
W4A4 |
|
✅ Converges |
|
QA-GRPO |
W4A16 |
|
✅ Smoke tested on MoE |
|
QA-GRPO real quantization rollout |
W4A16 |
|
✅ Converges |
|
The nvfp4_a16.yaml custom YAML enables NVFP4 e2m1 weight quantization (with dynamic e4m3 micro-block scales) and leaves activations unquantized; weights are still exercised through both Megatron training and vLLM generation. The nvfp4_a16_mlp_only.yaml recipe restricts W4A16 to MLP weights for real-quant rollout.
ModelOpt Layer Spec Toggle#
For QARL configs, try setting policy.disable_modelopt_layer_spec=true first.
This keeps ModelOpt quantization enabled while using the standard Megatron layer
specs instead of ModelOpt’s custom layer specs. This is usually faster and works
for most models, but it is not guaranteed for every architecture or recipe. If
you encounter errors with the standard Megatron layer specs, leave it unset or
set it to false to exercise ModelOpt’s Megatron layer-spec path.
Quantization-Aware GRPO (QA-GRPO)#
Configuration#
The QA-GRPO config extends the standard Megatron GRPO config by adding quantization parameters. See Verified Configurations for the status of W4A4 vs W4A16 on GRPO.
# examples/modelopt/qa_grpo_llama8b_megatron.v2.yaml
defaults: "../configs/grpo_math_8B_megatron.yaml"
policy:
quant_cfg: "examples/modelopt/quant_configs/nvfp4_a16.yaml"
quant_calib_data: "cnn_dailymail"
quant_calib_size: 512
quant_batch_size: 1
quant_sequence_length: 2048
generation:
quant_cfg: "examples/modelopt/quant_configs/nvfp4_a16.yaml"
Running QA-GRPO#
Single node (8 GPUs):
uv run examples/run_grpo.py \
--config examples/modelopt/qa_grpo_llama8b_megatron.v2.yaml \
policy.model_name=meta-llama/Llama-3.1-8B-Instruct
Via Slurm:
COMMAND="uv run examples/run_grpo.py \
--config examples/modelopt/qa_grpo_llama8b_megatron.v2.yaml \
policy.model_name=meta-llama/Llama-3.1-8B-Instruct \
checkpointing.checkpoint_dir=results/qa_grpo" \
CONTAINER=YOUR_CONTAINER \
MOUNTS="$PWD:$PWD" \
sbatch \
--nodes=1 \
--account=YOUR_ACCOUNT \
--job-name=qa-grpo \
--partition=YOUR_PARTITION \
--time=4:0:0 \
--gres=gpu:8 \
ray.sub
Real-Quant NVFP4 Rollout (W4A16)#
Real-quant rollout is intended for checking the deployment-style vLLM path during RL, not only the fake-quant training path. With policy.generation.real_quant: true, the Megatron policy worker exports ModelOpt QAT weights as packed NVFP4 tensors during refit, and the vLLM worker loads them into ModelOpt NVFP4 layers. This exercises vLLM’s real FP4 kernel path during rollout while the policy training worker remains a QAT model.
This path is validated for W4A16.
Minimal Configuration#
Start from a Megatron GRPO config and add the ModelOpt weight-only recipe to both the policy and generation sections:
policy:
quant_cfg: examples/modelopt/quant_configs/nvfp4_a16_mlp_only.yaml
generation:
backend: vllm
quant_cfg: examples/modelopt/quant_configs/nvfp4_a16_mlp_only.yaml
real_quant: true
The ready-to-run 2-node DAPO long-context recipe is:
examples/configs/recipes/llm/grpo-qwen3-8b-base-dapo-2n8g-long-megatron-qa-nvfp4-w4a16.yaml
For a BF16 baseline, copy the recipe, remove policy.quant_cfg,
policy.generation.quant_cfg, and policy.generation.real_quant, and use distinct
checkpoint and log directories.
Running the Example#
From the repository root inside the NeMo RL container:
uv run --extra mcore --extra modelopt --extra vllm \
examples/run_grpo.py \
--config examples/configs/recipes/llm/grpo-qwen3-8b-base-dapo-2n8g-long-megatron-qa-nvfp4-w4a16.yaml
For Slurm, wrap the same command in ray.sub as shown in Running QA-GRPO. Keep the W4A16 and BF16 runs separate and use distinct checkpoint directories.
Checkpoints and Fresh Starts#
Real-quant rollout is sensitive to stale Megatron conversion checkpoints. For a clean first-step comparison, move aside or remove both the training checkpoint and the converted Megatron checkpoint before launching:
# Training checkpoint: matches `checkpointing.checkpoint_dir` in your config.
mv checkpoints/grpo-qwen3-8b-base-dapo-2n-long-w4a16 checkpoints/grpo-qwen3-8b-base-dapo-2n-long-w4a16.old
# Converted Megatron checkpoint: under `$NRL_MEGATRON_CHECKPOINT_DIR` if set,
# else `$HF_HOME/nemo_rl` or `~/.cache/huggingface/nemo_rl`. The subdirectory is
# named after the HF model; see "Megatron Checkpoint Directory" below.
MEGATRON_CKPT_ROOT="${NRL_MEGATRON_CHECKPOINT_DIR:-${HF_HOME:-$HOME/.cache/huggingface}/nemo_rl}"
mv "$MEGATRON_CKPT_ROOT/<hf-model-subdir>" "$MEGATRON_CKPT_ROOT/<hf-model-subdir>.old"
If NRL_MEGATRON_CHECKPOINT_DIR is set, clear the subdirectory used by the run. On first startup, the log should show that iteration 0 was saved or loaded from a freshly generated conversion checkpoint.
For long runs on queues with short wall times, enable periodic checkpointing and submit dependency jobs with afterany so the next job can resume from the checkpoint written by the previous job.
Log Checks#
A healthy W4A16 real-rollout run should include these lines or equivalent vLLM logs:
quantization=modelopt_fp4
Using NvFp4LinearBackend.MARLIN for NVFP4 GEMM
MegatronQuantPolicyWorker[rank=0]: Packed ... groups of tensors
It should not include:
Using rollout logprobs
negative scales
CUDA error: invalid argument
For an initial sanity check, compare the first Generation KL Error with the BF16 baseline. They should be close on step 1. A substantially larger W4A16 first-step KL usually means the rollout model does not match the policy model, the run reused a stale checkpoint, or the real-quant export/refit path did not load the expected tensors.
Troubleshooting#
Symptom |
Likely Cause |
Action |
|---|---|---|
vLLM does not log |
|
Check the YAML under |
|
The run is bypassing policy/reference logprob computation |
Do not use rollout logprobs for real-quant validation |
First-step W4A16 |
Stale converted Megatron checkpoint or refit/export mismatch |
Clear checkpoints and rerun; confirm packed tensors are streamed |
|
Invalid or stale NVFP4 scale tensors reached vLLM |
Clear checkpoints and verify |
CUDA invalid argument during refit or generation |
vLLM consumed malformed packed tensors or stale IPC state |
Restart from a fresh job and inspect the first real-quant refit logs |
Quantization-Aware Distillation (On-Policy QAD)#
QAD combines on-policy distillation with quantization. The student model is quantized while the teacher remains in full precision, allowing the student to recover accuracy lost from quantization through knowledge distillation.
Configuration#
# examples/modelopt/qa_distillation_math_megatron.yaml
defaults: "../configs/distillation_math_megatron.yaml"
policy:
quant_cfg: "NVFP4_DEFAULT_CFG"
quant_calib_data: "cnn_dailymail"
quant_calib_size: 512
quant_batch_size: 1
quant_sequence_length: 2048
generation:
quant_cfg: "NVFP4_DEFAULT_CFG"
Running QAD#
uv run examples/run_distillation.py \
--config examples/modelopt/qa_distillation_math_megatron.yaml \
policy.model_name=Qwen/Qwen3-1.7B \
teacher.model_name=Qwen/Qwen3-1.7B
Quantization Parameters#
These parameters are added under the policy section:
Parameter |
Description |
|---|---|
|
Quantization config. Accepts: a built-in ModelOpt config name (e.g. |
|
Dataset name used for calibration. See the ModelOpt PTQ examples for supported datasets. |
|
Number of samples for the calibration pass |
|
Batch size during calibration |
|
Sequence length for calibration data |
The policy.generation.quant_cfg should match policy.quant_cfg to ensure consistent quantization between training and generation.
Generation-specific parameters are added under policy.generation:
Parameter |
Description |
|---|---|
|
Quantization config used by the vLLM generation worker. For QARL, this should normally match |
|
When |
|
Optional list of vLLM parameter name patterns that should stay in native dtype during real-quant rollout. If omitted, NeMo RL uses the default ModelOpt NVFP4 ignore set for sensitive layers such as attention and output heads. |
Megatron Checkpoint Directory#
On first run, the HF model is automatically converted to a Megatron checkpoint. By default, this checkpoint is saved under $HF_HOME/nemo_rl (or ~/.cache/huggingface/nemo_rl if HF_HOME is not set). To control where the converted checkpoint is stored — for example, to keep it alongside your experiment outputs — set the NRL_MEGATRON_CHECKPOINT_DIR environment variable:
export NRL_MEGATRON_CHECKPOINT_DIR="/path/to/your/megatron/checkpoints"
Differences from FP8 Training#
QARL (via ModelOpt) and NeMo RL’s built-in FP8 training (via TransformerEngine) serve different purposes:
TransformerEngine FP8 focuses on speeding up pre-training and fine-tuning using real quantization. It replaces linear layers with FP8-native implementations that compute directly in reduced precision for throughput gains.
ModelOpt QARL focuses on recovering accuracy under quantization using quantization-aware training. The policy forward pass uses quantized weights and, depending on the recipe, quantized activations while the backward pass uses full-precision gradients, so the model learns to be robust to quantization error. vLLM generation can run fake-quantized layers for W4A8/W4A16 recipes. W4A16 experiments can also use real ModelOpt NVFP4 kernels.
Supported Quantization Formats#
Weight quantization: per-tensor, per-channel, and block-wise formats are all supported. In fake-quant rollout, weights are pre-folded on the policy (Megatron) side before transfer to vLLM. In W4A16 real-quant rollout, weights are packed as NVFP4 tensors and streamed with their scale tensors.
Input (activation) quantization: only per-tensor is supported. The input quantizer amax is synced to vLLM as a per-tensor scalar.
Exporting Checkpoints#
After quantization-aware training, the Megatron checkpoint contains BF16 weights alongside quantization metadata (amax values, scales). To export a trained checkpoint to a fully quantized HuggingFace format (with real low-precision weights), use the Megatron-Bridge export tool. The exported checkpoint is ready for deployment with inference engines like vLLM or TensorRT-LLM.
From within the NeMo RL container:
cd /opt/nemo-rl
PYTHONPATH=$PWD/3rdparty/Megatron-Bridge-workspace/Megatron-Bridge/3rdparty/Megatron-LM:${PYTHONPATH:-} \
uv run --extra mcore --extra modelopt \
torchrun --nproc_per_node <pipeline-parallel-size> \
examples/modelopt/export_quantized_to_hf.py \
--hf-model-id <hf-model-name-or-path> \
--megatron-load-path <path-to-megatron-checkpoint>/policy/weights \
--export-dir <output-hf-directory> \
--tp 1 --pp <pipeline-parallel-size>
examples/modelopt/export_quantized_to_hf.pyis a thin wrapper aroundMegatron-Bridge/examples/quantization/export.py. All CLI flags pass through to the upstream script unchanged.--hf-model-idshould point to the original (pre-training) HuggingFace model so that the exporter knows the model architecture and tokenizer.The
PYTHONPATHprefix exposes Megatron-LM’smegatron.trainingto the bridge script.--tp 1is required: modelopt currently does not support TP>1 at export time. Training at TP>1 is fine; the bridge re-shards on load viamp_overrides.--ppcan be >1 for large models that don’t fit on one GPU.--nproc_per_nodemust equal--pp(since--tpis fixed at 1).
Limitations#
Generation: Currently only vLLM is supported for generation.
DTensor backend: Quantization support for the DTensor policy worker is not yet implemented.
Real-quant rollout: W4A16 real rollout is supported for dense vLLM ModelOpt NVFP4 layers.
Input quantization: Only per-tensor input (activation) quantization is supported.
Model support: Dense Transformer, MoE (Mixture of Experts), and hybrid MoE/Mamba models are supported on the Megatron policy + vLLM generation path when Megatron-Bridge and ModelOpt support the model architecture and quantization recipe. MoE/Mamba support is currently covered by smoke-tested example configs rather than broad convergence guarantees.