Weight Refit: Choosing a Transport#
Weight refit copies updated policy weights into the rollout model. Choose the
topology first, then select one transport with
policy.generation.refit_transport.
Pick a Transport#
Topology |
|
Transport |
Use when |
|---|---|---|---|
|
|
CUDA IPC |
Policy and rollout workers share GPUs; vLLM uses ZMQ plus CUDA IPC handles and SGLang uses Ray CUDA IPC. |
|
|
MCore wake-reshard |
Megatron policy and generation share GPUs. |
|
|
NCCL broadcast |
vLLM and Megatron use the default full-weight collective; SGLang uses its own weight-update group. |
|
|
MCore native refit |
Megatron policy sends directly to Megatron generation through an MCore copy service. |
|
|
NCCL reshard |
Provides the best performance for large models (>100s B models). |
|
|
Sparse delta over ZeroMQ |
The link is bandwidth-limited and workers can reach a relay over TCP. |
|
|
Sparse delta through S3 |
Workers communicate through shared object storage. |
|
|
NIXL checkpoint engine |
The cluster has a fast UCX/RDMA fabric for full-weight refit. |
null is the default. For non-colocated Megatron generation it selects the
packed collective; Megatron’s native refit must be selected explicitly with
mcore. The sparse transports read only refit_cfg.sparse; NIXL reads only
refit_cfg.nixl. Because one selector chooses the transport, sparse delta and
NIXL cannot both be active.
vLLM Reload API#
For non-colocated vLLM using the default NCCL transport, set
policy.generation.vllm_cfg.refit_with_reload_api: true to install refitted
weights through vLLM’s native reload_weights API. The flag is off by default.
Use this path when vLLM’s post-load processing is required after every refit,
for example flashinfer_trtllm MoE models where process_weights_after_loading
reshapes expert weight buffers.
Early measurements from this PR show that reload refit can be faster for Llama
8B and Qwen3 32B, while Qwen3 30B-A3B was about 44% slower. MXFP8 reload refit
was about 2x slower in the measured run and used higher peak memory. Leave
additional KV-cache headroom for MXFP8, for example by lowering
policy.generation.vllm_cfg.gpu_memory_utilization.
The reload API path currently supports only policy.generation.backend: vllm,
policy.generation.colocated.enabled: false, and
policy.generation.refit_transport: null. It is explicitly unsupported with
nccl_reshard, ModelOpt quantization (policy.generation.quant_cfg), sparse
refit transports, and checkpoint-engine/NIXL transports. Eagle/MTP draft
weights refitted from the trainer are not supported yet; use the legacy loader
for those cases.
Constraints#
Transport |
Generation backend |
Policy backend |
Quantization and MoE |
|---|---|---|---|
Colocated CUDA IPC |
vLLM or SGLang |
DTensor or Megatron |
Uses the generation backend’s standard loader. |
NCCL |
vLLM or Megatron |
DTensor or Megatron |
Uses the standard full-weight loader. |
NCCL + reload API |
vLLM |
DTensor or Megatron |
Default non-colocated NCCL only. Rejects colocated refit, |
MCore native refit |
Megatron |
Megatron |
Supports MCore’s configured copy-service backend. |
SGLang NCCL weight-update group |
SGLang |
Megatron |
Trainer rank 0 broadcasts finalized HF tensors to the engine leaders. |
NCCL reshard |
vLLM or Megatron |
Megatron |
Supports BF16 and the documented FP8/MXFP8 combinations; see the NCCL Reshard Refit design. |
Sparse delta |
vLLM |
Megatron |
BF16/FP16, unquantized rollout only. |
NIXL, full weights |
vLLM |
DTensor or Megatron |
Supports the standard full-weight FP8 loader. DTensor FP8 KV-cache scale transfer is not yet supported. |
NIXL, sharded experts |
vLLM |
DTensor or Megatron |
Unquantized BF16/FP16 Triton MoE only; FP8/MXFP8 and dynamic expert placement are rejected. |
Non-colocated SGLang generation is supported with a Megatron policy and
refit_transport: null; it creates its own NCCL weight-update group. The NIXL
restrictions are on the generation backend; both Megatron and DTensor policy
workers can send weights. Sparse delta is currently limited to GRPO. NIXL is
initialized by the GRPO and distillation setup paths; PPO currently requires
colocated generation.
Minimal Configuration#
Colocated vLLM and SGLang refit need no transport configuration:
policy:
generation:
colocated:
enabled: true
refit_transport: null
For non-colocated NCCL, change the topology and leave the selector unset:
policy:
generation:
colocated:
enabled: false
refit_transport: null
For native MCore refit, select it explicitly:
policy:
generation:
backend: megatron
refit_transport: mcore
mcore_generation_config:
refit_backend: nccl # gloo | nccl | nccl_m2n (nvshmem is broken; see #3646)
For NCCL reshard with Megatron policy training and vLLM generation:
policy:
generation:
colocated:
enabled: false
refit_transport: nccl_reshard
For sparse delta, select one data plane and configure its scope:
policy:
generation:
colocated:
enabled: false
refit_transport: vllm_zmq_sparse # or vllm_s3_sparse
refit_cfg:
sparse:
delta_compression:
encoding: xor
storage:
s3_bucket: null # required for vllm_s3_sparse
For NIXL, select the checkpoint engine and configure its scope:
policy:
generation:
colocated:
enabled: false
refit_transport: nixl
refit_cfg:
nixl:
update_weights_bucket_memory_ratio: 0.05
device: cuda
backend_name: UCX
release_after_refit: false
shard_expert_weights: false
Learn More#
NCCL Reshard Refit describes its requirements, architecture, and shard-to-shard transfer.
Sparse Delta Refit explains baseline, compression, ZeroMQ, and S3 behavior.
Checkpoint-Engine Refit covers NIXL setup, performance tuning, FP8, sharded experts, and fault tolerance.
Checkpoint Engines describes the checkpoint-engine protocol and implementation.
Training and Generation Backends summarizes backend compatibility.