nemo_rl.weight_sync.factory#
Factory for creating WeightSynchronizer instances.
Selects the appropriate weight synchronizer based on the deployment topology (colocated vs. non-colocated) and the generation backend:
Megatron -> Megatron reshard synchronizer
vLLM colocated -> IPC (ZMQ + CUDA IPC handles)
vLLM non-colocated -> NCCL collective
SGLang colocated -> Ray CUDA-IPC bucket transfer
SGLang non-colocated -> NCCL broadcast over SGLang’s own weight-update group
Module Contents#
Functions#
Create the appropriate WeightSynchronizer for the given deployment. |
API#
- nemo_rl.weight_sync.factory.create_weight_synchronizer(
- policy: Any,
- generation: Any,
- generation_backend: str,
- colocated: bool,
- train_cluster: Optional[Any] = None,
- inference_cluster: Optional[Any] = None,
- refit_buffer_size_gb: Optional[float | int] = None,
- refit_timeout_s: Optional[float] = None,
Create the appropriate WeightSynchronizer for the given deployment.
- Parameters:
policy – Policy object (ColocatablePolicyInterface).
generation – Generation object (GenerationInterface).
generation_backend – Name of the generation backend (“vllm”, “sglang”, “megatron”, or “dynamo”).
colocated – Whether policy and generation share the same GPUs.
train_cluster – RayVirtualCluster for training workers. Required for non-colocated deployments except SGLang, which owns its own group.
inference_cluster – RayVirtualCluster for inference workers. Same requirement as
train_cluster.refit_buffer_size_gb – Optional fixed buffer size for weight staging.
refit_timeout_s – Deadline for one refit collective, after which each participating worker aborts its own communicator so the controller can rebuild over the survivors. None disarms it.
- Returns:
A WeightSynchronizer instance appropriate for the deployment topology.
- Raises:
NotImplementedError – If the requested configuration is not supported.
ValueError – If required arguments are missing.