Deterministic Training#
Deterministic training guarantees that two runs with identical inputs produce identical outputs at every step. Useful for debugging regressions and for reproducibility studies.
Pass --deterministic-mode to any Megatron training entry point (e.g. pretrain_gpt.py):
python pretrain_gpt.py \
--deterministic-mode \
<other args ...>
When enabled, Megatron applies the env vars and config overrides below via megatron.training.determinism.apply_determinism_to_args (called from validate_args).
Environment variables#
Each variable may be set by the launcher or left unset. If set, the value must be one that has been validated as deterministic — anything else fails hard with an assertion. If unset, apply_determinism_env fills the canonical default (except MAMBA_DETERMINISTIC, which the Mamba SSM helper auto-detects from torch.are_deterministic_algorithms_enabled()). Must be set before the first cuBLAS / Transformer Engine call — apply_determinism_to_args runs early in validate_args to guarantee this.
Variable |
Accepted values (or unset) |
Default filled if unset |
Reason |
|---|---|---|---|
|
subset of |
|
Conservative default — |
|
|
|
Forces Transformer Engine to use deterministic algorithms |
|
|
|
Deterministic cuBLAS workspace (both sizes are reproducible per NVIDIA docs; |
|
any string starting with |
(none — SSM auto-detects) |
Mamba SSM auto-follows |
If you override NCCL_ALGO, the value must be a subset of {Ring, CollnetDirect, CollnetChain, ^NVLS}. Tree is intentionally excluded: its intra-node chain reduction order is not user-controllable, and the inter-node tree topology can vary across runs without a pinned topology file, so it cannot be vouched for as bit-exact across stacks. ^NVLS is accepted (banning NVLS is a legitimate user choice on hardware that exposes it); the user is responsible for ensuring whatever NCCL falls back to is deterministic on their environment.
Config requirements#
Checked against the parsed args Namespace in apply_determinism_to_args. Incompatible options are rejected with an explicit error rather than silently flipped off — you must disable them yourself so the run matches the config you asked for:
Flag |
Behavior under |
|---|---|
|
Must be off — asserted (fused CE is non-deterministic); drop the flag yourself |
|
Must be off — asserted (the overlap path is not bit-exact); drop the flag yourself |
|
Set to |
Flash attention is permitted: Transformer Engine’s flash-attention backend is deterministic when NVTE_ALLOW_NONDETERMINISTIC_ALGO=0 (see the Transformer Engine docs).
Verifying determinism#
The bit-exact correctness suite lives at tests/unit_tests/determinism/correctness/. It parametrizes over model presets (GPT-like, Llama-like, Hybrid/Mamba) × parallelism cells (TP, PP, VPP, EP, FSDP, and composites) and asserts that two runs of the same configuration produce bit-identical outputs and gradients. FP8 / FP4 recipes (tensorwise, delayed, mxfp8, nvfp4) are covered by tests/unit_tests/determinism/correctness/test_fp8_determinism.py; the Blackwell-only recipes are capability-skipped on Hopper.
The cost of --deterministic-mode is measured outside pytest by an nsys-driven per-NVTX-range breakdown: tests/performance_tests/shell_test_utils/determinism/run_nsys_breakdown.sh wraps any training entry point (e.g. pretrain_gpt.py --profile) under nsys for a det-vs-nondet comparison, and tests/performance_tests/shell_test_utils/determinism/print_nsys_leaderboard.py joins the two CSVs into a side-by-side table. The CI invocation lives at tests/test_utils/recipes/h100/determinism-perf.yaml.