Advanced Features#
Guides for Megatron Core training and inference features.
- CUDA Graph
- Fine-Grained Activation Offloading
- Mixture of Experts
- Megatron Core MoE
- Megatron-FSDP
- Distributed Optimizer
- Weighted Checkpoint Merge
- Optimizer CPU Offload
- MoE Paged Stash
- Tokenizers
- Megatron Energon
- Megatron RL
- Megatron Core Inference User Guide
- Table of Contents
- What Megatron Inference Is For
- Supported Model Architectures
- Supported Features
- Basic Usage: The High-Level API
- Async Scheduling
- Streaming
- OpenAI-Compatible HTTP Server
- Weight Refit and Resharding for RL
- Multimodal (Vision-Language) Inference
- Disaggregated Prefill and Decode
- Customizing the Pipeline
- Examples Directory
- Known Limitations
- Roadmap and Future Work
- Additional Resources