Hy-MT2-30B-A3B

View as Markdown

Hy-MT2-30B-A3B is Tencent’s translation Mixture-of-Experts language model with 30B total parameters and 3B activated per token. It features 48 transformer layers (layer 0 dense, layers 1-47 MoE), 128 routed experts plus 1 shared expert with top-8 sigmoid routing, Grouped Query Attention (32 Q / 4 KV heads), per-head QK RMSNorm, RoPE, and an in-forward fp32 upcast on the language-model head (enable_lm_head_fp32). It supports a 256K context window.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune Hy-MT2-30B-A3B

From the repository root, run:

uv run automodel --nproc-per-node=8 examples/llm_finetune/hy_mt2/hy_mt2_30b_a3b_sft.yaml

Refer to the NeMo AutoModel Installation Guide and LLM Fine-Tuning Guide.

Choose a Workflow

GoalStart Here
Supervised fine-tuning (SFT) - Hy-MT2-30B-A3B with FSDP2 + EP8 + fp32 LM headUse hy_mt2_30b_a3b_sft.yaml.

Model Reference

Model Architecture

PropertyValue
TaskText Generation (MoE, translation)
ArchitectureHyMT2ForCausalLM
Parameters30B total / 3B activated
Hugging Face Organizationtencent

Available Models

ModelHF ID
Hy-MT2-30B-A3Btencent/Hy-MT2-30B-A3B