Hy3-preview

View as Markdown

Hy3-preview is a 295B Mixture-of-Experts language model from Tencent. It features 80 transformer layers (layer 0 dense, layers 1-79 MoE), 192 routed experts plus 1 shared expert with top-8 sigmoid routing, Grouped Query Attention (64 Q / 8 KV heads), per-head QK RMSNorm, RoPE, and an e_score_correction_bias gate buffer for expert-load correction. It supports a 256K context window.

Set up NeMo AutoModel with the latest container or follow the installation instructions.

Fine-Tune Hy3-preview

From the repository root, run:

uv run automodel --nproc-per-node=8 examples/llm_finetune/hy_v3/hy3_preview_deepep.yaml

See the NeMo AutoModel Installation Guide and LLM Fine-Tuning Guide.

Choose a Workflow

GoalStart Here
Supervised fine-tuning (SFT) - Hy3-preview with DeepEPUse hy3_preview_deepep.yaml.

Model Reference

Model Architecture

PropertyValue
TaskText Generation (MoE)
ArchitectureHYV3ForCausalLM
Parameters295B total
Hugging Face Organizationtencent

Available Models

ModelHF ID
Hy3-previewtencent/Hy3-preview