> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# Hy3-preview

> Use Hy3-preview with NeMo AutoModel for language model fine-tuning, with documented checkpoints, runnable recipes, setup guidance, and model reference details.

[Hy3-preview](https://huggingface.co/tencent/Hy3-preview) is a 295B Mixture-of-Experts language model from Tencent. It features 80 transformer layers (layer 0 dense, layers 1-79 MoE), 192 routed experts plus 1 shared expert with top-8 sigmoid routing, Grouped Query Attention (64 Q / 8 KV heads), per-head QK RMSNorm, RoPE, and an `e_score_correction_bias` gate buffer for expert-load correction. It supports a 256K context window.

Set up NeMo AutoModel with the [latest container](https://catalog.ngc.nvidia.com/orgs/nvidia/containers/nemo-automodel) or follow the [installation instructions](/get-started/installation).

## Fine-Tune Hy3-preview

From the repository root, run:

```bash
uv run automodel --nproc-per-node=8 examples/llm_finetune/hy_v3/hy3_preview_deepep.yaml
```

See the [NeMo AutoModel Installation Guide](/get-started/installation) and [LLM Fine-Tuning Guide](/recipes-e2e-examples/sft-peft).

## Choose a Workflow

| Goal                                                   | Start Here                                                                                                                               |
| ------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------- |
| Supervised fine-tuning (SFT) - Hy3-preview with DeepEP | Use [hy3\_preview\_deepep.yaml](https://github.com/NVIDIA-NeMo/Automodel/blob/main/examples/llm_finetune/hy_v3/hy3_preview_deepep.yaml). |

## Model Reference

### Model Architecture

| Property                  | Value                                     |
| ------------------------- | ----------------------------------------- |
| Task                      | Text Generation (MoE)                     |
| Architecture              | `HYV3ForCausalLM`                         |
| Parameters                | 295B total                                |
| Hugging Face Organization | [tencent](https://huggingface.co/tencent) |

### Available Models

| Model       | HF ID                                                               |
| ----------- | ------------------------------------------------------------------- |
| Hy3-preview | [`tencent/Hy3-preview`](https://huggingface.co/tencent/Hy3-preview) |

## Related Resources

* [LLM Fine-Tuning Guide](/recipes-e2e-examples/sft-peft)
* [Large MoE Fine-Tuning Guide](/recipes-e2e-examples/large-moe-fine-tuning)