> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/automodel/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/automodel/_mcp/server.

# nemo_automodel.components.models.gpt_oss.rope_utils

## Module Contents

### Classes

| Name                                                                                      | Description |
| ----------------------------------------------------------------------------------------- | ----------- |
| [`RotaryEmbedding`](#nemo_automodel-components-models-gpt_oss-rope_utils-RotaryEmbedding) | -           |

### Functions

| Name                                                                                                          | Description                                       |
| ------------------------------------------------------------------------------------------------------------- | ------------------------------------------------- |
| [`apply_rotary_emb`](#nemo_automodel-components-models-gpt_oss-rope_utils-apply_rotary_emb)                   | Apply rotary embeddings to input tensor.          |
| [`apply_rotary_emb_qk`](#nemo_automodel-components-models-gpt_oss-rope_utils-apply_rotary_emb_qk)             | Apply rotary embeddings to query and key tensors. |
| [`position_ids_to_freqs_cis`](#nemo_automodel-components-models-gpt_oss-rope_utils-position_ids_to_freqs_cis) | -                                                 |

### API

```python
class nemo_automodel.components.models.gpt_oss.rope_utils.RotaryEmbedding(
    head_dim: int,
    base: int,
    dtype: torch.dtype,
    initial_context_length: int = 4096,
    scaling_factor: float = 1.0,
    ntk_alpha: float = 1.0,
    ntk_beta: float = 32.0,
    partial_rotary_factor: float = 1.0,
    device: torch.device | None = None
)
```

**Bases:** `Module`

**`rotary_dim`** `= int(head_dim * partial_rotary_factor)`

---

```python
nemo_automodel.components.models.gpt_oss.rope_utils.RotaryEmbedding._compute_concentration_and_inv_freq() -> torch.Tensor
```

See YaRN paper: [https://arxiv.org/abs/2309.00071](https://arxiv.org/abs/2309.00071)

Uses rotary\_dim instead of head\_dim to support partial rotary embeddings.

```python
nemo_automodel.components.models.gpt_oss.rope_utils.RotaryEmbedding._compute_cos_sin(
    num_tokens: int
)
```

```python
nemo_automodel.components.models.gpt_oss.rope_utils.RotaryEmbedding.forward(
    query: torch.Tensor,
    key: torch.Tensor
) -> tuple[torch.Tensor, torch.Tensor]
```

```python
nemo_automodel.components.models.gpt_oss.rope_utils.apply_rotary_emb(
    x: torch.Tensor,
    cos: torch.Tensor,
    sin: torch.Tensor
) -> torch.Tensor
```

Apply rotary embeddings to input tensor.

If cos/sin have fewer dimensions than x (due to partial\_rotary\_factor \< 1.0),
only the first rotary\_dim dimensions of x are rotated, and the rest are passed through.

**Parameters:**

**`x`** `torch.Tensor`

Input tensor (..., head\_dim)

---

**`cos`** `torch.Tensor`

Cosine tensor (..., rotary\_dim // 2)

---

**`sin`** `torch.Tensor`

Sine tensor (..., rotary\_dim // 2)

---

```python
nemo_automodel.components.models.gpt_oss.rope_utils.apply_rotary_emb_qk(
    q: torch.Tensor,
    k: torch.Tensor,
    freqs_cis: torch.Tensor,
    format: str = 'bshd',
    rope_fusion: bool = True,
    cu_seqlens: torch.Tensor | None = None,
    concentration: float | None = None,
    cp_size: int = 1,
    cp_rank: int = 0
) -> tuple[torch.Tensor, torch.Tensor]
```

Apply rotary embeddings to query and key tensors.

**Parameters:**

**`q`** `torch.Tensor`

Query tensor.

---

**`k`** `torch.Tensor`

Key tensor.

---

**`freqs_cis`** `torch.Tensor`

Frequency tensor. Format depends on rope\_fusion:

* If rope\_fusion=True: \[angles, angles] for TE fused rope
* If rope\_fusion=False: \[cos, sin] with concentration applied

---

**`format`** `str` — default: 'bshd'

QKV format ("bshd" or "thd").

---

**`rope_fusion`** `bool` — default: True

If True, use TE fused rope. If False, use non-fused rope.

---

**`cu_seqlens`** `torch.Tensor | None` — default: None

Cumulative sequence lengths for variable-length sequences.

---

**`cp_size`** `int` — default: 1

Context parallelism size.

---

**`cp_rank`** `int` — default: 0

Context parallelism rank.

---

**Returns:** `tuple[torch.Tensor, torch.Tensor]`

Tuple of (q, k) with rotary embeddings applied.

```python
nemo_automodel.components.models.gpt_oss.rope_utils.position_ids_to_freqs_cis(
    rotary_emb: nemo_automodel.components.models.gpt_oss.rope_utils.RotaryEmbedding,
    position_ids: torch.Tensor,
    qkv_format: str = 'bshd',
    for_fused_rope: bool = True,
    cp_size: int = 1
) -> torch.Tensor
```