RoPE (Rotary Position Embedding)
RoPE (Rotary Position Embedding)
RoPE is a first-class backend operation (CUDNN_BACKEND_OPERATION_ROPE_FWD_DESCRIPTOR / _BWD_DESCRIPTOR), not an SDPA attribute. The backend op-graph contains RoPE(Q) + RoPE(K) + SDPA(Q', K', V); the SDPA fusion engine pattern-matches and runs the RoPE kernels around the attention kernel as a single execution unit.
The kernel takes raw freqs (angles), not pre-computed cos+sin. It computes sincosf(freqs) once per position into shared memory, amortized across heads. One tensor instead of two; matches the TE/Mcore convention where RotaryEmbedding already caches raw angles.
Math
Non-interleaved (halved) rotation. Head dim D is split into a leading nope segment of size D − rope_dim (scaled passthrough) and a trailing rope segment of size rope_dim:
where x_lo, x_hi are the two halves of the rope segment and α = output_scale. The same scale folds into both rotated halves and the passthrough so the entire output is uniformly scaled.
The backward op is rotation by −θ (transpose of the forward rotation) — implemented as the same kernel with sin_sign = −1. cos is unchanged.
Forward graph
User builds:
Q_rot and K_rot are real (user-bound) outputs, not workspace, so they survive into the backward graph as inputs to SDPA bwd.
Backward graph
dfreqs is not produced — frequencies are constant during training.
User builds:
Folding attn_scale into RoPE
Standard scaled attention multiplies QKᵀ by α = 1/√D inside the softmax. Because RoPE is linear, that scale can be hoisted into the RoPE Q output without changing the math:
The fwd kernel applies output_scale uniformly to both rotated halves and the nope-dim passthrough so every position of the Q output carries the α factor. SDPA then runs with attn_scale = 1.
Backward: where α actually flows
This is subtle. Standard scaled SDPA bwd has α between the softmax bwd and the matmul bwds:
α appears explicitly in both dQ' and dK'.
In the folded case (Q_sdpa = α·Q', K_sdpa = K', attn_scale = 1), S = Q_sdpa · K_sdpa^T has the same value (α moved across the dot product), so P, O, dS, dV are unchanged. But the engine no longer multiplies by α between softmax bwd and the matmuls:
So K is already correct out of SDPA bwd; only Q needs RoPE_bwd to re-introduce the α factor.
General rule
The fold setting on each side is the same in fwd and bwd. The bwd correctness then drops out of the chain rule mechanically: whatever output_scale you used on RoPE_fwd_X, use the same on RoPE_bwd_X.
Why the asymmetry of the Q-only row “just works” for K bwd: the SDPA bwd computes dK_sdpa = dS^T · Q_sdpa, and Q_sdpa already carries α — so K’s gradient gets α “for free” without an explicit scale on RoPE_bwd_K.
Why bother
One fewer pass over Q/K data: the α multiply happens during the same load that does the rotation, instead of being a separate op. Also unblocks future TE-block fusion where attention scaling needs to be a property of the Q/K producer, not the matmul consumer.
Folding inv_ln2 (extension)
The same fold knob can absorb the 1/ln(2) constant the softmax kernel uses to bridge exp ↔ exp2. GPU hardware only has exp2, so the kernel internally does:
— a per-element inv_ln2 multiply against the score matrix S (the largest intermediate tensor in attention). If the producer pre-scales by inv_ln2 instead, the kernel can use exp2 directly. The chain reasoning:
— same correct softmax, but the multiply moved from O(BHS_qS_{kv}) (per-S-cell) to O(BHS_qD) (per-Q-element) and is amortized into the RoPE load.
Note: softmax is not scale-invariant — softmax(c·x) ≠ softmax(x) in general. So attn_scale = 1/√D stays mandatory (it controls the softmax shape). inv_ln2 is different: it’s a base-conversion artifact, removable only because the kernel switches exp ↔ exp2 to match.
Bwd contract
The trick has two halves that move together. The kernel needs an softmax_input_is_log2 flag that toggles both:
- Fwd: skip the internal
S *= inv_ln2; runexp_2(S - \max)directly. - Bwd: emit
dS_sdpa = ∂L/∂(kernel's S)— the gradient w.r.t. the already-log2-scaled S, not the natural-log S.
Math for the bwd half. With P_i = \exp_2(S_i - m) / Z, the chain rule gives
which propagates to
So the bwd output is bigger by a factor of ln 2 than today’s kernel. That extra factor cancels with the user’s output_scale = α · inv_ln2 on RoPE_bwd:
Same on K via Q_sdpa. Net effect: no change to the FE-side rule. RoPE_bwd uses the same output_scale as RoPE_fwd, whatever you folded into it. The kernel-internal ln 2 factor cancels with the inv_ln2 you put in output_scale, automatically, by the chain rule.
API
Python:
Support matrix
freqs is [S, 1, 1, rope_dim] in f32. For a partial-rotation config (e.g. DeepSeek-V3 MLA with qk_rope_head_dim = 64 inside D = 192), the user supplies the smaller freqs tensor and sets rope_dim = 64; the leading D − rope_dim positions become a scaled passthrough.