Under reviewPreprintActionable Interpretability Workshop · CoLM 2026

LayerRoPE

Dynamic Depth-wise Magnitude & Angular Superposition

TL;DRHidden-state norms, variance grow by orders of magnitude with depth, a “curse of depth” treated as a pathology. We find it encodes an emergent depth code in the norm weights γ, and make it explicit. LayerRoPE is a RoPE-style depth encoding in the residual stream. It improves compute scaling, depth scaling up to 512 layers, and widens the learning-rate basin. Surprisingly, it does this by learning to amplify residual variance (and regulating Block I/O), rather than dampening it.

{{ l.label }}
{{ hero.arcLbl.label }}
LLaMA-2 7B, one ray per layer. MDS projection of the MLP-norm weights γ: their direction drifts and their magnitude shifts with depth. {{ hero.read }}
3.4×
less compute to reach Pre-Norm’s 1.3B loss
512
layers, and the only method with strong convergence and near monotonic improvement with depth
3–10×
lower learning-rate sensitivity, at every scale and on both Pre/Peri-Norm backbones
<0.02%
extra FLOPs, with fewer parameters than the norm weights it replaces
1 Finding

Norm weights encode depth

A Transformer exposes exactly one learned gain per layer directly on the residual path: the RMSNorm weight γ. That makes it the most direct place for any depth-conditional code the network may wish to learn.
Across 16 open-weight LLMs from 9 families (dense, mixture-of-experts, parallel-branch and hybrid linear-attention; Pre-, Peri- and Post-Norm), we find that the MLP-block γ carries exactly such a depth-conditioned code:

Magnitude is depth-conditioned. The mean and upper percentiles of γ shift systematically with depth, rising or falling with clear depth conditioning in most models.
Direction is depth-conditioned. We observe that the normalized gain \gamma_\ell / \lVert \gamma_\ell \rVert also drifts steadily with depth. MDS projections, shown below, order the layers along a coherent structured arc, showing a clear directional drift with the layers. This is not a rigid rotation, however. See1

Figure 1
An emergent depth code in the norm weights
Model
(a) Magnitude with depth{{ f1.meta }}
γ value, per layer
{{ t.label }}
{{ t.base }}
{{ vio.tip.title }}
{{ t.text }}
Layer index · dashed line: mean{{ vio.note }}
(b) Direction with depth{{ f1.meta }}
MDS projection of per-layer γ, one ray per layer
{{ l.label }}
{{ rot.tip.title }}
{{ t.text }}
firstlast layer{{ rot.note }}
Figure 1 · {{ f1.name }}. (a) Per-layer distribution of MLP-block RMSNorm weights γ: the magnitude shifts systematically with depth, in some cases near-monotonically. (b) MDS projection of the same per-layer γ vectors, reflected so the first layer points right: successive layers occupy neighbouring positions, tracing a depth-ordered arc that shows a steady directional drift. Angles are measured in the 2-D projection around the layers’ centre, not between the γ vectors themselves. Of the 16 models we study, only the MLP norms of Qwen2.5, Qwen3.5 and Gemma show no clear depth code; attention-block norms carry little of it.
2 Method

LayeRoPE: Along the depth axis

Where RoPE encodes a token’s position by rotating pairs of query and key features, LayerRoPE encodes a block’s layer index by rotating and scaling a shared gain. Instead of storing L independent vectors γℓ, every layer’s gain is generated from one shared γ ∈ ℝᵈ per site and a single depth-conditioned complex factor:

γℓ = γ ⊙ c(ℓ), c(ℓ, j) = e^r(ℓ) · e^(iθ(ℓ, j))(1)
Magnitude r(ℓ) = α + β log(ℓ+1) Rotation θ(ℓ, j) = exp(α + β log(ℓ+1)) · b₀^(−2j/d)

LayerRoPE can thus be simply interpreted as a depth-conditioned scaling and rotation of γ in the complex plane. Reading the real gain as d/2 complex pairs \tilde\gamma_j = \gamma_{2j} + i\,\gamma_{2j+1}, the magnitude scales all pairs equally through a log-linear schedule (the shared trajectory found above), while the rotation spreads a per-layer base angle across pairs with RoPE’s fixed spectrum, b₀ being the base frequency.

LayerRoPE sits at four sites (the read and the write of every attention and MLP block), so it regulates how each block reads from and writes to the residual stream, rather than the stream itself. Four scalars per site replace O(Ld) per-layer gain parameters with O(d).

Figure 2 · Interactive
The depth code LayerRoPE learns
{{ t.label }}
shared γ
θ
Layer ℓ{{ dial.L }}of 24 · 1.3B model
Magnitude{{ dial.mag }}{{ dial.rel }}
Rotation{{ dial.ang }}highest-frequency pair
Across the rotation spectrum · pair j turns by θ(ℓ)·b₀^(−2j/d), b₀ = 100
{{ s.ang }}
{{ s.f }}
layer 0drag, or use ← →layer 23
Figure 2. Learned schedules from the trained 1.3B models. Left: one feature pair of the shared γ (dotted, at ×1) mapped to its gain at layer ℓ: scaled by e^r(ℓ) (radius, log scale) and rotated by θ(ℓ); faint rays trace every other layer. Reads shrink with depth while writes grow. Right: the same layer across the rotation spectrum.

Essentially free

4d + 16parameters replace 2Ld per-layer gains: 90K fewer than Pre-Norm and 188K fewer than Peri-Norm at 1.3B.
≤ 0.013%extra training FLOPs over Pre-Norm, ≈ 0 over Peri-Norm. No matrix multiplications; LayerRoPE is input-independent and computed once per forward pass.
One configfor every scale: depth slopes initialised at a fixed β, base frequency. Only the learning rate is tuned, as for every baseline.
LAYERROPE.PY · REFERENCE PYTORCH SNIPPET
import torch, torch.nn as nn

class LayerRoPE(nn.Module):
    """Depth-conditioned gains for one site (e.g. the attention read).
    One shared γ and four scalars replace L per-layer gain vectors."""

    def __init__(self, d, n_layers, base=100.0, beta_init=-0.5):
        super().__init__()
        self.gamma = nn.Parameter(torch.ones(d))  # shared γ
        self.a_mag = nn.Parameter(torch.zeros(()))
        self.b_mag = nn.Parameter(torch.tensor(beta_init))
        self.a_rot = nn.Parameter(torch.zeros(()))
        self.b_rot = nn.Parameter(torch.tensor(beta_init))
        j = torch.arange(d // 2)
        self.register_buffer("freq", base ** (-2 * j / d))    # RoPE spectrum
        self.register_buffer("log_l", torch.log(torch.arange(1.0, n_layers + 1)))

    def forward(self):                  # input-independent: once per pass
        r = self.a_mag + self.b_mag * self.log_l                 
        theta = torch.exp(self.a_rot + self.b_rot * self.log_l) 
        ang = theta[:, None] * self.freq                         
        c = torch.polar(torch.exp(r)[:, None].expand_as(ang), ang)
        g = torch.view_as_complex(self.gamma.view(-1, 2))        
        return torch.view_as_real(g * c).flatten(1)  # One LayerRoPE module per site
sites = [LayerRoPE(d, n_layers) for _ in range(4)]
attn_read, attn_write, mlp_read, mlp_write = (site() for site in sites)

# normalize(h) = h / RMS(h), i.e. RMSNorm without its own learned weight:
for layer, block in enumerate(blocks):
    h = h + attn_write[layer] * block.attn(normalize(h) * attn_read[layer])
    h = h + mlp_write[layer] * block.mlp(normalize(h) * mlp_read[layer])
Illustrative re-implementation of Eq. (1); see the paper for the exact configuration.
3 Results

Better scaling, deeper networks, wider basins

Scaling ladder · 58M → 1.3B

Lowest loss at every scale

We train LLaMA-style models from 58M to 1.3B parameters on C4 at 4× the Chinchilla budget, up to 107B tokens. Every method, from Pre-, Post- and Peri-Norm to Layer-Norm Scaling (LNS) and LayerRoPE, gets its own learning rate, swept until the optimum is bracketed and extrapolated to 1.3B; LayerRoPE’s own hyperparameters stay fixed across scales.

LayerRoPE attains the lowest loss at every scale with the steepest compute exponent.

Figure 3
Compute scaling on a 4× Chinchilla ladder
{{ g.label }}{{ g.note }}
Final C4 validation loss
{{ t.label }}
{{ t.base }}{{ t.exp }}
{{ l.label }}
{{ cmp.tip.title }}
{{ t.text }}
Training compute (FLOPs) · top: model size
Figure 3. Every method at its own tuned learning rate. Curves are a joint fit L(C) = E + A_m C^{-\alpha_m} with a shared irreducible loss E = 1.81; the legend gives each exponent \alpha_m. Post-Norm is not fit: its best 554M run plateaus at > 4 nats and its 1.3B run fails to converge. 
Table 1
Zero-shot downstream accuracy
MethodARC-eHellaS.PIQABoolQOBQAWino.LMB.MMLUAvg ↑LMB ppl↓
{{ r.name }} {{ c.v }}
Table 1. {{ tbl.note }} Bold: best per column; underlined: second best.
Depth · 48 → 512 layers

Strong convergence all the way to 512 layers

To isolate the curse of depth, we fix a narrow backbone (width 128) and vary only the number of layers, from 48 to 512 (18M–111M parameters), adding DeepNet, designed for very deep Transformers, as a baseline. All baselines diverge with depth (Post-Norm from 192 layers, Peri-Norm and LNS degrade beyond 96, and Pre-Norm stops improving after 256). DeepNet improves with depth, but remains well above LayerRoPE throughout.

LayerRoPE reaches the lowest loss, with a margin that grows with depth; an exhaustive per-depth learning-rate sweep also confirms our finding.

Figure 4
Depth scaling
{{ g.label }}
Final validation loss
{{ t.label }}
{{ t.base }}
{{ l.label }}
{{ depth.tip.title }}
{{ t.text }}
Number of layers (log scale)
Figure 4. Fixed width 128, 48–512 layers (17.8M–111.1M parameters), 2.6B C4 tokens. Shared LR: one peak rate of 10⁻³ for every method; points are the mean of two seeds, bars span both (single seed at 512). Tuned LR: each method at its best rate from a per-depth sweep (one seed). The upper strip collects diverged runs.
Stability · 58M, 368M, 1.3B

A wider learning-rate basin, and later divergence

Sweeping the peak learning rate over two to three orders of magnitude, with and without LayerRoPE on both backbones, changes the loss curve in three ways:

  • 01Better optimum
  • 02Wider basin
  • 03Later divergence
Figure 5
Learning-rate basins
{{ g.label }}
Final validation loss
{{ t.label }}
{{ t.base }}{{ t.exp }}
{{ l.label }}
{{ lr.tip.title }}
{{ t.text }}
Peak learning rate
LR sensitivity · lower is better
Mean excess loss over the swept rates (Wortsman et al., 2024). ● with LayerRoPE, ● without; the number is the reduction in LR sensitivity.
{{ s.scale }}
{{ s.ratio }}
00.71.4
Figure 5. Final validation loss against peak learning rate with (solid) and without (dotted) LayerRoPE, at 58M, 368M and 1.3B scales. Rings identify the tuned optima LR for each scale. Runs with losses too high for the panel, or diverged, are omitted. Right: LR sensitivity drops 3–10× in every setting.
Transfer · no tuning

Beyond language models: looped models and ViTs

With no tuning of their recipes, LayerRoPE carries over to two architectures far from the language-model ladder: a looped latent language model and Vision Transformers.

Looped language models

Parcae · FineWeb-Edu · T = 8

Parcae applies a recurrent block of layers T times to a latent state between an input and an output block. We train it at 140M, 370M and 770M parameters with its official recipe, beside a non-looped GPT of the same size, and apply LayerRoPE at every layer (each layer’s depth index is its position in the network, independent of the loop iteration), with magnitude slopes initialised at zero. Parcae + LayerRoPE is the best model on nearly every metric: lower validation perplexity, markedly lower Lambada perplexity at 140M and 370M, and higher average downstream accuracy.

Table 2 · Figure 6
Parcae with and without LayerRoPE
140M11.2B tokens
370M29.6B tokens
770M61.6B tokens
ModelVal ppl ↓LMB ppl ↓Avg ↑Val ppl ↓LMB ppl ↓Avg ↑Val ppl ↓LMB ppl ↓Avg ↑
GPT20.28105.629.915.0137.240.612.8020.847.0
Parcae18.6681.233.014.2232.543.412.4918.650.1
+ LayerRoPE18.3667.334.914.0528.044.012.0918.450.6
{{ g.label }}
Validation loss over training · 770M
{{ t.label }}
{{ t.base }}
{{ l.label }}
{{ loop.tip.title }}
{{ t.text }}
Training tokens
Table 2 · Figure 6. Above: FineWeb-Edu validation perplexity, Lambada perplexity, and average downstream accuracy (ARC-Easy, HellaSwag 0- and 10-shot, PIQA, Lambada, SQuAD, CoQA, BIG-bench QA WikiData; three few-shot seeds). GPT is a non-looped baseline with the same parameter count and recipe. Left: validation loss over training at 770M, omitting the first ~1B tokens; hover to compare.

Vision Transformers

ImageNet-1k · DeiT recipe

We train ViT-T and ViT-S on ImageNet-1k for 300 epochs with the DeiT recipe and apply LayerRoPE naively at every block, conditioning the gains γ of the pre-norm LayerNorms while keeping their centring and bias, with the magnitude slope initialised at zero. LayerRoPE improves both ViTs on every metric.

Table 3
Vision Transformers on ImageNet-1k
ViT-TDeiT · 300 epochs
MetricViT-T+ LayerRoPEΔ
Top-1 (%)72.2173.08+0.87
Top-5 (%)91.2691.62+0.36
Val. loss1.2111.171−0.040
ViT-SDeiT · 300 epochs
MetricViT-S+ LayerRoPEΔ
Top-1 (%)79.8080.38+0.58
Top-5 (%)94.8995.26+0.37
Val. loss0.8970.865−0.032
Table 3. Each metric is the best over training. Δ is the change from adding LayerRoPE.
4 Mechanism

A wider stream, not a narrower one

Prior remedies such as Layer-Norm Scaling and Peri-Norm damp the residual stream directly. LayerRoPE’s learned schedule does the opposite.

LayerRoPE does not shrink the residual pipe; it widens it, while regulating the blocks that read from and write to it.

Figure 7
Residual-stream variance through a 1.3B model
{{ t.base }}
{{ l.label }}
{{ vr.tip.title }}
{{ t.text }}
Layer (24 layers)
Figure 7. Variance of the residual stream after each layer, on 512 C4 validation samples. Ribbon height ∝ Var1/3\mathrm{Var}^{1/3}, color ∝ log⁡Var\log \mathrm{Var}, on one shared scale across both views; final values at right.

How? LayerRoPE lets the residual stream grow in variance and norm, and instead regulates it at the inputs and outputs of each computational block.
Measuring just the learned magnitude gains, we see:

Read

LayerRoPE learns to attenuate what each block reads from the residual stream.

Write

LayerRoPE learns to amplify each block's write backs into the residual stream.

Figure 8
Read less, write more: learned multipliers at 1.3B
Pre-Norm + LayerRoPE Peri-Norm + LayerRoPE Layer-Norm Scaling (fixed)
Read · block input multiplier
{{ t.label }}
{{ t.base }}
{{ l.label }}
Write · residual output multiplier
{{ t.label }}
{{ t.base }}
{{ l.label }}
{{ mult.roTitle }} {{ o.t }}
Figure 8. Learned magnitude multipliers e^α (ℓ+1)^β at each site (log scale), labelled with the learned slope β, against Layer-Norm Scaling’s fixed input multiplier. Both slopes start at −0.5; the write slopes turn positive. Hover either panel to compare a layer.

LayerRoPE’s behaviour, of residual-stream amplification with regulation of each block’s I/O, is therefore a learned property.

Takeaway

Regulate the blocks, not the stream.

Learning stability & depth scaling, our results suggest, depend less on how large the residual stream grows than on whether each block can scale what it reads and writes to its depth.

We find that LayerRoPE’s depth-conditioned I/O enables this regulation of the computational block, while amplifying the variance of the residual pipe.

Citation

BibTeX
@misc{srivastava2026layerrope,
  title         = {LayerRoPE: Dynamic Depth-wise Magnitude \& Angular Superposition},
  author        = {Srivastava, Shikhar and Kanan, Christopher},
  year          = {2026},
  eprint        = {2610.09179},
  archivePrefix = {arXiv},
  url           = {https://arxiv.org/abs/2610.09179}
}
Acknowledgements

This work was partly supported by NSF EFRI Award 2317706 and NSF CAREER Awards 2047556/2326491. The views and conclusions contained herein are those of the authors and should not be interpreted as representing any sponsor’s official policies or endorsements. We gratefully acknowledge use of the research computing resources of the Empire AI Consortium, Inc, with support from Empire State Development of the State of New York, the Simons Foundation, and the Secunda Family Foundation. This work was also supported in part by compute credits provided by Modal Labs through the Modal for Academics program.