RiT:在表示空间中使用原生扩散变换器已足够
RiT: Vanilla Diffusion Transformers Suffice in Representation Space
本研究探讨预训练表示空间在流匹配学习中的优势。比较像素、SD-VAE与DINOv2特征后发现,尽管像素与DINOv2的内在维度相近,但DINOv2在几何统计特性(如有效秩、协方差条件等)上表现更优,使回归过程更稳定。基于此,我们提出了表示图像变换器(RiT),它使用冻结的DINOv2特征,通过x-prediction目标训练一个原生扩散变换器。在ImageNet 256×256生成任务上,RiT性能优于参数量更多的DiT^DH-XL模型,且生成的常微分方程仅需少量步骤即可高效求解。
这篇论文没发明新架构,但通过剖析DINOv2特征的统计属性,证明简单结构在表示空间也能做出SOTA,对做图像生成的人来说是个省钱省参数的好思路。
Abstract
Flow matching with -prediction—regressing the clean data point rather than the ambient velocity—is known to exploit low-dimensional manifold structure effectively in pixel space [18]. We ask whether a pretrained representation space, while containing a low-dimensional data manifold of comparable intrinsic dimensionality, offers a distribution more favorable for flow-matching learning. Comparing pixel, SD-VAE, and DINOv2 features along four geometric axes, we find that pixel and DINOv2 share nearly identical intrinsic dimensionalities (both ) yet DINOv2 exhibits higher effective rank, better covariance conditioning, lower excess kurtosis, and lower on-manifold interpolation error; SD-VAE latents are consistently intermediate, indicating that the advantage stems from representation-learning objectives rather than mere compression. These statistical properties render the flow-matching regression well-conditioned and remove the need for the specialized prediction heads or Riemannian transport used by prior DINOv2 diffusion methods. We propose the Representation Image Transformer (RiT): a vanilla Diffusion Transformer trained by -prediction on frozen DINOv2 features, augmented only by a dimension-aware noise schedule and joint [CLS]-patch modeling. On ImageNet , RiT attains FID 1.45 without guidance and 1.14 with classifier-free guidance, outperforming DiT-XL with fewer parameters (676M vs. 839M). The resulting ODE is efficiently solvable at coarse discretizations: with classifier-free guidance, Heun steps already reach FID 2.0 and steps reach 1.25, without distillation or consistency training. Code at https://github.com/lezhang7/RiT.
1 Introduction
Flow matching [19, 7] learns a velocity field that transports Gaussian noise to data along linear paths. When data concentrates near a low-dimensional manifold, -prediction—parameterizing the network to output the clean data point rather than the ambient-space velocity—places the regression target on that manifold, as demonstrated by JiT [18] in pixel space. A natural question is whether a pretrained representation space, while containing a data manifold of comparable intrinsic dimensionality, offers a distribution more favorable for learning the flow-matching velocity field.
Comparing pixel, SD-VAE [23], and DINOv2 [21] features along four geometric axes, we find that pixel and DINOv2 share nearly identical intrinsic dimensionalities (both ) yet embed this manifold differently relative to . The pixel manifold is anisotropic, has strongly non-Gaussian per-coordinate marginals, and admits linear chords that traverse low-density regions. DINOv2 features exhibit near-isotropic variance, near-Gaussian per-coordinate marginals [38], and approximately on-manifold linear interpolants. These are marginal properties, not joint ones: DINOv2 features still concentrate on a -dimensional manifold, but each coordinate’s transport toward is short and well-conditioned.
Section 2 quantifies these gaps: DINOv2 attains higher effective rank, better covariance conditioning, lower excess kurtosis, and lower on-manifold interpolation error than pixels. SD-VAE latents fall consistently between the two, indicating that the advantage arises from representation-learning objectives rather than compression alone. These distributional advantages coexist with a DINOv2-specific pathology at off-manifold intermediate states: per-token LayerNorm pins , so linear flow-matching paths traverse ambient regions the encoder never outputs, and the -target at such acquires a large radial component.
The prevailing response to this radial ambiguity has been architectural. RAE [44] handles it with a specialized wide prediction head (DDT [37]) atop -prediction, alongside a ViT decoder that maps DINOv2 features back to pixels. Concurrent work [16] calls this phenomenon geometric interference and replaces the Euclidean transport with Riemannian Flow Matching on the norm-concentration sphere. Both modifications add complexity to either the architecture or the transport path.
We take a target-side alternative: -prediction. Under this parameterization, the network regresses , which lies on the data manifold by construction, so the radial ambiguity is resolved at the network’s output (which targets on the manifold) rather than at its input (where remains off-manifold). The reparameterization itself is not new [18]; its effectiveness here stems from the combination with DINOv2’s isotropic per-coordinate variance and near-Gaussian marginals, which render the denoising regression well-conditioned enough for a vanilla DiT. We instantiate this combination as the Representation Image Transformer (RiT) (Section 3): a vanilla Diffusion Transformer trained by -prediction flow matching in representation space, augmented by a dimension-aware noise schedule and joint [CLS]-patch modeling. As DiT operates on the SD-VAE latent space, RiT operates on a representation space provided by a frozen encoder–decoder; we use RAE’s [44] frozen DINOv2 encoder and ViT decoder. RiT thus models the high-dimensional DINOv2 feature distribution directly, without adapting the encoder for generation. On ImageNet , RiT attains FID 1.45 without guidance and 1.14 with classifier-free guidance, outperforming DiT-XL with fewer parameters. The resulting ODE converges in few Heun steps, yielding 5-step FID 2.0 and 10-step FID 1.25 (guided) without distillation or consistency training (Section 4.3).
2 The Geometry of Representation Spaces for Flow Matching
The manifold hypothesis—that data concentrates on a low-dimensional surface—holds regardless of representation. What differs across representations is how favorably this manifold is positioned relative to , and thus whether transport paths are short and the ODE is efficiently solvable in few steps. We characterize each representation along four complementary axes—intrinsic dimensionality (the manifold’s true degrees of freedom), effective rank (how uniformly variance is spread across directions), marginal Gaussianity (per-coordinate similarity to ), and on-manifold linear interpolation (whether linear chords stay near data)—each predicting a distinct mechanism that makes flow matching easier or harder. We quantify these on three spaces over 10,000 ImageNet images: (i) raw pixels (, ), (ii) DINOv2-Base features (, ), and (iii) SD-VAE latents (, )—the pretrained VAE used as the latent space of latent diffusion models [23]. Pixels and DINOv2 share the same ambient dimensionality, enabling direct geometric comparison; the inclusion of SD-VAE isolates the effect of representation-learning training (exemplified by DINOv2’s SSL) from generic compression.
Intrinsic dimensionality is the manifold’s true degrees of freedom—the number of independent directions needed to describe the data after stripping away ambient redundancy. Two spaces with comparable intrinsic dimensionality face manifolds of comparable underlying complexity; any difference in flow-matching learning difficulty must therefore arise from how the manifold is positioned rather than from its size. We use the TwoNN estimator [8], which recovers by maximum likelihood from the ratio of second- to first-nearest-neighbor distances under local uniformity (Appendix C). Bootstrapping over 10 independent subsamples of 5,000 points gives for pixels and for DINOv2—nearly identical, with the 1-dimension gap well within the combined estimator standard deviation. Both spaces therefore share essentially the same underlying manifold dimensionality; DINOv2’s advantage, by elimination, lies in how that manifold is embedded relative to the noise.
| Metric | Pixel | SD-VAE | DINOv2 |
|---|---|---|---|
| Mean | 0.923 | 0.467 | 0.114 |
| Median | 0.958 | 0.228 | 0.083 |
| 0.0% | 74.2% | 98.7% | |
| 70.6% | 86.7% | 99.7% |
Effective rank quantifies how uniformly variance is distributed across principal directions. It equals when all variance concentrates in a single direction (a thin needle in ) and equals the ambient dimension when variance spreads perfectly evenly (an isotropic ball). Since the flow-matching source is itself isotropic, higher effective rank on the data side translates to shorter, more uniform transport paths from noise to data. Concretely, with the normalized PCA eigenvalues [24]. Figure 1(a) plots per-component variance (log scale) and its cumulative: pixel’s first 50 components capture of total variance versus for DINOv2. The effective ranks are , , and for pixels, SD-VAE, and DINOv2 respectively—a gap between pixels and DINOv2. DINOv2’s per-token LayerNorm further fixes by construction, so this high effective rank concentrates features near an approximately isotropic shell of radius containing the -dim data manifold [38, 16].
Optimization conditioning reflects whether different variance directions can be learned in parallel during training: a well-conditioned regression converges along all directions at comparable rates, while a poorly-conditioned one over-fits high-variance directions while starving low-variance ones. Flow matching at time is implicitly such a regression, with effective covariance interpolating between at (pure noise) and the data covariance at (clean data); ill-conditioned therefore propagates into late-schedule training. Concretely, under a local Gaussian approximation , Ahamed et al. [1] show the regression covariance is ; we use the standard condition number as the diagnostic. Figure 1(b) plots across : both spaces start at near (where ) and grow monotonically toward as . At (representative of late-schedule fine-grained data-fitting), pixel space reaches while DINOv2 stays at —a gap, enabling all variance components to be learned at comparable rates. The same distributional proximity to also tightens the posterior , shrinking the irreducible variance of the per-pair velocity target—a distinct mechanism contributing to the faster convergence in §4.
Marginal Gaussianity measures how close each individual coordinate’s 1D distribution is to a Gaussian. The source is Gaussian along every axis, so closer-to-Gaussian per-coordinate marginals on the data side keep each dimension’s transport from noise to data short and well-behaved. We use the per-dimension excess kurtosis , which is for a Gaussian, positive for heavier-than-Gaussian (outlier-prone) tails, and negative for lighter tails. As shown in Table 1 and Figure 2, DINOv2 dimensions are markedly more Gaussian: 98.7% satisfy (vs. 74.2% for SD-VAE and 0% for pixels), with median lower than pixels and lower than SD-VAE. This captures marginal behavior only; the interpolation experiment below probes the joint geometry.
On-manifold linear interpolation. The previous three axes summarize variance per-direction; the final axis probes the joint geometry. Flow matching transports samples along straight paths , so if linear chords between data points themselves wander off the manifold, intermediate states will too, leaving the velocity target poorly defined. Cross-class image interpolation makes this concrete: pixel interpolation produces ghosting artifacts characteristic of paths crossing low-density voids, while DINOv2 interpolation yields smooth semantic transitions (Figure 3). We quantify this via a round-trip reconstruction error (full procedure in Appendix C): each intermediate frame—whether obtained by pixel blending or by linear interpolation in DINOv2 space followed by RAE decoding—is passed through the same DINOv2 encoder–RAE-decoder pipeline [44], and the MSE versus the input measures off-manifold distance. Because both conditions traverse the identical pipeline, the encoder–decoder reconstruction bias is shared; the remaining gap isolates whether the frame lies on or off the image manifold. Pixel frames incur higher error than DINOv2 frames ( vs. ); Figure 1(c) shows DINOv2 remains close throughout while pixel stays uniformly off-manifold.
Summary. Pixel and DINOv2 share nearly identical intrinsic dimensionalities (both ) yet DINOv2 is far better suited to flow-matching learning: higher effective rank, better covariance conditioning, lower excess kurtosis, and lower on-manifold interpolation error; SD-VAE is consistently intermediate, indicating the advantage arises from representation-learning objectives rather than compression alone. These properties predict that DDT heads, Riemannian transports, and wider backbones are not required for competitive performance—a prediction we validate in Sections 3–4 with a vanilla DiT and -prediction.
3 RiT: A Vanilla DiT for Representation-Space Diffusion
Guided by the geometry of Section 2, we instantiate the Representation Image Transformer (RiT) (Figure 4): a vanilla DiT backbone (SwiGLU [26], RMSNorm [43], 2D-RoPE [32], QK-normalization [11], plus in-context class tokens following JiT [18]), trained with a recipe tailored to DINOv2 features. We reuse RAE’s [44] frozen DINOv2-with-Registers encoder [21] and ViT decoder to move between pixels and features. The encoder yields patch tokens and a [CLS] token ; both are projected, concatenated, and jointly attended, with separate linear heads predicting and . Full details in Appendix D.
Flow matching preliminaries. Flow matching [19, 7] learns a velocity field that transports noise to data along straight paths , with , pure noise and clean data; the path’s time derivative is . The standard -prediction objective trains a network conditioned on timestep and class to regress this velocity:
| (1) |
Generation integrates the learned ODE from to via an Euler or Heun solver.
3.1 -Prediction on Standardized Features
Element-wise standardization. DINOv2’s per-token LayerNorm pins within each token, but leaves the cross-dataset per-channel variance heterogeneous ( spread across channels, §4). Before diffusion, we therefore standardize both patch tokens and the [CLS] token to zero mean and unit variance per element: (analogously for ), using statistics precomputed on the training set. This diagonal preconditioner [1] reduces the condition number of the data covariance and relaxes the near-constant-norm constraint that LayerNorm imposes on raw DINOv2 features [16]. The inverse transform is applied before decoding. Henceforth we use to denote the standardized feature. We find this step is a prerequisite rather than an optimization: training on raw DINOv2 features diverges entirely (Table 3).
-Prediction. As established in §2, DINOv2’s norm concentration means linear flow-matching paths traverse ambient regions the encoder never outputs; at such off-manifold , the -target acquires a radial component orthogonal to the data manifold—the geometric interference phenomenon diagnosed by Kumar and Patel [16], who resolve it with Riemannian Flow Matching [2] using SLERP paths. Under -prediction, the network must fit this radial component and therefore spends capacity on the norm direction rather than the tangential (along-manifold) direction. We resolve the same problem more simply, by changing the output parameterization.
Setting to the standardized DINOv2 feature, -prediction [18] instead outputs directly, with predicted velocity ; plugging this into the -prediction loss (1) yields
| (2) |
| 80 ep. | 200 ep. | 400 ep. | |
|---|---|---|---|
| -pred | 3.17 | 2.11 | 1.86 |
| -pred | 2.63 | 1.89 | 1.70 |
which is equivalent to the -prediction loss up to a reweighting (Appendix B). The - and -forms thus coincide as loss functionals, but impose different learning problems on the network, because the output parameterization determines which function is actually fit. Under -prediction, the network must fit : a target that depends on the off-manifold , diverges as near , and spans the full ambient space. Under -prediction, the network fits : a target that lies on the low-dimensional data manifold by construction and does not depend on explicitly. The chord through off-manifold persists at the network input; at the output, the target is confined to the data manifold rather than spanning the full ambient space. This manifold-targeting property is not itself DINOv2-specific [18]; what is unique here is the combination with DINOv2’s isotropic per-coordinate variance and near-Gaussian marginals, which together bound the target and smooth its dependence on (§2), so a vanilla DiT suffices. Table 2 confirms this empirically: with the same architecture, encoder, and noise schedule, -prediction consistently outperforms -prediction.
3.2 Joint CLS–Patch Modeling
A unique advantage of operating in representation space is direct access to the [CLS] token—a global semantic summary that encodes category, layout, and appearance complementary to local patch content. In standard latent diffusion on VAE features such a global token is not part of the latent representation itself; in representation space, it is intrinsic. We model [CLS] jointly with patches in the same diffusion process: is projected, prepended to the patch sequence, and participates in bidirectional self-attention, aggregating spatial evidence into a global context and broadcasting refined guidance back to local tokens. A separate linear head produces the [CLS] prediction , yielding an auxiliary -prediction loss (written in velocity form for symmetry with Eq. 1; equivalent to up to the same reweighting of Appendix B) and total objective . During training, [CLS] noise is sampled independently from patch noise to avoid a variance collapse ([CLS] is a single vector while patch noise is ). At inference, we advance [CLS] and patches jointly with Heun + classifier-free guidance under (optionally separate) guidance scales; we use the same scale for both in all reported experiments, but the mechanism permits decoupling. Only patch tokens are decoded while [CLS] is discarded. We also couple the two noise streams at initialization via (coupled noise)—a minor but consistent improvement at convergence (§4.3).
3.3 Dimension-Aware Noise Schedule
The SNR of is , but the effective per-token SNR scales with the per-token dimension [13]: for a -dimensional token, the noise magnitude grows as while the signal stays at unit scale, so higher- tokens need lower (more noise) to reach the same relative corruption. A DINOv2-Small token has , the per-pixel dimension of , so a pixel-space schedule undertrains on noisy states. Following RAE [44] and SD3 [7], we apply the dimension-dependent time shift with to and set . This pushes the median from to ( lower median SNR, Figure 5); §4 shows this closes a FID gap over the pixel-space logit-normal baseline (3.17 1.44 at 800 epochs).
4 Experiments
Setup. RiT-XL has 28 layers, hidden dimension 1152, 16 attention heads, and FFN expansion ratio , totaling 676M parameters. We train with a frozen DINOv2-Small encoder () and a pretrained RAE decoder [44] on ImageNet, using 8 H200 GPUs (12 min per epoch). We evaluate with FID-50K using class-balanced sampling (50 images per class) following RAE [44]. Full hyperparameters are in Appendix D.
| 200 ep | 400 ep | 800 ep | |
| Default (ours) | 1.83 | 1.67 | 1.44 |
| Preconditioning | |||
| w/o standardization† | 387.4 | 362.1 | 344.8 |
| Noise schedule | |||
| Logit-normal | 4.45 | 4.11 | 3.17 |
| CLS loss weight | |||
| w/o CLS | 1.89 | 1.70 | 1.63 |
| 1.86 | 1.64 | 1.50 | |
| Encoder | |||
| DINOv2-Base | 2.20 | 1.78 | 1.56 |
4.1 Convergence and Efficiency
Figure 6 compares RiT-XL against baselines. Against RAE-XL (DINOv2-S) [44]—a -prediction DiT-XL with the same encoder, decoder, and parameter count (676M) as RiT-XL, isolating the §3 design choices—RiT-XL leads at every epoch and reaches FID 1.45 at 800 ep ( better than 1.87). At 100 ep, RiT-XL already matches the larger RAE-XL (DINOv2-B) baseline at 720 ep ( speedup); at 200 ep it matches RAE-XL (DINOv2-S) at 800 ep ( speedup). RiT also surpasses the 800-ep FID of representation-alignment methods REPA [41] and REG [39] within 20–200 epochs. Concurrent RJF [16] tackles the same DINOv2 radial ambiguity via Riemannian Flow Matching on the norm-concentration sphere; at the matched 80-ep budget shown in Figure 6, RiT-XL reaches FID 2.48 (DINOv2-S) versus RJF’s 3.62 (DINOv2-B).
4.2 Every Recipe Choice Is Necessary
Ablations use the full recipe unless noted, varying one factor at a time; we report ImageNet FID-50K without guidance at Heun 50 steps.
Element-wise standardization. Raw DINOv2 features have heterogeneous per-channel variances ( range across channels). Training on raw features diverges: the loss oscillates and FID stays at random-init level () throughout training.
Noise schedule. The time-shift closes a FID gap over the original JiT logit-normal schedule (3.17 1.44 at 800 ep), confirming that reallocating training density toward higher noise is critical when per-token dimensionality grows (here vs. pixel’s ).
CLS token. Without CLS modeling (), FID plateaus at 1.63; reaches 1.44. Attention visualization (Appendix I) shows [CLS] aggregates coarse scene cues in early layers, integrates object–context relations in middle layers, and broadcasts refined guidance back in late layers.
Encoder size. Despite half the feature dimensionality, DINOv2-Small consistently beats DINOv2-Base (1.44 vs. 1.56 at 800 ep). Model capacity is not the bottleneck here: DINOv2-B has twice the ambient dimensionality at the same , so the denoiser must regress over a higher-dimensional target without a corresponding gain in underlying structure. The Section 2 analysis on DINOv2-Base is thus a conservative characterization—the Small manifold used in main experiments is at least as favorable and the regression task is easier.
4.3 Efficient ODE Convergence Enables Few-Step Generation
The four geometric properties established in Section 2—high effective rank, well-conditioned covariance, near-Gaussian marginals, and on-manifold linear interpolants—jointly predict that the noise-to-data ODE should be efficiently solvable in few Heun steps on DINOv2 features. We verify this directly by measuring Heun-solver truncation error in pixel space and show that it translates into order-of-magnitude gains under tight sampling budgets—without any distillation or consistency training.
| w/o guidance (CFG , Heun steps) | w/ guidance (CFG , Heun steps) | |||||||
|---|---|---|---|---|---|---|---|---|
| Schedule | 5 | 10 | 25 | 50 | 5 | 10 | 25 | 50 |
| Uniform | 12.8 / 12.7 | 6.88 / 6.76 | 2.34 / 2.29 | 1.61 / 1.58 | 10.8 / 10.7 | 6.19 / 6.12 | 1.93 / 1.90 | 1.30 / 1.28 |
| EDM | 2.37 / 2.34 | 1.61 / 1.58 | 1.49 / 1.47 | 1.47 / 1.46 | 2.01 / 1.99 | 1.33 / 1.32 | 1.17 / 1.15 | 1.16 / 1.14 |
| Cosine | 5.96 / 5.86 | 2.16 / 2.12 | 1.57 / 1.56 | 1.48 / 1.45 | 5.69 / 5.63 | 1.91 / 1.88 | 1.29 / 1.28 | 1.19 / 1.18 |
| Power-2 | 2.41 / 2.39 | 1.74 / 1.72 | 1.50 / 1.48 | 1.47 / 1.44 | 1.99 / 1.98 | 1.48 / 1.46 | 1.18 / 1.16 | 1.16 / 1.15 |
| Log-SNR | 4.56 / 4.51 | 1.96 / 1.95 | 1.51 / 1.50 | 1.46 / 1.44 | 3.79 / 3.78 | 1.73 / 1.74 | 1.23 / 1.22 | 1.18 / 1.17 |
| Time-shift | 2.44 / 2.38 | 1.59 / 1.58 | 1.47 / 1.45 | 1.46 / 1.44 | 1.99 / 1.99 | 1.27 / 1.25 | 1.15 / 1.14 | 1.15 / 1.14 |
Pixel-space truncation error measurement. For each model we generate 128 trajectories per space from matched pairs, run Heun sampling at plus a reference , decode every endpoint to pixel space (RiT through the frozen RAE decoder), and measure the Frobenius distance . Because each model is compared against its own -step reference, the metric isolates ODE convergence speed from the absolute quality of either endpoint: a model that happens to converge to a poor fixed point is not artifactually rewarded. JiT uses the schedule reported in its own paper.
RiT’s truncation error decays steeper than JiT’s. Figure 8 shows the averages. RiT’s truncation drops from to (), while JiT’s drops only (). At (RiT’s default), RiT is within Frobenius units of its -step endpoint—consistent with Table 4’s finding that RiT’s FID of already matches the converged . JiT at the same remains at Frobenius units and its FID is still far from converged (Figure 7: JiT-H’s -NFE FID is ). Under the 2nd-order Heun error bound , the gap in decay rate translates into a correspondingly smaller effective curvature for DINOv2’s marginal flow. This is the empirical counterpart of the four geometric properties quantified in Section 2: higher effective rank shortens the transport, Gaussianity matches the source distribution, tighter posterior lowers velocity-target variance, and on-manifold interpolants keep in well-defined regions—each independently predicting a smoother, easier-to-integrate velocity field.
Few-step generation. Figure 7 isolates the role of representation space under matched NFE. At 10 NFE, pixel-space JiT-H yields FID 26.2 and DINOv2-space DiT-XL yields 3.29, while RiT-XL reaches 2.38—an order-of-magnitude improvement over pixel space and a clear gain over the DDT-equipped DINOv2 baseline; the same ordering holds at 20 NFE ( for JiT-H, for DiT-XL, for RiT-XL). Within RiT itself (Table 4), 5 Heun steps with a time-shift schedule reach FID 2.44 without guidance and 1.99 with guidance, 10 steps reach 1.59 and 1.25, and 25 steps already match full convergence—surpassing the majority of prior VAE-latent baselines in Table 5, which typically require sampling steps.
Sampling schedule ablation. Table 4 compares six ODE time-discretization schedules (formal definitions in Appendix G). At steps, all non-uniform schedules converge to similar FID (1.43–1.45 w/o guidance, 1.14–1.15 w/ guidance), confirming the ODE is well-approximated. At 5 steps, the three concentrated schedules (EDM, power-2, time-shift: FID ) outperform uniform spacing (12.7) by , as they allocate more evaluations to the high-noise region where the velocity field varies most rapidly. Coupled noise (§3.2, independent / coupled cells in Table 4) shifts FID by uniformly across schedules and step counts; we adopt it in our main results for its small consistent gain but do not regard it as an essential component.
From geometry to empirics. The three mechanisms isolated in Section 2 map directly onto RiT’s gains. (i) The better covariance conditioning ( vs. ) lets all variance directions train at comparable rates. (ii) The tighter posterior shrinks the irreducible target variance and contributes to the convergence speedup (Figure 6). (iii) High effective rank (shorter transport), near-Gaussian marginals (smoother sourcedata interpolation), and on-manifold interpolants (no void-crossing) yield low effective curvature, verified by the faster truncation-error decay (Figure 8) and the few-step regime (Figure 7). Pixel latents fail on all three and SD-VAE on at least two, so these mechanisms are empirically independent; their joint materialization on DINOv2 explains why a vanilla M DiT surpasses M DDT-equipped baselines both at convergence and under tight sampling budgets.
| Method | AE | Epochs | #Params | w/o guidance | w/ guidance | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| FID | IS | Prec. | Rec. | FID | IS | Prec. | Rec. | ||||
| Pixel Diffusion | |||||||||||
| ADM [6] | – | 400 | 554M | 10.94 | 101.0 | 0.69 | 0.63 | 3.94 | 215.8 | 0.83 | 0.53 |
| PixelFlow-XL [3] | – | 320 | 677M | – | – | – | – | 1.98 | 282.1 | 0.81 | 0.60 |
| PixNerd-XL [36] | – | 320 | 700M | – | – | – | – | 1.93 | 298.0 | 0.80 | 0.60 |
| PixelDiT-XL [42] | – | 320 | 797M | – | – | – | – | 1.61 | 292.7 | 0.78 | 0.64 |
| JiT-G [18] | – | 600 | 2B | – | – | – | – | 1.82 | 292.6 | 0.79 | 0.62 |
| Latent Diffusion | |||||||||||
| DiT-XL [22] | SD-VAE | 1400 | 675M | 9.62 | 121.5 | 0.67 | 0.67 | 2.27 | 278.2 | 0.83 | 0.57 |
| SiT-XL [20] | SD-VAE | 1400 | 675M | 8.61 | 131.7 | 0.68 | 0.67 | 2.06 | 270.3 | 0.82 | 0.59 |
| MaskDiT [45] | SD-VAE | 1600 | 675M | 5.69 | 177.9 | 0.74 | 0.60 | 2.28 | 276.6 | 0.80 | 0.61 |
| MDTv2-XL [9] | SD-VAE | 1080 | 675M | – | – | – | – | 1.58 | 314.7 | 0.79 | 0.65 |
| REPA-XL [41] | SD-VAE | 800 | 675M | 5.78 | 158.3 | 0.70 | 0.67 | 1.29 | 306.3 | 0.79 | 0.64 |
| LightningDiT [40] | VA-VAE | 800 | 675M | 2.17 | 205.6 | 0.77 | 0.65 | 1.35 | 295.3 | 0.79 | 0.65 |
| DDT-XL [37] | SD-VAE | 400 | 675M | 6.27 | 154.7 | 0.68 | 0.69 | 1.26 | 310.6 | 0.79 | 0.65 |
| SVG-XL [27] | SVGTok | 1400 | 675M | 3.36 | 181.2 | – | – | 1.92 | 264.9 | – | – |
| REG-XL [39] | SD-VAE | 800 | 675M | 1.80 | - | - | - | 1.36 | 299.4 | 0.77 | 0.66 |
| RAE-XL [44] | RAE-DINOv2-S | 800 | 676M | 1.87 | 209.7 | 0.80 | 0.63 | 1.41 | 309.4 | 0.80 | 0.63 |
| RAE-XL[44] | RAE-DINOv2-B | 800 | 839M | 1.51 | 242.9 | 0.79 | 0.63 | 1.16 | 261.0 | 0.77 | 0.67 |
| FAE-XL [10] | FAE-DINOv2-G | 800 | 675M | 1.48 | 239.8 | 0.81 | 0.63 | 1.29 | 268.0 | 0.80 | 0.64 |
| RiT-XL (ours) | RAE-DINOv2-S | 800 | 676M | 1.45 | 231.6 | 0.81 | 0.62 | 1.14 | 299.7 | 0.80 | 0.63 |
4.4 Comparison with Prior Methods
Table 5 reports the full comparison. RiT’s unguided FID of 1.45 already surpasses the guided FID of every method without a representation encoder (PixelDiT-XL 1.61, JiT-G 1.82, MDTv2-XL 1.58, DiT-XL 2.27, SiT-XL 2.06)—the well-structured semantic space captures the distribution so faithfully that CFG becomes far less necessary. Among representation-based methods, RiT achieves the best unguided FID (1.45 vs. DiT-XL 1.51, FAE-DINOv2-G 1.48, RAE-XL 1.87) and the best guided FID (1.14 vs. DiT-XL 1.28, FAE 1.29, REPA-XL 1.29). The margin over FAE is narrow at the unguided level (1.45 vs. 1.48), but obtained under a substantially simpler setup: FAE uses the largest DINOv2 variant (DINOv2-G, ) and jointly fine-tunes its encoder for generation, whereas RiT uses the smallest variant (DINOv2-S, ) with the encoder entirely frozen—indicating that the geometric advantages of Section 2 are already accessible off-the-shelf, without encoder co-adaptation. More fundamentally, FAE and RiT address different problems: FAE eases generation by adapting the encoder to produce a more diffusion-friendly latent space (compressing DINOv2-G’s features to a generation latent), whereas RiT directly tackles modeling the existing high-dimensional representation distribution without altering the encoder. The two contributions are therefore largely orthogonal—encoder-side adaptation (FAE) and denoiser-side recipes (RiT) could in principle compose.
RiT uses uniformly smaller components. The denoiser is M parameters ( smaller than DiT-XL’s M, no DDT head); the encoder is DINOv2-S (, the smallest DINOv2 variant), versus DiT-XL’s DINOv2-B () and FAE’s DINOv2-G (). RiT also attains the highest unguided Precision () at competitive Recall (), confirming that an off-the-shelf DINOv2 encoder and a vanilla backbone are sufficient when the representation distribution is favorable for flow matching.
5 Related Work
Diffusion and flow matching for images. Diffusion and flow matching [12, 30, 19, 7] underpin modern image generation. Latent diffusion [23] compresses images via a VAE; DiT/SiT [22, 20] replace the U-Net with transformers. A critical design choice is the prediction target—, , or . JiT [18] showed that -prediction substantially outperforms the alternatives in pixel space by placing the target on the low-dimensional data manifold. Our work extends this insight by characterizing how a pretrained representation space embeds that manifold differently relative to , and shows that DINOv2 is especially well-suited to -prediction with a vanilla backbone.
Leveraging representations for generation. VA-VAE, EQ-VAE, and Diffusability shape autoencoder training for diffusion-friendly latents [40, 15, 28]. REPA [41] adds a DINOv2 alignment loss that is active only during training. REG [39] addresses this train–inference gap by entangling a DINOv2 [CLS] into the SD-VAE trajectory; RiT differs in operating natively in DINOv2 space where [CLS] is intrinsic. Neither REPA nor REG analyzes the manifold geometry, and both retain -prediction on SD-VAE latents. RAE [44] replaces the VAE with a DINOv2 encoder but adopts -prediction and needs a DDT head for the ill-conditioned velocity field; concurrent RJF [16] instead uses Riemannian Flow Matching with SLERP paths on the norm-concentration sphere. We show -prediction with element-wise standardization suffices to model DINOv2 features with a vanilla DiT—no architectural modification, no Riemannian reformulation.
Few-step generation and distillation. Progressive distillation [25], consistency models [31, 29], and rectified flow [19] reduce sampling cost by training a dedicated few-step student or straightening the teacher’s trajectories. These are orthogonal to RiT’s contribution: RiT shows that the base model itself already reaches competitive few-step FID in a geometry-friendly representation space, without any distillation or consistency loss, and remains a natural teacher for such methods.
Toward unified understanding–generation. Unified vision models [5, 33, 4, 34, 35] typically maintain separate encoders for perception (CLIP/DINOv2) and synthesis (SD-VAE), with task-specific architectural components bolted onto the generative side. RiT’s ability to generate competitively in DINOv2 space suggests a cleaner alternative: a single semantic representation and a single vanilla Transformer backbone—no DDT head, no Riemannian reformulation, no representation-alignment loss—can serve both tasks. The training speedup at matched encoder (§4) and competitive few-step sampling (FID 2.0 at 5 Heun steps, 1.25 at 10 steps) further reduce the practical cost of attaching a generative head to an existing perception stack, making DINOv2-space RiT a promising base for unified pipelines in which the same features drive classification, retrieval, and synthesis.
6 Conclusion
We presented RiT, a vanilla DiT trained with -prediction on frozen DINOv2 features that achieves FID 1.45 without guidance and 1.14 with classifier-free guidance on ImageNet using fewer denoiser parameters than DiT-XL (676M vs. 839M) and the smallest DINOv2 variant (DINOv2-S, ); it further supports few-step generation (guided FID 2.0 at 5 Heun steps, 1.25 at 10 steps) without distillation or consistency training. The Section 2 analysis indicates that representation-space diffusion becomes architecturally simpler whenever the feature distribution satisfies the four geometric axes we identify—high effective rank, well-conditioned covariance, near-Gaussian marginals, and on-manifold linear interpolants. Our results argue for a target-side reformulation (-prediction) over architecture-side (DDT heads) or transport-side (Riemannian flow matching) solutions whenever the representation’s distributional geometry is already favorable.
Acknowledgments
We thank the Mila IDT team and their technical support for maintaining the Mila compute cluster. We also acknowledge the material support of NVIDIA in the form of computational resources. Throughout this project, Aishwarya Agrawal received support from the Canada CIFAR AI Chair award.
References
- Ahamed et al. [2026] Shadab Ahamed, Eshed Gal, Simon Ghyselincks, Md Shahriar Rahim Siddiqui, Moshe Eliasof, and Eldad Haber. Preconditioned score and flow matching. arXiv preprint arXiv:2603.02337, 2026.
- Chen and Lipman [2023] Ricky TQ Chen and Yaron Lipman. Flow matching on general geometries. arXiv preprint arXiv:2302.03660, 2023.
- Chen et al. [2025a] Shoufa Chen, Chongjian Ge, Shilong Zhang, Peize Sun, and Ping Luo. Pixelflow: Pixel-space generative models with flow. arXiv preprint arXiv:2504.07963, 2025a.
- Chen et al. [2025b] Xiaokang Chen, Zhiyu Wu, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, and Chong Ruan. Janus-pro: Unified multimodal understanding and generation with data and model scaling. arXiv preprint arXiv:2501.17811, 2025b.
- Deng et al. [2025] Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, et al. Emerging properties in unified multimodal pretraining. arXiv preprint arXiv:2505.14683, 2025.
- Dhariwal and Nichol [2021] Prafulla Dhariwal and Alexander Nichol. Diffusion models beat GANs on image synthesis. In NeurIPS, 2021.
- Esser et al. [2024] Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first international conference on machine learning, 2024.
- Facco et al. [2017] Elena Facco, Maria d’Errico, Alex Rodriguez, and Alessandro Laio. Estimating the intrinsic dimension of datasets by a minimal neighborhood information. Scientific reports, 7(1):12140, 2017.
- Gao et al. [2023] Shanghua Gao, Pan Zhou, Ming-Ming Cheng, and Shuicheng Yan. Mdtv2: Masked diffusion transformer is a strong image synthesizer. arXiv preprint arXiv:2303.14389, 2023.
- Gao et al. [2025] Yuan Gao, Chen Chen, Tianrong Chen, and Jiatao Gu. One layer is enough: Adapting pretrained visual encoders for image generation. arXiv preprint arXiv:2512.07829, 2025.
- Henry et al. [2020] Alex Henry, Prudhvi Raj Dachapally, Shubham Shantaram Pawar, and Yuxuan Chen. Query-key normalization for transformers. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4246–4253, 2020.
- Ho et al. [2020] Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
- Hoogeboom et al. [2023] Emiel Hoogeboom, Jonathan Heek, and Tim Salimans. simple diffusion: End-to-end diffusion for high resolution images. In International Conference on Machine Learning, pages 13213–13232. PMLR, 2023.
- Karras et al. [2022] Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. Advances in neural information processing systems, 35:26565–26577, 2022.
- Kouzelis et al. [2025] Theodoros Kouzelis, Ioannis Kakogeorgiou, Spyros Gidaris, and Nikos Komodakis. Eq-vae: Equivariance regularized latent space for improved generative image modeling. arXiv preprint arXiv:2502.09509, 2025.
- Kumar and Patel [2026] Amandeep Kumar and Vishal M Patel. Learning on the manifold: Unlocking standard diffusion transformers with representation encoders. arXiv preprint arXiv:2602.10099, 2026.
- Levina and Bickel [2004] Elizaveta Levina and Peter Bickel. Maximum likelihood estimation of intrinsic dimension. Advances in neural information processing systems, 17, 2004.
- Li and He [2025] Tianhong Li and Kaiming He. Back to basics: Let denoising generative models denoise. arXiv preprint arXiv:2511.13720, 2025.
- Liu et al. [2022] Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003, 2022.
- Ma et al. [2024] Nanye Ma, Mark Goldstein, Michael S Albergo, Nicholas M Boffi, Eric Vanden-Eijnden, and Saining Xie. Sit: Exploring flow and diffusion-based generative models with scalable interpolant transformers. In European Conference on Computer Vision, pages 23–40. Springer, 2024.
- Oquab et al. [2024] Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. Transactions on Machine Learning Research, 2024.
- Peebles and Xie [2023] William Peebles and Saining Xie. Scalable diffusion models with transformers, 2023. URL https://arxiv.org/abs/2212.09748.
- Rombach et al. [2022] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022.
- Roy and Vetterli [2007] Olivier Roy and Martin Vetterli. The effective rank: A measure of effective dimensionality. In 2007 15th European signal processing conference, pages 606–610. IEEE, 2007.
- Salimans and Ho [2022] Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. arXiv preprint arXiv:2202.00512, 2022.
- Shazeer [2020] Noam Shazeer. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
- Shi et al. [2025] Minglei Shi, Haolin Wang, Wenzhao Zheng, Ziyang Yuan, Xiaoshi Wu, Xintao Wang, Pengfei Wan, Jie Zhou, and Jiwen Lu. Latent diffusion model without variational autoencoder. arXiv preprint arXiv:2510.15301, 2025.
- Skorokhodov et al. [2025] Ivan Skorokhodov, Sharath Girish, Benran Hu, Willi Menapace, Yanyu Li, Rameen Abdal, Sergey Tulyakov, and Aliaksandr Siarohin. Improving the diffusability of autoencoders. arXiv preprint arXiv:2502.14831, 2025.
- Song and Dhariwal [2023] Yang Song and Prafulla Dhariwal. Improved techniques for training consistency models. arXiv preprint arXiv:2310.14189, 2023.
- Song et al. [2020] Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020.
- Song et al. [2023] Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. Consistency models. 2023.
- Su et al. [2024] Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
- Team [2024] Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818, 2024.
- Tong et al. [2025] Shengbang Tong, David Fan, Jiachen Li, Yunyang Xiong, Xinlei Chen, Koustuv Sinha, Michael Rabbat, Yann LeCun, Saining Xie, and Zhuang Liu. Metamorph: Multimodal understanding and generation via instruction tuning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 17001–17012, 2025.
- Tong et al. [2026] Shengbang Tong, Boyang Zheng, Ziteng Wang, Bingda Tang, Nanye Ma, Ellis Brown, Jihan Yang, Rob Fergus, Yann LeCun, and Saining Xie. Scaling text-to-image diffusion transformers with representation autoencoders. arXiv preprint arXiv:2601.16208, 2026.
- Wang et al. [2025a] Shuai Wang, Ziteng Gao, Chenhui Zhu, Weilin Huang, and Limin Wang. Pixnerd: Pixel neural field diffusion. arXiv preprint arXiv:2507.23268, 2025a.
- Wang et al. [2025b] Shuai Wang, Zhi Tian, Weilin Huang, and Limin Wang. Ddt: Decoupled diffusion transformer, 2025b.
- Wang and Isola [2020] Tongzhou Wang and Phillip Isola. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In International conference on machine learning, pages 9929–9939. PMLR, 2020.
- Wu et al. [2025] Ge Wu, Shen Zhang, Ruijing Shi, Shanghua Gao, Zhenyuan Chen, Lei Wang, Zhaowei Chen, Hongcheng Gao, Yao Tang, Jian Yang, et al. Representation entanglement for generation: Training diffusion transformers is much easier than you think. arXiv preprint arXiv:2507.01467, 2025.
- Yao et al. [2025] Jingfeng Yao, Bin Yang, and Xinggang Wang. Reconstruction vs. generation: Taming optimization dilemma in latent diffusion models. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 15703–15712, 2025.
- Yu et al. [2024] Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. Representation alignment for generation: Training diffusion transformers is easier than you think. arXiv preprint arXiv:2410.06940, 2024.
- Yu et al. [2025] Yongsheng Yu, Wei Xiong, Weili Nie, Yichen Sheng, Shiqiu Liu, Anima Anandkumar, and Arash Vahdat. Pixeldit: Pixel diffusion transformers for image generation. arXiv preprint arXiv:2511.20645, 2025.
- Zhang and Sennrich [2019] Biao Zhang and Rico Sennrich. Root mean square layer normalization. Advances in neural information processing systems, 32, 2019.
- Zheng et al. [2025] Boyang Zheng, Nanye Ma, Shengbang Tong, and Saining Xie. Diffusion transformers with representation autoencoders. arXiv preprint arXiv:2510.11690, 2025.
- Zheng et al. [2023] Hongkai Zheng, Weili Nie, Arash Vahdat, and Anima Anandkumar. Fast training of diffusion models with masked transformers. TMLR, 2023.
Appendix A Limitations
DINOv2 encoder bias. RiT inherits the inductive biases of the frozen DINOv2 encoder. DINOv2’s SSL objective emphasizes semantic content over photometric detail, and prior work has observed weaker feature resolution on fine textures, thin structures, and small objects. These biases propagate directly into what RiT can generate, since the RAE decoder operates on the same features. Joint encoder fine-tuning (as in FAE [10]) could mitigate this, at the cost of the simpler frozen-encoder setup we advocate.
Class conditioning and resolution. All RiT results are class-conditional on ImageNet at . We have not evaluated text-to-image generation, higher resolutions (e.g., or ), or non-image modalities. The Section 2 geometric analysis was measured on ImageNet images at DINOv2-Base and DINOv2-Small scales; whether the same four axes persist at larger model/data scales or under text conditioning is left to future work.
Local-Gaussian assumption in the analysis. The covariance-conditioning diagnostic (Section 2) rests on a local Gaussian approximation ; the other three geometric axes are assumption-light but still report aggregate scalars that cannot rule out adversarial pockets of the manifold where the favorable properties fail. The flow-matching results empirically corroborate the geometric claims without establishing each axis as individually necessary for the observed efficiency gains.
Appendix B Equivalence of Velocity Loss and Reweighted -Prediction Loss
We show that the velocity MSE loss with -prediction parameterization (Eq. 2) is equivalent to a reweighted -prediction loss. Given the forward process , the target velocity is:
| (3) |
where the second equality follows from substituting . Under -prediction, the network outputs , and the predicted velocity is . Substituting both into the velocity loss:
| (4) |
Thus the velocity loss equals the -prediction loss reweighted by , which upweights the loss at high (near clean data). The two losses are therefore equivalent as functionals; what differs is the network’s parameterization, which determines the function actually fit (§3.1): -prediction makes the direct network output, so its regression target always lies on the data manifold, whereas -prediction asks the network to produce the ambient velocity , which depends on the off-manifold and diverges as .
Appendix C Manifold Analysis Details
This appendix provides formal definitions and implementation details for the manifold analysis metrics used in Section 2. All experiments are conducted on 10,000 randomly sampled ImageNet training images.
PCA spectrum and effective rank.
We fit PCA with 512 components on the flattened feature vectors of each space. Let denote the eigenvalues of the sample covariance matrix and the normalized eigenvalues. The effective rank [24] is defined as
| (5) |
It equals 1 when all variance concentrates in a single direction (maximally anisotropic) and equals when variance is perfectly uniform (maximally isotropic). For flow matching, higher effective rank means the data distribution is closer to the isotropic Gaussian source, requiring a less complex velocity field.
Intrinsic dimensionality (TwoNN).
The TwoNN estimator [8] estimates intrinsic dimensionality from the ratio of first- and second-nearest-neighbor distances. For each point , let and be the distances to its nearest and second-nearest neighbor, and . Under the assumption that data is locally uniform on a -dimensional manifold, the MLE estimator is
| (6) |
where is the number of valid samples with . We subsample 5,000 points and compute pairwise Euclidean distances in chunks to control memory.
Robustness of at .
At such high ambient dimensionality, nearest-neighbor distances concentrate and any single-run intrinsic-dimension estimate can be noisy. We bootstrap TwoNN over 10 independent subsamples of 5,000 points, reporting the sample mean and standard deviation:
| Space | (mean std, 10 bootstraps) |
|---|---|
| Pixel | |
| DINOv2 |
The pixel–DINOv2 gap of dimension is substantially below the combined standard deviation (, ), so the two estimates are statistically indistinguishable. Larger- variants of the MLE estimator [17] are known to suffer from upward bias at high ambient dimension [8] and do not share this convergence property; we therefore rely on TwoNN as the primary estimator. DINOv2’s advantage is not in manifold dimensionality but in the global geometry characterized by the other three axes of Section 2.
Marginal Gaussianity (excess kurtosis).
For each dimension , the excess kurtosis is
| (7) |
where and are the per-dimension mean and standard deviation. A Gaussian distribution has ; positive values indicate heavier tails, negative values indicate lighter tails. We report the median of across all dimensions as a scalar summary.
On-manifold interpolation score.
For each intermediate frame along a linear interpolation path, we measure how well it stays on the natural image manifold via reconstruction error under a unified pipeline: image encode(DINOv2) decode MSE versus the input image. This pipeline is applied identically to both pixel-space and DINOv2-space interpolation frames:
- •
Pixel interpolation: The intermediate frame is passed through encodedecode. Ghosting artifacts (off-manifold) produce high MSE because the encoder projects them to the nearest valid representation.
- •
DINOv2 interpolation: The intermediate representation is first decoded to a pixel-space frame, which is then passed through the same encodedecode pipeline.
By measuring both through the identical pipeline, the comparison is unbiased: the only difference is whether the frame was produced by pixel blending or DINOv2 latent interpolation. We average over 100 same-class pairs with 11 interpolation steps each.
Sanity check: encoder round-trip on DINOv2 interpolants.
To rule out a trivial explanation that the DINOv2 interpolation pipeline enjoys near-zero reconstruction error by construction, we additionally measure whether and are close in feature space. If the encoder merely re-projected arbitrary inputs to their nearest valid representation, the pixel-versus-DINOv2 MSE gap could be artifactually large. We report the average cosine similarity between and its re-encoded counterpart, which remains high throughout the interpolation path; the gap in Figure 1(c) therefore reflects the off-manifold position of pixel blends rather than a baseline asymmetry in how the encoder treats each input.
Appendix D Architecture and Hyperparameters
D.1 Model Architecture
RiT uses a modernized DiT backbone (following JiT [18]/LightningDiT [40]) operating on the spatial grid of DINOv2 features. Our main experiments use DINOv2-Small (); we also report DINOv2-Base () results in encoder ablations. Table 6 summarizes the model variants.
| Model | Layers | Hidden dim | Heads | FFN dim | Params |
|---|---|---|---|---|---|
| RiT-L | 24 | 1024 | 16 | 4096 | 458M |
| RiT-XL | 28 | 1152 | 16 | 4608 | 676M |
Each DiT block consists of:
- 1.
adaLN modulation: timestep and class embeddings are summed () and projected to per-layer scale/shift parameters via a shared SiLU–Linear layer.
- 2.
Multi-head self-attention with QK-normalization (RMSNorm on Q and K before attention) and VisionRoPE for 2D spatial position encoding. [CLS] and register tokens are excluded from RoPE.
- 3.
SwiGLU FFN: .
The final layer uses adaLN-modulated RMSNorm followed by a linear projection to output channels (384 for DINOv2-Small, 768 for DINOv2-Base). A separate linear head predicts the [CLS] token. The code supports attention and projection dropout applied only in the middle 50% of layers; our main RiT-XL training sets both rates to .
Following JiT [18], we inject 32 learnable in-context tokens at an intermediate layer (layer 8 for RiT-L). These tokens are initialized from the class embedding with added learnable positional embeddings, participate in self-attention for all subsequent layers, and are discarded before the final projection. They provide additional capacity for class-conditional generation without modifying the core DiT block.
D.2 Training Hyperparameters
| Hyperparameter | Value |
|---|---|
| Hardware | 8 NVIDIA H200 GPUs |
| Throughput | 12 minutes per epoch |
| Optimizer | AdamW (, ) |
| Base learning rate | (scaled by ) |
| LR schedule | Constant (after warmup) |
| Warmup epochs | 5 |
| Weight decay | 0.0 |
| Gradient clipping | (max norm) |
| Total epochs | 800 |
| Batch size | 1536 (8 192 per GPU) |
| EMA decay | 0.9999 / 0.9996 (dual tracking) |
| Label dropout | 0.1 |
| Attention / projection dropout | 0.0 (main configuration) |
| Noise schedule | Truncated logit-normal (, ) |
| Time shift | ( for ) |
| Noise scale | 1.0 |
| Epsilon clamp | 0.05 |
| CLS loss weight | 0.2 |
D.3 Sampling Hyperparameters
| Hyperparameter | Value |
|---|---|
| ODE solver | Heun (2nd order) |
| Number of steps | 25 |
| CFG scale (patches) | 3.7 |
| CFG scale (CLS) | 3.7 |
| CFG interval | |
| Epsilon clamp | 0.05 |
| Generation precision | FP32 |
| EMA model | 0.9999 |
D.4 Pseudocode
⬇
t=sample_logit_normal(B,shift=s)
eps=randn_like(z0)*sigma
zt=t*z0+(1-t)*eps
eps_cls=randn_like(z_cls)*sigma
zt_cls=t*z_cls+(1-t)*eps_cls
z0_hat,cls_hat=dit(zt,t,y,zt_cls)
v=(z0-zt)/(1-t).clamp_min(eps_t)
v_hat=(z0_hat-zt)/(1-t).clamp_min(eps_t)
v_cls=(z_cls-zt_cls)/(1-t).clamp_min(eps_t)
v_cls_hat=(cls_hat-zt_cls)/(1-t).clamp_min(eps_t)
loss=mse(v_hat,v)+lam*mse(v_cls_hat,v_cls)
loss.backward()
⬇
z=randn(B,C,H,W)*sigma
z_cls=randn(B,C)*sigma
dt=1.0/K
foriinrange(K):
t=i/K
z0_hat,cls_hat=dit(z,t,y,z_cls)
v=(z0_hat-z)/(1-t).clamp_min(eps_t)
v_cls=(cls_hat-z_cls)/(1-t).clamp_min(eps_t)
z=z+dt*v
z_cls=z_cls+dt*v_cls
x=rae_decoder(z)
Appendix E Encoder Size Ablation
See Section 2 (main text) for the encoder size ablation and Table 3 for the full ablation results. DINOv2-Small () consistently outperforms DINOv2-Base () despite having half the feature dimensionality, reaching FID 1.44 vs. 1.56 at 800 epochs. The lower-dimensional latent space ( vs. ) is easier for the denoiser to model, while DINOv2-Small still retains sufficient semantic information (TwoNN intrinsic dimensionality is comparable across encoder sizes).
Appendix F Uncurated Sample Grid
Figure 11 shows uncurated samples generated by RiT-XL (DINOv2-Small encoder, 760 training epochs) using the Heun sampler with 100 steps and classifier-free guidance scale 3.7. The samples span diverse ImageNet categories including animals, food, landscapes, vehicles, and plants, demonstrating RiT’s ability to produce high-fidelity, diverse images across a wide range of semantic categories.
Appendix G Sampling Schedule Analysis
Table 9 lists the six ODE time-discretization schedules evaluated in this work. Each schedule maps a normalized step index () to a timestep , where is pure noise and is clean data.
| Schedule | Formula |
|---|---|
| Uniform | |
| Cosine | |
| Log-SNR uniform | |
| EDM [14] | |
| Power-2 | |
| Time-shift |
Figure 12 visualizes these schedule functions and their effect on generation quality. Panel (a) shows the schedule functions : uniform distributes steps evenly, while EDM, power-2, and time-shift concentrate steps near (the high-noise end); cosine and log-SNR are denser near both endpoints. Panels (b) and (c) show the corresponding FID as a function of Heun step count. Schedules that allocate more evaluations to the high-noise regime—where the velocity field varies most rapidly—achieve dramatically better FID at low step counts (), while all non-uniform schedules converge at steps.
| w/o guidance (CFG , Heun steps) | w/ guidance (CFG , Heun steps) | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Schedule | 2 | 5 | 10 | 25 | 50 | 125 | 2 | 5 | 10 | 25 | 50 | 125 |
| Uniform | 23.67 / 23.75 | 12.84 / 12.72 | 6.88 / 6.76 | 2.34 / 2.29 | 1.61 / 1.58 | 1.47 / 1.45 | 20.67 / 20.63 | 10.80 / 10.68 | 6.19 / 6.12 | 1.93 / 1.90 | 1.30 / 1.28 | 1.16 / 1.15 |
| EDM | 11.78 / 11.70 | 2.37 / 2.34 | 1.61 / 1.58 | 1.49 / 1.47 | 1.47 / 1.46 | 1.47 / 1.45 | 10.43 / 10.34 | 2.01 / 1.99 | 1.33 / 1.32 | 1.17 / 1.15 | 1.16 / 1.14 | 1.14 / 1.13 |
| Cosine | 23.67 / 23.75 | 5.96 / 5.86 | 2.16 / 2.12 | 1.57 / 1.56 | 1.48 / 1.45 | 1.45 / 1.44 | 20.67 / 20.63 | 5.69 / 5.63 | 1.91 / 1.88 | 1.29 / 1.28 | 1.19 / 1.18 | 1.15 / 1.14 |
| Power-2 | 11.16 / 11.09 | 2.41 / 2.39 | 1.74 / 1.72 | 1.50 / 1.48 | 1.47 / 1.44 | 1.45 / 1.43 | 9.44 / 9.37 | 1.99 / 1.98 | 1.48 / 1.46 | 1.18 / 1.16 | 1.16 / 1.15 | 1.15 / 1.14 |
| Log-SNR | 23.67 / 23.75 | 4.56 / 4.51 | 1.96 / 1.95 | 1.51 / 1.50 | 1.46 / 1.44 | 1.45 / 1.44 | 20.67 / 20.63 | 3.79 / 3.78 | 1.73 / 1.74 | 1.23 / 1.22 | 1.18 / 1.17 | 1.15 / 1.14 |
| Time-shift | 14.05 / 14.03 | 2.44 / 2.38 | 1.59 / 1.58 | 1.47 / 1.45 | 1.46 / 1.44 | 1.45 / 1.43 | 8.59 / 8.64 | 1.99 / 1.99 | 1.27 / 1.25 | 1.15 / 1.14 | 1.15 / 1.14 | 1.15 / 1.14 |
Appendix H Random Samples
Figure 13 shows 192 randomly generated samples (8 per class, 24 classes) from RiT-XL without any curation or cherry-picking. Each row corresponds to a single ImageNet class. The model produces consistently high-quality and diverse samples across all categories.
Appendix I CLS–Patch Attention Analysis
Figure 14 visualizes the bidirectional attention between [CLS] and patch tokens across layers and timesteps. We observe a clear stage-wise communication pattern:
CLSPatch (left). In early layers, [CLS] attends broadly to salient foreground regions, aggregating coarse object cues. In middle layers, its attention expands to contextual/background regions, forming a global scene summary. In late layers, [CLS] re-focuses on semantically critical details (e.g., head and eyes), which strongly influence structural consistency and perceptual realism.
PatchCLS (right). Patch tokens increasingly query [CLS] at deeper layers, with the strongest reliance in semantically important regions, while low-information background patches rely less on it. These observations suggest that [CLS] acts as a global message hub: it collects distributed evidence, integrates object–context relations, and broadcasts refined global guidance back to patch tokens, improving object–background disentanglement and final generation quality.
来源:HuggingFace Daily Papers(社区热门论文) · arxiv.org