PixelDiT2

Representation-Grounded Pixel Diffusion Transformers

Yongsheng Yu1,2 Wei Xiong1† Yichen Sheng1 Shiqiu Liu1 Jiebo Luo2

1NVIDIA2University of Rochester

† Project Lead and Main Advising

TL;DR A frozen vision encoder reads the noisy image; its per-patch features ground a pixel diffusion transformer at every denoising step.

PixelDiT2 sample, ImageNet class: lesser panda
PixelDiT2 sample, ImageNet class: espresso
PixelDiT2 sample, ImageNet class: Rhodesian ridgeback
PixelDiT2 sample, ImageNet class: volcano
PixelDiT2 sample, ImageNet class: bottlecap
PixelDiT2 sample, ImageNet class: hot pot
PixelDiT2 sample, ImageNet class: fox squirrel
PixelDiT2 sample, ImageNet class: car mirror
PixelDiT2 sample, ImageNet class: Madagascar cat
PixelDiT2 sample, ImageNet class: bonnet
Class-conditional ImageNet 512×512 samples from PixelDiT2. The noise shows the forward process \(x_t = t\,x + (1-t)\,\varepsilon\) (\(t = 1\) clean, \(t = 0\) noise) over a 16-px patch grid, replayed when you hover, tap or focus a tile; it is not the sampling trajectory.
1.46
FID-50K lower is better
ImageNet 256×256 · 600 epochs Tab. 2, 14
1.48
FID-50K lower is better
ImageNet 512×512 · 680 epochs Tab. 3
4.25×
lower epoch budget
At 512×512, FID 1.78 at epoch 200 surpasses PixelDiT's 1.81 at epoch 850. Fig. 1

01 Motivation

Pixel diffusion learns two things at once

Pixel diffusion has narrowed the quality gap with latent diffusion, yet it still converges more slowly and lags in final quality. We argue that a key reason is the lack of an explicit representation prior.

Schematic · not data

ALatent diffusion

A pretrained autoencoder (AE) partly organizes the space before denoising. A decoder maps latents back to pixels.

BPixel diffusion e.g. PixelDiT, JiT

Starting from raw RGB, one network must learn denoising representations and pixel generation at the same time.

CPixelDiT2

A frozen encoder \(\E\) reads \(x_t\); the timestep-conditioned \(\Pg\) maps its tokens to grounding tokens \(\g\). No autoencoder, no decoder.

Where the representation comes from learned inside the DiT partly organized by the AE guided by grounding tokens \(\g\)
Where the representation job sits in each pipeline. The thumbnail is a PixelDiT2 sample; \(x_t = t\,x + (1-t)\,\varepsilon\) is the forward process applied to it, not a sampling trajectory. Intro · Fig. 1

Learning denoising representations and pixel generation in one network is a coupling that can make learning harder and may increase the training budget. PixelDiT2 decouples representation learning from pixel generation without an autoencoder or a latent reconstruction bottleneck: a frozen vision encoder provides explicit per-patch representation guidance throughout denoising, so the pixel diffusion transformer can focus more on pixel generation.

02 The idea

Read the noisy image with a frozen encoder

At every denoising step, a frozen vision encoder reads the noisy image \(x_t\); after a timestep-conditioned projection, its per-patch features modulate the pixel DiT through spatial AdaLN, in training and in sampling. We call this representation grounding; diffusion stays fully in pixel space.

Schematic · one denoising step · after Fig. 1

One denoising step

Hover, focus, or tap a part of the diagram for a one-line description.

  • pixels \(x_t\)
  • frozen vision encoder \(\E\), raw tokens \(\gbar\)
  • projection \(\Pg\), grounding tokens \(\g\)
  • timestep \(t\)
Token grids are drawn as 4×4 for readability; a 256 px image gives 16×16 tokens on both paths, since the encoder and the DiT share 16-px patches. The noisy input is drawn with the forward process \(x_t = t\,x + (1-t)\,\varepsilon\).

Frozen vision encoder \(\E\)

DINOv3, kept frozen in training and inference. It reads the noisy \(x_t\), not the clean image, and returns one token per 16-px patch on the same grid as the DiT.

S/16 at 256 px · B/16 at 512 px §3.2, Tab. 10

Timestep-conditioned projection \(\Pg\)

One DiT-style block with AdaLN on the timestep embedding. It maps raw tokens \(\gbar\) to grounding tokens \(\g\) and absorbs the encoder's drift on noisy inputs.

Noise-level adaptation is confined to \(\Pg\) Tab. 7, 8

Spatial AdaLN

\(\g\) is the AdaLN condition of the DiT blocks, applied element-wise on the shared patch grid, so each patch is modulated by its own token. It stays active at inference.

Lowest FID of the four injection mechanisms tested (H/16, epoch 600) Tab. 5

03 Method

Representation grounding, step by step

Steps 1–2 set up a plain single-path pixel DiT. Steps 3–7 add representation grounding, from the frozen encoder to the sampling loop.

  1. Step 1

    Pixel-space rectified flow

    PixelDiT2 denoises raw pixels. A noisy sample \(x_t\) lies on the straight line from noise \(\textcolor{#b5651d}{\varepsilon}\) to the clean image \(x\).

    \[x_t = \textcolor{#b5651d}{t}\,x + (1-\textcolor{#b5651d}{t})\,\textcolor{#b5651d}{\varepsilon},\qquad x_1 = x,\quad x_0 = \textcolor{#b5651d}{\varepsilon}\]
  2. The DiT cuts \(x_t\) into 16 × 16-pixel patches. A 256 × 256 image becomes \(L = 16 \times 16 = 256\) tokens.

  3. Step 2

    Predict the clean image, train on velocity

    The DiT reads \(x_t\), the timestep \(\textcolor{#b5651d}{t}\), and the class label \(y\), and predicts the clean image \(\hat x_\theta\). The loss is taken on the velocity \(v = x - \textcolor{#b5651d}{\varepsilon}\).

    \[\hat v_\theta = \frac{\hat x_\theta - x_t}{\max(1-\textcolor{#b5651d}{t},\ \tau)},\qquad \mathcal L_{\text{diff}} = \mathbb E\,\bigl\lVert \hat v_\theta - v \bigr\rVert_2^2\]

    \(\tau = 0.05\) guards against divergence as \(\textcolor{#b5651d}{t} \to 1\). Up to here, this is a single-path pixel DiT, similar to PixelDiT’s patch-level path, with x-prediction and a velocity loss following JiT.

  4. Step 3

    A frozen encoder reads the noisy image

    A pretrained DINOv3 encoder \(\textcolor{#3b73b9}{\mathcal E}\) reads \(x_t\), not the clean image, and stays frozen. Its patch size is also 16, so its \(L\) raw tokens align one-to-one with the DiT’s patch tokens.

    \[\textcolor{#3b73b9}{\bar{\mathbf g}_t} = \textcolor{#3b73b9}{\mathcal E}(x_t) \in \mathbb R^{L \times d_{\mathcal E}}\]

    DINOv3-S/16 at 256 px; DINOv3-B/16 at 512 px, where the grid is 32 × 32.

  5. Step 4

    The encoder never saw noise

    \(\textcolor{#3b73b9}{\mathcal E}\) was pretrained on clean images. As \(\textcolor{#b5651d}{t}\) moves away from 1, the raw encoder tokens drift from the clean-image features, and the drop accelerates below \(\textcolor{#b5651d}{t} = 0.7\).

    \[\cos\bigl(\textcolor{#3b73b9}{\bar{\mathbf g}_t},\ \textcolor{#3b73b9}{\bar{\mathbf g}_1}\bigr),\qquad \textcolor{#3b73b9}{\bar{\mathbf g}_1} = \textcolor{#3b73b9}{\mathcal E}(x)\]

    Scroll, or pick \(\textcolor{#b5651d}{t}\) on the stage. Readouts are the measured per-image mean cosine for DINOv3-S/16 Fig. 4a

  6. Step 5

    A timestep-conditioned projection

    Noise-level adaptation happens outside the encoder, in \(\textcolor{#336b00}{P_g}\), a single DiT-style block whose AdaLN reads the timestep embedding \(\textcolor{#b5651d}{e_t}\).

    \[\textcolor{#336b00}{\mathbf g_t} = \textcolor{#336b00}{P_g}\bigl(\textcolor{#3b73b9}{\bar{\mathbf g}_t},\ \textcolor{#b5651d}{t}\bigr) \in \mathbb R^{L\times D}\]

    Its output keeps cosine similarity above 0.77 to its own clean-input output at every \(\textcolor{#b5651d}{t}\), including pure noise. \(\textcolor{#336b00}{P_g}\) does not restore clean-image features; it emits a different, \(\textcolor{#b5651d}{t}\)-stable set of grounding tokens. Fig. 4a

  7. Step 6

    Spatial AdaLN

    Grounding tokens \(\textcolor{#336b00}{\mathbf g_t}\) share the DiT’s \(L\)-patch grid. Inspired by DDT and PixelDiT, \(\textcolor{#336b00}{\mathbf g_t}\) is used directly as the spatial AdaLN condition: every DiT token is modulated by the grounding token at its own position, in every block.

    \[h_i \leftarrow (1+\gamma_i)\odot \mathrm{LN}(h_i) + \beta_i,\quad (\gamma_i, \beta_i)\ \text{from}\ \textcolor{#336b00}{\mathbf g_{t,i}}\]

    Generic AdaLN form, shown as a schematic.

  8. Step 7

    The encoder stays in the loop at sampling

    Sampling integrates the flow ODE with a 50-step Heun solver. At every step, the frozen encoder re-reads the current \(x_t\).

    \[\frac{\mathrm d x_t}{\mathrm d t} = \hat v_\theta\bigl(x_t,\ \textcolor{#b5651d}{t},\ y;\ \textcolor{#336b00}{\mathbf g_t}\bigr),\qquad \textcolor{#336b00}{\mathbf g_t} = \textcolor{#336b00}{P_g}\bigl(\textcolor{#3b73b9}{\mathcal E}(x_t),\ \textcolor{#b5651d}{t}\bigr)\]

    For classifier-free guidance, the unconditional branch uses null grounding, which bypasses \(\textcolor{#3b73b9}{\mathcal E}\) and \(\textcolor{#336b00}{P_g}\). See Recipe.

    Schematic: the stage frames mix \(x\) and \(\textcolor{#b5651d}{\varepsilon}\); they are not sampler outputs.

Representation grounding

Frozen encoderDINOv3 reads the noisy \(x_t\) and is never updated.

Timestep-conditioned projectionOne DiT block absorbs the drift on noisy inputs.

Spatial AdaLNPer-patch, per-block modulation of the DiT.

Active in training and at every sampling step. Training also keeps the standard REPA loss, whose target is a separate DINOv2-B/14 on the clean image; see REPA.

04 Where representations enter

An encoder inside the sampling loop

Prior work uses pretrained features as training targets, as the space to be generated, or as variables generated alongside the image. PixelDiT2 runs the frozen encoder on the current noisy image at every step, in training and in sampling.

Select a row to compare training and sampling. Tab. 1 · §2.2
How pretrained visual representations enter diffusion (paper Tab. 1). The rows describe the formulations of the listed methods. The last column indicates whether the pretrained visual encoder itself is evaluated during sampling.
Method Role of pretrained features Diffused variables Encoder run
during sampling?
Representations as supervision
Feature-alignment targets Feature-alignment targets VAE latents ✗
Alignment for joint VAE–DiT training Alignment for joint VAE–DiT training VAE latents ✗
Representations as a generation space
Representation space for diffusion Representation space for diffusion Patch features ✗
Prior for a learned tokenizer Prior for a learned tokenizer Compressed feature latents ✗
Representations as additional generation targets
Jointly generated global semantics Jointly generated global semantics VAE latents and a [CLS] token ✗
Jointly generated spatial features Jointly generated spatial features Pixels and patch features ✗
Representations as spatial conditioning
Spatial conditioning from \(x_t\) Spatial conditioning from \(x_t\) Pixels ✓

The rows describe the formulations of the listed methods. A cross does not imply that pretrained knowledge is absent: it may be learned by the denoiser, define its generation space, or be represented by jointly generated features.

Schematic · training vs. sampling

PixelDiT2

Spatial conditioning from \(x_t\)

The features act as spatial conditioning; only the pixels have a diffusion trajectory.

Training
Sampling
Encoder at sampling
Evaluated every step
Diffused
Pixels
Decoder
None

Running the frozen encoder at every sampling step adds inference cost (Limitations).

05 Results

Faster convergence, lower final FID

Class-conditional ImageNet at 256×256 and 512×512, scored by FID-50K. PixelDiT2 samples with a 50-step Heun ODE solver and classifier-free guidance restricted to an interval.

Convergence at 512×512

PixelDiT2 against its predecessor PixelDiT, measured in training epochs.

ImageNet 512×512 · FID-50K

Epoch 200

1.78FID

below PixelDiT's 1.81 at 850 epochs

Epoch budget

\( \dfrac{\textcolor{#7c877a}{850}}{\textcolor{#76b900}{200}} = 4.25\times \)

lower budget to surpass PixelDiT's 850-epoch FID

Epoch 680

1.48FID

IS 295.7, lowest FID among the pixel-space methods in Tab. 3

ImageNet 512×512 convergence. Fig. 1 (bottom right)

Epochs count passes over the data, not compute. PixelDiT2-H has 1074M parameters including its frozen DINOv3-B/16 grounding encoder; PixelDiT-XL has 797M (Tab. 3).

System-level comparison

Class-conditional ImageNet with classifier-free guidance. The last column places every FID on one shared axis.

Class-conditional ImageNet 256×256.
Method Budget Params (M) FID↓ IS↑

* delayed grounding dropout: class dropout only until epoch 160, then independent class and grounding dropout (§09). Bold: best, underline: second best. Tab. 2

Trained with independent class and grounding dropout from initialization, PixelDiT2-H reaches FID 1.46 (IS 301.6) at 600 epochs. With delayed grounding dropout it reaches 1.48 at 480 epochs, improving on PixelDiT-XL (1.61 at 320 epochs) and JiT-G/16 (1.82 at 600 epochs, about 2B parameters).

Lower FIDs remain: SiD2 reports 1.38 (listed budget 1280), and the latent baselines in Tab. 2 reach 1.13 (RAE-XL) and 1.29 (REPA-XL). Those latent models use a separately trained tokenizer; PixelDiT2 does not.

TakeawayPixelDiT2 surpasses PixelDiT's 850-epoch FID at 512×512 by epoch 200. Its final FID is also lower at both resolutions: 1.46 (600 epochs) vs. 1.54 (800 epochs) at 256×256, and 1.48 (680) vs. 1.81 (850) at 512×512. The best latent-diffusion baselines remain lower. Fig. 1, Tab. 2, 3, 14

Scale and training length

All PixelDiT2 rows of the full 256×256 comparison.

Model scale at 200 epochs. FID falls from 5.29 (B/16, 165M) to 2.35 (L/16, 503M) and 1.82 (H/16, 1008M), at 91, 281, and 568 GFLOPs. Tab. 14
PixelDiT2-H/16 over training. Standard recipe: 1.82 → 1.59 → 1.52 → 1.46. Delayed grounding dropout trains without grounding dropout until epoch 160, then enables it: 1.66 → 1.50 → 1.48. Gray markers are reference pixel models. Tab. 14

Matched nominal budget against Latent Forcing

Latent Forcing jointly generates pixels and spatial representation features. PixelDiT2 instead encodes the current noisy image at each step.

  • H/16, from scratch
  • 200 epochs
  • batch 1024
  • 50K ImageNet-256 samples
  • Heun-50
  • own guidance sweep each

PixelDiT2 reaches FID 1.822 versus 2.287, with 1007.8M versus 1080.7M parameters. Latent Forcing has the higher IS: 295.7 versus 275.4. Tab. 12, App. B.2

06 Relation to REPA

Grounding complements REPA after early training

REPA uses a pretrained encoder as a training target. Representation grounding uses one as an input at every denoising step. PixelDiT2 keeps both, and in a matched-encoder study the combination gives the lowest FID from epoch 320 onward.

Schematic · two roles for a pretrained encoder

REPA

training only
Encoder reads
the clean image \(x\)
Its features are
an alignment target for projected DiT features
At sampling
the encoder is not evaluated

Representation grounding

training and sampling
Encoder reads
the noisy sample \(x_t\)
Its features are
a conditioning input, through \(P_g\) and spatial AdaLN
At sampling
the encoder runs at every step

PixelDiT2 keeps both paths. By default, the REPA target is a frozen DINOv2-B/14 on the clean image (loss weight 0.5, training only). The grounding encoder is a frozen DINOv3-S/16 at 256 px and DINOv3-B/16 at 512 px. Tab. 1, Tab. 10

A matched-encoder 2 × 2

Each path is switched on or off. Both active paths use DINOv3-S/16, removing encoder choice as a confound. Tab. 4

Grounding off Grounding on REPA off REPA on
Epoch

Early. At epoch 200, REPA alone and the combination are nearly identical, at FID 1.89 and 1.90. Grounding alone reaches 2.28, against 2.53 for the bare backbone.

Later. From epoch 320 onward, combining both gives the lowest FID at every evaluated epoch, improving from 1.70 to 1.59 by epoch 480. Over the same interval, REPA alone plateaus near 1.7, whereas grounding alone keeps improving from 2.08 to 1.83.

TakeawayOn the same backbone, grounding alone lowers FID at every evaluated epoch (2.53 → 2.28 at epoch 200, 2.19 → 1.83 at epoch 480). The two paths have different optimization profiles, and their complementarity emerges after the early training regime rather than appearing immediately. Tab. 4

07 What matters

Design choices, one at a time

Unless noted: PixelDiT2-H/16, ImageNet 256×256, 600 epochs, with separate sweeps over CFG scale and guidance interval.

FID-50K, lower is better. In each chart the hairline marks the default setting. Hover or tap a row to show that variant in the schematic.

Where should grounding tokens enter?

Tab. 5
Schematic
Spatial AdaLNFID 1.462 \(\g\) modulates every DiT block element-wise, patch by patch, on the shared 16×16 grid.

Four ways to feed the same grounding tokens \(\g\) into the pixel DiT.

TakeawaySpatial AdaLN gives the lowest FID of the four tested mechanisms.

The experiment does not by itself isolate which property of AdaLN causes the advantage.

One global token instead of a per-patch grid Tab. 11 · App. B.1

PixelDiT2-B/16, epoch 200, DINOv3-S/16. A REG-inspired baseline puts the encoder's [CLS] token in the in-context conditioning slot instead of modulating each patch. FID rises from 5.29 to 6.04.

Collapsing the grounding signal into one token erases the 16×16 spatial alignment that AdaLN exploits.

Should the encoder adapt to noisy inputs?

Tab. 7
Schematic
Frozen (default)FID 1.462 Off-the-shelf \(\E\). Noise-level adaptation is confined to the timestep-conditioned \(\Pg\).

\(\E\) was pretrained on clean images but reads noisy \(x_t\). Should the encoder side absorb this shift?

Joint
the adapted part is optimized together with the diffusion model.
Two-stage
the adapted part comes from a stage-1 conditioner checkpoint and stays frozen while \(\Pg\) trains.
Timestep adapter
a timestep-conditioned encoder-side module.
First DINO block
the encoder's first transformer block.

TakeawayAll four adapted variants are worse. In the evaluated setting, these results favor preserving the pretrained encoder and adapting its output downstream.

This does not establish that \(\Pg\) alone determines performance.

What should \(\Pg\) be?

Tab. 8
Schematic · 143M trainable in both
DiT block + AdaLN(\(e_t\))FID 5.29 \(\Pg\) is one DiT-style block modulated by the timestep embedding \(e_t\). The DiT has 12 layers.

PixelDiT2-B/16 at epoch 200. Replace the DiT-style \(\Pg\) with a linear projection and add one block to the DiT, so both models keep 143M trainable parameters.

TakeawayAt matched parameters, the timestep-conditioned DiT-style \(\Pg\) reaches FID 5.29, versus 6.79 for the linear projection. Its benefit is not interchangeable with placing the same capacity in the DiT.

Which frozen encoder?

Tab. 6 · Tab. 9 · Tab. 15
Schematic
DINOv3-S/16FID 1.462 Default grounding encoder at 256px.

Swap the frozen encoder \(\E\) on the grounding path, under the shared protocol.

TakeawayDINOv3-S/16 gives the lowest FID of the six. Larger DINOv3 encoders are worse at 256px under the shared evaluation at epoch 600.

Is \(\Pg\) the bottleneck for DINOv3-L? Tab. 9

A single-block \(\Pg\) may bottleneck higher-dimensional DINOv3-L features. Fix the grounding encoder to DINOv3-L/16 and widen or deepen \(\Pg\).

Schematic

Doubling the width lowers FID at both epochs (1.82 → 1.78, 1.73 → 1.64); four blocks help only at epoch 320 (1.69). Neither reaches the DINOv3-S/16 reference of 1.59 at epoch 320.

At 512px, the ordering changes with training Tab. 15 · App. C

PixelDiT2-H/16 at 512px. Grounding encoder DINOv3-B/16 or DINOv3-L/16; REPA target fixed to DINOv2-B/14.

DINOv3-L is lower at epochs 200, 320 and 400. At epoch 600, DINOv3-B is lower: 1.52 vs 1.54.

TakeawayLimited \(\Pg\) capacity may contribute to the negative S-to-L scaling at 256px, but does not fully explain it. In the evaluated 512px setting, the larger encoder speeds up early convergence without improving the later epoch. The results do not support a resolution-independent rule that smaller grounding encoders are always preferable.

Fewer tokens at 512px?

Tab. 16 · App. C
Illustration · DiT patch grid
Patch 16FID 1.543 ep. 400 32×32 = 1024 tokens. PixelDiT2-H/16, the 512px default.

Patch size 32 cuts the 512px token grid from 32×32 to 16×16. Can a larger DiT compensate?

TakeawayScaling H/32 to G/32 improves FID at every evaluated epoch, but G/32 stays behind H/16 at epoch 400: 1.625 vs 1.543. Replacing DINOv2-B with DINOv3-B as the H/32 grounding encoder gives a smaller, consistent gain.

Epoch budget. At matched H/32 scale, PixelDiT2 reaches 2.093 at epoch 200, below the authors' reproduction of JiT-H/32 at epoch 320 (2.321). The epoch budgets differ by 1.6×; this is not a wall-clock or iso-FLOP speedup.

08 Inside the model

What \(P_g\) learns, and what the DiT inherits

Two diagnostics: what \(P_g\) does across noise levels, and how grounding changes the DiT's own features.

a Feature stability Fig. 4a

\(P_g\) turns drifting features into stable grounding tokens

The drift is not specific to DINOv3. Per-image mean cosine similarity between \(\E(x_t)\) and \(\E(x_1)\) also collapses for DINOv2-B/14 and MAE-B/16 once the input becomes noisy, with the drop accelerating below \(t = 0.7\).

After the trained \(P_g(\cdot, t)\), the analogous similarity stays above 0.77 at every \(t\), including pure noise. \(P_g\) does not restore clean-image features; it emits a different, \(t\)-stable set of grounding tokens.

Cosine to the clean-input feature
t = 0.3
A PixelDiT2 sample blended with Gaussian noise at the selected t

Input

\(x_t = t\,x + (1-t)\,\varepsilon\)

forward process on a PixelDiT2 sample

raw \(\E(x_t)\)

cos 0.44

\(P_g(\E(x_t), t)\)

cos 0.81

Reading the grids. Schematic, DINOv3-S/16. Each arrow is one token; its faint line is the token's direction at \(t = 1\). The spread is set so that the mean cosine across the grid equals the measured value at the selected \(t\). Drag the slider or click the chart; it snaps to the seven measured \(t\).

b Linear probing Fig. 4b

Grounding makes the DiT's own features more linearly separable

Following REPA, linear probes read the per-block hidden states of PixelDiT2-B/16 at epoch 200, which uses frozen DINOv3-S/16, \(P_g\), and REPA. The baseline is an iso-architecture REPA-only model, JiT-B/16 + REPA, without grounding tokens.

Inputs are clean. Every probe uses the null class token, so it cannot read the label off the in-context conditioning slot.

TakeawayWith grounding, probe accuracy exceeds the REPA-only baseline at every block: +38.5 points at block 0 (Fig. 4b) and a stable margin of about +8 to +10 points in the deeper half.

Both curves peak at the in-context-injection block, \(i = 4\). Past it, absolute accuracy drops in both because the null-class probe moves deeper layers off-distribution; the relative gap is the meaningful comparison.

Linear-probe accuracy per DiT block, PixelDiT2-B/16 at epoch 200, clean inputs, null class token. Values read from Fig. 4b. Fig. 4b

09 Training and sampling recipe

Drop the grounding, too

Classifier-free guidance (CFG) needs an unconditional branch. Under representation grounding, replacing the class label with a null token is not enough to build one.

A null class is not a null condition

During training, the class token is replaced by a null token \(\varnothing\) with probability 0.1, as in standard CFG. The grounding tokens \(\g\), however, are computed from \(x_t\) and remain class-discriminative. A branch that drops only the class token can therefore still receive class information. §3.2

Schematic · one network evaluation · DiT timestep input not drawn

Training

Each sample draws two independent Bernoulli masks with \(p = 0.1\): one for the class token, one for the grounding tokens. A dropped grounding uses the null grounding tokens instead of \(\Pg(\E(x_t), t)\). §3.2, §4

Inference

The unconditional CFG branch uses the null class token and the same null grounding tokens, matching the training distribution. The two branches combine in the standard CFG form, with guidance applied only inside an interval of \(t\):

\[\hat v = \hat v_{\varnothing} + w\,\bigl(\hat v_{y} - \hat v_{\varnothing}\bigr)\]
Two independent masks per sample schematic
A random draw for 50 samples, each mask dropped with probability 0.1. Class only, grounding only, and both can be dropped.

Grounding dropout and late-stage FID

The null-grounding branch also changes optimization. Three H/16 runs on ImageNet-256 use different grounding-dropout schedules; class dropout is active in all three. §4.2, Fig. 3

grounding retained class and grounding dropout grounding dropout enabled at epoch 160
With independent class and grounding dropout from initialization, FID decreases monotonically from 1.59 (epoch 320) to 1.46 (epoch 600). Retaining grounding reaches 1.54 at epoch 500, then rises to 1.555 at epoch 600, a pattern consistent with mild late-stage overfitting. Fig. 3 left
Training starts with class dropout while grounding is retained; independent class and grounding dropout begins at epoch 160 . FID reaches 1.66, 1.50, and 1.48 at epochs 200, 320, and 480, versus 1.75, 1.59, and 1.54 with grounding retained. Fig. 3 right; overlay: Fig. 3 left, Tab. 14

TakeawayGrounding dropout plays two empirical roles: it matches the null-grounding distribution that CFG requires, and it is associated with steadier late-stage FID. It need not start at initialization.

A single delayed run suggests a curriculum: first learn the grounded denoising function, then spend a shorter phase on the matched null-grounding branch. GLIDE is an earlier precedent; it introduced text-condition dropout only during fine-tuning. §4.2, Fig. 3 right

Unless stated otherwise, the paper applies class and grounding dropout from the start of training. The delayed run is reported separately (PixelDiT2-H*, FID 1.48 at epoch 480, Tab. 2).

Sampling: Heun-50, guidance on an interval

The guidance scale \(w\) and the interval are swept per backbone scale and epoch. §4, Tab. 10

Tab. 13 · H/16 · epoch 600 · Heun-50 · FID-50K
settingw = 2.4, [0.125, 0.900]
FID-50K1.4622
IS301.6

Stronger guidance raises IS and FID together. Extending the interval's upper bound from 0.9 to 0.975 raises FID from 1.4622 to 1.6202, while IS moves from 301.6 to 303.2. The lowest FID in these one-dimensional sweeps, 1.4622, is at \(w = 2.4\) on \([0.125, 0.9]\), the headline configuration. App. B.3, Tab. 13

Sampler at matched compute

\(N\) Heun steps cost \(2N\) network evaluations (NFE); \(N\) Euler steps cost \(N\). At NFE = 100, Heun-50 reaches FID 1.50 and Euler-100 reaches 1.54. At NFE = 50 the order reverses: Heun-25 1.76, Euler-50 1.72. IS favors Euler-100, 300.8 vs. 292.2.

The paper reports Heun-50 because it is FID-optimal at matched compute. App. B.4

This sweep uses H/16 at epoch 320, trained with delayed grounding dropout, and \(w = 2.4\) on \([0.1, 0.9]\). The CFG factor of two is paid by both samplers and omitted from NFE.

Each tick is one network evaluation. Heun ticks come in predictor and corrector pairs.
Setup Tab. 10

Model configurations

Training

Sampling and evaluation

The grounding encoder runs on \(x_t\) in training and sampling. Unless stated otherwise, the DiT, \(\Pg\), and the REPA head are optimized jointly from scratch under one optimizer. The 512 px component counts are measured at epoch 680.

10 Samples

Selected and uncurated samples

Class-conditional ImageNet samples from PixelDiT2: selected sets at 512×512 (Fig. 1) and 256×256 (PixelDiT2-H/16, Fig. 2), and uncurated sets at 256×256 (PixelDiT2-H/16, Figs. 6–7). Select any image to enlarge it.

\(\varepsilon\) \(x\) t = 1.00

The slider applies the forward process to finished samples (\(t = 1\) clean, \(t = 0\) pure noise). Tiles enter the same way. Neither is the model's sampling trajectory.

11 Limitations

What remains open

01

The full gap to the best latent baselines is not closed

On ImageNet 256×256, the best latent baselines still reach lower FID than PixelDiT2-H: RAE-XL 1.13 and REPA-E 1.15, against 1.46. Tab. 14

Selected latent-diffusion baselines from Tab. 14.

02

The frozen encoder runs at every step

Grounding evaluates \(\E\) on \(x_t\) at every sampling step, which adds inference cost: DINOv3-S/16 (21.6M parameters) at 256 px and DINOv3-B/16 (85.7M) at 512 px. Tab. 10

PixelDiT2-H/16; parameter counts in millions. The REPA target, DINOv2-B/14, is training-only and is not counted.

Future work Address both limitations, and apply the architecture to scaled text-to-image generation.

We hope our findings encourage future pixel-space generative models to treat representations not only as auxiliary training targets but as active components of the denoising computation.

Citation

BibTeX

@article{yu2026pixeldit2,
  title   = {PixelDiT2: Representation-Grounded Pixel Diffusion Transformers},
  author  = {Yu, Yongsheng and Xiong, Wei and Sheng, Yichen and Liu, Shiqiu and Luo, Jiebo},
  journal = {arXiv preprint arXiv:2609.24919},
  year    = {2026}
}

PixelDiT2 builds on PixelDiT: project page · arXiv:2511.20645.