NeurIPS 2026
Each pair: OBS-Diff (left) vs. Ours (right) at the same sparsity. Hover to pause.
OBS-style pruning such as OBS-Diff minimizes the average reconstruction error, so an error on the subject costs the same as an error on the background. At high sparsity, the subject is what breaks. We weight that error with a spatial importance map.
OBS-Diff · uniform objective
$\|\hat W X - W X\|_2^2$


Motorcycle loses its shape


Ours · importance-aware objective
$\|A \odot (\hat W X - W X)\|_2^2$


Subject preserved (red boxes)
Structured pruning of PixArt-Σ. Prompt: "A black Honda motorcycle parked in front of a garage."
We propose importance-aware pruning for diffusion models, a training-free framework that prioritizes preserving parameters critical to semantically salient image regions. To do so, we incorporate spatial importance maps, derived from conditioning signals or model attention, into the pruning objective. This produces parameter rankings aligned with perceptual relevance rather than uniform reconstruction error. On the MS-COCO dataset, our approach consistently retains subject fidelity and structural correctness at high compression ratios where conventional pruning causes visible degradation. These results show that content-aware objectives are key to perceptually faithful compression of generative models.
OBS ranks each weight by its saliency $S_q = w_q^2 / 2[H^{-1}]_{qq}$, with the Hessian $H$ built from layer activations. This score is already an approximation, so we correct it with information about what the image should preserve: a spatial importance map re-weights the activations before the Hessian is accumulated.
Take the magnitude of the CFG response, averaged over channels and min-max normalized:
$$M_t = \big|\epsilon_\theta(x_t,t,y) - \epsilon_\theta(x_t,t,\varnothing)\big|$$
$$A_t=\lambda M_t$$
Canny edges or detector boxes can replace $M_t$.
Errors in important regions count more:
$$\min_{\hat W}\sum_t \alpha_t\big\|A_t\odot(\hat W X_t - W X_t)\big\|_2^2$$
Linear layers act on channels, so $A_t$ commutes with $W$.
With filtered activations $X'_t = A_t\odot X_t$:
$$H_{\text{imp}} = 2\sum_t \alpha_t\,\mathbb{E}\big[X'_t X_t'^{\top}\big]$$
The OBS closed-form weight update is unchanged.
Calibrated on 1,000 GCC3M prompts and evaluated on 1,000 MS-COCO 2017 prompts. The gap grows with sparsity.
| Sparsity | Method | CLIP (↑) | ImageReward (↑) | MUSIQ (↑) |
|---|---|---|---|---|
| 0% | Dense | 32.26 | 0.98 | 72.68 |
| 30% | OBS-Diff | 32.25 | 0.96 | 72.53 |
| Ours (CFG) | 32.24 | 0.98 | 72.41 | |
| 40% | OBS-Diff | 32.30 | 0.92 | 71.02 |
| Ours (CFG) | 32.35 | 0.95 | 71.48 | |
| 45% | OBS-Diff | 32.25 | 0.84 | 68.76 |
| Ours (CFG) | 32.30 | 0.87 | 69.86 | |
| 50% | OBS-Diff | 32.16 | 0.71 | 65.98 |
| Ours (CFG) | 32.20 | 0.76 | 66.30 |
Best value per sparsity level (or category) in green.
Application · Flexible importance guidance
The pruning formulation never changes. The only input you choose is the importance map $A_t$, which says where reconstruction errors should cost more. Changing the signal behind $A_t$ changes what the pruned model preserves: prompt content, whole objects, or fine structure.
Takeaway: the importance signal is a plug-in. Canny makes the method usable when CFG maps are unavailable or inconvenient to extract, and adding a detector steers pruning toward object-centric preservation. Every variant beats OBS-Diff on PixArt-Σ at all tested sparsities (Other Signals tab above).
Outputs of our pruned models at their highest tested sparsity. Click any image to compare against OBS-Diff at every sparsity level. The last row shows category-targeted pruning, where calibration uses only prompts that contain a chosen category.
@inproceedings{lam2026importance,
title = {Importance-Aware OBS Pruning for Diffusion Models},
author = {Lam, Ba-Thinh and Das, Srijan and Le, Hieu},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
year = {2026},
eprint = {2607.20048},
archivePrefix = {arXiv}
}