A General Framework for Inference-time Scaling and Steering of Diffusion Models

Raghav SinghalZachary HorvitzRyan TeehanMengye RenZhou YuKathleen McKeownRajesh Ranganath

article2025ICML139 citations

Presents Feynman-Kac steering, a training-free framework that resamples particle trajectories at intermediate generation steps using arbitrary reward functions, allowing smaller diffusion models to outperform fine-tuned models and larger architectures in sample quality and prompt fidelity.

Listen

Diffusion-based generative models have achieved strong performance across images and text, yet ensuring that outputs consistently adhere to user prompts or safety preferences remains a major challenge. Standard solutions—such as fine-tuning models with preference datasets, gradient-based guidance, or generating multiple outputs and picking the best (best-of-n sampling)—are computationally expensive, inflexible, or limited to differentiable settings. The article introduces and evaluates Feynman-Kac (FK) steering, an inference-time framework designed to steer diffusion models toward user-defined reward functions without requiring model retraining.

To achieve this, FK steering tracks multiple candidate generation trajectories (called particles) and iteratively scores and resamples them during intermediate diffusion steps. Paths with higher likelihood of achieving high final rewards are multiplied, while low-reward paths are pruned. The researchers evaluated this approach across standard image generation models (Stable Diffusion variants) and text diffusion models (continuous and discrete), testing alignment metrics, text fluency, and rare-attribute steering such as toxicity red-teaming.

Key findings demonstrate that FK steering consistently outperforms conventional fine-tuning and inference baselines with minimal compute overhead. With as few as two particles, FK steering of base text-to-image models surpassed models fine-tuned specifically on preference alignment benchmarks. In addition, steering a smaller 0.8-billion parameter model achieved higher prompt fidelity and human preference scores than a substantially larger 2.6-billion parameter model while requiring less wall-clock time (9.1 seconds versus 11.5 seconds). In text diffusion tasks, FK steering lowered perplexity, improved grammatical acceptability, and dramatically enhanced rare-attribute detection—boosting toxicity detection rates from under 1% up to 64.7% for automated safety testing.

These results show that scaling compute during inference offers a viable and often superior alternative to costly model retraining. Organizations can maintain a single, frozen base model and apply diverse, plug-and-play reward functions dynamically at deployment, reducing training costs and improving adaptability. For teams implementing generative systems, the article recommends adopting particle-based inference steering when fine-tuning is impractical and tuning steering hyperparameters—such as resampling schedules and temperature—to balance output quality against sample diversity. Because highly aggressive steering can reduce output variety, future work should focus on optimizing reward estimators and testing broader production workflows.

No sufficiently relevant recommendations were found.

Cover for A General Framework for Inference-time Scaling and Steering of Diffusion Models

Abstract

Diffusion models have demonstrated remarkable performance in generative modeling, but generating samples with specific desiderata remains challenging. Existing solutions — such as fine-tuning, best-of-n sampling, and gradient-based guidance — are expensive, inefficient, or limited in applicability. In this work, we introduce Feynman-Kac (FK) steering, which applies Feynman-Kac interacting particle systems to the inference-time steering of diffusion models with arbitrary reward functions. FK steering works by generating multiple trajectories, called particles, and resampling particles at intermediate steps based on scores computed using functions called potentials. Potentials are defined using rewards for intermediate states and are chosen such that a high score indicates the particle will yield a high-reward sample. We explore various choices of potentials, rewards, and samplers. Steering text-to-image models with a human preference reward, we find that FK steering outperforms fine-tuned models with just 2 particles. Moreover, FK steering a 0.8B parameter model outperforms a 2.6B model, achieving state-of-the-art performance on prompt fidelity. We also steer text diffusion models with rewards for text quality and rare attributes such as toxicity, and find that FK steering generates lower perplexity text and enables gradient-free control. Overall, inference-time scaling and steering of diffusion models, even training-free, provides significant quality and controllability benefits. Code available here.

Citation

MLA
Singhal, R., et al. “A General Framework for Inference-time Scaling and Steering of Diffusion Models”. arXiv, 2025, http://arxiv.org/abs/2501.06848v5.
APA
Singhal, R., Horvitz, Z., Teehan, R., Ren, M., Yu, Z., McKeown, K., & Ranganath, R. (2025). A General Framework for Inference-time Scaling and Steering of Diffusion Models. arXiv. http://arxiv.org/abs/2501.06848v5
Chicago
Singhal, R., Z. Horvitz, R. Teehan, et al. 2025. “A General Framework for Inference-time Scaling and Steering of Diffusion Models”. arXiv. http://arxiv.org/abs/2501.06848v5.
Harvard
Singhal, R. et al. (2025) “A General Framework for Inference-time Scaling and Steering of Diffusion Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2501.06848v5.
Vancouver
1. Singhal R, Horvitz Z, Teehan R, Ren M, Yu Z, McKeown K, Ranganath R (2025) A General Framework for Inference-time Scaling and Steering of Diffusion Models. arXiv

BibTeX

@article{singhal2025general,
  title = {A General Framework for Inference-time Scaling and Steering of Diffusion Models},
  author = {Singhal, Raghav and Horvitz, Zachary and Teehan, Ryan and Ren, Mengye and Yu, Zhou and McKeown, Kathleen and Ranganath, Rajesh},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2501.06848v5},
  eprint = {2501.06848}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/