Generating Images of Rare Concepts Using Pre-trained Diffusion Models
Dvir SamuelRami Ben-AriSimon RavivNir DarshanGal Chechik
Proposes SeedSelect, a method that accurately generates rare and complex visual concepts from pre-trained diffusion models without fine-tuning by optimizing the initial noise seed using only a handful of reference images.
Modern text-to-image generative models produce impressive visual content but frequently fail when asked to generate uncommon concepts or structurally complex objects such as hands. An analysis reveals that roughly 25% of ImageNet categories are poorly rendered by standard diffusion models, primarily because web-scale training datasets are heavily imbalanced. For categories appearing less than 10,000 times during training, the generated images are accurately classified only about 50% of the time, limiting the reliability of generative artificial intelligence across specialized business and scientific domains.
The article evaluates whether these under-represented concepts remain accessible within pre-trained diffusion models without requiring computationally expensive model retraining or fine-tuning. The authors introduce SeedSelect, an optimization method that identifies starting noise seeds that correctly prompt the pre-trained model to generate rare and structurally complex visual concepts using only a small set of 3 to 5 reference images.
The authors conducted empirical evaluations using public foundation models on benchmark datasets including ImageNet, CUB (birds), and iNaturalist (species). The SeedSelect technique searches the input noise space at generation time by optimizing a joint objective: semantic similarity calculated with a visual-language encoder and appearance consistency evaluated using the model's image encoder. The approach was evaluated against state-of-the-art baselines via automated classification, standardized image quality metrics, human evaluation studies, and downstream few-shot classification tasks.
The analysis yielded several key findings. First, SeedSelect significantly outperformed competing methods in generation accuracy, with human evaluators preferring SeedSelect images over fine-tuned models by 3.4 times on CUB, 5.0 times on iNaturalist, and 4.3 times on rare ImageNet classes. Second, the method maintained baseline visual realism (an FID score of 6.5 compared to 6.4 for standard Stable Diffusion) and preserved sample diversity, avoiding the quality drops seen in fine-tuning alternatives (FID of 10.2). Third, for difficult structural generations like human hands, human judges found SeedSelect images approximately 4.5 times more prompt-aligned and 4 times more realistic than vanilla outputs. Finally, synthetic images produced by SeedSelect achieved state-of-the-art results when used as semantic data augmentations to train visual recognition classifiers.
These findings demonstrate that pre-trained diffusion models retain the semantic knowledge needed to render rare concepts, but standard random noise sampling fails to activate these representations. SeedSelect provides an efficient, low-cost way to utilize existing foundation models for niche applications. By eliminating the need for full-model fine-tuning, organizations can reduce training timelines and computing costs from hours per concept to a few minutes of initialization followed by rapid, second-scale image synthesis.
Organizations utilizing generative models for domain-specific applications, rare-class visual generation, or dataset enrichment should adopt seed-optimization strategies over resource-heavy model fine-tuning. For multi-image production pipelines, teams should implement the authors' bootstrapping strategy to reduce generation times from minutes to seconds per image. However, practitioners should note that the method is prompt-specific, does not capture the artistic style of reference images, and shows reduced efficacy on extremely scarce concepts represented by only a handful of web samples.
- Paper: An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion, Rinon Gal et al. (2022). Introduces Textual Inversion for personalizing pre-trained diffusion models using 3–5 reference images, establishing the core baseline and problem setting that SeedSelect aims to improve without optimization-based pseudo-tokens.
- Paper: DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation, Nataniel Ruiz et al. (2023). Presents DreamBooth, a primary fine-tuning baseline compared against SeedSelect for few-shot subject-driven generation from 3–5 casual images.
- Paper: Multi-Concept Customization of Text-to-Image Diffusion, Nupur Kumari et al. (2022). Develops parameter-efficient cross-attention tuning for multi-concept customization, serving as a key benchmark method evaluated against test-time seed optimization.
- Paper: SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations, Chenlin Meng et al. (2022). Pioneers test-time noise-guided image generation and editing via stochastic differential equations without model fine-tuning, providing essential intuition for input noise manipulation.
- Paper: Classifier-Free Diffusion Guidance, Jonathan Ho et al. (2022). Introduces classifier-free guidance, the core sampling mechanism utilized by modern pre-trained latent diffusion models upon which SeedSelect operates.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Establishes the foundational mathematical framework of denoising diffusion probabilistic models and iterative reverse sampling from starting noise.
- Paper: Initno: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization, Xiefan Guo et al. (2024). Explores initial noise optimization to resolve semantic omission and attribute blending in pre-trained text-to-image models without retraining, directly extending test-time latent noise steering.
- Paper: DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation, Hong Chen et al. (2024). Advances subject-driven customization by disentangling identity from background and pose variations in reference images during tuning.
- Paper: DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image Editing, Chong Mou et al. (2024). Applies localized test-time noise control and deterministic rollback sampling to enhance flexibility and precision during diffusion-based image editing.
