Generating Images of Rare Concepts Using Pre-trained Diffusion Models

Dvir SamuelRami Ben-AriSimon RavivNir DarshanGal Chechik

article2024AAAI91 citations

Proposes SeedSelect, a method that accurately generates rare and complex visual concepts from pre-trained diffusion models without fine-tuning by optimizing the initial noise seed using only a handful of reference images.

Listen

Modern text-to-image generative models produce impressive visual content but frequently fail when asked to generate uncommon concepts or structurally complex objects such as hands. An analysis reveals that roughly 25% of ImageNet categories are poorly rendered by standard diffusion models, primarily because web-scale training datasets are heavily imbalanced. For categories appearing less than 10,000 times during training, the generated images are accurately classified only about 50% of the time, limiting the reliability of generative artificial intelligence across specialized business and scientific domains.

The article evaluates whether these under-represented concepts remain accessible within pre-trained diffusion models without requiring computationally expensive model retraining or fine-tuning. The authors introduce SeedSelect, an optimization method that identifies starting noise seeds that correctly prompt the pre-trained model to generate rare and structurally complex visual concepts using only a small set of 3 to 5 reference images.

The authors conducted empirical evaluations using public foundation models on benchmark datasets including ImageNet, CUB (birds), and iNaturalist (species). The SeedSelect technique searches the input noise space at generation time by optimizing a joint objective: semantic similarity calculated with a visual-language encoder and appearance consistency evaluated using the model's image encoder. The approach was evaluated against state-of-the-art baselines via automated classification, standardized image quality metrics, human evaluation studies, and downstream few-shot classification tasks.

The analysis yielded several key findings. First, SeedSelect significantly outperformed competing methods in generation accuracy, with human evaluators preferring SeedSelect images over fine-tuned models by 3.4 times on CUB, 5.0 times on iNaturalist, and 4.3 times on rare ImageNet classes. Second, the method maintained baseline visual realism (an FID score of 6.5 compared to 6.4 for standard Stable Diffusion) and preserved sample diversity, avoiding the quality drops seen in fine-tuning alternatives (FID of 10.2). Third, for difficult structural generations like human hands, human judges found SeedSelect images approximately 4.5 times more prompt-aligned and 4 times more realistic than vanilla outputs. Finally, synthetic images produced by SeedSelect achieved state-of-the-art results when used as semantic data augmentations to train visual recognition classifiers.

These findings demonstrate that pre-trained diffusion models retain the semantic knowledge needed to render rare concepts, but standard random noise sampling fails to activate these representations. SeedSelect provides an efficient, low-cost way to utilize existing foundation models for niche applications. By eliminating the need for full-model fine-tuning, organizations can reduce training timelines and computing costs from hours per concept to a few minutes of initialization followed by rapid, second-scale image synthesis.

Organizations utilizing generative models for domain-specific applications, rare-class visual generation, or dataset enrichment should adopt seed-optimization strategies over resource-heavy model fine-tuning. For multi-image production pipelines, teams should implement the authors' bootstrapping strategy to reduce generation times from minutes to seconds per image. However, practitioners should note that the method is prompt-specific, does not capture the artistic style of reference images, and shows reduced efficacy on extremely scarce concepts represented by only a handful of web samples.

arXiv: 2304.14530
Cover for Generating Images of Rare Concepts Using Pre-trained Diffusion Models

Abstract

Text-to-image diffusion models can synthesize high quality images, but they have various limitations. Here we highlight a common failure mode of these models, namely, generating uncommon concepts and structured concepts like hand palms. We show that their limitation is partly due to the long-tail nature of their training data: web-crawled data sets are strongly unbalanced, causing models to under-represent concepts from the tail of the distribution. We characterize the effect of unbalanced training data on text-to-image models and offer a remedy. We show that rare concepts can be correctly generated by carefully selecting suitable generation seeds in the noise space, using a small reference set of images, a technique that we call SeedSelect. SeedSelect does not require retraining or finetuning the diffusion model. We assess the faithfulness, quality and diversity of SeedSelect in creating rare objects and generating complex formations like hand images, and find it consistently achieves superior performance. We further show the advantage of SeedSelect in semantic data augmentation. Generating semantically appropriate images can successfully improve performance in few-shot recognition benchmarks, for classes from the head and from the tail of the training data of diffusion models.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Motivating Analysis
  • 4 Notations and Definitions
  • 5 Our Approach: Seed Select
  • 5.1 Improving Speed and Quality
  • 6 Experiments
  • 6.1 Rare Concept Generation
  • 6.2 Hand Generation
  • 6.3 Synthetic Data for Few-Shot Recognition
  • 7 Discussion and Limitations
  • References

Knowls

  1. Knowl 1 — Hypothesis of Rare-Concept Representation in Latent Noise Space

    assumption

    Web-crawled training datasets for text-to-image diffusion models (such as LAION-2B) exhibit extreme class imbalance, where tail concepts appear infrequently (e.g., fewer than 10k10\text{k} times). For frequent ("head") concepts, diffusion models learn to map broad regions of the standard high-dimensional Gaussian noise space N(0,I)\mathcal{N}(0, I) to semantically correct images. For rare ("tail") concepts, the model learns valid generative mappings only within small, restricted subregions of the noise space. Consequently, unguided random noise seeds zT∼N(0,I)z_T \sim \mathcal{N}(0, I) at inference time fall outside the viable generative distribution for rare concepts, resulting in semantic corruption or incorrect concept generation. However, because rare concepts were observed thousands of times during training, their visual representations are retained in the pre-trained weights and can be recovered by identifying in-distribution noise seeds zTz_T using test-time optimization guided by a small reference set.

  2. Knowl 2 — SeedSelect Latent Seed Optimization Objective

    model/method

    Given a target rare concept text prompt yy, a pre-trained latent diffusion model with frozen parameters (comprising a VAE encoder E\mathcal{E}, a VAE decoder D\mathcal{D}, and a denoising network ϵθ\epsilon_\theta), and a reference set of kk real images {I1,…,Ik}\{I^1, \dots, I^k\} of the concept (k≈3–5k \approx 3\text{--}5), SeedSelect optimizes the initial noise tensor zTGz_T^G without modifying any model weights. The objective function balances semantic fidelity and natural visual appearance:

    LTotal=λLSemantic+(1−λ)LAppearance\mathcal{L}_{\text{Total}} = \lambda \mathcal{L}_{\text{Semantic}} + (1 - \lambda) \mathcal{L}_{\text{Appearance}}

    where λ∈[0,1]\lambda \in [0, 1] is a trade-off hyperparameter.

    The semantic loss enforces semantic correspondence by computing the Euclidean distance between the CLIP image feature vG=CLIP(IG)v^G = \text{CLIP}(I^G) of the generated image IG=D(z0G)I^G = \mathcal{D}(z_0^G) and the centroid of the reference CLIP feature embeddings μv=1k∑i=1kCLIP(Ii)\mu_v = \frac{1}{k}\sum_{i=1}^k \text{CLIP}(I^i):

    LSemantic=∥μv−vG∥2\mathcal{L}_{\text{Semantic}} = \|\mu_v - v^G\|_2

    The natural appearance loss enforces structural consistency by measuring the mean squared error between the spatial latent z0Gz_0^G (obtained at the end of the denoising trajectory) and the VAE-encoded reference latents zi=E(Ii)z^i = \mathcal{E}(I^i):

    LAppearance=1k∑i=1k∥zi−z0G∥22\mathcal{L}_{\text{Appearance}} = \frac{1}{k} \sum_{i=1}^k \|z^i - z_0^G\|_2^2

    Optimization updates zTGz_T^G via gradient backpropagation through the denoising process and terminates when LTotal\mathcal{L}_{\text{Total}} plateaus or increases for more than 3 iterations.

  3. Knowl 3 — Supervised Contrastive Semantic Loss for Multi-Class Optimization

    equation

    When generating images across a predefined set of candidate classes CC, SeedSelect replaces the single-centroid Euclidean distance with a supervised contrastive loss in CLIP embedding space. This pulls the generated image representation vGv^G toward the target class centroid μvc\mu_v^c while pushing it away from centroids of other classes μvc′\mu_v^{c'}:

    LSemantic=−log⁡exp⁡(−∥μvc−vG∥2)∑c′∈Cexp⁡(−∥μvc′−vG∥2)\mathcal{L}_{\text{Semantic}} = -\log \frac{\exp\left(-\|\mu_v^c - v^G\|_2\right)}{\sum_{c' \in C} \exp\left(-\|\mu_v^{c'} - v^G\|_2\right)}

    where c∈Cc \in C is the target class of the generated image, vG=CLIP(IG)v^G = \text{CLIP}(I^G) is the CLIP embedding of the generated image IGI^G, and μvc=1kc∑i=1kcCLIP(Ic,i)\mu_v^c = \frac{1}{k_c}\sum_{i=1}^{k_c} \text{CLIP}(I^{c, i}) is the empirical mean feature vector of the reference images for class cc.

  4. Knowl 4 — Stabilized and Bootstrapped SeedSelect Optimization

    algorithm

    To improve convergence stability and accelerate batch image synthesis, SeedSelect incorporates two algorithmic strategies: (1) multi-step loss aggregation over the final tt denoising steps (t=2t=2), and (2) reference set bootstrapping.

    Input: Text prompt yy, reference set I={I1,…,Ik}I = \{I^1, \dots, I^k\}, pre-trained latent diffusion model with encoder E\mathcal{E} and decoder D\mathcal{D}, CLIP image encoder, number of subsets MM, subset size mm, stabilization horizon t=2t=2, step size η\eta, trade-off λ\lambda.
    Output: Set of generated images G\mathcal{G}.
    Compute reference latents zi=E(Ii)z^i = \mathcal{E}(I^i) and embeddings vi=CLIP(Ii)v^i = \text{CLIP}(I^i) for all i∈{1,…,k}i \in \{1, \dots, k\}
    Compute full-set centroid μv=1k∑i=1kvi\mu_v = \frac{1}{k} \sum_{i=1}^k v^i
    Initialize noise seed zTG∼N(0,I)z_T^G \sim \mathcal{N}(0, I)
    // Phase 1: Global Seed Warm-up
    while stopping criterion not met do
        Run denoising from zTGz_T^G to obtain trajectory latents ztG,…,z0Gz_t^G, \dots, z_0^G and image IG=D(z0G)I^G = \mathcal{D}(z_0^G)
        Compute LSemantic=∑i=0t∥μv−CLIP(D(ziG))∥2\mathcal{L}_{\text{Semantic}} = \sum_{i=0}^t \|\mu_v - \text{CLIP}(\mathcal{D}(z_i^G))\|_2
        Compute LAppearance=1k∑i=1k∥zi−z0G∥22\mathcal{L}_{\text{Appearance}} = \frac{1}{k} \sum_{i=1}^k \|z^i - z_0^G\|_2^2
        Compute LTotal=λLSemantic+(1−λ)LAppearance\mathcal{L}_{\text{Total}} = \lambda \mathcal{L}_{\text{Semantic}} + (1 - \lambda) \mathcal{L}_{\text{Appearance}}
        Update zTG←zTG−η∇zTGLTotalz_T^G \leftarrow z_T^G - \eta \nabla_{z_T^G} \mathcal{L}_{\text{Total}}
    end while
    Initialize G←∅\mathcal{G} \leftarrow \emptyset
    // Phase 2: Bootstrapped Generation
    for j=1j = 1 to MM do
        Sample subset Sj⊂IS_j \subset I with ∣Sj∣=m|S_j| = m
        Compute subset centroid μvSj=1m∑Is∈SjCLIP(Is)\mu_v^{S_j} = \frac{1}{m} \sum_{I^s \in S_j} \text{CLIP}(I^s)
        Initialize zTSj←zTGz_T^{S_j} \leftarrow z_T^G
        Optimize zTSjz_T^{S_j} for a small number of iterations using SjS_j
        Generate ISj=D(Denoise(zTSj,y))I^{S_j} = \mathcal{D}(\text{Denoise}(z_T^{S_j}, y))
        G←G∪{ISj}\mathcal{G} \leftarrow \mathcal{G} \cup \{I^{S_j}\}
    end for
    return G\mathcal{G}

    This bootstrapping procedure reduces the per-image generation time on an NVIDIA A100 GPU from 1--4 minutes to 1--2 seconds while maintaining sample diversity.

  5. Knowl 5 — Faithfulness Degradation on Long-Tail Concepts in Diffusion Models

    empirical result

    When evaluating image generation across all 1000 ImageNet classes sorted by their frequency in the LAION-2B dataset (using 100 Stable Diffusion v2.1 generations per class and a pre-trained MaxViT classifier achieving 88.2% top-1 accuracy on real ImageNet test data), generation accuracy drops sharply for tail classes. While head classes exhibit high accuracy, classes in the bottom frequency quartile (<10k<10\text{k} training occurrences in LAION-2B) achieve an average accuracy of only 52.2% with vanilla Stable Diffusion. On fine-grained tail benchmarks (CUB-200 and iNaturalist), pre-trained classifiers (achieving 93.1% and 83.8% test accuracy on real data, respectively) similarly identify severe accuracy degradation in vanilla Stable Diffusion and fine-tuned baselines. SeedSelect increases the average accuracy on the ImageNet tail quartile from 52.2% to 76.1% and achieves higher classification accuracy across all class frequency percentiles across ImageNet, CUB, and iNaturalist.

  6. Knowl 6 — Realism and Sample Diversity Comparison

    data/table

    Image quality and sample diversity were measured on 50,000 generated images versus 50,000 real ImageNet test images. Image quality was evaluated via Fr'echet Inception Distance (FID), while diversity and distribution coverage were assessed via Number of Statistically Different Bins (NDB), Precision, Recall, Fidelity, and Diversity metrics.

    Method FID ↓\downarrow NDB ↓\downarrow Precision ↑\uparrow Recall ↑\uparrow Fidelity ↑\uparrow Diversity ↑\uparrow
    SD (Stable Diffusion v2.1) 6.4 2.48 0.79 0.20 0.85 0.37
    SD + RG + RF 6.9 2.60 0.70 0.15 0.79 0.28
    Finetuned SD 10.2 2.90 0.61 0.11 0.71 0.20
    SeedSelect 6.5 2.52 0.77 0.18 0.83 0.36

    Fine-tuning model parameters on few-shot samples (Finetuned SD) or steering diffusion latents with real reference noise (SD + RG + RF) deteriorates FID and causes mode collapse (indicated by higher NDB and lower Recall/Diversity). Because SeedSelect optimizes only the initial random noise vector zTGz_T^G for a frozen model, it maintains the realism (FID 6.5 vs. 6.4) and diversity profile of vanilla Stable Diffusion.

  7. Knowl 7 — Human Evaluation of Generated Rare Concepts and Hand Structures

    data/table

    Human preference evaluations were conducted using a 2-alternative forced choice (2AFC) protocol. For rare concepts, raters evaluated 10 generated images per class across 30 CUB classes, 30 iNaturalist classes, and 90 ImageNet classes stratified into Many (>1M>1\text{M} occurrences in LAION-2B), Med (10k10\text{k}--1M1\text{M}), and Few (<10k<10\text{k}) categories, comparing SeedSelect against Finetuned SD. For hand generation, raters evaluated whether images matched the prompt and appeared realistic across 5 hand prompts comparing SeedSelect against vanilla Stable Diffusion.

    Dataset / Metric Finetuned SD (%) SeedSelect (%) Neither (%)
    ImageNet Many (>1M>1\text{M}) 48.01±1.0148.01 \pm 1.01 50.12±1.0050.12 \pm 1.00 1.87±1.121.87 \pm 1.12
    ImageNet Med (10k10\text{k}–1M1\text{M}) 41.55±1.5541.55 \pm 1.55 55.33±1.4255.33 \pm 1.42 3.12±1.483.12 \pm 1.48
    ImageNet Few (<10k<10\text{k}) 15.84±2.2915.84 \pm 2.29 69.08±2.4669.08 \pm 2.46 15.08±2.2215.08 \pm 2.22
    CUB All 20.18±2.3120.18 \pm 2.31 68.98±2.7168.98 \pm 2.71 10.84±3.1110.84 \pm 3.11
    iNaturalist All 14.45±2.7714.45 \pm 2.77 72.44±2.1372.44 \pm 2.13 13.11±2.7913.11 \pm 2.79
    Hand Generation Stable Diffusion (%) SeedSelect (%) Neither (%)
    Matches prompt 16.22±2.616.22 \pm 2.6 70.21±2.670.21 \pm 2.6 13.57±4.113.57 \pm 4.1
    Looks realistic 16.19±5.816.19 \pm 5.8 62.46±6.562.46 \pm 6.5 21.35±7.221.35 \pm 7.2

    SeedSelect is preferred over Finetuned SD by a margin of ×4.3\times 4.3 on the ImageNet tail, ×3.4\times 3.4 on CUB, and ×5.0\times 5.0 on iNaturalist. On hand generation, SeedSelect achieves a ∼×4.5\sim \times 4.5 preference gain for prompt alignment and ∼×4.0\sim \times 4.0 gain for anatomical realism compared to vanilla Stable Diffusion.

  8. Knowl 8 — Few-Shot Recognition via SeedSelect Semantic Augmentation

    empirical result

    SeedSelect can generate synthetic training data to enhance few-shot image classification. Given N∈{1,2,4,8,16}N \in \{1, 2, 4, 8, 16\} real training samples per class on ImageNet and CUB, 800 synthetic images are generated per class using SeedSelect and combined with real samples in a mix-training regime to fine-tune a pre-trained CLIP-RN50 (ResNet-50) classifier. Across all shot levels, CLIP classifiers fine-tuned on SeedSelect augmentations consistently outperform Zero-shot CLIP, CooP prompt learning, Tip-Adapter residual feature adaptation, Classifier Tuning with vanilla Stable Diffusion (CT w. SD), and Classifier Tuning with Textual Inversion (CT w. Textual Inversion). Substantial improvements are maintained even in the extreme 1-shot setting.

  9. Knowl 9 — Limitations of SeedSelect

    limitation

    SeedSelect presents three key operational limitations: (1) Reference Style Invariance: SeedSelect is biased toward the pre-trained diffusion model's natural photographic distribution and struggles to replicate non-photorealistic styles present in reference images (e.g., providing sketch references of dogs yields natural photographic dogs rather than sketch-styled generations). (2) Prompt Dependency: An optimized initial latent noise tensor zTGz_T^G is conditioned on the specific text prompt yy used during optimization and does not directly transfer or generalize to modified or novel prompts. (3) Sensitivity to Extreme Data Scarcity: For concepts that are exceptionally sparse in the pre-training corpus (having only a few isolated occurrences in LAION-2B), the internal generative representation is insufficiently formed, preventing SeedSelect from generating high-quality images.

Coverage note — None was omitted; all key contributions including long-tail characterization, optimization objective, algorithmic acceleration, empirical results across datasets, human studies, few-shot recognition, and limitations are fully covered.

References

  1. 1.Avrahami, O.; Hayes, T.; Gafni, O.; Gupta, S.; Taigman, Y.; Parikh, D.; Lischinski, D.; Fried, O.; and Yin, X. 2023. Spa-Text: Spatio-Textual Representation for Controllable Image Generation. CVPR.
  2. 2.Azizi, S.; Kornblith, S.; Saharia, C.; Norouzi, M.; and Fleet, D. J. 2023. Synthetic Data from Diffusion Models Improves ImageNet Classification. ArXiv.
  3. 3.Balaji, Y.; Nah, S.; Huang, X.; Vahdat, A.; Song, J.; Kreis, K.; Aittala, M.; Aila, T.; Laine, S.; Catanzaro, B.; et al. 2022. ediffi: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324.
  4. 4.Chefer, H.; Alaluf, Y.; Vinker, Y.; Wolf, L.; and Cohen-Or, D. 2023. Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models. SIGGRAPH.
  5. 5.Chou, P.-Y.; Kao, Y.-Y.; and Lin, C.-H. 2023. Fine-grained Visual Classification with High-temperature Refinement and Background Suppression. ArXiv.
  6. 6.Deng, J.; Dong, W.; Socher, R.; Li, L.; Kai Li; and Li Fei-Fei. 2009. ImageNet: A large-scale hierarchical image database. In cvpr.
  7. 7.Dhariwal, P.; and Nichol, A. 2021. Diffusion models beat gans on image synthesis. NeurIPS.
  8. 8.Efron, B. 1992. Bootstrap methods: another look at the jackknife. In Breakthroughs in statistics: Methodology and distribution.
  9. 9.Feng, W.; He, X.; Fu, T.-J.; Jampani, V.; Akula, A.; Narayana, P.; Basu, S.; Wang, X. E.; and Wang, W. Y. 2023. Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis. ICLR.
  10. 10.Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A. H.; Chechik, G.; and Cohen-Or, D. 2023a. An image is worth one word: Personalizing text-to-image generation using textual inversion. ICLR.
  11. 11.Gal, R.; Arar, M.; Atzmon, Y.; Bermano, A. H.; Chechik, G.; and Cohen-Or, D. 2023b. Designing an Encoder for Fast Personalization of Text-to-Image Models. arXiv preprint arXiv:2302.12228.
  12. 12.He, R.; Sun, S.; Yu, X.; Xue, C.; Zhang, W.; Torr, P. H. S.; Bai, S.; and Qi, X. 2023. Is synthetic data from generative models ready for image recognition? ICLR.
  13. 13.Ho, J.; and Salimans, T. 2021. Classifier-free diffusion guidance. NeurIPS workshop on Deep Generative Models and Downstream Applications.
  14. 14.Karras, T.; Aittala, M.; Aila, T.; and Laine, S. 2022. Elucidating the design space of diffusion-based generative models. NeurIPS.
  15. 15.Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; and Krishnan, D. 2020. Supervised contrastive learning. NeurIPS.
  16. 16.Lian, L.; Li, B.; Yala, A.; and Darrell, T. 2023. LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models. arXiv preprint arXiv:2305.13655.
  17. 17.Liu, N.; Li, S.; Du, Y.; Torralba, A.; and Tenenbaum, J. B. 2022. Compositional visual generation with composable diffusion models. In ECCV.
  18. 18.Liu, V.; and Chilton, L. B. 2022. Design guidelines for prompt engineering text-to-image generative models. In ACM SIGCHI.
  19. 19.Marcus, G.; Davis, E.; and Aaronson, S. 2022. A very preliminary analysis of dall-e 2. arXiv preprint arXiv:2204.13807.
  20. 20.Naeem, M. F.; Oh, S. J.; Uh, Y.; Choi, Y.; and Yoo, J. 2020. Reliable Fidelity and Diversity Metrics for Generative Models. Proceedings of Machine Learning Research. PMLR.
  21. 21.Nichol, A.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; and Chen, M. 2022. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. NeurIPS.
  22. 22.Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125.
  23. 23.Rassin, R.; Hirsch, E.; Glickman, D.; Ravfogel, S.; Goldberg, Y.; and Chechik, G. 2023. Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map Alignment. ArXiv.
  24. 24.Richardson, E.; and Weiss, Y. 2018. On gans and gmms. NeurIPS.
  25. 25.Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022. High-resolution image synthesis with latent diffusion models. In CVPR.
  26. 26.Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; and Aberman, K. 2023. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. CVPR.
  27. 27.Ryali, C. K.; Hu, Y.-T.; Bolya, D.; Wei, C.; Fan, H.; Huang, P.-Y. B.; Aggarwal, V.; Chowdhury, A.; Poursaeed, O.; Hoffman, J.; Malik, J.; Li, Y.; and Feichtenhofer, C. 2023. Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles. ICML.
  28. 28.Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E.; Ghasemipour, S. K. S.; Ayan, B. K.; Mahdavi, S. S.; Lopes, R. G.; et al. 2022. Photorealistic text-to-image diffusion models with deep language understanding. NeurIPS.
  29. 29.Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. 2022. Laion-5b: An open large-scale dataset for training next generation image-text models. NeurIPS.
  30. 30.Tewel, Y.; Gal, R.; Chechik, G.; and Atzmon, Y. 2023. Key-Locked Rank One Editing for Text-to-Image Personalization. SIGGRAPH.
  31. 31.Thorndike, R. L. 1953. Who belongs in the family? Psychometrika.
  32. 32.Tu, Z.; Talebi, H.; Zhang, H.; Yang, F.; Milanfar, P.; Bovik, A.; and Li, Y. 2022. MaxViT: Multi-Axis Vision Transformer. ECCV.
  33. 33.Van Horn, G.; Mac Aodha, O.; Song, Y.; Cui, Y.; Sun, C.; Shepard, A.; Adam, H.; Perona, P.; and Belongie, S. 2018. The inaturalist species classification and detection dataset. In CVPR.
  34. 34.Wah, C.; Branson, S.; Welinder, P.; Perona, P.; and Belongie, S. 2011. The Caltech-UCSD Birds-200-2011 Dataset. Technical Report CNS-TR-2011-001, California Institute of Technology.
  35. 35.Wang, Z. J.; Montoya, E.; Munechika, D.; Yang, H.; Hoover, B.; and Chau, D. H. 2022. DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models. arXiv preprint arXiv:2210.14896.
  36. 36.Zhang, L.; and Agrawala, M. 2023. Adding Conditional Control to Text-to-Image Diffusion Models. arXiv preprint arXiv:2302.05543.
  37. 37.Zhang, R.; Fang, R.; Zhang, W.; Gao, P.; Li, K.; Dai, J.; Qiao, Y. J.; and Li, H. 2022. Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling. ECCV.
  38. 38.Zhang, Y.; Kang, B.; Hooi, B.; Yan, S.; and Feng, J. 2021. Deep long-tailed learning: A survey. arXiv preprint arXiv:2110.04596.
  39. 39.Zhou, K.; Yang, J.; Loy, C. C.; and Liu, Z. 2021. Learning to Prompt for Vision-Language Models. International Journal of Computer Vision.

Citation

MLA
Samuel, D., et al. “Generating Images of Rare Concepts Using Pre-trained Diffusion Models”. arXiv, 2023, http://arxiv.org/abs/2304.14530v3.
APA
Samuel, D., Ben-Ari, R., Raviv, S., Darshan, N., & Chechik, G. (2023). Generating images of rare concepts using pre-trained diffusion models. arXiv. http://arxiv.org/abs/2304.14530v3
Chicago
Samuel, D., R. Ben-Ari, S. Raviv, N. Darshan, and G. Chechik. 2023. “Generating Images of Rare Concepts Using Pre-trained Diffusion Models”. arXiv. http://arxiv.org/abs/2304.14530v3.
Harvard
Samuel, D. et al. (2023) “Generating images of rare concepts using pre-trained diffusion models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.14530v3.
Vancouver
1. Samuel D, Ben-Ari R, Raviv S, Darshan N, Chechik G (2023) Generating images of rare concepts using pre-trained diffusion models. arXiv

BibTeX

@article{samuel2023generating,
  title = {Generating images of rare concepts using pre-trained diffusion models},
  author = {Samuel, Dvir and Ben-Ari, Rami and Raviv, Simon and Darshan, Nir and Chechik, Gal},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.14530v3},
  eprint = {2304.14530}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF