A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image Models

James Urquhart AllinghamJie RenMichael W. DusenberryXiuye GuYin CuiDustin TranJeremiah Zhe LiuBalaji Lakshminarayanan

article2023ICML71 citations

Proposes a bias-corrected zero-shot prompt weighting algorithm that automatically scores and ensembles prompts for text-image models without needing labeled validation data or manual prompt engineering.

Listen

Modern text-image models can classify images into novel categories without task-specific training, a capability known as zero-shot classification. To achieve high accuracy, however, these systems depend heavily on prompt engineering—the process of selecting descriptive text templates to accompany class names. Hand-crafting and tuning these prompt sets across different visual domains is labor-intensive and typically requires access to labeled validation data, which undermines the practical, out-of-the-box utility of zero-shot artificial intelligence.

The article demonstrates an automated method called Zero-shot Prompt Ensembling (ZPE) to score, weight, and select effective prompts from a large candidate pool for specific downstream tasks without using labeled validation data or model retraining.

The researchers developed a scoring framework that evaluates how well candidate prompts align with unlabeled target images. To prevent common scoring failures, they introduced an optimization-free normalization technique that adjusts raw scores against reference data from pre-training and test distributions, followed by softmax weighting to suppress unhelpful prompts. The approach was evaluated across 16 benchmark image datasets—including ImageNet, four robustness variants, and 11 fine-grained classification tasks—using established vision-language architectures such as CLIP and LiT with candidate pools containing up to 426 prompts.

The evaluation produced several key findings. First, naive prompt scoring based purely on confidence produces significant distortions because models over-score prompts containing frequent pre-training words or incidental concepts present in background images. Second, normalizing scores against general image distributions successfully eliminates these biases. Third, ZPE consistently outperformed both uniform prompt ensembling and extensively hand-tuned prompt baselines. On average across all 16 datasets, ZPE prompt selection achieved 67.73% accuracy on CLIP (compared to 67.29% for hand-crafted prompts and 66.06% for equal-weight pools) and 77.38% on LiT (compared to 76.51% for hand-crafted and 74.74% for equal-weight pools). Finally, the approach proved highly efficient, requiring as few as 5,000 reference images and a small fraction of unlabeled test images to calculate robust prompt weights.

These findings indicate that organizations can bypass costly, manual prompt design and remove the dependency on labeled validation data when deploying zero-shot vision systems. Because ZPE operates without iterative optimization or complex hyperparameter sweeps, it provides an interpretable, plug-and-play enhancement that lowers deployment costs and reduces operational timelines without introducing additional training risks.

Organizations deploying vision-language models should consider replacing manual prompt curation with automated scoring pipelines using ZPE weighted averaging or threshold-based prompt selection. Practitioners should maintain a diverse, domain-relevant pool of candidate prompts to maximize performance. Further engineering efforts should explore per-image prompt dynamic weighting and prompt-combination scoring to capture additional accuracy gains.

Confidence in these findings is strong across diverse benchmarks and model architectures. However, decision-makers should note that ZPE performance remains bounded by the diversity and baseline quality of the initial prompt candidate pool, and narrow fine-grained domains require adequate domain-specific templates to fully realize the method's accuracy benefits.

Allingham et al (2023).pdf

No sufficiently relevant recommendations were found.

Cover for A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image Models

Table of Contents

  • 1. Introduction
  • 2. Background
  • 3. Zero-shot weighted prompt ensembling
  • 3.1. A simple baseline - max logit scoring
  • 3.2. Tackling frequency biases via logit normalization
  • 3.3. Handling long-tails via softmax weighting
  • 3.4. Prompt selection
  • 4. Experimental Evaluation
  • 4.1. Creating a pool of prompts
  • 4.2. ZPE weighted average
  • 4.3. ZPE prompt selection
  • 4.4. Ablation studies and sensitivity analyses
  • 4.4.1. NORMALIZATION SCHEMES ABLATION
  • 4.4.2. WEIGHTING SCHEMES ABLATION
  • 4.4.3. MODEL ARCHITECTURE SENSITIVITY
  • 4.4.4. OTHER ABLATIONS AND SENSITIVITY ANALYSES
  • 5. Related Work
  • 6. Discussion and Conclusion
  • Acknowledgements
  • References
  • A. Additional Experimental Results
  • A.1. Sensitivity to the size of the pool set
  • A.2. Sensitivity to the number of random images
  • A.3. Sensitivity of the number of test images for prompt score estimation
  • A.4. Initial investigation into per-example scoring
  • B. Additional Experimental Details
  • B.1. Dataset Details
  • B.2. Implementation Details
  • C. Per-dataset Prompt Scores
  • D. The Prompt Pool

Knowls

  1. Knowl 1 — Bias-corrected maximum-logit prompt scoring

    model/method

    For a pool of PP prompt templates and a classification task with CC classes, let xn∈RDx_n\in\mathbb{R}^D be the embedding of unlabeled target image nn, and let tp,c∈RDt_{p,c}\in\mathbb{R}^D be the text embedding for prompt pp composed with class name cc. Define the target and reference mean logits for each prompt and class as μp,ctarget=N−1∑n=1Nxn⊤tp,c\mu^{\mathrm{target}}_{p,c}=N^{-1}\sum_{n=1}^{N}x_n^\top t_{p,c} and μp,cref=N0−1∑j=1N0rj⊤tp,c\mu^{\mathrm{ref}}_{p,c}=N_0^{-1}\sum_{j=1}^{N_0}r_j^\top t_{p,c}, where rj∈RDr_j\in\mathbb{R}^D are reference-image embeddings. The corrected logit for target image nn is ℓn,p,c=xn⊤tp,c−12(μp,cref+μp,ctarget)\ell_{n,p,c}=x_n^\top t_{p,c}-\tfrac12(\mu^{\mathrm{ref}}_{p,c}+\mu^{\mathrm{target}}_{p,c}), and prompt pp receives score sp=N−1∑n=1Nmax⁡1≤c≤Cℓn,p,cs_p=N^{-1}\sum_{n=1}^{N}\max_{1\le c\le C}\ell_{n,p,c}. Thus scoring centers each prompt’s class logits using both a broad reference-image distribution and the unlabeled target-image distribution before taking the maximum over classes and averaging over target images. The reference images address high logits associated with frequent pretraining words; target-image centering reduces scores from concepts common in the target images but unrelated to the classes. Because the text-image model’s pretraining data were unavailable, the experiments used the first 20,000 LAION400M images as the reference set; performance was nearly unchanged using 5,000 or 10,000 reference images, and using 10% rather than all target images to estimate the target means.

  2. Knowl 2 — Softmax-weighted zero-shot prompt ensemble

    model/method

    Given a text-image classifier, a set of PP prompt templates, class-specific text embeddings tp,ct_{p,c}, and corrected prompt scores sps_p, the method assigns each prompt the weight wp=exp⁡(sp)/∑q=1Pexp⁡(sq)w_p=\exp(s_p)/\sum_{q=1}^{P}\exp(s_q). For an image embedding xx, it predicts the class maximizing ∑p=1Pwpx⊤tp,c\sum_{p=1}^{P}w_p x^\top t_{p,c}. Equivalently, this is a softmax-weighted average of the prompt-specific class logits; an overall factor of 1/P1/P does not change the predicted class. Scores and weights are computed without labeled validation examples and without optimization. In the reported implementation, the target-image means used to score prompts are estimated from unlabeled images in the target dataset, so the procedure uses the target image collection but not its labels.

  3. Knowl 3 — Robust outlier-based prompt selection

    model/method

    For prompt scores s1,…,sPs_1,\ldots,s_P, the selection procedure assumes that most templates in a large pool are irrelevant to the target task and that useful templates are score outliers. It calculates the pool median m=median⁡p(sp)m=\operatorname{median}_p(s_p) and median absolute deviation d=median⁡p(∣sp−m∣)d=\operatorname{median}_p(|s_p-m|), then assigns each prompt the robust standardized score zp=(sp−m)/dz_p=(s_p-m)/d. A prompt is retained when zp>τz_p>\tau; the classifier ensembles the retained prompts’ logits. This relative threshold avoids setting a cutoff in the task-dependent units of raw prompt scores. In the experiments, τ=0.5\tau=0.5 was used for ImageNet and its variants and τ=2.0\tau=2.0 for the fine-grained datasets; these values were selected by sweeping candidate thresholds and choosing the best average classification performance across datasets.

  4. Knowl 4 — Why uncorrected maximum-logit scores can misrank prompts

    empirical result

    The naive score for a prompt—its maximum class logit averaged over target images—can be high even when the prompt is poorly matched to the classification task. The paper identifies two sources: words frequent in pretraining can elicit high logits across images, and prompts can describe concepts frequent in target images but unrelated to the labels. For example, prompts containing “person” ranked highly for ImageNet and Sun397 even though “person” was not the class concept of interest; Sun397 scene images often contain people despite being labeled by locations. This shows why a raw maximum logit is not a reliable measure of prompt suitability and motivates centering against both reference and target image means.

  5. Knowl 5 — ZPE improves average accuracy over prompt baselines

    empirical result

    The evaluation used a pool of 247 unique templates assembled from existing prompt sets and tested on ImageNet, four ImageNet variants, and 11 fine-grained classification datasets. The reported averages cover all 16 datasets and are zero-shot accuracies in percent. For CLIP ViT-B/16, the equal-average pool achieved 66.06%, the manually tuned hand-crafted prompt ensemble 67.29%, ZPE softmax weighting 67.44%, and ZPE prompt selection 67.73%. For LiT ViT-L/16, the corresponding results were 74.74%, 76.51%, 76.79%, and 77.38%. On ImageNet alone, CLIP ViT-B/16 achieved 67.59% with the equal-average pool, 68.31% with hand-crafted prompts, 68.56% with ZPE weighting, and 68.60% with ZPE selection; LiT ViT-L/16 achieved 77.49%, 78.55%, 78.90%, and 79.26%, respectively. ZPE outperformed the equal-average pool and hand-crafted baseline on these aggregate comparisons while avoiding task-specific manual prompt tuning.

  6. Knowl 6 — Prompt selection transfers across text-image architectures

    empirical result

    Across eight evaluated CLIP and LiT architectures, ZPE prompt selection improved the average accuracy over all 16 evaluation datasets relative to both the equal-average pool and the hand-crafted prompt ensemble. The triples below give equal-average pool, hand-crafted ensemble, and ZPE prompt-selection accuracy, respectively, in percent: CLIP ResNet-50, 52.71/55.15/55.46; CLIP ResNet-101, 57.04/58.90/59.21; CLIP ViT-B/32, 61.02/62.76/63.18; CLIP ViT-B/16, 66.06/67.29/67.73; CLIP ViT-L/14, 73.82/75.85/76.13; LiT ViT-B/32, 64.94/66.33/67.58; LiT ViT-B/16, 68.89/70.94/71.71; and LiT ViT-L/16, 74.74/76.51/77.38. The hand-crafted prompts were designed for CLIP, whereas ZPE scoring was applied to each model’s prompt pool without model-specific manual tuning.

  7. Knowl 7 — Ablation supports combining reference and target centering

    empirical result

    For CLIP ViT-B/16, the normalization ablation reports average accuracy over all 16 datasets. Under softmax-weighted averaging, no normalization achieved 66.92%, reference-only centering EpretrainE_{\mathrm{pretrain}} 67.42%, target-only centering EtestE_{\mathrm{test}} 67.00%, the variant averaging over both reference images and classes Epretrain∗E^{*}_{\mathrm{pretrain}} 67.17%, and the proposed combination of reference- and target-image centering 67.44%. Under prompt selection, the corresponding values were 67.15%, 67.69%, 67.09%, 67.39%, and 67.73%. The combined correction performed best in both ensemble modes, while reference-only centering was the strongest single correction; target-only centering alone was less effective than no normalization in these aggregate comparisons.

  8. Knowl 8 — Softmax weighting suppresses the long tail of weak prompts

    empirical result

    Prompt scores over a large pool have a long tail: a few templates score highly while many weak templates receive small but nonzero scores. On CLIP ViT-B/16, average accuracy across all 16 datasets for the weighted ensemble was 66.18% using raw scores as weights, 67.30% using scores raised to the tenth power, and 67.44% using softmax-normalized scores. For prompt selection, the same weighting comparisons yielded 67.70%, 67.72%, and 67.73%. Thus softmax weighting gave the best aggregate result and particularly helped the weighted-average ensemble, where many individually small weights could otherwise accumulate.

  9. Knowl 9 — Prompt-pool size and quality affect performance

    empirical result

    In a CLIP ViT-B/16 evaluation averaged over all 16 datasets, the original 247-template pool produced 67.44% accuracy with softmax weighting and 67.73% with prompt selection. An 80-template pool produced 67.13% and 67.26%, respectively; an expanded 426-template pool, which added 179 templates generated by ChatGPT, produced 67.30% and 67.54%. Applying ZPE weights to the hand-crafted prompts increased their equal-average result from 67.29% to 67.43%. The added templates did not improve on the 247-template pool, indicating that pool quality as well as size matters; nevertheless, ZPE selection with the 426-template pool remained above the hand-crafted equal-average result.

  10. Knowl 10 — Dataset-level independent scoring leaves room for improvement

    limitation

    ZPE requires a large, varied pool of high-quality prompt templates and scores each template independently, so it does not model complementary combinations of prompts. Its reported scores are dataset-level rather than image-specific, even though a template useful for one image may not suit another. An exploratory CLIP ViT-B/16 comparison found that per-example scoring of the pool increased average accuracy over all 16 datasets from 67.44% to 67.60% and fine-grained-dataset average accuracy from 70.71% to 71.01%, but reduced ImageNet accuracy from 68.56% to 67.97%; it was therefore not uniformly better. The paper presents per-example scoring and scoring prompt combinations as directions for further work, not as established improvements.

Coverage note — Detailed per-dataset prompt rankings and every individual benchmark score are omitted because the aggregate results and the diagnosed prompt-bias examples capture the main contribution without reproducing lengthy dataset-specific listings.

References

  1. 1.Allingham, J. U., Wenzel, F., Mariet, Z. E., Mustafa, B., Puigcerver, J., Houlsby, N., Jerfel, G., Fortuin, V., Lakshminarayanan, B., Snoek, J., Tran, D., Ruiz, C. R., and Jenatton, R. Sparse MoEs meet efficient ensembles. TMLR, 2022.
  2. 2.Best, W. R. and Cowper, D. C. The ratio of observed-to-expected mortality as a quality of care indicator in non-surgical va patients. Medical care, 32(4):390–400, 1994.
  3. 3.Beyer, L., Zhai, X., and Kolesnikov, A. Big Vision. https://github.com/google-research/big_vision, 2022.
  4. 4.Bossard, L., Guillaumin, M., and Van Gool, L. Food-101 – mining discriminative components with random forests. In ECCV, 2014.
  5. 5.Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q. JAX: composable transformations of Python+NumPy programs, 2018. URL http://github.com/google/jax.
  6. 6.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), NeurIPS, 2020.
  7. 7.Casella, G. and Berger, R. L. Statistical inference. Cengage Learning, 2021.
  8. 8.Cheng, G., Han, J., and Lu, X. Remote sensing image scene classification: Benchmark and state of the art. Proceedings of the IEEE, 105(10):1865–1883, Oct 2017.
  9. 9.Cherti, M., Beaumont, R., Wightman, R., Wortsman, M., Ilharco, G., Gordon, C., Schuhmann, C., Schmidt, L., and Jitsev, J. Reproducible scaling laws for contrastive language-image learning. CoRR, abs/2212.07143, 2022.
  10. 10.Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A. Describing textures in the wild. In CVPR, pp. 3606–3613, 2014.
  11. 11.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR. OpenReview.net, 2021.
  12. 12.Fei-Fei, L., Fergus, R., and Perona, P. Learning generative visual models from few training examples: An incremental Bayesian approach tested on 101 object categories. Computer Vision and Pattern Recognition Workshop, 2004.
  13. 13.Galvan-Turner, V. B., Chang, J., Ziogas, A., and Bristow, R. E. Observed-to-expected ratio for adherence to treatment guidelines as a quality of care indicator for ovarian cancer. Gynecologic oncology, 139(3):495–499, 2015.
  14. 14.Gao, T., Fisch, A., and Chen, D. Making pre-trained language models better few-shot learners. In Zong, C., Xia, F., Li, W., and Navigli, R. (eds.), ACL/IJCNLP, pp. 3816–3830. Association for Computational Linguistics, 2021.
  15. 15.Ge, Y., Ren, J., Wang, Y., Gallagher, A., Yang, M.-H., Itti, L., Adam, H., Lakshminarayanan, B., and Zhao, J. Improving zero-shot generalization and robustness of multi-modal models. arXiv preprint arXiv:2212.01758, 2022.
  16. 16.Heek, J., Levskaya, A., Oliver, A., Ritter, M., Rondepierre, B., Steiner, A., and van Zee, M. Flax: A neural network library and ecosystem for JAX, 2020. URL http://github.com/google/flax.
  17. 17.Helber, P., Bischke, B., Dengel, A., and Borth, D. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification, 2017.
  18. 18.Hendrycks, D., Basart, S., Mazeika, M., Mostajabi, M., Steinhardt, J., and Song, D. Scaling out-of-distribution detection for real-world settings. arXiv preprint arXiv:1911.11132, 2019.
  19. 19.Hendrycks, D., Basart, S., Mu, N., Kadavath, S., Wang, F., Dorundo, E., Desai, R., Zhu, T., Parajuli, S., Guo, M., Song, D., Steinhardt, J., and Gilmer, J. The many faces of robustness: A critical analysis of out-of-distribution generalization. In ICCV, pp. 8320–8329. IEEE, 2021a.
  20. 20.Hendrycks, D., Zhao, K., Basart, S., Steinhardt, J., and Song, D. Natural adversarial examples. In CVPR, pp. 15262–15271. Computer Vision Foundation/IEEE, 2021b.
  21. 21.Jia, C., Yang, Y., Xia, Y., Chen, Y., Parekh, Z., Pham, H., Le, Q. V., Sung, Y., Li, Z., and Duerig, T. Scaling up visual and vision-language representation learning with noisy text supervision. 139:4904–4916, 2021.
  22. 22.Jones, K. S. A statistical interpretation of term specificity and its application in retrieval. Journal of documentation, 1972.
  23. 23.King, G. Unifying political methodology: The likelihood theory of statistical inference. Cambridge University Press, 1989.
  24. 24.Krause, J., Stark, M., Deng, J., and Fei-Fei, L. 3d object representations for fine-grained categorization. In Proceedings of the IEEE international conference on computer vision workshops, pp. 554–561, 2013.
  25. 25.Krizhevsky, A. Learning multiple layers of features from tiny images. Technical report, 2009.
  26. 26.MacQueen, J. Classification and analysis of multivariate observations. In 5th Berkeley Symp. Math. Statist. Probability, pp. 281–297, 1967.
  27. 27.Nado, Z., Band, N., Collier, M., Djolonga, J., Dusenberry, M., Farquhar, S., Filos, A., Havasi, M., Jenatton, R., Jerfel, G., Liu, J., Mariet, Z., Nixon, J., Padhy, S., Ren, J., Rudner, T., Wen, Y., Wenzel, F., Murphy, K., Sculley, D., Lakshminarayanan, B., Snoek, J., Gal, Y., and Tran, D. Uncertainty Baselines: Benchmarks for uncertainty & robustness in deep learning. arXiv preprint arXiv:2106.04015, 2021.
  28. 28.Nilsback, M.-E. and Zisserman, A. Automated flower classification over a large number of classes. In Proceedings of the Indian Conference on Computer Vision, Graphics and Image Processing, Dec 2008.
  29. 29.OpenAI. ChatGPT: Optimizing language models for dialogue (January 9th release). https://openai.com/blog/chatgpt/, 2022.
  30. 30.Parkhi, O. M., Vedaldi, A., Zisserman, A., and Jawahar, C. V. Cats and dogs. In IEEE Conference on Computer Vision and Pattern Recognition, 2012.
  31. 31.Pham, H., Dai, Z., Ghiasi, G., Liu, H., Yu, A. W., Luong, M., Tan, M., and Le, Q. V. Combined scaling for zero-shot transfer learning. CoRR, abs/2111.10050, 2021.
  32. 32.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. 139:8748–8763, 2021.
  33. 33.Recht, B., Roelofs, R., Schmidt, L., and Shankar, V. Do imageNet classifiers generalize to imageNet? In Chaudhuri, K. and Salakhutdinov, R. (eds.), ICML, volume 97 of Proceedings of Machine Learning Research, pp. 5389–5400. PMLR, 2019.
  34. 34.Ren, J., Liu, P. J., Fertig, E., Snoek, J., Poplin, R., Depristo, M., Dillon, J., and Lakshminarayanan, B. Likelihood ratios for out-of-distribution detection. NeurIPS, 32, 2019.
  35. 35.Ren, J., Fort, S., Liu, J., Roy, A. G., Padhy, S., and Lakshminarayanan, B. A simple fix to mahalanobis distance for improving near-ood detection. arXiv preprint arXiv:2106.09022, 2021.
  36. 36.Ren, J., Luo, J., Zhao, Y., Krishna, K., Saleh, M., Lakshminarayanan, B., and Liu, P. J. Out-of-distribution detection and selective generation for conditional language models. ICLR, 2022.
  37. 37.Rousseeuw, P. J. and Croux, C. Alternatives to the median absolute deviation. Journal of the American Statistical association, 88(424):1273–1283, 1993.
  38. 38.Rubin, O., Herzig, J., and Berant, J. Learning to retrieve prompts for in-context learning. In Carpuat, M., de Marneffe, M., and Ruíz, I. V. M. (eds.), NAACL.
  39. 39.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M. S., Berg, A. C., and Fei-Fei, L. ImageNet large scale visual recognition challenge. Int. J. Comput. Vis., 115(3):211–252, 2015.
  40. 40.Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A. LAION-400M: open dataset of CLIP-filtered 400 million image-text pairs. CoRR, abs/2111.02114, 2021.
  41. 41.Shin, T., Razeghi, Y., IV, R. L. L., Wallace, E., and Singh, S. AutoPrompt: Eliciting knowledge from language models with automatically generated prompts. In Webber, B., Cohn, T., He, Y., and Liu, Y. (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pp. 4222–4235. Association for Computational Linguistics, 2020.
  42. 42.Shu, M., Nie, W., Huang, D., Yu, Z., Goldstein, T., Anandkumar, A., and Xiao, C. Test-time prompt tuning for zero-shot generalization in vision-language models. CoRR, abs/2209.07511, 2022. doi: 10.48550/arXiv.2209.07511.
  43. 43.Sorensen, T., Robinson, J., Rytting, C. M., Shaw, A. G., Rogers, K. J., Delorey, A. P., Khalil, M., Fulda, N., and Wingate, D. An information-theoretic approach to prompt engineering without ground truth labels. In Muresan, S., Nakov, P., and Villavicencio, A. (eds.), ACL, pp. 819–862. Association for Computational Linguistics, 2022.
  44. 44.Wang, H., Ge, S., Lipton, Z. C., and Xing, E. P. Learning robust global representations by penalizing local predictive power. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alche-Buc, F., Fox, E. B., and Garnett, R. (eds.), NeurIPS, pp. 10506–10518, 2019.
  45. 45.Xiao, J., Hays, J., Ehinger, K. A., Oliva, A., and Torralba, A. SUN database: Large-scale scene recognition from abbey to zoo. In CVPR, pp. 3485–3492, June 2010.
  46. 46.Zhai, X., Wang, X., Mustafa, B., Steiner, A., Keysers, D., Kolesnikov, A., and Beyer, L. Lit: Zero-shot transfer with locked-image text tuning. pp. 18102–18112, 2022.
  47. 47.Zhou, K., Yang, J., Loy, C. C., and Liu, Z. Conditional prompt learning for vision-language models. In CVPR, pp. 16816–16825, 2022a.
  48. 48.Zhou, K., Yang, J., Loy, C. C., and Liu, Z. Learning to prompt for vision-language models. Int. J. Comput. Vis., 130(9):2337–2348, 2022b.
  49. 49.Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J. Large language models are human-level prompt engineers. CoRR, abs/2211.01910, 2022c.

Citation

MLA
Allingham, J. U., et al. “A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image Models”. International Conference on Machine Learning, vol. 202, 2023, pp. 547–68, https://proceedings.mlr.press/v202/allingham23a.html.
APA
Allingham, J. U., Ren, J., Dusenberry, M. W., Gu, X., Cui, Y., Tran, D., Liu, J. Z., & Lakshminarayanan, B. (2023). A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image Models. International Conference on Machine Learning, 202, 547–568. https://proceedings.mlr.press/v202/allingham23a.html
Chicago
Allingham, J. U., J. Ren, M. W. Dusenberry, et al. 2023. “A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image Models”. International Conference on Machine Learning 202: 547–68. https://proceedings.mlr.press/v202/allingham23a.html.
Harvard
Allingham, J.U. et al. (2023) “A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image Models”, International Conference on Machine Learning. PMLR, pp. 547–568. Available at: https://proceedings.mlr.press/v202/allingham23a.html.
Vancouver
1. Allingham JU, Ren J, Dusenberry MW, Gu X, Cui Y, Tran D, Liu JZ, Lakshminarayanan B (2023) A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image Models. In: International Conference on Machine Learning. PMLR, pp 547–568

BibTeX

@InProceedings{pmlr-v202-allingham23a,
  title = 	 {A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image Models},
  author =       {Allingham, James Urquhart and Ren, Jie and Dusenberry, Michael W and Gu, Xiuye and Cui, Yin and Tran, Dustin and Liu, Jeremiah Zhe and Lakshminarayanan, Balaji},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {547--568},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/allingham23a/allingham23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/allingham23a.html},
  abstract = 	 {Contrastively trained text-image models have the remarkable ability to perform zero-shot classification, that is, classifying previously unseen images into categories that the model has never been explicitly trained to identify. However, these zero-shot classifiers need prompt engineering to achieve high accuracy. Prompt engineering typically requires hand-crafting a set of prompts for individual downstream tasks. In this work, we aim to automate this prompt engineering and improve zero-shot accuracy through prompt ensembling. In particular, we ask “Given a large pool of prompts, can we automatically score the prompts and ensemble those that are most suitable for a particular downstream dataset, without needing access to labeled validation data?". We demonstrate that this is possible. In doing so, we identify several pathologies in a naive prompt scoring method where the score can be easily overconfident due to biases in pre-training and test data, and we propose a novel prompt scoring method that corrects for the biases. Using our proposed scoring method to create a weighted average prompt ensemble, our method overall outperforms equal average ensemble, as well as hand-crafted prompts, on ImageNet, 4 of its variants, and 11 fine-grained classification benchmarks. while being fully automatic, optimization-free, and not requiring access to labeled validation data.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/