Automatic image annotation and retrieval using cross-media relevance models

J. JeonV. LavrenkoR. Manmatha

article2003SIGIR1,353 citationsSIGIR Test of Time Award Honorable Mention

Presents a cross-media relevance model that learns joint distributions of visual region clusters and keywords to automatically annotate and retrieve images, doubling the mean precision of previous machine translation approaches.

Listen

Organizations managing large visual libraries have traditionally relied on manual image annotation to enable keyword-based search. However, manual tagging is expensive, labor-intensive, and prone to human error and inconsistency, while traditional content-based retrieval systems require technical queries (such as color or texture) that non-specialists find difficult to use. To overcome these barriers, the article evaluates a statistical framework known as the Cross-Media Relevance Model (CMRM), which automatically annotates images with keywords and performs ranked image retrieval based on natural text queries.

The evaluated approach models images by first segmenting them into visual regions and clustering these regions into discrete visual tokens termed "blobs." Using a standardized benchmark dataset of 5,000 stock photos spanning 371 vocabulary words and 500 visual blobs, the model learns the joint statistical distribution between words and visual features from training examples. This framework is applied in three configurations: a fixed-length annotation model (FACMRM) that assigns the top keywords to an image; a probabilistic annotation model (PACMRM) that scores images using language modeling; and a direct retrieval model (DRCMRM) that maps text queries into visual blob distributions to rank uncaptioned images.

The findings show that the proposed relevance models dramatically outperform existing automatic annotation methods. On a common 70-query benchmark, the fixed-length model achieved a mean precision of 0.33 and a mean recall of 0.37, roughly doubling the performance of state-of-the-art translation models (0.14 precision, 0.24 recall) and improving nearly fivefold over basic co-occurrence models (0.07 precision, 0.11 recall). When evaluated on the top 49 query words, mean precision reached 0.41 compared to 0.20 for the translation baseline. For ranked image search, the direct retrieval model consistently outperformed the probabilistic annotation model across queries of varying length, achieving average precision scores above 0.20 on complex, multi-word queries.

These results demonstrate that formal information retrieval models can significantly reduce the costs and operational bottlenecks associated with manual indexing while improving search reliability. Furthermore, the model proved capable of surfacing relevant images even when original human annotations were incomplete or flawed. Decision-makers should consider implementing cross-media relevance modeling to automate cataloging and quality-check human-curated collections, giving preference to direct retrieval architectures when deploying ranked multi-word search functionality.

Nevertheless, several technical limitations warrant measured confidence. The visual segmentation and clustering process remains imperfect, occasionally assigning identical visual features to semantically unrelated objects. Future development should focus on testing larger and more diverse datasets, incorporating advanced continuous visual features, and extending the model from isolated keywords to full natural language captions.

Cover for Automatic image annotation and retrieval using cross-media relevance models

Abstract

Libraries have traditionally used manual image annotation for indexing and then later retrieving their image collections. However, manual image annotation is an expensive and labor intensive procedure and hence there has been great interest in coming up with automatic ways to retrieve images based on content. Here, we propose an automatic approach to annotating and retrieving images based on a training set of images. We assume that regions in an image can be described using a small vocabulary of blobs. Blobs are generated from image features using clustering. Given a training set of images with annotations, we show that probabilistic models allow us to predict the probability of generating a word given the blobs in an image. This may be used to automatically annotate and retrieve images given a word as a query. We show that relevance models allow us to derive these probabilities in a natural way. Experiments show that the annotation performance of this cross-media relevance model is almost six times as good (in terms of mean precision) than a model based on word-blob co-occurrence model and twice as good as a state of the art model derived from machine translation. Our approach shows the usefulness of using formal information retrieval models for the task of image annotation and retrieval.

Table of Contents

  • 1. INTRODUCTION
  • 2. RELATED WORK
  • 3. DISCRETE FEATURES IN IMAGES
  • 4. CROSS-MEDIA RELEVANCE MODELS
  • 4.1 A Model of Image Annotation
  • 4.1.1 Using the model for Image Annotation
  • 4.2 Two Models of Image Retrieval
  • 4.2.1 Annotation-based Retrieval Model
  • 4.2.2 Direct Retrieval Model (DRCMRM)
  • 5. EXPERIMENTAL RESULTS
  • 5.1 Dataset
  • 5.2 Automatic Image Annotation
  • 5.2.1 Finding model parameters
  • 5.2.2 Model Comparison
  • 5.3 Evaluation of Ranked Retrieval
  • 5.4 Illustrative Examples
  • 6. CONCLUSIONS AND FUTURE WORK
  • 7. ACKNOWLEDGMENTS
  • 8. REFERENCES

Knowls

  1. Knowl 1 — Cross-Media Relevance Model Joint Probability Formulation

    model/method

    The Cross-Media Relevance Model (CMRM) estimates the joint distribution of visual features (discrete image regions called blobs) and textual keywords (words) without requiring a one-to-one alignment between individual blobs and words. Let an unannotated image II be represented by a set of mm discrete blobs I={b1,…,bm}I = \{b_1, \dots, b_m\}, and let ww be a word from a vocabulary VV.

    The probability P(w∣I)P(w \mid I) of generating word ww given the observed blobs of II is approximated by conditioning on the blob sequence: P(w∣I)≈P(w∣b1,…,bm)=P(w,b1,…,bm)P(b1,…,bm)P(w \mid I) \approx P(w \mid b_1, \dots, b_m) = \frac{P(w, b_1, \dots, b_m)}{P(b_1, \dots, b_m)}

    Given a training collection TT of annotated images, where each training image J∈TJ \in T possesses a dual representation J={b1,…,bmJ;w1,…,wnJ}J = \{b_1, \dots, b_{m_J}; w_1, \dots, w_{n_J}\}, the joint probability is computed as the expectation over all training images J∈TJ \in T: P(w,b1,…,bm)=∑J∈TP(J)P(w,b1,…,bm∣J)P(w, b_1, \dots, b_m) = \sum_{J \in T} P(J) P(w, b_1, \dots, b_m \mid J)

    Assuming that occurrences of word ww and blobs b1,…,bmb_1, \dots, b_m are conditionally independent given training image JJ (treating each image as an urn containing both words and blobs), the joint distribution factorizes as: P(w,b1,…,bm)=∑J∈TP(J)P(w∣J)∏i=1mP(bi∣J)P(w, b_1, \dots, b_m) = \sum_{J \in T} P(J) P(w \mid J) \prod_{i=1}^m P(b_i \mid J)

    where the prior probability P(J)P(J) is assumed to be uniform across all training images, P(J)=1∣T∣P(J) = \frac{1}{|T|}.

  2. Knowl 2 — Smoothed Estimation of Word and Blob Probabilities in CMRM

    equation

    In the Cross-Media Relevance Model (CMRM), the probability of drawing a word ww or a blob bb from a specific training image J∈TJ \in T is computed using linear interpolation between the image-level maximum-likelihood estimate and the background collection-level frequency:

    P(w∣J)=(1−αJ)#(w,J)∣J∣+αJ#(w,T)∣T∣P(w \mid J) = (1 - \alpha_J) \frac{\#(w, J)}{|J|} + \alpha_J \frac{\#(w, T)}{|T|}

    P(b∣J)=(1−βJ)#(b,J)∣J∣+βJ#(b,T)∣T∣P(b \mid J) = (1 - \beta_J) \frac{\#(b, J)}{|J|} + \beta_J \frac{\#(b, T)}{|T|}

    where:

    • #(w,J)\#(w, J) is the number of times word ww occurs in the caption of image JJ (typically 0 or 1).
    • #(w,T)\#(w, T) is the cumulative count of word ww across all captions in the training collection TT.
    • #(b,J)\#(b, J) is the count of regions in image JJ labeled with discrete blob identifier bb.
    • #(b,T)\#(b, T) is the cumulative occurrences of blob bb across all training images in TT.
    • ∣J∣|J| is the aggregate count of all words and blobs in image JJ.
    • ∣T∣|T| is the total aggregate count of all words and blobs in the training set TT.
    • αJ∈[0,1]\alpha_J \in [0, 1] is the word smoothing hyperparameter (tuned to αJ=0.1\alpha_J = 0.1).
    • βJ∈[0,1]\beta_J \in [0, 1] is the blob smoothing hyperparameter (tuned to βJ=0.9\beta_J = 0.9).

    Distinct smoothing parameters are used because words follow a Zipfian distribution, whereas visual blob clusters have a much more uniform occurrence pattern.

  3. Knowl 3 — Direct-Retrieval Cross-Media Relevance Model (DRCMRM)

    model/method

    The Direct-Retrieval Cross-Media Relevance Model (DRCMRM) translates a multi-word textual query Q=w1…wkQ = w_1 \dots w_k into a distribution over the visual blob vocabulary BB, ranking unannotated images I={b1,…,bm}I = \{b_1, \dots, b_m\} directly by distribution similarity without producing intermediate text annotations.

    1. Query Relevance Model Estimation: Assuming query QQ is sampled from an underlying relevance model P(⋅∣Q)P(\cdot \mid Q), the probability of observing a blob b∈Bb \in B is estimated as: P(b∣Q)≈P(b∣w1,…,wk)=P(b,w1,…,wk)P(w1,…,wk)P(b \mid Q) \approx P(b \mid w_1, \dots, w_k) = \frac{P(b, w_1, \dots, w_k)}{P(w_1, \dots, w_k)} where the joint probability is estimated over the training set TT: P(b,w1,…,wk)=∑J∈TP(J)P(b∣J)∏i=1kP(wi∣J)P(b, w_1, \dots, w_k) = \sum_{J \in T} P(J) P(b \mid J) \prod_{i=1}^k P(w_i \mid J) with uniform prior P(J)=1∣T∣P(J) = \frac{1}{|T|}.

    2. Image Ranking via Negative KL Divergence: Test images II are ranked by the negative Kullback-Leibler (KL) divergence between the query model P(⋅∣Q)P(\cdot \mid Q) and the image blob model P(⋅∣I)P(\cdot \mid I): −KL(Q∥I)=∑b∈BP(b∣Q)log⁡P(b∣I)P(b∣Q)-KL(Q \parallel I) = \sum_{b \in B} P(b \mid Q) \log \frac{P(b \mid I)}{P(b \mid Q)} where P(b∣I)P(b \mid I) is the smoothed estimate of blob bb in test image II.

  4. Knowl 4 — Fixed Annotation-Based Cross-Media Relevance Model (FACMRM)

    model/method

    The Fixed Annotation-based Cross-Media Relevance Model (FACMRM) assigns a discrete set of keywords to an unannotated image I={b1,…,bm}I = \{b_1, \dots, b_m\}:

    1. Compute the posterior probability P(w∣I)≈P(w∣b1,…,bm)P(w \mid I) \approx P(w \mid b_1, \dots, b_m) for every word ww in the vocabulary VV using the Cross-Media Relevance Model joint distribution: P(w∣b1,…,bm)∝∑J∈TP(J)P(w∣J)∏i=1mP(bi∣J)P(w \mid b_1, \dots, b_m) \propto \sum_{J \in T} P(J) P(w \mid J) \prod_{i=1}^m P(b_i \mid J)

    2. Rank all vocabulary words by descending order of P(w∣I)P(w \mid I).

    3. Select the top NN words (typically N∈{3,4,5}N \in \{3, 4, 5\}) to form the discrete caption for image II.

    FACMRM produces fixed-length annotations suitable for human inspection and coordination-level keyword lookup, but because it discards word probability weights and produces uniform caption lengths, it is not well-suited for ranked retrieval with document-length normalization.

  5. Knowl 5 — Probabilistic Annotation-Based Cross-Media Relevance Model (PACMRM)

    model/method

    The Probabilistic Annotation-based Cross-Media Relevance Model (PACMRM) performs ranked retrieval by treating each unannotated image I={b1,…,bm}I = \{b_1, \dots, b_m\} as a complete language model distribution P(⋅∣I)P(\cdot \mid I) over all words w∈Vw \in V, where P(w∣I)≈P(w∣b1,…,bm)P(w \mid I) \approx P(w \mid b_1, \dots, b_m) is computed via the Cross-Media Relevance Model.

    Given a text query Q=w1…wkQ = w_1 \dots w_k, images are ranked by the query likelihood under the assumption of independent and identically distributed term generation from the image model: P(Q∣I)=∏j=1kP(wj∣I)P(Q \mid I) = \prod_{j=1}^k P(w_j \mid I)

    This continuous scoring eliminates the need to select an arbitrary annotation length cutoff and enables fine-grained ranking for multi-word queries.

  6. Knowl 6 — Benchmark Setup for Image Annotation and Retrieval on Corel Stock Photos

    experimental setup

    The experimental evaluation uses a benchmark dataset of 5,000 images sampled from 50 Corel Stock Photo CDs (100 images per CD):

    • Visual Blob Vocabulary: Each image is segmented using Normalized Cuts. Regions exceeding a minimum size threshold are parameterized with 33 color, texture, shape, and position features. KK-means clustering with K=500K = 500 clusters over training regions defines a discrete vocabulary of 500 blobs. Each image contains 1 to 10 blobs.
    • Annotation Vocabulary: Each image is associated with 1 to 5 manual keywords, resulting in a vocabulary of 371 distinct words.
    • Data Splits:
      • The dataset is initially split into 4,000 training images, 500 evaluation images (used to optimize smoothing parameters αJ=0.1\alpha_J = 0.1 and βJ=0.9\beta_J = 0.9 via F-measure), and 500 test images.
      • Final models are trained on 4,500 images (4,000 training + 500 evaluation) and evaluated on the 500 held-out test images.
  7. Knowl 7 — Automatic Image Annotation Performance Comparison

    empirical result

    Automatic image annotation was evaluated on 500 test images (assigning 5 keywords per image for fixed models) comparing the Fixed Annotation-based Cross-Media Relevance Model (FACMRM), the Co-occurrence Model, and the Machine Translation Model (Brown Model 2):

    • 70-Query Set Evaluation (the union of all single-word queries retrieving at least one relevant image across the three systems, out of 263 total single-word vocabulary queries):

      • Co-occurrence Model: Mean Precision = 0.07, Mean Recall = 0.11 (retrieved ≥1\ge 1 relevant image on 19 queries)
      • Translation Model: Mean Precision = 0.14, Mean Recall = 0.24 (retrieved ≥1\ge 1 relevant image on 49 queries)
      • FACMRM: Mean Precision = 0.33, Mean Recall = 0.37 (retrieved ≥1\ge 1 relevant image on 66 queries)
    • Top-49 Query Subset (restricted to queries where the Translation Model retrieved at least one relevant document):

      • Translation Model: Mean Precision = 0.20, Mean Recall = 0.34
      • FACMRM: Mean Precision = 0.41, Mean Recall = 0.49

    FACMRM doubles the mean precision of the Translation Model, achieves over four-fold precision improvement over the Co-occurrence Model, and substantially increases recall and vocabulary coverage.

  8. Knowl 8 — Ranked Retrieval Performance Across Query Lengths

    data/table

    Ranked retrieval performance was evaluated on the 500 test images across all single- and multi-word query combinations present in the testing set (excluding combinations appearing only once). Ground-truth relevance requires an image to contain all query words in its manual annotations. Performance is measured by non-interpolated Average Precision (AveP).

    Query length 1 word 2 words 3 words 4 words
    Number of queries 179 386 178 24
    Relevant images 1675 1647 542 67
    AveP (PACMRM) 0.1501 0.1419 0.1730 0.2364
    AveP (DRCMRM) 0.1697 0.1642 0.2030 0.2765

    The Direct-Retrieval Cross-Media Relevance Model (DRCMRM) consistently outperforms the Probabilistic Annotation-based Model (PACMRM) across all query lengths. The improvements are statistically significant at the 5% confidence level under the Wilcoxon test for 1-, 2-, and 3-word queries (the 4-word query set lacked statistical significance due to the small sample size of 24 queries). Retrieval accuracy increases as query length increases from 2 to 4 words.

  9. Knowl 9 — Effect of Manual Annotation Errors on Model Performance

    empirical result

    Evaluation of single-word queries revealed that omissions and errors in the original manual annotations degrade model metrics, but correcting ground-truth labels allows the Cross-Media Relevance Model to capture semantic concepts effectively:

    Ground-Truth Labels Sun Recall Sun Precision Sunset Recall Sunset Precision
    Original Manual 0.60 0.46 0.00 0.00
    Corrected Manual 0.50 0.71 0.57 0.21

    Under original annotations, "sunset" scored 0.00 recall and precision because images visibly depicting sunsets lacked the keyword label in both training and test sets. After manual re-labeling and re-training, "sunset" recall reached 0.57 and precision reached 0.21, while "sun" precision improved from 0.46 to 0.71.

Coverage note — Detailed specifications of the 33 low-level region features (color, texture, shape) and Normalized Cuts segmentation procedures were omitted as they are standard components adopted directly from Duygulu et al. (2002).

References

  1. 1.K. Barnard, P. Duygulu, N. de Freitas, D. Forsyth, D. Blei, and M. I. Jordan. Matching words and pictures. Journal of Machine Learning Research, 3:1107–1135, 2003.
  2. 2.K. Barnard and D. Forsyth. Learning the semantics of words and pictures. In International Conference on Computer Vision, Vol.2, pages 408-415, 2001.
  3. 3.D. Blei, Michael, and M. I. Jordan. Modeling annotated data. To appear in the Proceedings of the 26th annual international ACM SIGIR conference
  4. 4.Berger, A. and Lafferty, J. Information retrieval as statistical translation. In Proceedings of the 22nd annual international ACM SIGIR conference, pages 222–229, 1999.
  5. 5.P. Brown, S. D. Pietra, V. D. Pietra, and R. Mercer. The mathematics of statistical machine translation: Parameter estimation. In Computational Linguistics, 19(2):263-311, 1993.
  6. 6.W. B. Croft. Combining Approaches to Information Retrieval, in Advances in Information Retrieval ed. W. B. Croft, Kluwer Academic Publishers, Boston, MA.
  7. 7.C. Carson, M. Thomas, S. Belongie, J. M. Hellerstein, and J. Malik. Blobworld: A system for region-based image indexing and retrieval. In Third International Conference on Visual Information Systems, Lecture Notes in Computer Science, 1614, pages 509-516, 1999.
  8. 8.M. Das and R. Manmatha and E. M. Riseman, Indexing Flowers by Color Names using Domain Knowledge-driven Segmentation, IEEE Intelligent Systems, 14(5):24–33, 1999.
  9. 9.P. Duygulu, K. Barnard, N. de Freitas, and D. Forsyth. Object recognition as machine translation: Learning a lexicon for a fixed image vocabulary. In Seventh European Conference on Computer Vision, pages 97-112, 2002.
  10. 10.D. Forsyth and J. Ponce, Computer Vision: A Modern Approach Prentice Hall, 2003
  11. 11.D. Hiemstra Using Language Models for Information Retrieval. PhD dissertation, University of Twente, Enschede, The Netherlands, 2001.
  12. 12.J. M. Ponte, and W. B. Croft, A language modeling approach to information retrieval. Proceedings of the 21st annual international ACM SIGIR Conference, pages 275–281, 1998.
  13. 13.V. Lavrenko and W. Croft. Relevance-based language models. Proceedings of the 24th annual international ACM SIGIR conference, pages 120-127, 2001.
  14. 14.V. Lavrenko, M. Choquette, and W. Croft. Cross-lingual relevance models. Proceedings of the 25th annual international ACM SIGIR conference, pages 175–182, 2002.
  15. 15.Y. Mori, H. Takahashi, and R. Oka. Image-to-word transformation based on dividing and vector quantizing images with words. In MISRM’99 First International Workshop on Multimedia Intelligent Storage and Retrieval Management, 1999.
  16. 16.R. W. Picard and T. P. Minka”, Vision Texture for Annotation, In Multimedia Systems, 3(1):3–14, 1995.
  17. 17.J. Shi and J. Malik. Normalized cuts and image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(8):888–905, 2000.
  18. 18.J. Lafferty and C. Zhai. Document language models, query models, and risk minimization for information retrieval, Proceedings of the 24th annual international ACM SIGIR Conference, pages 111-119, 2001.

Citation

MLA
Jeon, J., et al. “Automatic Image Annotation and Retrieval Using Cross-media Relevance Models”. Proceedings of the 26th Annual International ACM SIGIR Conference on Research and Development in Informaion Retrieval, 2003, pp. 119–26, https://doi.org/10.1145/860435.860459.
APA
Jeon, J., Lavrenko, V., & Manmatha, R. (2003). Automatic image annotation and retrieval using cross-media relevance models. Proceedings of the 26th Annual International ACM SIGIR Conference on Research and Development in Informaion Retrieval, 119–126. https://doi.org/10.1145/860435.860459
Chicago
Jeon, J., V. Lavrenko, and R. Manmatha. 2003. “Automatic Image Annotation and Retrieval Using Cross-media Relevance Models”. Proceedings of the 26th Annual International ACM SIGIR Conference on Research and Development in Informaion Retrieval, 119–26. https://doi.org/10.1145/860435.860459.
Harvard
Jeon, J., Lavrenko, V. and Manmatha, R. (2003) “Automatic image annotation and retrieval using cross-media relevance models”, Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval. ACM, pp. 119–126. Available at: https://doi.org/10.1145/860435.860459.
Vancouver
1. Jeon J, Lavrenko V, Manmatha R (2003) Automatic image annotation and retrieval using cross-media relevance models. In: Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval. ACM, pp 119–126

BibTeX

@inproceedings{Jeon_2003, series={SIGIR03}, title={Automatic image annotation and retrieval using cross-media relevance models}, url={http://dx.doi.org/10.1145/860435.860459}, DOI={10.1145/860435.860459}, booktitle={Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval}, publisher={ACM}, author={Jeon, J. and Lavrenko, V. and Manmatha, R.}, year={2003}, month=July, pages={119–126}, collection={SIGIR03} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF