Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics

Micah HodoshPeter YoungJulia Hockenmaier

article2013JAIR1,471 citations

Establishes a unified ranking framework and a benchmark of 8,000 images with multiple descriptive captions to evaluate sentence-based image description and retrieval independently of text generation challenges.

Listen

Efficiently searching and describing the billions of images across digital platforms remains a major challenge because standard systems rely heavily on surrounding, often irrelevant text. Most research has framed automatic image captioning as a natural language generation task, which complicates evaluation by blending image understanding with linguistic fluency and fails to address the practically vital task of sentence-based image search. The article addresses this challenge by framing image description as a ranking task, mapping images and natural language sentences into a shared space to evaluate both sentence-based image annotation and sentence-based image search within a single framework.

The article demonstrates that sentence-based image understanding can be effectively modeled as a ranking problem using minimally supervised models and establishes a standardized benchmark for comparative evaluation. To conduct this evaluation, the authors constructed the Flickr 8K dataset, which pairs 8,000 images depicting people and animals in action with five crowdsourced conceptual descriptions each. Using a split of 6,000 training, 1,000 development, and 1,000 test images, the authors evaluated 30 distinct systems, focusing on Kernel Canonical Correlation Analysis—a machine learning method that finds correlated projections between two feature spaces—and nearest-neighbor baselines. They coupled low-level visual features (color, texture, and shape) with text representations incorporating word order and lexical similarities derived from alignments and text corpora.

The findings show that ranking-based Kernel Canonical Correlation Analysis substantially outperforms nearest-neighbor baselines across both image search and annotation tasks. Incorporating word-order sequences alongside corpus- and alignment-based lexical similarities delivered the highest performance, raising retrieval precision significantly over basic word-matching models. For example, the best-performing model placed a relevant caption in the top 10 results for 49.1% of images and a relevant image in the top 10 for 48.5% of caption queries, whereas a random baseline achieves suitable descriptions only about 1.5% of the time. Crucially, the authors found that standard automated natural language generation metrics (such as unigram precision and recall scores against reference captions) correlate poorly with human judgments when candidates differ from the references, while ranking metrics evaluating the top 5 or 10 positions correlate strongly with comprehensive human evaluations.

These results demonstrate that complex, detector-based object representations and full natural language generation pipelines are not required to achieve meaningful cross-modal image understanding. By isolating semantic matching from surface-level text generation, organizations can benchmark and improve multimodal retrieval systems more transparently and cost-effectively. Furthermore, the findings caution practitioners against relying on automated text-generation metrics to assess conceptual accuracy in cross-modal AI systems.

Organizations developing multimodal search and captioning capabilities should adopt ranking-based evaluation frameworks and utilize training sets with multiple human-authored conceptual descriptions per image. Future technical efforts should explore combining these minimally supervised representations with richer visual detectors and deeper syntactic parsing, while also testing the ranking methodology on larger, more varied visual collections.

The findings are supported by high inter-annotator agreement among human evaluators and statistically rigorous testing across 30 system configurations. However, readers should note that the training process for Kernel Canonical Correlation Analysis requires storing large kernel matrices in memory, which may present scaling constraints on massive datasets without alternative optimization techniques.

Cover for Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics

Abstract

The ability to associate images with natural language sentences that describe what is depicted in them is a hallmark of image understanding, and a prerequisite for applications such as sentence-based image search. In analogy to image search, we propose to frame sentence-based image annotation as the task of ranking a given pool of captions. We introduce a new benchmark collection for sentence-based image description and search, consisting of 8,000 images that are each paired with five different captions which provide clear descriptions of the salient entities and events. We introduce a number of systems that perform quite well on this task, even though they are only based on features that can be obtained with minimal supervision. Our results clearly indicate the importance of training on multiple captions per image, and of capturing syntactic (word order-based) and semantic features of these captions. We also perform an in-depth comparison of human and automatic evaluation metrics for this task, and propose strategies for collecting human judgments cheaply and on a very large scale, allowing us to augment our collection with additional relevance judgments of which captions describe which image. Our analysis shows that metrics that consider the ranked list of results for each query image or sentence are significantly more robust than metrics that are based on a single response per query. Moreover, our study suggests that the evaluation of ranking-based image description systems may be fully automated.

Table of Contents

  • 1. Introduction
  • 1.1 Related Work
  • 1.2 Our Approach
  • 1.3 Contributions and Outline of this Paper
  • 2. A New Data Set for Image Description
  • 2.1 What Do We Mean by Image Description?
  • 2.2 The Need for New Data Sets
  • 2.3 Our Data Sets
  • 2.3.1 The PASCAL VOC-2008 Data Set
  • 2.3.2 The Flickr 8K Data Set
  • 3. Systems for Sentence-Based Image Description
  • 3.1 Nearest-Neighbor Search for Image Description
  • 3.2 Kernel Canonical Correlation Analysis for Image Description
  • 3.2.1 Kernel Canonical Correlation Analysis (KCCA)
  • 3.2.2 Using KCCA to Associate Images and Sentences
  • 3.3 Image Kernels
  • 3.3.1 The Histogram Kernel ( K Histo )
  • 3.3.2 The Pyramid Kernel K Py
  • 3.4 Basic Text Kernels
  • 3.4.1 The Bag of Words Kernel ( BoW )
  • 3.4.2 The Tag Rank Kernel ( TagRank )
  • 3.4.3 The Trigram Kernel ( Tri )
  • 3.5 Extending the Trigram Kernel with Lexical Similarities
  • 3.5.1 String Kernels with Lexical Similarities
  • 3.5.2 The Lin Similarity Kernel ( σ Lin )
  • 3.5.3 Distributional Similarity ( σ D C )
  • 3.5.4 Alignment-Based Similarity ( σ A )
  • 3.5.5 Comparing the Similarity Metrics (Figure 3)
  • 3.5.6 Combining Different Similarities
  • 4. Evaluation Procedures and Metrics for Image Description
  • 4.1 Experimental Setup
  • 4.1.1 The Data
  • 4.1.2 The Tasks
  • 4.1.3 The Systems
  • 4.2 Metrics for the Quality of Individual Image-Caption Pairs
  • 4.2.1 Human Evaluation with Graded 'Expert' Judgments
  • 4.2.2 Automatic Evaluation with Bleu and Rouge
  • 4.3 Metrics for the Large-Scale Evaluation of Image Description Systems
  • 4.3.1 Recall and Median Rank of the Original Item
  • 4.3.2 Collecting Binary Relevance Judgments on a Large Scale
  • 4.3.3 Large-Scale Evaluation with Relevance Judgments
  • 4.3.4 Measuring the Impact of Linguistic Features (Table 6)
  • 4.3.5 Can Human Evaluations Be Approximated by Automatic Techniques?
  • 5. Summary of Contributions and Conclusions
  • 5.1 The Advantages of Framing Image Description as a Ranking Task
  • 5.2 Our Data Set
  • 5.3 Our Models
  • 5.4 Evaluating Ranking-Based Image Description Systems
  • 5.5 Implications for the Evaluation of Caption Generation Systems
  • Acknowledgments
  • Appendix A. Agreement Between Approximate Metrics and Expert Human Judgments
  • Appendix B. Performance of All Systems
  • References

Knowls

  1. Knowl 1 — Bidirectional Ranking Formulation for Sentence-Based Image Annotation and Retrieval

    model/method

    Sentence-based image annotation (description) and sentence-based image search (retrieval) are unified into a bidirectional ranking problem over a candidate pool of sentences Scand\mathcal{S}_{\text{cand}} and a candidate pool of images Icand\mathcal{I}_{\text{cand}} using an affinity function f(i,s)f(i, s) that measures the degree of association between an image ii and a sentence ss.

    Given a query caption sq∈Scands_q \in \mathcal{S}_{\text{cand}}, image search identifies the image i∗∈Icandi^* \in \mathcal{I}_{\text{cand}} that maximizes f(i,sq)f(i, s_q): i∗=arg⁡max⁡i∈Icandf(i,sq)i^* = \arg\max_{i \in \mathcal{I}_{\text{cand}}} f(i, s_q)

    Conversely, given a query image iq∈Icandi_q \in \mathcal{I}_{\text{cand}}, image annotation selects the sentence s∗∈Scands^* \in \mathcal{S}_{\text{cand}} that maximizes f(iq,s)f(i_q, s): s∗=arg⁡max⁡s∈Scandf(iq,s)s^* = \arg\max_{s \in \mathcal{S}_{\text{cand}}} f(i_q, s)

    This ranking formulation allows both tasks to be evaluated symmetrically on unseen test pools Itest\mathcal{I}_{\text{test}} and Stest\mathcal{S}_{\text{test}} that are disjoint from the training data Dtrain=(Itrain,Strain)\mathcal{D}_{\text{train}} = (\mathcal{I}_{\text{train}}, \mathcal{S}_{\text{train}}).

    A baseline nearest-neighbor (NN) affinity function uses unimodal similarity functions fIf_I (for images) and fSf_S (for text) to route queries via the closest training pair in Dtrain\mathcal{D}_{\text{train}}: fNN(i,sq)=fI(iNN,i)where ⟨iNN,sNN⟩=arg⁡max⁡⟨it,st⟩∈DtrainfS(sq,st)f_{\text{NN}}(i, s_q) = f_I(i^{\text{NN}}, i) \quad \text{where } \langle i^{\text{NN}}, s^{\text{NN}} \rangle = \arg\max_{\langle i^t, s^t \rangle \in \mathcal{D}_{\text{train}}} f_S(s_q, s^t) fNN(iq,s)=fS(sNN,s)where ⟨iNN,sNN⟩=arg⁡max⁡⟨it,st⟩∈DtrainfI(iq,it)f_{\text{NN}}(i_q, s) = f_S(s^{\text{NN}}, s) \quad \text{where } \langle i^{\text{NN}}, s^{\text{NN}} \rangle = \arg\max_{\langle i^t, s^t \rangle \in \mathcal{D}_{\text{train}}} f_I(i_q, i^t)

  2. Knowl 2 — Cross-Modal Kernel Canonical Correlation Analysis (KCCA) for Image-Sentence Mapping

    model/method

    Kernel Canonical Correlation Analysis (KCCA) maps images i∈Ii \in \mathcal{I} and sentences s∈Ss \in \mathcal{S} from two different feature spaces into a shared latent semantic space Z\mathcal{Z} where corresponding images and captions are maximally correlated.

    Given training pairs Dtrain={(it,st)}t=1M\mathcal{D}_{\text{train}} = \{(i_t, s_t)\}_{t=1}^M, image kernel matrix KX[i,j]=⟨ϕX(ii),ϕX(ij)⟩K_X[i, j] = \langle \phi_X(i_i), \phi_X(i_j) \rangle and sentence kernel matrix KY[i,j]=⟨ϕY(si),ϕY(sj)⟩K_Y[i, j] = \langle \phi_Y(s_i), \phi_Y(s_j) \rangle, KCCA optimizes projection weights (α∗,β∗)(\alpha^*, \beta^*) according to: (α∗,β∗)=arg⁡max⁡α,βα′KXKYβ(α′KX2α+κα′KXα)(β′KY2β+κβ′KYβ)(\alpha^*, \beta^*) = \arg\max_{\alpha, \beta} \frac{\alpha' K_X K_Y \beta}{\sqrt{(\alpha' K_X^2 \alpha + \kappa \alpha' K_X \alpha)(\beta' K_Y^2 \beta + \kappa \beta' K_Y \beta)}} where κ\kappa is a regularization parameter that penalizes solution norms to prevent overfitting.

    This optimization is cast as the generalized eigenproblem: (KX+κI)−1KY(KY+κI)−1KXα=λ2α(K_X + \kappa I)^{-1} K_Y (K_Y + \kappa I)^{-1} K_X \alpha = \lambda^2 \alpha and solved via partial Gram-Schmidt orthogonalization.

    For a new image ii and sentence ss, kernel evaluation vectors KI(i)=[KI(i1,i),…,KI(iM,i)]′K_I(i) = [K_I(i_1, i), \dots, K_I(i_M, i)]' and KS(s)=[KS(s1,s),…,KS(sM,s)]′K_S(s) = [K_S(s_1, s), \dots, K_S(s_M, s)]' are projected into Z\mathcal{Z}. The cross-modal affinity function is defined as their cosine similarity in Z\mathcal{Z}: fKCCA(i,s)=sim(α∗KI(i),β∗KS(s))f_{\text{KCCA}}(i, s) = \text{sim}(\alpha^* K_I(i), \beta^* K_S(s))

    Hyperparameters (regularization parameter κ∈{0.1,0.5,1,5}\kappa \in \{0.1, 0.5, 1, 5\} and dimensionality n∈[10,6000]n \in [10, 6000] of the projection) are tuned on development data by selecting the top five settings for Recall@1, top five for Recall@5, and top five for Recall@10, combining all 15 rankings using Borda counts.

  3. Knowl 3 — Flickr 8K Dataset for Conceptual Image Description

    experimental setup

    The Flickr 8K dataset consists of 8,092 Flickr images depicting people and animals (primarily dogs) performing actions across diverse scenes. Unlike web captions or news captions that provide non-visual context, personal commentary, or proper nouns, Flickr 8K focuses on concrete conceptual descriptions of visible entities, activities, and scenes.

    Each image is annotated with five independent, descriptive captions collected via Amazon Mechanical Turk from qualified US workers who passed grammar and spelling qualification tests. Captions average 11.8 words in length, and 79% contain action verbs beyond static verbs like sit, stand, wear, or look.

    The dataset is partitioned into three disjoint sets:

    1. Training set Dtrain\mathcal{D}_{\text{train}}: 6,000 images, each paired with 5 captions (30,000 image-caption pairs).
    2. Development set Ddev\mathcal{D}_{\text{dev}}: 1,000 images, each paired with 1 arbitrarily selected caption.
    3. Test set Dtest\mathcal{D}_{\text{test}}: 1,000 images, each paired with 1 arbitrarily selected caption.

    Captions are preprocessed with spellchecking, compound normalization (e.g., t-shirt), stop word removal, and lemmatization.

  4. Knowl 4 — Pixel-Based Visual Histogram and Spatial Pyramid Matching Kernels

    model/method

    Visual representations are constructed from low-level pixel descriptors without using pre-trained object detectors. Three descriptor types are extracted and clustered into discrete visual vocabularies via K-means on 200 PASCAL 2008 images: color in CIELAB coordinates (128 visual words), texture filter responses centered on pixels (256 visual words), and SIFT edge/shape descriptors (256 visual words).

    Two image kernels combine these features:

    1. Histogram Kernel (KHistoK^{\text{Histo}}): Image xix_i is represented as a normalized feature histogram HiH_i. The intersection of histograms over vocabulary VV is: K(xi,xj)=∑v=1Vmin⁡(Hi(v),Hj(v))K(x_i, x_j) = \sum_{v=1}^V \min(H_i(v), H_j(v)) The composite kernel is the pp-th power (p∈{2,3}p \in \{2, 3\}) of the averaged color (CC), texture (TT), and SIFT (SS) kernels: KHisto(xi,xj)=[13∑F∈{C,S,T}KFHisto(xi,xj)]pK^{\text{Histo}}(x_i, x_j) = \left[ \frac{1}{3} \sum_{F \in \{C, S, T\}} K_F^{\text{Histo}}(x_i, x_j) \right]^p

    2. Spatial Pyramid Matching Kernel (KPyK^{\text{Py}}): Partitions an image at scale levels l∈{0,1,2}l \in \{0, 1, 2\} into grids of Cl=2l×2lC_l = 2^l \times 2^l cells (C0=1,C1=4,C2=16C_0=1, C_1=4, C_2=16). The cell histogram intersection sum at level ll is Iijl=∑c=0Cl−1∑v=1Vmin⁡(Hic(v),Hjc(v))I_{ij}^l = \sum_{c=0}^{C_l-1} \sum_{v=1}^V \min(H_{ic}(v), H_{jc}(v)). The pyramid kernel weights finer scales higher: KPy(xi,xj)=12LIij0+∑l=1L12L−l+1IijlK^{\text{Py}}(x_i, x_j) = \frac{1}{2^L} I_{ij}^0 + \sum_{l=1}^L \frac{1}{2^{L-l+1}} I_{ij}^l Individual kernels KCPy,KTPy,KSPyK_C^{\text{Py}}, K_T^{\text{Py}}, K_S^{\text{Py}} are averaged and exponentiated to power pp.

  5. Knowl 5 — Trigram Subsequence Text Kernel for Image Captions

    model/method

    The trigram kernel (KTriK_{\text{Tri}}) is a truncated string kernel that captures word order dependencies (such as head-modifier pairs and subject-verb-object triples) across ordered word subsequences w=w1…wkw = w_1 \dots w_k of length k≤3k \le 3, without requiring full syntactic parses.

    Let Ms,wM_{s, w} be the set of substrings in sentence ss that start with w1w_1, end with wkw_k, and contain the ordered sequence ww: Ms,w={(i,j)∣w=w1…wk∈si…sj,si=w1,sj=wk}M_{s, w} = \{(i, j) \mid w = w_1 \dots w_k \in s_i \dots s_j, s_i = w_1, s_j = w_k\}

    With match parameter λm=0.5\lambda_m = 0.5, non-penalized gap parameter λg=1\lambda_g = 1, and match count ms,w=∣Ms,w∣m_{s, w} = |M_{s, w}|, the base trigram kernel is: KTri(s,s′)=∑w:k≤3ms,wms′,wλm2l(w)K_{\text{Tri}}(s, s') = \sum_{w: k \le 3} m_{s, w} m_{s', w} \lambda_m^{2 l(w)} where l(w)l(w) is the length of subsequence ww. Kernel values are normalized by the geometric mean K(s,s)K(s′,s′)\sqrt{K(s, s) K(s', s')}.

    When augmented with Inverse Document Frequency (IDF) weights λwi=log⁡∣Dtrain∣∣Dtrain(wi)∣+1\lambda_{w_i} = \log \frac{|\mathcal{D}_{\text{train}}|}{|\mathcal{D}_{\text{train}}(w_i)| + 1} where λw=∏k=ijλwk\lambda_w = \prod_{k=i}^j \lambda_{w_k}, the idf\sqrt{\text{idf}}-weighted trigram kernel (KTriidfK_{\text{Tri}\sqrt{\text{idf}}}) is: KTriidf(s,s′)=∑w:k≤3λwms,wms′,wλm2l(w)K_{\text{Tri}\sqrt{\text{idf}}}(s, s') = \sum_{w: k \le 3} \lambda_w m_{s, w} m_{s', w} \lambda_m^{2 l(w)}

  6. Knowl 6 — Alignment-Based Lexical Similarity ($\sigma_A$) for Parallel Captions

    model/method

    To account for lexical variation across descriptions of the same scene, an alignment-based lexical similarity metric σA\sigma_A is trained using statistical machine translation alignment across pairs of captions written for the same image in Dtrain\mathcal{D}_{\text{train}}.

    Nouns and verbs are processed separately to prevent cross-part-of-speech noise:

    1. Noun and verb vocabularies are extracted (words occurring ≥5\ge 5 times and tagged as noun ≥50%\ge 50\% of the time, or tagged as verb ≥25\ge 25 times and ≥25%\ge 25\% of occurrences).
    2. IBM Alignment Models 1–2 are trained using GIZA++ over all pairs of noun-only and verb-only captions describing identical images, yielding translation probabilities Pn(wi∣w)P_n(w_i \mid w) for nouns and Pv(wi∣w)P_v(w_i \mid w) for verbs.
    3. A word ww is represented as a similarity vector w⃗A\vec{w}_A over vocabulary words wiw_i, weighted by the relative frequencies of ww being tagged as noun (Pn(w)P_n(w)) or verb (Pv(w)P_v(w)): w⃗A(i)=Pn(wi∣w)Pn(w)+Pv(wi∣w)Pv(w)\vec{w}_A(i) = P_n(w_i \mid w) P_n(w) + P_v(w_i \mid w) P_v(w) Entries smaller than 0.050.05 are set to zero.

    The word kernel between words ww and w′w' is the cosine similarity κA(w,w′)=cos⁡(w⃗A,w⃗A′)\kappa_A(w, w') = \cos(\vec{w}_A, \vec{w}'_A). For a subsequence ww of length ll, the sequence similarity is σA(w,w′)=∏i=1lκA(wi,wi′)\sigma_A(w, w') = \prod_{i=1}^l \kappa_A(w_i, w'_i), which replaces exact matches in the trigram text kernel.

  7. Knowl 7 — Lexical and Distributional Extensions to Text Kernels

    model/method

    Text kernels are extended to capture partial semantic matches by mapping each word ww to a vector w⃗S\vec{w}_S in an NN-dimensional vocabulary space where w⃗S(i)=simS(w,wi)\vec{w}_S(i) = \text{sim}_S(w, w_i). The word kernel is defined as κS(w,w′)=cos⁡(w⃗S,w⃗S′)\kappa_S(w, w') = \cos(\vec{w}_S, \vec{w}'_S), and sequence similarity for length ll is σS(w,w′)=∏i=1lκS(wi,wi′)\sigma_S(w, w') = \prod_{i=1}^l \kappa_S(w_i, w'_i).

    The resulting idf\sqrt{\text{idf}}-weighted sequence kernel is: KS(s,s′)=∑w∑w′∈σS(w)λwλw′ms,wms′,w′λm2l(w)σS(w′,w)K_S(s, s') = \sum_w \sum_{w' \in \sigma_S(w)} \sqrt{\lambda_w \lambda_{w'}} m_{s, w} m_{s', w'} \lambda_m^{2 l(w)} \sigma_S(w', w)

    Three lexical similarity sources simS\text{sim}_S are defined:

    1. WordNet Lin similarity (σLin\sigma_{\text{Lin}}): Uses first noun senses in WordNet 3.0 and information content of the lowest common subsumer LCS(si,sj)\text{LCS}(s_i, s_j): simLin(si,sj)=2log⁡P(LCS(si,sj))log⁡P(si)+log⁡P(sj)\text{sim}_{\text{Lin}}(s_i, s_j) = \frac{2 \log P(\text{LCS}(s_i, s_j))}{\log P(s_i) + \log P(s_j)}
    2. Distributional similarity (σDC\sigma_{D_C}): Non-negative pointwise mutual information (PMI) computed over corpus CC: w⃗DC(i)=max⁡(0,log⁡2PC(w,wi)PC(w)PC(wi))\vec{w}_{D_C}(i) = \max\left(0, \log_2 \frac{P_C(w, w_i)}{P_C(w) P_C(w_i)}\right) Evaluated on the training captions (DicD_{\text{ic}}, threshold <0.01< 0.01 zeroed) and the British National Corpus (DBNCD_{\text{BNC}}).
    3. Combined kernels: Averaged distributional kernel κDBNC,ic=12(κDBNC+κDic)\kappa_{D_{\text{BNC}}, \text{ic}} = \frac{1}{2}(\kappa_{D_{\text{BNC}}} + \kappa_{D_{\text{ic}}}), and hybrid alignment-distributional kernel κD+A(w,w′)=max⁡(κD(w,w′),κA(w,w′))\kappa_{D+A}(w, w') = \max(\kappa_D(w, w'), \kappa_A(w, w')).
  8. Knowl 8 — Ablation Analysis of Linguistic Features on Cross-Modal R-Precision

    data/table

    The impact of adding Inverse Document Frequency (IDF) weighting (ii), alignment-based similarities (aa), and distributional similarities (dd) to the base trigram model (Tri5) was evaluated on the Flickr 8K test set using R-precision (the percentage of relevant items in the top rir_i responses, where rir_i is the number of relevant items for query qiq_i).

    Tri5 +IDF +Align +AlignIDF
    Ann. Search Ann. Search Ann. Search Ann. Search
    Tri5 11.6 11.0 12.5 11.3 13.4 12.3 13.4 13.2
    +DBNCD_{\text{BNC}} 12.7 12.1 12.9 12.2 13.2 12.8 12.9 12.9
    +DicD_{\text{ic}} 12.7 12.8 12.8 13.1 13.0 12.8 13.3 13.4
    +DBNC+icD_{\text{BNC+ic}} 12.5 12.7 13.3 13.0 13.4 13.2 13.7 13.4

    Key empirical findings from this ablation:

    • Any model incorporating lexical similarities significantly outperforms the base Tri5 model on both annotation (p<0.0001p < 0.0001) and search (p≤0.02p \le 0.02).
    • Alignment-based similarity (+Align+Align) provides substantial gains over base Tri5, increasing annotation R-precision from 11.6% to 13.4% and search from 11.0% to 12.3%.
    • The full semantic model Tri5Sem (Tri5 + Align & IDF + DBNC+icD_{\text{BNC+ic}}) achieves the highest overall R-precision (13.7% annotation, 13.4% search).
    • WordNet Lin similarity (extTri5Lin ext{Tri5}_{\text{Lin}}) yielded 11.7% on annotation and 10.7% on search, failing to improve performance because WordNet hypernym relations conflate visually distinct concepts (e.g., grouping swimming and football as sports).
  9. Knowl 9 — Cross-Modal Performance Comparison Across Retrieval and Annotation Systems

    data/table

    System performance was measured on 1,000 Flickr 8K test queries for image annotation and image search using automated rank metrics (Recall@kk, Median Rank rr) and crowdsourced human relevance metrics (Success Rate S@kk, R-precision).

    Image Annotation (Rank of Gold) Image Retrieval (Rank of Gold) Human Metrics
    Model R@1 R@5 R@10 Med. rr R@1 R@5 R@10 Med. rr S@10 (Ann) R-prec (Tot)
    NN 2.5 7.6 9.7 251.0 2.5 4.7 7.2 272.0 20.2 4.5
    BoW1 4.8 13.5 19.7 64.0 4.5 14.3 20.8 67.0 39.7 10.1
    BoW5 6.2 17.1 24.3 58.0 5.8 16.7 23.6 60.0 42.7 10.8
    TagRank 6.0 17.0 23.8 56.0 5.4 17.4 24.3 52.5 42.9 11.1
    Tri5 7.1 17.2 23.7 53.0 6.0 17.8 26.2 55.0 43.4 11.3
    Tri5Sem 8.3 21.6 30.3 34.0 7.6 20.7 30.1 38.0 49.1 13.5

    Key comparative findings:

    • All KCCA models significantly outperform nearest-neighbor baselines (p<0.001p < 0.001), reducing median rank from >250>250 to <65<65.
    • Training on five captions per image (BoW5) substantially improves over training on a single caption (BoW1), increasing R@10 from 19.7% to 24.3% in annotation.
    • Tri5Sem achieves the best performance across all rank and human metrics, returning a relevant caption in the top 10 for 49.1% of images and a relevant image in the top 10 for 48.5% of queries.
    • S@1 scores are at least double R@1 scores (e.g., Tri5Sem S@1 is 16.6% vs. R@1 of 8.3%), demonstrating that highest-ranked predictions are frequently valid alternative descriptions not in the original single-caption gold pairing.
  10. Knowl 10 — Inadequacy of BLEU/ROUGE and Robustness of Automated Ranking Metrics for Image Description

    empirical result

    An evaluation of metrics comparing unigram BLEU-1 and ROUGE-1 against 4-point graded expert human judgments ({1,2,3,4}\{1, 2, 3, 4\}, Krippendorff's α=0.81\alpha = 0.81) and crowdsourced binary relevance judgments revealed fundamental limitations of standard n-gram overlap metrics:

    1. BLEU and ROUGE show high agreement with expert judgments only when the candidate pool contains the exact original reference caption (κ=0.72\kappa = 0.72 for BLEU, κ=0.54\kappa = 0.54 for ROUGE). When candidate captions are human-written but disjoint from the reference set (simulating novel generation), agreement drops sharply to κ=0.36\kappa = 0.36--0.520.52 for BLEU and κ=0.42\kappa = 0.42--0.510.51 for ROUGE, demonstrating that BLEU and ROUGE are unreliable for assessing semantic content quality.
    2. BLEU with stemming and stop word removal (BLEUpre\text{BLEU}_{\text{pre}}) at threshold ≥0.25\ge 0.25 effectively filters out 86.0% of all 1,000×1,0001,000 \times 1,000 test image-caption pairs while eliminating only 3.5% of pairs with expert scores ≥3.0\ge 3.0.
    3. System rankings produced by automatic ranking metrics based on the gold item (Recall@5, Recall@10, and Median Rank) correlate strongly with system rankings obtained from crowdsourced human relevance judgments (Success Rate S@k and R-precision) across 30 systems, with Spearman's ρ≥0.92\rho \ge 0.92 and Kendall's τ≥0.76\tau \ge 0.76 for annotation and ρ≥0.96,τ≥0.87\rho \ge 0.96, \tau \ge 0.87 for search. Conversely, R@1 correlates poorly with human rankings (τ=0.68\tau = 0.68). This confirms that rank-based evaluation metrics considering lists of results allow fully automated, reliable evaluation of image description systems.

Coverage note — None was omitted; all key contributions—including the ranking formulation, the Flickr 8K benchmark, visual and text kernels (BoW, TagRank, Trigram, Alignment, Distributional, Lin), KCCA methodology, feature ablation results, comparative system evaluation, and metric correlation analyses—are fully represented.

References

  1. 1.Artstein, R., & Poesio, M. (2008). Inter-coder agreement for computational linguistics. Computational Linguistics, 34 (4), 555–596.
  2. 2.Bach, F. R., & Jordan, M. I. (2002). Kernel independent component analysis. Journal of Machine Learning Research, 3, 1–48.
  3. 3.Barnard, K., Duygulu, P., Forsyth, D., Freitas, N. D., Blei, D. M., & Jordan, M. I. (2003). Matching words and pictures. Journal of Machine Learning Research, 3, 1107–1135.
  4. 4.Blei, D. M., & Jordan, M. I. (2003). Modeling annotated data. In SIGIR 2003: Proceedings of the 26th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 127–134, Toronto, Ontario, Canada.
  5. 5.Bloehdorn, S., Basili, R., Cammisa, M., & Moschitti, A. (2006). Semantic kernels for text classification based on topological measures of feature similarity. In Proceedings of the 6th IEEE International Conference on Data Mining (ICDM 2006), pp. 808–812, Hong Kong, China.
  6. 6.BNC Consortium (2007). The British National Corpus, version 3 (BNC XML edition). http://www.natcorp.ox.ac.uk.
  7. 7.Brown, P. F., Pietra, V. J. D., Pietra, S. A. D., & Mercer, R. L. (1993). The mathematics of statistical machine translation: parameter estimation. Computational Linguistics, 19 (2), 263–311.
  8. 8.Callison-Burch, C., Osborne, M., & Koehn, P. (2006). Re-evaluation the role of bleu in machine translation research. In Proceedings of the 11th Conference of the European Chapter of the Association for Computational Linguistics (EACL), pp. 249–256, Trento, Italy.
  9. 9.Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20 (1), 37–46.
  10. 10.Croce, D., Moschitti, A., & Basili, R. (2011). Structured lexical similarity via convolution kernels on dependency trees. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 1034–1046, Edinburgh, UK.
  11. 11.Dale, R., & White, M. (Eds.). (2007). Workshop on Shared Tasks and Comparative Evaluation in Natural Language Generation: Position Papers, Arlington, VA, USA.
  12. 12.Datta, R., Joshi, D., Li, J., & Wang, J. Z. (2008). Image retrieval: Ideas, influences, and trends of the new age. ACM Computing Surveys, 40 (2), 5:1–5:60.
  13. 13.Deschacht, K., & Moens, M.-F. (2007). Text analysis for automatic image annotation. In Proceedings of the 45th Annual Meeting of the Association of Computational Linguistics (ACL), pp. 1000–1007, Prague, Czech Republic.
  14. 14.Dietterich, T. G. (1998). Approximate statistical tests for comparing supervised classification learning algorithms. Neural Computation, 10 (7), 1895–1923.
  15. 15.Everingham, M., Gool, L. V., Williams, C., Winn, J., & Zisserman, A. (2008). The PASCAL Visual Object Classes Challenge 2008 (VOC2008) Results. http://www.pascal-network.org/challenges/VOC/voc2008/workshop/.
  16. 16.Farhadi, A., Hejrati, M., Sadeghi, M. A., Young, P., Rashtchian, C., Hockenmaier, J., & Forsyth, D. (2010). Every picture tells a story: Generating sentences from images. In Proceedings of the European Conference on Computer Vision (ECCV), Part IV, pp. 15–29, Heraklion, Greece.
  17. 17.Fellbaum, C. (1998). WordNet: An Electronic Lexical Database. Bradford Books.
  18. 18.Felzenszwalb, P., McAllester, D., & Ramanan, D. (2008). A discriminatively trained, multiscale, deformable part model. In Proceedings of the 2008 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1–8, Anchorage, AK, USA.
  19. 19.Feng, Y., & Lapata, M. (2008). Automatic image annotation using auxiliary text information. In Proceedings of the 46th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies (ACL-08: HLT), pp. 272–280, Columbus, OH, USA.
  20. 20.Feng, Y., & Lapata, M. (2010). How many words is a picture worth? automatic caption generation for news images. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics (ACL), pp. 1239–1249, Uppsala, Sweden.
  21. 21.Fisher, R. A. (1935). The Design of Experiments. Olyver and Boyd, Edinburgh, UK.
  22. 22.Grangier, D., & Bengio, S. (2008). A discriminative kernel-based approach to rank images from text queries. IEEE Transactions on Pattern Analysis and Machine Intelligence, 30, 1371–1384.
  23. 23.Grice, H. P. (1975). Logic and conversation. In Davidson, D., & Harman, G. H. (Eds.), The Logic of Grammar, pp. 64–75. Dickenson Publishing Co., Encino, CA, USA.
  24. 24.Grubinger, M., Clough, P., Müller, H., & Deselaers, T. (2006). The IAPR benchmark: A new evaluation resource for visual information systems. In OntoImage 2006, Workshop on Language Resources for Content-based Image Retrieval during LREC 2006, pp. 13–23, Genoa, Italy.
  25. 25.Gupta, A., Verma, Y., & Jawahar, C. (2012). Choosing linguistics over vision to describe images. In Proceedings of the Twenty-Sixth AAAI Conference on Artificial Intelligence, Toronto, Ontario, Canada.
  26. 26.Hardoon, D. R., Saunders, C., Szedmak, S., & Shawe-Taylor, J. (2006). A correlation approach for automatic image annotation. In Li, X., Zaïane, O. R., & Li, Z.-H. (Eds.), Advanced Data Mining and Applications, Vol. 4093 of Lecture Notes in Computer Science, pp. 681–692. Springer Berlin Heidelberg.
  27. 27.Hardoon, D. R., Szedmak, S. R., & Shawe-Taylor, J. R. (2004). Canonical correlation analysis: An overview with application to learning methods. Neural Computation, 16, 2639–2664.
  28. 28.Hotelling, H. (1936). Relations between two sets of variates. Biometrika, 28 (3/4), 321–377.
  29. 29.Hwang, S., & Grauman, K. (2012). Learning the relative importance of objects from tagged images for retrieval and cross-modal search. International Journal of Computer Vision, 100 (2), 134–153.
  30. 30.Jaimes, A., Jaimes, R., & Chang, S.-F. (2000). A conceptual framework for indexing visual information at multiple levels. In Internet Imaging 2000, Vol. 3964 of Proceedings of SPIE, pp. 2–15, San Jose, CA, USA.
  31. 31.Jurafsky, D., & Martin, J. H. (2008). Speech and Language Processing (2nd edition). Prentice Hall.
  32. 32.Krippendorff, K. (2004). Content analysis: An introduction to its methodology. Sage.
  33. 33.Kulkarni, G., Premraj, V., Dhar, S., Li, S., Choi, Y., Berg, A. C., & Berg, T. L. (2011). Baby talk: Understanding and generating simple image descriptions. In Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1601–1608.
  34. 34.Kuznetsova, P., Ordonez, V., Berg, A., Berg, T., & Choi, Y. (2012). Collective generation of natural image descriptions. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 359–368, Jeju Island, Korea.
  35. 35.Lavrenko, V., Manmatha, R., & Jeon, J. (2004). A model for learning the semantics of pictures. In Thrun, S., Saul, L., & Schölkopf, B. (Eds.), Advances in Neural Information Processing Systems 16, Cambridge, MA, USA.
  36. 36.Lazebnik, S., Schmid, C., & Ponce, J. (2009). Spatial pyramid matching. In S. Dickinson, A. Leonardis, B. S., & Tarr, M. (Eds.), Object Categorization: Computer and Human Vision Perspectives, chap. 21, pp. 401–415. Cambridge University Press.
  37. 37.Li, S., Kulkarni, G., Berg, T. L., Berg, A. C., & Choi, Y. (2011). Composing simple image descriptions using web-scale n-grams. In Proceedings of the Fifteenth Conference on Computational Natural Language Learning (CoNLL), pp. 220–228, Portland, OR, USA.
  38. 38.Lin, C.-Y. (2004). Rouge: A package for automatic evaluation of summaries. In Marie-Francine Moens, S. S. (Ed.), Text Summarization Branches Out: Proceedings of the ACL-04 Workshop, pp. 74–81, Barcelona, Spain.
  39. 39.Lin, C.-Y., & Hovy, E. H. (2003). Automatic evaluation of summaries using n-gram cooccurrence statistics. In Proceedings of the 2003 Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics (HLT-NAACL), pp. 71–78, Edmonton, AB, Canada.
  40. 40.Lin, D. (1998). An information-theoretic definition of similarity. In Proceedings of the Fifteenth International Conference on Machine Learning (ICML), pp. 296–304, Madison, WI, USA.
  41. 41.Lowe, D. G. (2004). Distinctive image features from scale-invariant keypoints. International Journal of Computer Vision, 60 (2), 91–110.
  42. 42.Makadia, A., Pavlovic, V., & Kumar, S. (2010). Baselines for image annotation. International Journal of Computer Vision, 90 (1), 88–105.
  43. 43.Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press.
  44. 44.Mitchell, M., Dodge, J., Goyal, A., Yamaguchi, K., Stratos, K., Han, X., Mensch, A., Berg, A., Berg, T., & Daume III, H. (2012). Midge: Generating image descriptions from computer vision detections. In Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics (EACL), pp. 747–756, Avignon, France.
  45. 45.Moschitti, A. (2009). Syntactic and semantic kernels for short text pair categorization. In Proceedings of the 12th Conference of the European Chapter of the Association for Computational Linguistics (EACL), pp. 576–584, Athens, Greece.
  46. 46.Moschitti, A., Pighin, D., & Basili, R. (2008). Tree kernels for semantic role labeling. Computational Linguistics, 34 (2), 193–224.
  47. 47.Och, F. J., & Ney, H. (2003). A systematic comparison of various statistical alignment models. Computational Linguistics, 29 (1), 19–51.
  48. 48.Ordonez, V., Kulkarni, G., & Berg, T. L. (2011). Im2text: Describing images using 1 million captioned photographs. In Advances in Neural Information Processing Systems 24, pp. 1143–1151.
  49. 49.Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). Bleu: a method for automatic evaluation of machine translation. In Proceedings of 40th Annual Meeting of the Association for Computational Linguistics (ACL), pp. 311–318, Philadelphia, PA, USA.
  50. 50.Popescu, A., Tsikrika, T., & Kludas, J. (2010). Overview of the Wikipedia retrieval task at ImageCLEF 2010. In CLEF (Notebook Papers/LABs/Workshops), Padua, Italy.
  51. 51.Porter, M. F. (1980). An algorithm for suffix stripping. Program, 14 (3), 130–137.
  52. 52.Rashtchian, C., Young, P., Hodosh, M., & Hockenmaier, J. (2010). Collecting image annotations using Amazon’s Mechanical Turk. In NAACL Workshop on Creating Speech and Language Data With Amazon’s Mechanical Turk, pp. 139–147, Los Angeles, CA, USA.
  53. 53.Rasiwasia, N., Pereira, J. C., Coviello, E., Doyle, G., Lanckriet, G. R., Levy, R., & Vasconcelos, N. (2010). A new approach to cross-modal multimedia retrieval. In Proceedings of the International Conference on Multimedia (MM), pp. 251–260, New York, NY, USA.
  54. 54.Reiter, E., & Belz, A. (2009). An investigation into the validity of some metrics for automatically evaluating natural language generation systems. Computational Linguistics, 35 (4), 529–558.
  55. 55.Shatford, S. (1986). Analyzing the subject of a picture: A theoretical approach. Cataloging & Classification Quarterly, 6, 39–62.
  56. 56.Shawe-Taylor, J., & Cristianini, N. (2004). Kernel Methods for Pattern Analysis. Cambridge University Press.
  57. 57.Smucker, M. D., Allan, J., & Carterette, B. (2007). A comparison of statistical significance tests for information retrieval evaluation. In Proceedings of the Sixteenth ACM Conference on Information and Knowledge Management (CIKM), pp. 623–632, Lisbon, Portugal.
  58. 58.Socher, R., & Li, F.-F. (2010). Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora. In Proceedings of the 2010 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 966–973, San Francisco, CA, USA.
  59. 59.van Erp, M., & Schomaker, L. (2000). Variants of the Borda count method for combining ranked classifier hypotheses. In Proceedings of the Seventh International Workshop on Frontiers in Handwriting Recognition (IWFHR), pp. 443–452, Nijmegen, Netherlands.
  60. 60.Varma, M., & Zisserman, A. (2005). A statistical approach to texture classification from single images. International Journal of Computer Vision, 62, 61–81.
  61. 61.Vedaldi, A., & Fulkerson, B. (2008). VLFeat: An open and portable library of computer vision algorithms. http://www.vlfeat.org/.
  62. 62.Weston, J., Bengio, S., & Usunier, N. (2010). Large scale image annotation: learning to rank with joint word-image embeddings. Machine Learning, 81 (1), 21–35.
  63. 63.Yang, Y., Teo, C., Daume III, H., & Aloimonos, Y. (2011). Corpus-guided sentence generation of natural images. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 444–454, Edinburgh, UK.

Citation

MLA
Hodosh, M., et al. “Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics”. Journal of Artificial Intelligence Research, vol. 47, 2013, pp. 853–99, https://doi.org/10.1613/jair.3994.
APA
Hodosh, M., Young, P., & Hockenmaier, J. (2013). Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics. Journal of Artificial Intelligence Research, 47, 853–899. https://doi.org/10.1613/jair.3994
Chicago
Hodosh, M., P. Young, and J. Hockenmaier. 2013. “Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics”. Journal of Artificial Intelligence Research 47: 853–99. https://doi.org/10.1613/jair.3994.
Harvard
Hodosh, M., Young, P. and Hockenmaier, J. (2013) “Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics”, Journal of Artificial Intelligence Research, 47, pp. 853–899. Available at: https://doi.org/10.1613/jair.3994.
Vancouver
1. Hodosh M, Young P, Hockenmaier J (2013) Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics. Journal of Artificial Intelligence Research 47:853–899

BibTeX

@article{Hodosh_2013, title={Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics}, volume={47}, ISSN={1076-9757}, url={http://dx.doi.org/10.1613/jair.3994}, DOI={10.1613/jair.3994}, journal={Journal of Artificial Intelligence Research}, publisher={AI Access Foundation}, author={Hodosh, M. and Young, P. and Hockenmaier, J.}, year={2013}, month=Aug, pages={853–899} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/