DeeCap: Dynamic Early Exiting for Efficient Image Captioning

Zhengcong FeiXu YanShuhui WangQi Tian

article2022CVPR66 citations

Proposes DeeCap, an efficient image captioning framework that uses imitation learning to approximate deep layer representations from shallow features, enabling dynamic early exiting in Transformer decoders to achieve a 4x inference speed-up with minimal accuracy loss.

Listen

Modern artificial intelligence systems for image captioning rely heavily on deep neural networks that translate visual information into natural language. While these architectures generate high-quality descriptions, their substantial computational requirements cause significant latency, making them difficult and costly to deploy in real-time, resource-constrained environments. Conventional acceleration techniques, such as standard early exiting—which terminates sentence processing in shallow layers when confidence appears high—often fail because initial layers lack the rich semantic information needed for accurate multimodal description.

The article introduces and evaluates DeeCap, an early-exiting framework designed to accelerate image captioning without compromising description quality. The objective is to demonstrate that lightweight imitation learning can predict higher-level semantic features from shallow network layers, enabling faster, reliable early exits during inference.

To test this approach, the researchers conducted extensive empirical experiments using standard benchmark datasets, specifically Microsoft COCO and Flickr30k. They evaluated DeeCap against both complete, standard captioning networks and alternative acceleration methods across standard caption quality metrics—such as BLEU-4 and CIDEr—as well as human evaluation studies and computational speed-up ratios.

The findings show that DeeCap achieves an approximate fourfold (4.35×) speed-up while maintaining caption quality that is nearly identical to fully computed models (achieving a CIDEr score of 129.0 versus 129.5 for the complete baseline). Standard early exiting without deep feature approximation experienced sharp performance drops at higher speeds, whereas DeeCap retained more than two-thirds of the lost performance margin. Furthermore, human evaluations revealed that 82.0% of DeeCap's generated captions passed a human distinction test, compared to only 61.3% for standard early exiting baselines. Combining shallow and imitated deep representations via a gating mechanism also resolved common text errors, such as repetitive phrasing and incomplete sentences.

These results demonstrate that organizations can reduce inference computational costs and server latency by roughly 75% without retraining separate models for different deployment environments. Because DeeCap allows dynamic adjustment of the speed-accuracy threshold at runtime, engineering teams gain operational flexibility to adapt to varying server loads or edge-computing constraints with minimal performance risk.

Organizations deploying image captioning at scale should consider piloting dynamic early-exiting architectures, particularly utilizing feature concatenation for combining layer representations, to lower infrastructure overhead. Before full-scale implementation, teams should evaluate DeeCap within their specific production hardware pipelines to establish latency thresholds that align with their operational latency and caption quality standards.

Confidence in these findings is supported by consistent results across standard academic benchmarks, the online Microsoft COCO evaluation server, and human evaluations. However, practitioners should exercise caution regarding performance on specialized or out-of-domain imagery, as testing was limited to standard academic datasets and evaluation relies on pre-extracted visual feature backbones.

Cover for DeeCap: Dynamic Early Exiting for Efficient Image Captioning

Abstract

Both accuracy and efficiency are crucial for image captioning in real-world scenarios. Although Transformer-based models have gained significant improved captioning performance, their computational cost is very high. A feasible way to reduce the time complexity is to exit the prediction early in internal decoding layers without passing the entire model. However, it is not straightforward to devise early exiting into image captioning due to the following issues. On one hand, the representation in shallow layers lacks high-level semantic and sufficient cross-modal fusion information for accurate prediction. On the other hand, the exiting decisions made by internal classifiers are unreliable sometimes. To solve these issues, we propose DeeCap framework for efficient image captioning, which dynamically selects proper-sized decoding layers from a global perspective to exit early. The key to successful early exiting lies in the specially designed imitation learning mechanism, which predicts the deep layer activation with shallow layer features. By deliberately merging the imitation learning into the whole image captioning architecture, the imitated deep layer representation can mitigate the loss brought by the missing of actual deep layers when early exiting is undertaken, resulting in significant reduction in calculation cost with small sacrifice of accuracy. Experiments on the MS COCO and Flickr30k datasets demonstrate the DeeCap can achieve competitive performances with 4× speed-up. Code is available at: https://github.com/feizc/DeeCap.

Table of Contents

  • 1. Introduction
  • 2. Investigations on Early Exiting
  • 2.1. Are Shallow Representations Sufficient?
  • 2.2. Are Internal Classifiers Reliable?
  • 3. Methodology
  • 3.1. Deep Representations Imitation
  • 3.2. Multi-Level Representations Fusion
  • 3.3. Gate Decision Mechanism
  • 3.4. Training and Inference
  • 4. Experiments
  • 4.1. Experimental Preparation
  • 4.2. Overall Results
  • 4.3. Model Analysis
  • 4.4. Case Study
  • 4.5. Human Evaluation
  • 5. Related Works
  • 6. Conclusion
  • Acknowledgment
  • References

Knowls

  1. Knowl 1 — Deep Representation Imitation Mechanism

    model/method

    In Transformer-based image captioning decoders with NN layers, early exiting at an intermediate layer m<Nm < N traditionally discards the high-level semantic and cross-modal representations learned in deeper layers k>mk > m. DeeCap approximates these uncomputed deep representations directly from the current shallow hidden state hm∈Rdh_m \in \mathbb{R}^d using layer-specific Multi-Layer Perceptron (MLP) networks.

    For any target deep layer k∈{m+1,…,N}k \in \{m+1, \dots, N\}, the kk-th imitation network generates an approximated representation h^km\hat{h}_k^m:

    h^km=MLPk(hm)\hat{h}_k^m = \text{MLP}_k(h_m)

    To align the approximated representation with the ground-truth hidden representation hkh_k obtained during full forward propagation, the discrepancy is measured via cosine distance:

    Cos-Sim(hk,h^km)=1−h^km⋅hk∥h^km∥2∥hk∥2\text{Cos-Sim}(h_k, \hat{h}_k^m) = 1 - \frac{\hat{h}_k^m \cdot h_k}{\|\hat{h}_k^m\|_2 \|h_k\|_2}

    where ∥⋅∥2\|\cdot\|_2 denotes the Euclidean (L2L_2) norm. Aggregating over all possible intermediate exit layers m∈{2,…,N−1}m \in \{2, \dots, N-1\} and all downstream target layers k∈{m+1,…,N}k \in \{m+1, \dots, N\}, the total deep imitation objective Limit\mathcal{L}_{imit} is:

    Limit=1N−1∑m=2N1N−m∑k=m+1NCos-Sim(hk,h^km)\mathcal{L}_{imit} = \frac{1}{N-1} \sum_{m=2}^N \frac{1}{N-m} \sum_{k=m+1}^N \text{Cos-Sim}(h_k, \hat{h}_k^m)

  2. Knowl 2 — Multi-Level Representation Fusion and Dynamic Gate Decision Mechanism

    model/method

    At decoder layer mm, the model collects all computed shallow hidden states {h1,…,hm}\{h_1, \dots, h_m\} and all approximated deep hidden states {h^m+1m,…,h^Nm}\{\hat{h}_{m+1}^m, \dots, \hat{h}_N^m\}. These collections are aggregated into shallow and deep summaries via a fusion operator g(⋅)g(\cdot):

    hshallow=g({h1,…,hm})h_{shallow} = g(\{h_1, \dots, h_m\})

    hdeep=g({h^m+1m,…,h^Nm})h_{deep} = g(\{\hat{h}_{m+1}^m, \dots, \hat{h}_N^m\})

    The operator g(⋅)g(\cdot) is implemented via sequence-dimension concatenation followed by a linear compression layer, which outperforms averaging, attention-pooling, and recurrent aggregation.

    Because computed shallow features and approximated deep features have different levels of certainty, a dynamic gating mechanism learns a weighting factor α∈[0,1]\alpha \in [0, 1]:

    α=σ(MLP([hshallow,hdeep]))\alpha = \sigma(\text{MLP}([h_{shallow}, h_{deep}]))

    zm=αhshallow+(1−α)hdeepz_m = \alpha h_{shallow} + (1 - \alpha) h_{deep}

    where [⋅,⋅][\cdot, \cdot] denotes vector concatenation, σ(⋅)\sigma(\cdot) is the sigmoid activation function, and MLP\text{MLP} is a multi-layer perceptron. The resulting representation zmz_m is fed to the mm-th internal classifier to produce the output token probability distribution pm=softmax(zm)p_m = \text{softmax}(z_m).

  3. Knowl 3 — DeeCap Multi-Task Objective, Layer Re-weighting, and Progressive Freezing

    model/method

    The DeeCap training objective combines layer-wise cross-entropy loss Lce\mathcal{L}_{ce} across all NN decoder layers with the deep imitation loss Limit\mathcal{L}_{imit} using a balancing hyperparameter λ∈[0,1]\lambda \in [0, 1]:

    L=λLce+(1−λ)Limit\mathcal{L} = \lambda \mathcal{L}_{ce} + (1 - \lambda) \mathcal{L}_{imit}

    where

    Lce=−∑m=1Nwm∑yi∈Vyilog⁡(pm(yi))\mathcal{L}_{ce} = - \sum_{m=1}^N w_m \sum_{y_i \in V} y_i \log(p_m(y_i))

    VV is the target vocabulary, yiy_i is the one-hot ground-truth token indicator, and pm(yi)p_m(y_i) is the predicted probability from the mm-th layer classifier.

    Training incorporates two specific modifications:

    1. Layer Loss Re-weighting: Because shallow layers receive updates from all downstream training signals, their cross-entropy loss is re-weighted by layer depth mm:

    wm=m∑k=1Nkw_m = \frac{m}{\sum_{k=1}^N k}

    1. Progressive Layer Freezing: To prevent destroying well-trained Transformer representations during early-exiting fine-tuning, the parameters of decoder layer mm are frozen with probability pmp_m, where pmp_m decreases linearly from 1.01.0 at layer 11 to 0.00.0 at layer NN.
  4. Knowl 4 — False Confidence Score (FCS) Metric for Early Exiting Classifiers

    definition

    The False Confidence Score (FCS) quantifies how reliably an intermediate classifier's prediction confidence reflects token difficulty. Let context Ci=(x,y<i)C_i = (x, y_{<i}) denote image xx and previously generated prefix y<iy_{<i}.

    • Token Difficulty d(yi)∈{0,1}d(y_i) \in \{0, 1\}: d(yi)=1d(y_i) = 1 if the model cannot generate ground-truth token yiy_i correctly under context CiC_i (difficult token), and d(yi)=0d(y_i) = 0 if it generates it correctly (easy token).
    • Prediction Confidence c(yi)∈[0,1]c(y_i) \in [0, 1]: the predicted probability assigned to ground-truth token yiy_i under context CiC_i.

    The pairwise False Confidence function FC(yi,yj)\text{FC}(y_i, y_j) measures whether a difficult token is incorrectly assigned higher confidence than an easier token:

    FC(yi,yj)={0if d(yi)>d(yj) and c(yi)<c(yj)1otherwise\text{FC}(y_i, y_j) = \begin{cases} 0 & \text{if } d(y_i) > d(y_j) \text{ and } c(y_i) < c(y_j) \\ 1 & \text{otherwise} \end{cases}

    All LL context-token pairs in the evaluation dataset are sorted in ascending order of confidence such that c(yi)<c(yj)c(y_i) < c(y_j) for all i<ji < j. The normalized False Confidence Score is computed as:

    FCS=1−1Q∑i=2L∑j=1i−1FC(yi,yj)\text{FCS} = 1 - \frac{1}{Q} \sum_{i=2}^L \sum_{j=1}^{i-1} \text{FC}(y_i, y_j)

    where Q=12L(L−1)Q = \frac{1}{2} L(L-1) normalizes FCS∈[0,1]\text{FCS} \in [0, 1]. Higher FCS values indicate that the classifier's confidence scores properly prioritize easier tokens over harder ones, yielding more dependable early exit decisions.

  5. Knowl 5 — Entropy-Based Dynamic Early Exiting Inference Algorithm

    algorithm

    During inference, DeeCap generates captions autoregressively. At each decoding step, forward propagation evaluates each decoder layer sequentially until the entropy of the current layer's prediction falls below a user-selected threshold τ\tau.

    Input: Visual features from encoder, prefix tokens y<iy_{<i}, number of decoder layers NN, entropy threshold τ\tau
    Output: Generated token yiy_i
    h0←Embedding(y<i)h_0 \leftarrow \text{Embedding}(y_{<i})
    for m=1m = 1 to NN do
        hm←DecoderLayerm(hm−1)h_m \leftarrow \text{DecoderLayer}_m(h_{m-1})
        for k=m+1k = m+1 to NN do
            h^km←MLPk(hm)\hat{h}_k^m \leftarrow \text{MLP}_k(h_m)
        end for
        hshallow←g({h1,…,hm})h_{shallow} \leftarrow g(\{h_1, \dots, h_m\})
        hdeep←g({h^m+1m,…,h^Nm})h_{deep} \leftarrow g(\{\hat{h}_{m+1}^m, \dots, \hat{h}_N^m\})
        α←σ(MLP([hshallow,hdeep]))\alpha \leftarrow \sigma(\text{MLP}([h_{shallow}, h_{deep}]))
        zm←αhshallow+(1−α)hdeepz_m \leftarrow \alpha h_{shallow} + (1 - \alpha) h_{deep}
        pm←softmax(zm)p_m \leftarrow \text{softmax}(z_m)
        e(pm)←−∑v∈Vpm(v)log⁡pm(v)e(p_m) \leftarrow -\sum_{v \in V} p_m(v) \log p_m(v)
        if e(pm)<τe(p_m) < \tau or m==Nm == N then
            yi←arg⁡max⁡v∈Vpm(v)y_i \leftarrow \arg\max_{v \in V} p_m(v)
            return yiy_i
        end if
    end for

    The effective computational speedup ratio across a dataset is computed by comparing the number of executed layers against full execution:

    SpeedUp=∑m=1NN×wm∑m=1Nm×wm\text{SpeedUp} = \frac{\sum_{m=1}^N N \times w^m}{\sum_{m=1}^N m \times w^m}

    where wmw^m is the count of tokens that exited at decoder layer mm.

  6. Knowl 6 — Performance Comparison on MS COCO Karpathy Test Split

    data/table

    The table compares DeeCap against autoregressive baselines, non-autoregressive fast models, and vanilla early exiting (TF-EE) on the MS COCO Karpathy test split. BLEU-1 (B-1), BLEU-4 (B-4), METEOR (M), ROUGE-L (R), CIDEr (C), and SPICE (S) are reported as percentages (%).

    Models BLEU-1 BLEU-4 METEOR ROUGE CIDEr SPICE SpeedUp
    Autoregressive Image Captioning models
    NIC-v2 - 32.1 25.7 - 99.8 - -
    Up-Down 79.8 36.3 27.7 56.9 120.1 21.4 -
    AoANet 80.2 38.9 29.2 58.8 129.8 22.4 -
    M2-T 80.8 39.1 29.2 58.6 131.2 22.6 -
    TF-Complete 80.2 38.8 29.0 58.3 129.5 22.7 1.00×\times
    Non-Autoregressive Image Captioning models
    MNIC 75.4 30.9 27.5 55.6 108.1 21.0 2.80×\times
    FNIC - 36.2 27.1 55.3 115.7 20.2 8.15×\times
    MIR - 32.5 27.2 55.4 109.5 20.6 1.56×\times
    CMAL 80.3 37.3 28.1 58.0 124.0 21.8 13.90×\times
    IBM 77.2 36.6 27.8 56.2 113.2 20.9 3.06×\times
    SAIC 80.3 38.4 29.0 58.1 127.1 21.9 3.42×\times
    Early Exiting-based Image Captioning models
    TF-EE 79.8 37.2 28.2 57.7 126.3 21.8 4.54×\times
    DeeCap 80.1 38.7 29.1 58.1 129.0 22.5 4.35×\times

    DeeCap obtains 38.738.7 BLEU-4 and 129.0129.0 CIDEr at a 4.35×4.35\times speedup, matching the quality of the full 6-layer Transformer (TF-Complete: 38.838.8 BLEU-4, 129.5129.5 CIDEr) while significantly outperforming vanilla early exiting (TF-EE: 37.237.2 BLEU-4, 126.3126.3 CIDEr at 4.54×4.54\times speedup).

  7. Knowl 7 — Ablation on Deep Representation Imitation Across Speedup Ratios

    data/table

    The ablation evaluates the impact of removing the deep information imitation mechanism (-w/o Deep Info.) from DeeCap on the MS COCO Karpathy offline test set across two speedup ratios (2×2\times and 4×4\times).

    Methods FCS (↑\uparrow) B-4 C
    DeeCap (×2\times 2) 80.12 38.9 129.5
    -w/o Deep Info. 78.67 38.5 128.3
    DeeCap (×4\times 4) 82.40 38.7 129.0
    -w/o Deep Info. 79.55 38.2 127.8

    Incorporating imitated deep representations improves decision reliability (FCS increases by +1.45+1.45 at 2×2\times and by +2.85+2.85 at 4×4\times) and caption quality (CIDEr increases by +1.2+1.2 at 2×2\times and by +1.2+1.2 at 4×4\times). The benefit is larger at higher acceleration (4×4\times), where tokens exit at shallower layers that otherwise lack high-level semantic context.

  8. Knowl 8 — Ablation of Multi-Level Hidden State Fusion Strategies

    data/table

    Comparison of four representation fusion operators g(⋅)g(\cdot) for aggregating multi-level hidden states on the MS COCO offline test set:

    • Average: Direct element-wise mean across hidden states.
    • Concatenation: Concatenation along the sequence dimension followed by a linear projection.
    • Attention-Pooling: Weighted sum using the last hidden state as attention query.
    • SeqNN: Feeding multi-level states sequentially into an LSTM and taking the final hidden state.
    Methods B-4 M R C S
    Average 38.3 28.8 57.7 127.3 21.9
    Concatenation 38.7 29.1 58.1 129.0 22.5
    Attention-Pooling 38.6 29.0 57.9 128.7 22.3
    SeqNN 38.5 29.0 58.0 129.0 22.3

    Concatenation achieves the highest scores across all metrics (BLEU-4 38.7, METEOR 29.1, ROUGE-L 58.1, CIDEr 129.0, SPICE 22.5), indicating that linear interaction across all historical layer states produces the most discriminative combined feature.

  9. Knowl 9 — Performance on Online MS COCO Test Server

    data/table

    Leaderboard evaluation results on the official MS COCO online test server with 5 reference captions (c5) and 40 reference captions (c40). Models marked with ∗* are ensemble models.

    BLEU-1 BLEU-2 BLEU-3 BLEU-4 METEOR ROUGE-L CIDEr
    Models c5 c40 c5 c40 c5 c40 c5 c40 c5 c40 c5 c40 c5 c40
    Up-Down∗^* 80.2 95.2 64.1 88.8 49.1 79.4 36.9 68.5 27.6 36.7 57.1 72.4 117.9 120.5
    AoANet∗^* 81.0 95.0 65.8 89.6 51.4 81.3 39.4 71.2 29.1 38.5 58.9 74.5 126.9 129.6
    M2-T∗^* 81.6 96.0 66.4 90.8 51.8 82.7 39.7 72.8 29.4 39.0 59.2 74.8 129.3 132.1
    CMAL 79.8 94.3 63.8 87.2 48.8 77.2 36.8 66.1 27.9 36.4 57.6 72.0 119.3 121.2
    DeeCap 80.5 95.1 65.2 89.1 50.3 80.0 38.1 69.5 28.0 37.0 58.4 73.5 121.4 124.4

    DeeCap outperforms the top-performing non-autoregressive acceleration baseline (CMAL) across all metrics, achieving an improvement of +2.1+2.1 CIDEr points on c5 (121.4121.4 vs 119.3119.3) and +3.2+3.2 CIDEr points on c40 (124.4124.4 vs 121.2121.2).

  10. Knowl 10 — Human Evaluation and Turing Test Comparison

    empirical result

    In a human evaluation study on 300 randomly selected images from the MS COCO test set, eight evaluators conducted a Turing test: each worker was presented with an image paired with a caption generated by human annotators, the proposed DeeCap model, or the vanilla early-exiting model (TF-EE), and was asked whether the caption was produced by a human or an automated system.

    The percentage of captions judged to be human-generated was:

    • Human Ground Truth: 91.7%91.7\%
    • DeeCap: 82.0%82.0\%
    • Vanilla Early Exiting (TF-EE): 61.3%61.3\%

    DeeCap shows a +20.7%+20.7\% absolute improvement over vanilla early exiting, demonstrating that deep feature approximation effectively prevents common early-exit errors such as incomplete syntax and word repetition.

Coverage note — None was omitted; all primary methodological components, formulations, training strategies, and empirical evaluation results were captured.

References

  1. 1.Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. Spice: Semantic propositional image caption evaluation. In Proc. ECCV, pages 382–398, 2016.
  2. 2.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. Bottom-up and top-down attention for image captioning and visual question answering. In Proc. IEEE CVPR, pages 6077–6080, 2018.
  3. 3.Alexandre Attia and Sharone Dayan. Global overview of imitation learning. arXiv preprint arXiv:1801.06503, 2018.
  4. 4.Shuang Bai and Shan An. A survey on automatic image caption generation. Neurocomputing, 311:291–304, 2018.
  5. 5.Ali Furkan Biten, Lluis Gomez, and Dimosthenis Karatzas. Let there be a clock on the beach: Reducing object hallucination in image captioning. arXiv preprint arXiv:2110.01705, 2021.
  6. 6.Long Chen, Zhihong Jiang, Jun Xiao, and Wei Liu. Human-like controllable image captioning with verb-specific semantic roles. In Proc. IEEE CVPR, pages 16846–16856, 2021.
  7. 7.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015.
  8. 8.Xinlei Chen and C Lawrence Zitnick. Mind’s eye: A recurrent visual representation for image caption generation. In Proc. IEEE CVPR, pages 2422–2431, 2015.
  9. 9.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Learning universal image-text representations. arXiv preprint arXiv:1909.11740, 2019.
  10. 10.Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, and Rita Cucchiara. Meshed-memory transformer for image captioning. In Proc. IEEE CVPR, pages 10578–10587, 2020.
  11. 11.Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli. Depth-adaptive transformer. In Proc. ICLR, pages 1–14, 2020.
  12. 12.Zhengcong Fei. Fast image caption generation with position alignment. arXiv preprint arXiv:1912.06365, 2019.
  13. 13.Zhengcong Fei. Iterative back modification for faster image captioning. In Proc. ACM MM, pages 3182–3190, 2020.
  14. 14.Zhengcong Fei. Memory-augmented image captioning. In Proc. AAAI, pages 2–9, 2021.
  15. 15.Zhengcong Fei. Partially non-autoregressive image captioning. In Proc. AAAI, volume 35, pages 1309–1316, 2021.
  16. 16.Junlong Gao, Xi Meng, Shiqi Wang, Xia Li, Shanshe Wang, Siwei Ma, and Wen Gao. Masked non-autoregressive image captioning. arXiv preprint arXiv:1906.00717, 2019.
  17. 17.Amir Ghodrati, Babak Ehteshami Bejnordi, and Amirhossein Habibian. Frameexit: Conditional early exiting for efficient video recognition. In Proc. IEEE CVPR, pages 15608–15618, 2021.
  18. 18.Alex Graves. Adaptive computation time for recurrent neural networks. arXiv preprint arXiv:1603.08983, 2016.
  19. 19.Jiatao Gu, James Bradbury, Caiming Xiong, Victor O. K. Li, and Richard Socher. Non-autoregressive neural machine translation. In Proc. ICLR, 2018.
  20. 20.Longteng Guo, Jing Liu, Xinxin Zhu, Xingjian He, Jie Jiang, and Hanqing Lu. Non-autoregressive image captioning with counterfactuals-critical multi-agent learning. arXiv preprint arXiv:2005.04690, 2020.
  21. 21.Gao Huang, Danlu Chen, Tianhong Li, Felix Wu, Laurens van der Maaten, and Kilian Q Weinberger. Multi-scale dense networks for resource efficient image classification. arXiv preprint arXiv:1703.09844, 2017.
  22. 22.Lun Huang, Wenmin Wang, Jie Chen, and Xiao-Yong Wei. Attention on attention for image captioning. In Proc. IEEE ICCV, pages 4634–4643, 2019.
  23. 23.Lee Jason, Mansimov Elman, Graham Neubig, and Cho Kyunghyun. Deterministic non-autoregressive neural sequence modeling by iterative refinement. In Proc. EMNLP, pages 1138–1149, 2018.
  24. 24.Andrej Karpathy and Li Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In Proc. IEEE CVPR, pages 3128–3137, 2015.
  25. 25.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  26. 26.Alon Lavie and Abhaya Agarwal. Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgments. In Proc. ACL Workshop, pages 228–231, 2007.
  27. 27.Chen-Yu Lee, Saining Xie, Patrick Gallagher, Zhengyou Zhang, and Zhuowen Tu. Deeply-supervised nets. In Artificial intelligence and statistics, pages 562–570. PMLR, 2015.
  28. 28.Guang Li, Linchao Zhu, Ping Liu, and Yi Yang. Entangled transformer for image captioning. In Proc. IEEE ICCV, pages 8928–8937, 2019.
  29. 29.Lei Li, Yankai Lin, Deli Chen, Shuhuai Ren, Peng Li, Jie Zhou, and Xu Sun. Cascadebert: Accelerating inference of pre-trained language models via calibrated complete models cascade. arXiv preprint arXiv:2012.14682, 2020.
  30. 30.Xiaonan Li, Yunfan Shao, Tianxiang Sun, Hang Yan, Xipeng Qiu, and Xuanjing Huang. Accelerating bert inference for sequence labeling via early-exit. arXiv preprint arXiv:2105.13878, 2021.
  31. 31.Kaiyuan Liao, Yi Zhang, Xuancheng Ren, Qi Su, Xu Sun, and Bin He. A global past-future early exit method for accelerating inference of pre-trained language models. In Proc. NAACL, pages 2013–2023, 2021.
  32. 32.Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Proc. ACL Workshops, pages 74–81, 2004.
  33. 33.Fenglin Liu, Xuancheng Ren, Xian Wu, Shen Ge, Wei Fan, Yuexian Zou, and Xu Sun. Prophet attention: Predicting attention with future attention for improved image captioning. In Proc. NIPS, 2021.
  34. 34.Weijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao, Haotang Deng, and Qi Ju. Fastbert: a self-distilling bert with adaptive inference time. In Proc. ACL, pages 6035–6044, 2020.
  35. 35.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021.
  36. 36.Yunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao, Yongjian Wu, Feiyue Huang, Chia-Wen Lin, and Rongrong Ji. Dual-level collaborative transformer for image captioning. In Proc. AAAI, volume 35, pages 2286–2293, 2021.
  37. 37.Yingwei Pan, Ting Yao, Yehao Li, and Tao Mei. X-linear attention networks for image captioning. In Proc. IEEE CVPR, pages 10971–10980, 2020.
  38. 38.Kishore Papineni, Salim Roukos, Todd Ward, and Wei Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Proc. ACL, pages 311–318, 2002.
  39. 39.Emmanouil Antonios Platanios, Otilia Stretcu, Graham Neubig, Barnabás Póczos, and Tom Mitchell. Competence-based curriculum learning for neural machine translation. In Proc. NACCL, pages 1162–1172, 2019.
  40. 40.Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In Proc. IEEE ICCV, pages 2641–2649, 2015.
  41. 41.Steven J. Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel. Self-critical sequence training for image captioning. In Proc. IEEE CVPR, pages 1179–1195, 2017.
  42. 42.Stéphane Ross, Geoffrey Gordon, and Drew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In Proc. ICAIS, pages 627–635. JMLR Workshop and Conference Proceedings, 2011.
  43. 43.Stefan Schaal. Learning from demonstration. Proc. NIPS, 9, 1996.
  44. 44.Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, and Noah A Smith. The right tool for the job: Matching model and instance complexities. In Proc. ACL, 2020.
  45. 45.Luca Soldaini and Alessandro Moschitti. The cascade transformer: an application for efficient answer sentence selection. In Proc. ACL, pages 5697–5708, 2020.
  46. 46.Tianxiang Sun, Yunhua Zhou, Xiangyang Liu, Xinyu Zhang, Hao Jiang, Zhao Cao, Xuanjing Huang, and Xipeng Qiu. Early exiting with ensemble internal classifiers. arXiv preprint arXiv:2105.13792, 2021.
  47. 47.Surat Teerapittayanon, Bradley McDanel, and Hsiang-Tsung Kung. Branchynet: Fast inference via early exiting from deep neural networks. In Proc. ICPR, pages 2464–2469. IEEE, 2016.
  48. 48.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proc. NIPS, pages 5998–6008, 2017.
  49. 49.Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evaluation. In Proc. IEEE CVPR, pages 4566–4575, 2015.
  50. 50.Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. Show and tell: A neural image caption generator. In Proc. IEEE CVPR, pages 3156–3164, 2015.
  51. 51.Bingzhen Wei, Mingxuan Wang, Hao Zhou, Junyang Lin, Jun Xie, and Xu Sun. Imitation learning for nonautoregressive neural machine translation. arXiv preprint arXiv:1906.02041, 2019.
  52. 52.Yijun Xiao and William Yang Wang. On hallucination and predictive uncertainty in conditional language generation. arXiv preprint arXiv:2103.15025, 2021.
  53. 53.Ji Xin, Rodrigo Nogueira, Yaoliang Yu, and Jimmy Lin. Early exiting bert for efficient document ranking. In Proceedings of SustaiNLP: Workshop on Simple and Efficient Natural Language Processing, pages 83–88, 2020.
  54. 54.Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. Deebert: Dynamic early exiting for accelerating bert inference. In Proc. ACL, pages 2246–2251, 2020.
  55. 55.Guanghui Xu, Shuaicheng Niu, Mingkui Tan, Yucheng Luo, Qing Du, and Qi Wu. Towards accurate text-based image captioning with content diversity exploration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12637–12646, 2021.
  56. 56.Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio. Show, attend and tell: Neural image caption generation with visual attention. In Proc. ICML, pages 2048–2057, 2015.
  57. 57.Xu Yan, Zhengcong Fei, Zekang Li, Shuhui Wang, Qingming Huang, and Qi Tian. Semi-autoregressive image captioning. In Proceedings of the 29th ACM International Conference on Multimedia, pages 2708–2716, 2021.
  58. 58.Bang Yang, Fenglin Liu, Can Zhang, and Yuexian Zou. Non-autoregressive coarse-to-fine video captioning. arXiv preprint arXiv:1911.12018, 2019.
  59. 59.Le Yang, Yizeng Han, Xi Chen, Shiji Song, Jifeng Dai, and Gao Huang. Resolution adaptive networks for efficient inference. In Proc. IEEE CVPR, pages 2369–2378, 2020.
  60. 60.Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei. Exploring visual relationship for image captioning. In Proc. ECCV, pages 684–699, 2018.
  61. 61.Xuying Zhang, Xiaoshuai Sun, Yunpeng Luo, Jiayi Ji, Yiyi Zhou, Yongjian Wu, Feiyue Huang, and Rongrong Ji. Rstnet: Captioning with adaptive attention on visual and nonvisual words. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15465–15474, 2021.
  62. 62.Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J Corso, and Jianfeng Gao. Unified vision-language pre-training for image captioning and vqa. arXiv preprint arXiv:1909.11059, 2019.
  63. 63.Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei. Bert loses patience: Fast and robust inference with early exit. Proc. NIPS, 33, 2020.
  64. 64.Yuanen Zhou, Yong Zhang, Zhenzhen Hu, and Meng Wang. Semi-autoregressive transformer for image captioning. In Proc. ICCV Workshop, pages 3139–3143, 2021.

Citation

MLA
Fei, Z., et al. “DeeCap: Dynamic Early Exiting for Efficient Image Captioning”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 12206–16, https://doi.org/10.1109/CVPR52688.2022.01190.
APA
Fei, Z., Yan, X., Wang, S., & Tian, Q. (2022). DeeCap: Dynamic Early Exiting for Efficient Image Captioning. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12206–12216. https://doi.org/10.1109/CVPR52688.2022.01190
Chicago
Fei, Z., X. Yan, S. Wang, and Q. Tian. 2022. “DeeCap: Dynamic Early Exiting for Efficient Image Captioning”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12206–16. https://doi.org/10.1109/CVPR52688.2022.01190.
Harvard
Fei, Z. et al. (2022) “DeeCap: Dynamic Early Exiting for Efficient Image Captioning”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 12206–12216. Available at: https://doi.org/10.1109/CVPR52688.2022.01190.
Vancouver
1. Fei Z, Yan X, Wang S, Tian Q (2022) DeeCap: Dynamic Early Exiting for Efficient Image Captioning. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 12206–12216

BibTeX

@inproceedings{Fei_2022, title={DeeCap: Dynamic Early Exiting for Efficient Image Captioning}, url={http://dx.doi.org/10.1109/CVPR52688.2022.01190}, DOI={10.1109/cvpr52688.2022.01190}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Fei, Zhengcong and Yan, Xu and Wang, Shuhui and Tian, Qi}, year={2022}, month=June, pages={12206–12216} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE