DeeCap: Dynamic Early Exiting for Efficient Image Captioning
Zhengcong FeiXu YanShuhui WangQi Tian
Proposes DeeCap, an efficient image captioning framework that uses imitation learning to approximate deep layer representations from shallow features, enabling dynamic early exiting in Transformer decoders to achieve a 4x inference speed-up with minimal accuracy loss.
Modern artificial intelligence systems for image captioning rely heavily on deep neural networks that translate visual information into natural language. While these architectures generate high-quality descriptions, their substantial computational requirements cause significant latency, making them difficult and costly to deploy in real-time, resource-constrained environments. Conventional acceleration techniques, such as standard early exiting—which terminates sentence processing in shallow layers when confidence appears high—often fail because initial layers lack the rich semantic information needed for accurate multimodal description.
The article introduces and evaluates DeeCap, an early-exiting framework designed to accelerate image captioning without compromising description quality. The objective is to demonstrate that lightweight imitation learning can predict higher-level semantic features from shallow network layers, enabling faster, reliable early exits during inference.
To test this approach, the researchers conducted extensive empirical experiments using standard benchmark datasets, specifically Microsoft COCO and Flickr30k. They evaluated DeeCap against both complete, standard captioning networks and alternative acceleration methods across standard caption quality metrics—such as BLEU-4 and CIDEr—as well as human evaluation studies and computational speed-up ratios.
The findings show that DeeCap achieves an approximate fourfold (4.35×) speed-up while maintaining caption quality that is nearly identical to fully computed models (achieving a CIDEr score of 129.0 versus 129.5 for the complete baseline). Standard early exiting without deep feature approximation experienced sharp performance drops at higher speeds, whereas DeeCap retained more than two-thirds of the lost performance margin. Furthermore, human evaluations revealed that 82.0% of DeeCap's generated captions passed a human distinction test, compared to only 61.3% for standard early exiting baselines. Combining shallow and imitated deep representations via a gating mechanism also resolved common text errors, such as repetitive phrasing and incomplete sentences.
These results demonstrate that organizations can reduce inference computational costs and server latency by roughly 75% without retraining separate models for different deployment environments. Because DeeCap allows dynamic adjustment of the speed-accuracy threshold at runtime, engineering teams gain operational flexibility to adapt to varying server loads or edge-computing constraints with minimal performance risk.
Organizations deploying image captioning at scale should consider piloting dynamic early-exiting architectures, particularly utilizing feature concatenation for combining layer representations, to lower infrastructure overhead. Before full-scale implementation, teams should evaluate DeeCap within their specific production hardware pipelines to establish latency thresholds that align with their operational latency and caption quality standards.
Confidence in these findings is supported by consistent results across standard academic benchmarks, the online Microsoft COCO evaluation server, and human evaluations. However, practitioners should exercise caution regarding performance on specialized or out-of-domain imagery, as testing was limited to standard academic datasets and evaluation relies on pre-extracted visual feature backbones.
- Paper: BranchyNet: Fast inference via early exiting from deep neural networks, Surat Teerapittayanon et al. (2016). This paper establishes the foundational early-exiting paradigm for deep networks, introducing multi-branch architectures that DeeCap adapts specifically for multimodal image captioning.
- Paper: Show and tell: A neural image caption generator, Oriol Vinyals et al. (2015). This seminal work introduces the end-to-end neural encoder-decoder paradigm for image captioning, providing the baseline translation architecture that DeeCap accelerates.
- Paper: Show, Attend and Tell: Neural Image Caption Generation with Visual Attention, Kelvin Xu et al. (2015). This foundational paper presents spatial visual attention mechanisms for image captioning, establishing the multimodal feature-extraction pipeline leveraged by DeeCap's early exit gates.
- Paper: Self-Critical Sequence Training for Image Captioning, Steven J. Rennie et al. (2016). This work introduces sequence-level metric optimization using reinforcement learning, which defines the standard training objectives and CIDEr benchmark evaluations utilized in DeeCap.
- Paper: CIDEr: Consensus-based image description evaluation, Ramakrishna Vedantam et al. (2014). This paper defines the CIDEr consensus metric, the primary evaluation standard used throughout DeeCap to measure caption quality preservation.
- Paper: Knowing When to Look: Adaptive Attention via a Visual Sentinel for Image Captioning, Jiasen Lu et al. (2016). This work introduces adaptive gating between visual and linguistic states during caption generation, providing conceptual grounding for DeeCap's imitation gating mechanism.
- Paper: Confident Adaptive Language Modeling, Tal Schuster et al. (2022). This paper extends dynamic early exiting to autoregressive language generation with distribution-free risk control guarantees across transformer layers.
- Paper: Speculative Decoding with Big Little Decoder, Sehoon Kim et al. (2023). This work explores speculative decoding with hierarchical decoders as an alternative paradigm to early exiting for accelerating sequence generation.
- Paper: Fast Inference from Transformers via Speculative Decoding, Yaniv Leviathan et al. (2023). This work formalizes speculative decoding to accelerate transformer generation, presenting a complementary approach to DeeCap for reducing autoregressive inference latency.
- Paper: Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time, Zichang Liu et al. (2023). This study investigates contextual sparsity prediction to skip intermediate parameters dynamically at inference time, generalizing dynamic efficiency concepts beyond layer-level early exits.
- Paper: Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration, Yuhang Han et al. (2026). This paper develops training-free visual token reduction to accelerate multimodal language models during inference, addressing computational bottlenecks complementary to early exiting.
