Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge

Pierre L. DogninIgor MelnykYoussef MrouehInkit PadhiMattia RigottiJarret RossYair SchiffRichard A. YoungBrian Belgodere

article2022JAIR65 citations

Presents the winning multimodal framework from the VizWiz 2020 Challenge, integrating rotation-invariant optical character recognition, object detection, and a copy-mechanism transformer to generate task-oriented captions from real-world photos taken by visually impaired users.

Listen

Traditional automated image captioning models are largely trained on curated datasets that produce generic scene descriptions, such as labeling an object simply as a bottle. These standard systems fail when deployed as assistive technologies for visually impaired users, who take everyday photos with mobile devices and require practical, goal-oriented information, such as reading a medicine bottle label or identifying an ingredient. Compounding this challenge, images captured by visually impaired individuals frequently suffer from quality defects like blur, poor lighting, occlusions, and severe tilt or rotation.

The article details the design, implementation, and evaluation of a winning multimodal artificial intelligence architecture engineered specifically for task-oriented assistive image captioning on the VizWiz dataset. The core objective is to evaluate how combining robust image feature extraction, rotation-invariant text reading, object detection, and dynamic vocabulary generation improves caption quality and practical utility for visually impaired users.

The researchers developed a multimodal Transformer pipeline that simultaneously processes three distinct streams of information: raw image features extracted using a network pre-trained on billions of mobile phone images, detected objects, and optical character recognition used to read text. Because over 28% of the dataset images have rotation flaws, the optical character recognition pipeline analyzes images across four 90-degree rotations and selects the orientation yielding the highest number of valid dictionary words. A copy mechanism was integrated to allow the model to directly lift out-of-vocabulary words, such as specific brand names, from the recognized text and object categories straight into the generated description. The system was trained on over 23,000 images and evaluated across competitive benchmark splits comprising 8,000 real-world test images.

The evaluation revealed several key findings regarding assistive caption performance. First, the complete ensemble system secured first place in the VizWiz 2020 Challenge, achieving a top score of 81.04 on the benchmark evaluation metric. Second, incorporating text reading produced the single largest performance leap; on images containing legible text, the system scored 91.60 compared to the second-place model's 77.78, representing an improvement of nearly 18%. Third, the integration of a dynamic copy mechanism significantly enhanced individual model performance, allowing the generator to accurately output previously unseen product names and specific details. Finally, the model demonstrated strong robustness across image difficulties, outperforming all baseline and competing architectures on easy, medium, and hard quality subsets.

These findings demonstrate that assistive vision systems must move beyond generic visual labeling toward specialized, goal-oriented recognition that prioritizes reading printed text and identifying functional objects. By capturing specific names, instructions, and environmental cues, the technology substantially mitigates safety risks for blind users, such as misidentifying medication or consumer goods. The researchers successfully packaged this pipeline into a real-time cloud-based web application paired with text-to-speech synthesis, proving that advanced multimodal sequence models can be practically deployed with low latency for cross-device assistive use.

To build on these results, decision-makers and engineering teams should invest in conversational, two-way visual dialogue capabilities rather than static one-shot captioners. An interactive feedback loop would allow the system to detect unreadable or occluded inputs and prompt the user to retake the picture at a better angle. Organizations should also explore integrating supplementary sensor data, including spatial mapping, color detection, and location coordinates, to improve situational navigation and scene awareness.

Confidence in these findings is high given the extensive ablation experiments and competitive benchmark validation. However, primary limitations remain: the system still experiences performance drops on severely degraded or blurry images where text cannot be recovered, and its vocabulary is bounded by the precision of the underlying text and object detection modules. Deployers should exercise caution in high-risk environments until interactive verification systems are established.

  • Paper: Towards VQA Models That Can Read, Amanpreet Singh et al. (2019). Its OCR-and-copy approach to text-dependent visual questions provides the direct groundwork for understanding why the source reads image text and copies unseen names into captions.
  • Paper: VizWiz Grand Challenge: Answering Visual Questions from Blind People, Danna Gurari et al. (2018). This earlier VizWiz dataset and challenge establish the real-world blind-user data and task context that the source’s 2020 captioning system builds on.

No sufficiently relevant recommendations were found.

Cover for Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge

Abstract

Image captioning has recently demonstrated impressive progress largely owing to the introduction of neural network algorithms trained on curated dataset like MS-COCO. Often work in this field is motivated by the promise of deployment of captioning systems in practical applications. However, the scarcity of data and contexts in many competition datasets renders the utility of systems trained on these datasets limited as an assistive technology in real-world settings, such as helping visually impaired people navigate and accomplish everyday tasks. This gap motivated the introduction of the novel VizWiz dataset, which consists of images taken by the visually impaired and captions that have useful, task-oriented information. In an attempt to help the machine learning computer vision field realize its promise of producing technologies that have positive social impact, the curators of the VizWiz dataset host several competitions, including one for image captioning. This work details the theory and engineering from our winning submission to the 2020 captioning competition. Our work provides a step towards improved assistive image captioning systems.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Model
  • 3.1 Image Feature Extraction with ResNeXt
  • 3.2 Incorporating OCR and Objects
  • 3.2.1 Incorporating OCR
  • 3.2.2 Dictionary-Guided Rotation-Invariant OCR
  • 3.2.3 Object Detection
  • 4. Multimodal Transformer
  • 4.1 Regular Transformer (without copy mechanism)
  • 4.2 Copy Transformer and Dynamic Vocabulary
  • 5. Experiments
  • 5.1 Dataset
  • 5.2 Training Details
  • 5.3 Evaluation on EvalAI
  • 5.4 Post Processing of Captions
  • 5.5 Ablation Studies
  • 5.6 Competition Results
  • 6. Discussion of Goal-Oriented Captioning versus Generic Captioning
  • 7. Image Captioning As an Assistive Technology : Real-Time Demo
  • 8. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Three-modality transformer for goal-oriented captioning

    model/method

    The assistive captioner combines image appearance, detected text, and detected objects so captions can include task-relevant details rather than relying on visual features alone. A ResNeXt extractor supplies 196 image-region vectors of dimension 2048. The OCR and object-detection branches supply, respectively, up to 20 words and up to 10 object labels, embedded as 300-dimensional fastText vectors. Linear projections map all three streams into a shared 512-dimensional space; concatenation produces a 226-position multimodal sequence for a Transformer encoder. A Transformer decoder generates the caption. This architecture lets the decoder use image context together with the text and object information that can identify products or other scene details.

  2. Knowl 2 — Dictionary-guided OCR over four image orientations

    algorithm

    The OCR procedure addresses the fact that images in the VizWiz training and validation sets have rotation issues in approximately 28.5% and 28.3% of cases, respectively. It evaluates the four right-angle orientations and selects the OCR output with the most tokens recognized by the fastText vocabulary as intelligible. Detected text regions are then ranked by bounding-box area and OCR confidence in descending order; at most 20 OCR words are passed onward as the text modality. Thus the method provides rotation robustness over the finite set of orientations 0∘0^\circ, 90∘90^\circ, 180∘180^\circ, and 270∘270^\circ, rather than claiming invariance to arbitrary rotations.

    Input: Image I, OCR recognizer, fastText vocabulary V
    Output: Selected OCR text tokens
    For each angle r in [0, 90, 180, 270]:
        Rotate I by r degrees
        Run OCR on the rotated image
        Tokenize the recognized text with fastText
        Count tokens that are in V
    Select the orientation with the largest count of tokens in V
    Rank that orientation's detected text regions by bounding-box area and OCR confidence, descending
    Return at most 20 ranked OCR words
  3. Knowl 3 — Copy mechanism creates a dynamic output vocabulary

    equation

    The decoder can generate tokens detected by OCR or object detection even when those tokens are absent from its ordinary caption vocabulary. At each generation step, the system expands the output vocabulary with the current image's OCR and object tokens, masks OCR or object logits for tokens not detected in that image, concatenates the remaining logits with the caption-vocabulary logits, and applies a softmax. For a candidate token yty_t, the probability is

    p(yt∣I,y<t)=exp⁡(ϕIMG(yt))+exp⁡(ϕOCR(yt))+exp⁡(ϕOBJ(yt))Z.p(y_t\mid I,y_{<t})=\frac{\exp(\phi_{\mathrm{IMG}}(y_t))+\exp(\phi_{\mathrm{OCR}}(y_t))+\exp(\phi_{\mathrm{OBJ}}(y_t))}{Z}.

    Here II is the input image, y<ty_{<t} is the sequence of previously generated tokens, and ϕIMG\phi_{\mathrm{IMG}}, ϕOCR\phi_{\mathrm{OCR}}, and ϕOBJ\phi_{\mathrm{OBJ}} are unnormalized scores from the caption, OCR, and object channels. An OCR or object score is set to −∞-\infty when the token is absent from that channel's detections; ZZ is the sum of the resulting exponentiated scores over the expanded output vocabulary. This permits copying out-of-vocabulary words read from labels or supplied by the object detector.

  4. Knowl 4 — Winning ensemble achieved the highest reported CIDEr

    empirical result

    The winning system combined 90 models using all three modalities, with both copy-enabled and non-copy models, followed by caption post-processing. On the test-challenge set of 8,000 images it scored 27.44 BLEU4, 22.25 METEOR, 50.20 ROUGE-L, 81.04 CIDEr, and 17.00 SPICE; the paper reports 81.04 as its highest CIDEr score. The test-dev results below show how copy and ensembling behaved: single copy models exceeded single non-copy models on all listed metrics, but the 40-model non-copy ensemble outperformed the 40-model copy ensemble. The mixed 90-model ensemble had the best test-dev CIDEr among these ensemble configurations.

    Evaluation splitCopy settingModel setBLEU4METEORROUGE-LCIDErSPICE
    test-devCopySingle23.6620.3447.3464.4814.81
    test-devCopyEnsemble (40)25.5021.4449.1972.5716.28
    test-devNon CopySingle22.1019.7146.3660.9014.70
    test-devNon CopyEnsemble (40)27.4322.2050.1078.8017.35
    test-devCopy + Non CopyEnsemble (90)27.3122.1350.0280.3817.10
    test-challengeCopy + Non CopyEnsemble (90)27.4422.2550.2081.0417.00
  5. Knowl 5 — Training, ensembling, and competition-time caption selection

    experimental setup

    Experiments used the VizWiz Captions dataset: 23,431 training images with 117,155 captions, 7,750 validation images with 138,750 captions, and 8,000 test images with 40,000 captions. Caption preprocessing removed newline, carriage-return, and tab characters and punctuation, lowercased the text, and applied the BERT tokenizer, replacing words outside its dictionary with <unk>. Captions plus detected OCR and object tokens yielded 35,555 unique observed training tokens.

    Models were first trained with cross-entropy for 10 epochs using batches of 80 and Adam with (β1,β2)=(0.9,0.98)(\beta_1,\beta_2)=(0.9,0.98). The learning rate was warmed for 2,000 minibatch iterations with a factor of 1 and then decayed proportional to 1/i1/\sqrt{i}, where ii is the iteration number. Self-Critical Sequence Training (SCST) fine-tuned selected cross-entropy checkpoints to optimize CIDEr: sampled captions received reward relative to a greedy-decoding baseline, and samples below the baseline were suppressed. SCST used batches of 80, carried over the Adam state, set the step number to 50,000 plus the final cross-entropy step, and ran for a randomly selected 15–40 additional epochs. Ensembles averaged model probabilities for each vocabulary token at every decoding step and selected the most likely token.

    Competition post-processing considered five candidates from high-CIDEr systems. It de-tokenized generated text and ranked candidates using self-BERT (favoring a candidate semantically similar to the other candidates) and OCR-token overlap (favoring a caption containing more recognized image text). The authors used these candidate-ranking steps for competition submissions after observing validation CIDEr improvements.

  6. Knowl 6 — ResNeXt features improved CIDEr over ResNet-101 features

    empirical result

    The authors report a 10-point CIDEr improvement from using ResNeXt rather than ResNet-101 features, with other conditions held equal. They hypothesize that the gain comes from the closer match between phone photographs in VizWiz and the Instagram phone images used to pretrain ResNeXt. Their extractor is the 101-layer ResNeXt, using its 99th layer, cardinality 32, and bottleneck dimension 8. Images are kept at their original size when they meet the minimum-size requirement; otherwise the undersized axis is increased to 320 pixels while preserving aspect ratio. Two-dimensional adaptive max pooling produces a 14×14×204814\times14\times2048 feature map, reshaped into 196 vectors of dimension 2048 for the multimodal Transformer.

  7. Knowl 7 — Object detections provide a separate semantic input stream

    model/method

    The captioner uses EfficientDet without modifications or VizWiz-specific adaptation. The detector recognizes the 80 object categories used by MS-COCO and returns labels with confidence scores. The system retains detections with confidence greater than 0.25, orders them by confidence, and supplies at most 10 object labels to the captioner. Each label is embedded with fastText and passed as its own modality. This choice favors the detector's accuracy as a reliable input, rather than the broader category coverage of alternatives the authors considered.

  8. Knowl 8 — OCR gave the largest gains in the modality ablation

    empirical result

    The authors evaluated ensembles of 20 models without the copy mechanism on the test-dev split, comparing image-only features with image features augmented by OCR or object detection. Each table entry gives the metric before and after post-processing. Adding OCR produced the strongest scores across all five metrics; object detection gave smaller improvements over image-only features. Post-processing also improved CIDEr by 2.46–3.06 points across the three settings.

    ModalitiesBLEU4METEORROUGE-LCIDErSPICE
    IMG only25.11 / 25.9721.01 / 21.2248.22 / 48.5566.8 / 69.4115.83 / 16.04
    IMG + OCR25.88 / 26.8321.49 / 21.7349.17 / 49.5472.45 / 75.5116.23 / 16.47
    IMG + OBJ25.39 / 26.2121.02 / 21.2148.33 / 48.6366.97 / 69.4315.91 / 16.07

    Within each metric, values are reported as before / after post-processing. All ensembles used varying Transformer layer counts, hidden dimensions, and random seeds.

  9. Knowl 9 — Performance advantage varied with text presence and image difficulty

    data/table

    On the test-challenge split, the winning captioner was compared with other competition systems using CIDEr, stratified by whether images contained text and by image difficulty. It led the listed systems on images with text and in all three difficulty groups, while SRC-B scored slightly higher on images without text. The organizers defined hard images as the 25% for which the maximum CIDEr across submissions was below 50, easy images as those with a minimum CIDEr of 100 across submissions (6.73% of images), and the remaining images as medium.

    ConditionOursSRC-BaburnsLittlePandaiimBaseline
    With text91.6077.7867.8864.7664.1764.81
    No text59.8060.4354.9050.3749.749.01
    Easy88.7677.4868.1564.7464.0764.52
    Medium66.7163.2656.6751.9451.4351.27
    Hard41.3139.8735.3033.8333.3931.84
  10. Knowl 10 — Cloud-based real-time assistive captioning demo

    model/method

    The authors implemented a cross-device progressive web application that accepts an image from a phone camera, webcam, or local upload. In the cloud, a first GPU runs the ResNeXt feature extractor, object detector, and OCR; their outputs are combined and sent to a second GPU, where the multimodal Transformer generates a caption using greedy decoding. The caption, recognized text, and detected objects are sent to the Watson Text-to-Speech API, which returns an MP3 for playback. The application also displays the image, caption, detected objects, and OCR words. This demo demonstrates an end-to-end path from image capture to spoken, task-oriented description.

Coverage note — The indoor and outdoor qualitative caption examples are omitted because they are illustrative case studies rather than a systematic result; their OCR, out-of-vocabulary copying, and object-detection functions are captured in the method and quantitative knowls.

References

  1. 1.Anderson, P., Fernando, B., Johnson, M., & Gould, S. (2016). Spice: Semantic propositional image caption evaluation. In European conference on computer vision (pp. 382–398).
  2. 2.Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., . . . Zhang, L. (2018a). Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the ieee/cvf conference on computer vision and pattern recognition (cvpr).
  3. 3.Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., . . . Zhang, L. (2018b). Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the ieee conference on computer vision and pattern recognition (pp. 6077–6086).
  4. 4.Baek, J., Kim, G., Lee, J., Park, S., Han, D., Yun, S., . . . Lee, H. (2019). What is wrong with scene text recognition model comparisons? dataset and model analysis. In International conference on computer vision (iccv).
  5. 5.Baek, Y., Lee, B., Han, D., Yun, S., & Lee, H. (2019). Character region awareness for text detection. In Proceedings of the ieee conference on computer vision and pattern recognition (pp. 9365–9374).
  6. 6.Banerjee, S., & Lavie, A. (2005). Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization (Vol. 29, pp. 65–72).
  7. 7.Bigham, J. P., Jayant, C., Ji, H., Little, G., Miller, A., Miller, R. C., . . . Yeh, T. (2010). Vizwiz: Nearly real-time answers to visual questions. In Proceedings of the 23nd annual acm symposium on user interface software and technology (p. 333–342). New York, NY, USA: Association for Computing Machinery.
  8. 8.Bojanowski, P., Grave, E., Joulin, A., & Mikolov, T. (2016). Enriching word vectors with subword information. arXiv preprint arXiv:1607.04606 .
  9. 9.Brady, E., Morris, M. R., Zhong, Y., White, S., & Bigham, J. P. (2013, April). Visual challenges in the everyday lives of blind people. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (pp. 2117–2126). Association for Computing Machinery.
  10. 10.Bujimalla, S., Subedar, M., & Tickoo, O. (2020). B-SCST: Bayesian Self-Critical Sequence Training for Image Captioning. arXiv preprint arXiv:2004.02435 .
  11. 11.Burns, A., Saenko, K., & Bryan, P. (2020). Feature Refinement for Common Sense Captioning. Boston University, ”https://ivc.ischool.utexas.edu/~yz9244/VizWiz workshop/videos/aburns-oral.mp4”.
  12. 12.Chiu, T.-Y., Zhao, Y., & Gurari, D. (2020, June). Assessing image quality issues for real-world problems. In Proceedings of the ieee/cvf conference on computer vision and pattern recognition (cvpr).
  13. 13.Dai, B., Lin, D., Urtasun, R., & Fidler, S. (2017). Towards diverse and natural image descriptions via a conditional GAN. International Conference on Computer Vision (ICCV).
  14. 14.Dognin, P., Melnyk, I., Mroueh, Y., Padhi, I., Rigotti, M., Ross, J., & Schiff, Y. (2020). Alleviating noisy data in image captioning with cooperative distillation. In Vizwiz grand challenge workshop, cvpr. (https://ivc.ischool.utexas.edu/~yz9244/VizWiz workshop/poster pdf/vizwiz workshop-co-distill.pdf)
  15. 15.Dognin, P., Melnyk, I., Mroueh, Y., Ross, J., & Sercu, T. (2019). Adversarial semantic alignment for improved image captions. In Proceedings of the ieee conference on computer vision and pattern recognition (pp. 10463–10471).
  16. 16.Fisch, A., Lee, K., Chang, M.-W., Clark, J. H., & Barzilay, R. (2020). CapWAP: Captioning with a Purpose. ACL 2020 . Retrieved 2020-12-01, from http://arxiv.org/abs/2011.04264
  17. 17.Gan, Z., & Lin, K. (2020). Box Features Masking and Bayesian Self-Critical Sequence Training for VizWiz Captions Challenge. Samsung Research China-Beijing, ”https://ivc.ischool.utexas.edu/~yz9244/VizWiz workshop/videos/SRC-B VCLab.mp4”.
  18. 18.Gleason, C., Pavel, A., McCamey, E., Low, C., Carrington, P., Kitani, K. M., & Bigham, J. P. (2020). Twitter a11y: A browser extension to make twitter images accessible. In Proceedings of the 2020 chi conference on human factors in computing systems (p. 1–12). Association for Computing Machinery.
  19. 19.Grinberg, M. (2018). Flask web development: developing web applications with python. ” O’Reilly Media, Inc.”.
  20. 20.Gu, J., Lu, Z., Li, H., & Li, V. O. (2016, August). Incorporating copying mechanism in sequence-to-sequence learning. In Proceedings of the 54th annual meeting of the association for computational linguistics (volume 1: Long papers) (pp. 1631–1640). Berlin, Germany: Association for Computational Linguistics. Retrieved from https://www.aclweb.org/anthology/P16-1154 doi: 10.18653/v1/P16-1154
  21. 21.Guinness, D., Cutrell, E., & Morris, M. R. (2018). Caption crawler: Enabling reusable alternative text descriptions using reverse image search. In Proceedings of the 2018 chi conference on human factors in computing systems (p. 1–11). Association for Computing Machinery.
  22. 22.Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., . . . Bigham, J. P. (2018, June). Vizwiz grand challenge: Answering visual questions from blind people. In Proceedings of the ieee conference on computer vision and pattern recognition (cvpr).
  23. 23.Gurari, D., & Zhao, Y. (2020). Vizwiz Grand Challenge:Caption Images Taken by People Who Are Blind. ”https://ivc.ischool.utexas.edu/~yz9244/VizWiz workshop/slides/zhao-talk.pdf”.
  24. 24.Gurari, D., Zhao, Y., Zhang, M., & Bhattacharya, N. (2020). Captioning images taken by people who are blind. In A. Vedaldi, H. Bischof, T. Brox, & J.-M. Frahm (Eds.), Computer vision – eccv 2020 (pp. 417–434). Springer International Publishing.
  25. 25.He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  26. 26.Huang, L., Wang, W., Chen, J., & Wei, X.-Y. (2019). Attention on Attention for Image Captioning. arXiv preprint arXiv:1908.06954 .
  27. 27.IBM Research. (2021). Watson Text to Speech. ”https://www.ibm.com/cloud/watson-text-to-speech”. (Accessed: 2021-06-21)
  28. 28.IBM Research and CMU. (2021). NavCog AI Suite case. ”https://www.cs.cmu.edu/~NavCog/navcog.html”. (Accessed: 2021-06-21)
  29. 29.Kacorri, H., Kitani, K. M., Bigham, J. P., & Asakawa, C. (2017). People with visual impairment training personal object recognizers: Feasibility and challenges. In Proceedings of the 2017 chi conference on human factors in computing systems (pp. 5839–5849).
  30. 30.Kant, Y., Batra, D., Anderson, P., Schwing, A., Parikh, D., Lu, J., & Agrawal, H. (2020). Spatially Aware Multimodal Transformers for TextVQA. arXiv preprint arXiv:2007.12146 .
  31. 31.Karpathy, A., & Li, F.-F. (2015). Deep visual-semantic alignments for generating image descriptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  32. 32.Lin, C.-Y. (2004, July). ROUGE: A package for automatic evaluation of summaries. In Text summarization branches out: Proceedings of the acl-04 workshop (pp. 74–81). Association for Computational Linguistics.
  33. 33.Lin, T., Maire, M., Belongie, S. J., Bourdev, L. D., Girshick, R. B., Hays, J., . . . Zitnick, C. L. (2014). Microsoft COCO: Common Objects in Context. EECV .
  34. 34.Liu, S., Zhu, Z., Ye, N., Guadarrama, S., & Murphy, K. (2017). Improved image captioning via policy gradient optimization of spider. In International conference on computer vision (iccv).
  35. 35.Lu, J., Yang, J., Batra, D., & Parikh, D. (2018). Neural baby talk. In Proceedings of the ieee conference on computer vision and pattern recognition (cvpr) (p. 7219-7228). IEEE Computer Society.
  36. 36.Luo, R., Vered, G., Bracha, L., Chechik, G., & Shakhnarovich, G. (2019). Winning Google Conceptual captions challenge. https://ttic.uchicago.edu/~rluo/files/ConceptualWorkshopSlides.pdf.
  37. 37.Mahajan, D., Girshick, R., Ramanathan, V., He, K., Paluri, M., Li, Y., . . . van der Maaten, L. (2018). Exploring the limits of weakly supervised pretraining. Retrieved from http://arxiv.org/abs/1805.00932
  38. 38.Microsoft AI. (2021). Seeing AI Mobile App. ”https://www.microsoft.com/en-us/ai/seeing-ai”. (Accessed: 2021-06-21)
  39. 39.Morris, M. R., Johnson, J., Bennett, C. L., & Cutrell, E. (2018). Rich representations of visual content for screen reader users. In Proceedings of the 2018 chi conference on human factors in computing systems (p. 1–11). Association for Computing Machinery.
  40. 40.Mroueh, Y., Voinea, S., & Poggio, T. A. (2015). Learning with group invariant features: A kernel perspective. In Advances in neural information processing systems 28.
  41. 41.OrCam. (2021). OrCam Wearable Device. ”https://www.orcam.com/en/about/”. (Accessed: 2021-06-21)
  42. 42.Pan, Y., Yao, T., Li, Y., & Mei, T. (2020). X-Linear Attention Networks for Image Captioning. In Proceedings of the ieee/cvf conference on computer vision and pattern recognition (pp. 10971–10980).
  43. 43.Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). Bleu: a method for automatic evaluation of machine translation. In Proceedings of the association for computational linguistics (pp. 311–318).
  44. 44.Ranzato, M., Chopra, S., Auli, M., & Zaremba, W. (2015). Sequence level training with recurrent neural networks. ICLR.
  45. 45.Redmon, J., & Farhadi, A. (2016). YOLO9000: Better, Faster, Stronger.
  46. 46.Rennie, S. J., Marcheret, E., Mroueh, Y., Ross, J., & Goel, V. (n.d.). Self-critical sequence training for image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), year = 2017,.
  47. 47.See, A., Liu, P. J., & Manning, C. D. (2017). Get to the point: Summarization with pointer-generator networks. In Proceedings of the 55th annual meeting of the association for computational linguistics (volume 1: Long papers) (pp. 1073–1083).
  48. 48.Sharma, P., Ding, N., Goodman, S., & Soricut, R. (2018). Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of acl.
  49. 49.Shetty, R., Rohrbach, M., Hendricks, L. A., Fritz, M., & Schiele, B. (2017). Speaking the same language: Matching machine to human captions by adversarial training. arXiv preprint arXiv:1703.10476 .
  50. 50.Simons, R. N., Gurari, D., & Fleischmann, K. R. (2020, October). ”I Hope This Is Helpful”: Understanding Crowdworkers’ Challenges and Motivations for an Image Description Task. Proc. ACM Hum.-Comput. Interact., 4 (CSCW2).
  51. 51.Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., . . . Rohrbach, M. (2019, June). Towards VQA Models That Can Read. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  52. 52.Stangl, A., Morris, M. R., & Gurari, D. (2020). ”person, shoes, tree. is the person naked?” what people with vision impairments want in image descriptions. In Proceedings of the 2020 chi conference on human factors in computing systems (p. 1–13). New York, NY, USA: Association for Computing Machinery.
  53. 53.Tan, M., Pang, R., & Le, Q. V. (2020). Efficientdet: Scalable and efficient object detection.
  54. 54.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., . . . Polosukhin, I. (2017). Attention is all you need. In Advances in neural information processing systems 30 (pp. 5998–6008).
  55. 55.Vedantam, R., Lawrence Zitnick, C., & Parikh, D. (2015). Cider: Consensus-based image description evaluation. In Proceedings of the ieee conference on computer vision and pattern recognition (pp. 4566–4575).
  56. 56.Vedantam, R., Zitnick, C. L., & Parikh, D. (2015). Cider: Consensus-based image description evaluation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Retrieved from http://dblp.uni-trier.de/db/conf/cvpr/cvpr2015.html#VedantamZP15
  57. 57.Vinyals, O., Toshev, A., Bengio, S., & Erhan, D. (2015). Show and Tell: A Neural Image Caption Generator. Pattern Analysis and Machine Intelligence, IEEE Transactions on.
  58. 58.Vinyals, O., Toshev, A., Bengio, S., & Erhan, D. (2017, Apr). Show and Tell: Lessons Learned from the 2015 MSCOCO Image Captioning Challenge. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39 (4), 652–663. Retrieved from http://dx.doi.org/10.1109/TPAMI.2016.2587640 doi: 10.1109/tpami.2016.2587640
  59. 59.Williams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning, 8 (3-4), 229–256.
  60. 60.Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A. C., Salakhutdinov, R., . . . Bengio, Y. (2015). Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. In Icml.
  61. 61.Xu, K., Ba, J., Kiros, R., Courville, A., Salakhutdinov, R., Zemel, R., & Bengio, Y. (2015). Show, attend and tell: Neural image caption generation with visual attention. arXiv preprint arXiv:1502.03044 .
  62. 62.Yao, T., Pan, Y., Li, Y., & Mei, T. (2017). Incorporating copying mechanism in image captioning for learning novel objects. In Proceedings of the ieee conference on computer vision and pattern recognition (pp. 6580–6588).
  63. 63.Zeng, X., Wang, Y., Chiu, T.-Y., Bhattacharya, N., & Gurari, D. (2020, October). Vision skills needed to answer visual questions. Proc. ACM Hum.-Comput. Interact., 4 (CSCW2). Retrieved from https://doi.org/10.1145/3415220 doi: 10.1145/3415220
  64. 64.Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2019). Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675 .
  65. 65.Zhu, Y., Lu, S., Zheng, L., Guo, J., Zhang, W., Wang, J., & Yu, Y. (2018). Texygen: A benchmarking platform for text generation models. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval (pp. 1097–1100).

Citation

MLA
Dognin, P., et al. “Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge”. Journal of Artificial Intelligence Research, vol. 73, 2022, pp. 437–59, https://doi.org/10.1613/JAIR.1.13113.
APA
Dognin, P., Melnyk, I., Mroueh, Y., Padhi, I., Rigotti, M., Ross, J., Schiff, Y., Young, R. A., & Belgodere, B. (2022). Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge. Journal of Artificial Intelligence Research, 73, 437–459. https://doi.org/10.1613/JAIR.1.13113
Chicago
Dognin, P., I. Melnyk, Y. Mroueh, et al. 2022. “Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge”. Journal of Artificial Intelligence Research 73: 437–59. https://doi.org/10.1613/JAIR.1.13113.
Harvard
Dognin, P. et al. (2022) “Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge”, Journal of Artificial Intelligence Research, 73, pp. 437–459. Available at: https://doi.org/10.1613/JAIR.1.13113.
Vancouver
1. Dognin P, Melnyk I, Mroueh Y, Padhi I, Rigotti M, Ross J, Schiff Y, Young RA, Belgodere B (2022) Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge. Journal of Artificial Intelligence Research 73:437–459

BibTeX

@article{Dognin_2022, title={Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge}, volume={73}, ISSN={1076-9757}, url={http://dx.doi.org/10.1613/JAIR.1.13113}, DOI={10.1613/jair.1.13113}, journal={Journal of Artificial Intelligence Research}, publisher={AI Access Foundation}, author={Dognin, Pierre and Melnyk, Igor and Mroueh, Youssef and Padhi, Inkit and Rigotti, Mattia and Ross, Jarret and Schiff, Yair and Young, Richard A. and Belgodere, Brian}, year={2022}, month=Jan, pages={437–459} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/