VizWiz Grand Challenge: Answering Visual Questions from Blind People

Danna GurariQing LiAbigale J. StanglAnhong GuoChi LinKristen GraumanJiebo LuoJeffrey P. Bigham

article2018CVPR1,315 citations

Introduces VizWiz, a dataset of over 31,000 real-world visual questions from blind users that challenges visual question answering models to handle conversational queries, imperfect mobile photos, and unanswerable prompts in genuine assistive settings.

Listen

Visual question answering technology holds significant promise for assisting blind individuals with daily tasks, such as reading packaging or identifying personal items. However, existing automated systems have predominantly been developed using synthetic, artificially curated datasets that fail to reflect authentic, real-world user needs. Human-powered assistance services bridge this gap today, but they introduce major concerns regarding high operational costs, multi-minute response latencies, poor scalability, and privacy risks when sensitive information is shared.

The article introduces and evaluates VizWiz, the first goal-oriented visual question answering dataset collected directly from blind users in everyday settings. The primary objective is to evaluate how well state-of-the-art vision and language algorithms answer natural visual questions and determine whether a given visual question can be answered at all.

To construct the dataset, the researchers sourced over 72,000 authentic visual questions captured via mobile devices and implemented a rigorous two-stage anonymization and filtering protocol. This process eliminated roughly 31% of the candidate data to protect user privacy and remove personally identifiable information, resulting in a finalized benchmark of 31,173 visual questions with ten crowdsourced answers each. The researchers then benchmarked multiple leading computer vision algorithms against this dataset, testing models pre-trained on standard benchmarks, fine-tuned models, and models trained entirely from scratch.

The analysis yielded several critical findings regarding real-world data and algorithmic capabilities. First, existing top-performing models failed abruptly on authentic data, achieving only around 14% accuracy due to significant vocabulary mismatches; only 824 of the top 3,000 answers in VizWiz overlap with standard benchmark answer sets. Second, fine-tuning or training models directly on the dataset improved accuracy to roughly 47%, yet a substantial performance gap remains when compared to the 75% human accuracy baseline. Third, 28.6% of real-world visual questions are completely unanswerable due to poor focus, bad lighting, or framing errors. Finally, algorithms designed to predict whether a question is answerable performed best when combining image and question data, achieving an average precision of 71.7% compared to the 30.6% baseline of prior caption-matching approaches.

These findings demonstrate that automated tools trained purely on standard benchmark images are not viable for deployment in real assistive applications without substantial domain-specific adaptation. Real-world systems must handle conversational, spoken queries and heavily degraded images while reliably detecting unanswerable queries to prevent erroneous answers that could compromise user safety.

Organizations developing assistive computer vision systems should integrate answerability detection mechanisms into their processing pipelines and prioritize fine-tuning models on authentic, domain-specific data. Systems should also provide immediate feedback to users when an image is unanswerable due to poor lighting or blur, prompting a clearer recapture before attempting an automated answer.

While the findings demonstrate high confidence based on extensive crowdsourced verification and standard machine learning metrics, certain limitations remain. The stringent privacy-filtering process deliberately removed many complex scenes and low-quality images containing sensitive text, leaving some gap between the public benchmark and the full spectrum of private, in-the-wild user interactions.

arXiv: 1802.08218
Cover for VizWiz Grand Challenge: Answering Visual Questions from Blind People

Abstract

The study of algorithms to automatically answer visual questions currently is motivated by visual question answering (VQA) datasets constructed in artificial VQA settings. We propose VizWiz, the first goal-oriented VQA dataset arising from a natural VQA setting. VizWiz consists of over 31,000 visual questions originating from blind people who each took a picture using a mobile phone and recorded a spoken question about it, together with 10 crowdsourced answers per visual question. VizWiz differs from the many existing VQA datasets because (1) images are captured by blind photographers and so are often poor quality, (2) questions are spoken and so are more conversational, and (3) often visual questions cannot be answered. Evaluation of modern algorithms for answering visual questions and deciding if a visual question is answerable reveals that VizWiz is a challenging dataset. We introduce this dataset to encourage a larger community to develop more generalized algorithms that can assist blind people.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 VizWiz: Dataset Creation
  • 3.1 Visual Question Collection Analysis
  • 3.2 Anonymizing and Filtering Visual Questions
  • 3.3 Collecting Answers
  • 4 VizWiz: Dataset Analysis
  • 4.1 Analysis of Questions
  • 4.2 Analysis of Images
  • 4.3 Analysis of Answers
  • 5 VizWiz Benchmarking
  • 5.1 Visual Question Answering
  • 5.2 Visual Question Answerability
  • 6 Conclusions
  • References

Knowls

  1. Knowl 1 — VizWiz Dataset Specification

    definition

    The VizWiz dataset is a goal-oriented visual question answering (VQA) dataset consisting of 31,17331,173 visual questions collected from real blind users in natural settings, paired with 10 crowdsourced answers per visual question.

    Each visual question originated from a user taking a photograph with a mobile phone and recording a spoken natural language question about the captured image. The dataset is partitioned into:

    • Training set: 20,00020,000 visual questions (~64.2%)
    • Validation set: 3,1733,173 visual questions (~10.2%)
    • Test set: 8,0008,000 visual questions (~25.7%)

    VizWiz differs from existing VQA datasets across three key dimensions:

    1. Image Quality: The images are taken by visually impaired photographers, frequently resulting in poor lighting, improper focus, extreme blur, and off-center or missing subjects of interest.
    2. Question Modality: Questions were spoken rather than typed, introducing conversational phrasing (e.g., greetings, polite requests), background noise, and audio recording artifacts (such as cut-off sentence starts or trailing speech).
    3. Visual Answerability: Because askers cannot verify the quality or framing of their photos, 28.63%28.63\% of the visual questions cannot be answered from the provided image due to severe image degradation or missing visual content.
  2. Knowl 2 — Privacy and Safety Filtering Protocol for Natural VQA Data

    model/method

    Because visual questions gathered in the wild from blind users can reveal sensitive information, a multi-stage anonymization and filtering protocol was developed to protect user privacy before dataset release.

    Anonymization

    • Voice Removal: Spoken audio recordings were transcribed into text by crowd workers, followed by spell-checking.
    • Metadata Stripping: All images were re-encoded using lossless compression to remove embedded EXIF and geolocation metadata.

    Vulnerability Taxonomy and Two-Stage Filtering

    Candidate questions (48,16948,169 total) were evaluated against five vulnerability categories:

    1. Personally-Identifying Information (PII): Human faces, full names, home addresses, bank account numbers, credit card numbers, or medical prescriptions.
    2. Location: Identifiable addressed mail or identifiable local business details.
    3. Indecent Content: Nudity or profanity.
    4. Suspicious Complex Scenes: Scenes where PII might be present but cannot be definitively ruled out.
    5. Suspicious Low-Quality Images: Blurry or over/underexposed images that could reveal PII under digital image enhancement.

    Filtering was executed in two successive human review stages:

    • Round 1 (Crowd Review): Amazon Mechanical Turk workers flagged and removed 4,6264,626 visual questions containing PII.
    • Round 2 (Domain Expert Review): In-house experts reviewed the remaining dataset, removing an additional 2,6932,693 questions (895895 for PII, 377377 for location, 5555 for indecent content, 725725 for complex scenes, 578578 for low-quality scenes, and 6363 for other concerns). An additional 7,4777,477 instances with missing questions (fewer than two words) were removed.

    In total, 14,79614,796 visual questions (~30.7%30.7\%) were filtered out, leaving the final set of 31,17331,173 visual questions.

  3. Knowl 3 — Crowdsourced Answer Collection and Normalization Protocol

    model/method

    To collect ground-truth answers for algorithm training and evaluation on VizWiz, a crowdsourcing and normalization pipeline was established:

    Annotation Interface and Rules

    • 10 independent answers were collected per visual question on Amazon Mechanical Turk from US workers who had completed at least 500 tasks with a ≥95%\ge 95\% approval rating and possessed an Adult Qualification.
    • Workers were instructed to return brief phrases rather than full sentences and were informed that images were captured by blind individuals.
    • Two explicit unanswerable tags were provided to allow fine-grained categorization:
      • "Unsuitable Image": Used when an image is too degraded in quality to answer the question (e.g., entirely white, entirely black, or heavily blurred).
      • "Unanswerable": Used when the image quality is acceptable but the question cannot be answered from the visible content.

    Text Normalization

    Answers were standardized by converting all text to lowercase, mapping numerical words to digits, stripping punctuation and articles ("a", "an", "the"), and removing filler words ("it", "is", "its", "and", "&", "with", "there", "are", "of", "or"). Spelling errors were corrected using an automated ensemble of the Enchant library and Norvig's frequency-based spell checker, with manual verification to preserve genuine non-standard text such as CAPTCHA strings.

  4. Knowl 4 — Linguistic, Visual, and Answerability Characteristics of VizWiz

    empirical result

    Analysis of the 31,17331,173 visual questions in VizWiz reveals distinct distributions compared to synthetic and web-curated VQA datasets:

    Linguistic Characteristics of Questions

    • Length: The median question length is 5 words (mean: 6.68 words; 25th percentile: 4 words; 75th percentile: 7 words).
    • Starter Diversity: 27.88%27.88\% of questions start with a rare first word (defined as occurring in <5%<5\% of all questions), compared to 13.4%13.4\% in VQA 2.0. This is driven by conversational starters (e.g., "Hi", "Please", "Can you") and speech-recording truncations.
    • Frequent Inquiries: The single most common question is "What is this?". Questions predominantly begin with "What", which broadens the space of plausible answers compared to categorical "Is..." or "How many..." formulations.

    Visual Characteristics

    • The pixel-wise average across all images in VizWiz yields a uniform gray image with no salient spatial structure, confirming the absence of photographer framing bias (e.g., center bias).
    • 28.0%28.0\% of all images were marked as "Unsuitable Image" by at least two independent crowd workers.

    Answer Distributions and Human Agreement

    • Vocabulary and Overlap: VizWiz contains 58,78958,789 unique answers. Only 824 of the top 3,000 most frequent answers in VizWiz overlap with the top 3,000 answers of VQA 2.0.
    • Answer Length: Answers have a median length of 1.0 word and a mean of 1.66 words (67.32%67.32\% 1 word, 20.74%20.74\% 2 words, 8.24%8.24\% 3 words, 3.52%3.52\% 4 words, and 0.01%>40.01\% >4 words).
    • Unanswerability Rate: 28.63%28.63\% of visual questions are unanswerable (defined as at least 5 out of 10 workers submitting "Unanswerable" or "Unsuitable Image").
    • Human Consensus: Using exact string matching, 97.7%97.7\% of visual questions have agreement among independent annotators: 72.83%72.83\% have >3>3 workers agreeing on the most common answer, 15.5%15.5\% have exactly 3 agreeing, and 9.67%9.67\% have exactly 2 agreeing.
  5. Knowl 5 — Performance of Visual Question Answering Models on VizWiz

    data/table

    Nine VQA models were benchmarked on the VizWiz test set (8,0008,000 visual questions) across four evaluation metrics: VQA Accuracy, CIDEr, BLEU-4, and METEOR. The architectures include question-image feature fusion without attention (Q+I), spatial image attention (Q+I+A), and bottom-up and top-down attention (Q+I+BUA). Models were evaluated under three regimes: pre-trained on VQA 2.0 as-is, fine-tuned on VizWiz (FT), and trained from scratch on VizWiz (VizWiz).

    Method Acc CIDEr BLEU4 METEOR
    Q+I (as-is) 0.137 0.224 0.000 0.078
    Q+I+A (as-is) 0.145 0.237 0.000 0.082
    Q+I+BUA (as-is) 0.134 0.226 0.000 0.077
    FT [Q+I] 0.466 0.675 0.314 0.297
    FT [Q+I+A] 0.469 0.691 0.351 0.299
    FT [Q+I+BUA] 0.475 0.713 0.359 0.309
    VizWiz [Q+I] 0.465 0.654 0.353 0.298
    VizWiz [Q+I+A] 0.469 0.661 0.356 0.302
    VizWiz [Q+I+BUA] 0.469 0.675 0.396 0.306

    Models pre-trained on VQA 2.0 perform poorly on VizWiz (extaccuracy≈0.134–0.145 ext{accuracy} \approx 0.134\text{--}0.145) due to severe answer distribution shift (only 824 of the top 3,000 VizWiz answers exist in VQA 2.0). Fine-tuning or training directly on VizWiz improves accuracy to 0.465–0.4750.465\text{--}0.475, but all models remain far below human performance (0.7500.750 accuracy). Attention mechanisms provide only marginal improvements because blind user images frequently contain few distinct objects or exhibit severe blur.

  6. Knowl 6 — VQA Accuracy Breakdown by Answer Type and Cross-Dataset Generalization

    data/table

    Evaluating VQA models across specific answer types on the VizWiz test set reveals where performance gains and failures occur. The four answer categories in VizWiz are Yes/No (4.80%4.80\%), Number (1.69%1.69\%), Unanswerable (34.60%34.60\%), and Other (58.91%58.91\%).

    Method Yes/No Number Unanswerable Other
    Q+I (as-is) 0.598 0.045 0.070 0.142
    Q+I+A (as-is) 0.605 0.068 0.071 0.155
    Q+I+BUA (as-is) 0.582 0.071 0.060 0.143
    FT [Q+I] 0.675 0.220 0.781 0.275
    FT [Q+I+A] 0.681 0.213 0.770 0.287
    FT [Q+I+BUA] 0.669 0.220 0.776 0.294
    VizWiz [Q+I] 0.597 0.262 0.805 0.264
    VizWiz [Q+I+A] 0.608 0.218 0.802 0.274
    VizWiz [Q+I+BUA] 0.596 0.210 0.805 0.273

    Training on VizWiz provides the largest gains on Unanswerable visual questions (increasing from ≈0.060\approx 0.060 to 0.8050.805) and moderate gains on Yes/No questions (reaching 0.6810.681). Gains remain low on Number (max 0.2620.262) and Other (max 0.2940.294), where models struggle with text-reading (e.g., CAPTCHAs, expiration dates, cooking directions) and fine-grained visual description.

    Cross-Dataset Performance on VQA 2.0

    When models trained on VizWiz are tested on the VQA 2.0 test set, performance drops significantly, demonstrating a pronounced domain shift:

    Model All Yes/No Number Other
    FT [Q+I] 0.300 0.612 0.094 0.079
    FT [Q+I+A] 0.318 0.601 0.163 0.110
    FT [Q+I+BUA] 0.304 0.595 0.082 0.105
    VizWiz [Q+I] 0.218 0.461 0.074 0.042
    VizWiz [Q+I+A] 0.228 0.465 0.131 0.049
    VizWiz [Q+I+BUA] 0.219 0.453 0.083 0.048
  7. Knowl 7 — Visual Question Answerability Prediction Benchmarking

    data/table

    Visual question answerability prediction evaluates the ability of a model to classify whether a given visual question cannot be answered from the provided image. Eight baseline approaches and feature ablations were evaluated on the VizWiz test set using Average Precision (AP) and average F1 score:

    • Q+C [30]: Pre-trained question-relevance model comparing an LSTM question encoding to an LSTM caption encoding generated via NeuralTalk2 on MSCOCO, as-is, fine-tuned (FT), and trained from scratch (VizWiz).
    • VQA [18]: Uses the output probability of the "unanswerable" class from the top-performing VQA classifier.
    • Q: A one-layer LSTM question encoder fed into a softmax classifier.
    • C: A one-layer LSTM image caption encoder fed into a softmax classifier.
    • I: ResNet-152 pool5 image features fed into a softmax classifier.
    • Q+I: Concatenation of LSTM question features and ResNet-152 image features fed into a softmax classifier.
    Model Average Precision (AP) Average F1 Score
    Q+C [30] (as-is) 0.306 0.383
    FT [30] 0.561 0.542
    VizWiz [30] 0.605 0.549
    VQA [18] 0.560 0.569
    Q (Question only) 0.490 0.233
    C (Caption only) 0.464 0.270
    I (Image only) 0.640 0.518
    Q+I (Question + Image) 0.717 0.648

    The image modality alone (I, extAP=0.640 ext{AP} = 0.640) provides the strongest single predictive signal, substantially outperforming question-only (Q, extAP=0.490 ext{AP} = 0.490) and caption-only (C, extAP=0.464 ext{AP} = 0.464) features. This indicates that low image quality (e.g., severe blur, extreme lighting, occluding fingers) is the dominant cause of unanswerability in VizWiz. Combining visual and question representations (Q+I) achieves the highest performance (extAP=0.717 ext{AP} = 0.717, extF1=0.648 ext{F1} = 0.648).

Coverage note — None was omitted; all contributed dataset construction details, privacy filtering protocols, empirical dataset statistics, VQA benchmark results, answer-type breakdowns, cross-dataset transfer tests, and answerability prediction benchmarks are covered.

References

  1. 1.Be my eyes. http://www.bemyeyes.org/. 2
  2. 2.http://camfindapp.com/. 1
  3. 3.http://www.taptapseeapp.com/. 1, 3
  4. 4.D. Adams, L. Morales, and S. Kurniawan. A qualitative study to support a blind photography mobile application. In International Conference on PErvasive Technologies Related to Assistive Environments, page 25. ACM, 2013. 1
  5. 5.T. Ahmed, R. Hoyle, K. Connelly, D. Crandall, and A. Kapadia. Privacy concerns and behaviors of people with visual impairments. In ACM Conference on Human Factors in Computing Systems, pages 3523–3532. ACM Conference on Human Factors in Computing Systems (CHI), 2015. 1, 4
  6. 6.P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang. Bottom-up and top-down attention for image captioning and visual question answering. arXiv preprint arXiv:1707.07998, 2017. 7, 8, 12, 14
  7. 7.J. Andreas, M. Rohrbach, T. Darrell, and D. Klein. Neural module networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 39–48, 2016. 1, 2, 3, 4
  8. 8.S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh. VQA: Visual Question Answering. In IEEE International Conference on Computer Vision (ICCV), pages 2425–2433, 2015. 1, 2, 3, 4, 5, 6, 7, 9, 11
  9. 9.J. P. Bigham, C. Jayant, H. Ji, G. Little, A. Miller, R. C. Miller, R. Miller, A. Tatarowicz, B. White, S. White, and T. Yeh. Vizwiz: Nearly real-time answers to visual questions. In ACM symposium on User interface software and technology (UIST), pages 333–342, 2010. 1, 2, 3
  10. 10.J. P. Bigham, C. Jayant, A. Miller, B. White, and T. Yeh. Vizwiz:: Locateit-enabling blind people to locate objects in their environment. In Computer Vision and Pattern Recognition Workshops (CVPRW), 2010 IEEE Computer Society Conference on, pages 65–72. IEEE, 2010. 2
  11. 11.E. Brady, M. R. Morris, Y. Zhong, S. White, and J. P. Bigham. Visual challenges in the everyday lives of blind people. In ACM Conference on Human Factors in Computing Systems (CHI), pages 2117–2126, 2013. 1, 3
  12. 12.M. A. Burton, E. Brady, R. Brewer, C. Neylan, J. P. Bigham, and A. Hurst. Crowdsourcing subjective fashion advice using VizWiz: Challenges and opportunities. In ACM SIGACCESS conference on Computers and accessibility (ASSETS), pages 135–142, 2012. 1, 2
  13. 13.X. Chen, H. Fang, T. Lin, R. Vedantam, S. K. Gupta, P. Dollár, and C. L. Zitnick. Microsoft coco captions: Data collection and evaluation server. CoRR, abs/1504.00325, 2015. 1, 2, 7
  14. 14.A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. F. Moura, D. Parikh, and D. Batra. Visual dialog. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 6, 7
  15. 15.J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei. ImageNet: A large-scale hierarchical image database. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 248–255. IEEE, 2009. 4, 6
  16. 16.D. Elliott and F. Keller. Image description using visual dependency representations. In Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1292–1302, 2013. 7
  17. 17.H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu. Are you talking to a machine? dataset and methods for multilingual image question answering. In arXiv preprint arXiv:1505.05612, 2015. 1, 3, 4
  18. 18.Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. arXiv preprint arXiv:1612.00837, 2016. 1, 3, 4, 7, 8, 12, 14
  19. 19.S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning, pages 448–456, 2015. 12
  20. 20.C. Jayant, H. Ji, S. White, and J. P. Bigham. Supporting blind photography. In ASSETS, 2011. 3
  21. 21.J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. L. Zitnick, and R. Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. arXiv preprint arXiv:1612.06890, 2016. 1, 2, 3, 4
  22. 22.K. Kafle and C. Kanan. An analysis of visual question answering algorithms. arXiv preprint arXiv:1703.09684, 2017. 1, 3, 4, 6, 7
  23. 23.A. Karpathy and L. Fei-Fei. Deep visual-semantic alignments for generating image descriptions. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3128–3137, 2015. 8
  24. 24.V. Kazemi and A. Elqursh. Show, ask, attend, and answer: A strong baseline for visual question answering. arXiv preprint arXiv:1704.03162, 2017. 1, 7, 8, 12, 14
  25. 25.D. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 12
  26. 26.R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L. Li, D. A. Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International Journal of Computer Vision, 123(1):32–73, 2017. 1, 3, 4
  27. 27.W. S. Lasecki, P. Thiha, Y. Zhong, E. Brady, and J. P. Bigham. Answering visual questions with conversational crowd assistants. In ACM SIGACCESS Conference on Computers and Accessibility (ASSETS), number 18, pages 1– 8, 2013. 1, 2
  28. 28.T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. L. Zitnick. Microsoft COCO: Common objects in context. In IEEE European Conference on Computer Vision (ECCV), pages 740–755, 2014. 1, 2, 3, 4, 8
  29. 29.H. MacLeod, C. L. Bennett, M. R. Morris, and E. Cutrell. Understanding blind people’s experiences with computergenerated captions of social media images. In ACM Conference on Human Factors in Computing Systems (CHI), pages 5988–5999. ACM, 2017. 1
  30. 30.A. Mahendru, V. Prabhu, A. Mohapatra, D. Batra, and S. Lee. The promise of premise: Harnessing question premises in visual question answering. arXiv preprint arXiv:1705.00601, 2017. 1, 3, 6, 8, 14
  31. 31.M. Malinowski and M. Fritz. A multi-world approach to question answering about real-world scenes based on uncertain input. In Advances in Neural Information Processing Systems (NIPS), pages 1682–1690, 2014. 1, 3, 4
  32. 32.K. Papineni, S. Roukos, T. Ward, and W. Zhu. BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics (ACL), pages 311–318. Association for Computational Linguistics, 2002. 7
  33. 33.G. Patterson and J. Hays. Sun attribute database: Discovering, annotating, and recognizing scene attributes. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2751–2758. IEEE, 2012. 1, 2
  34. 34.A. Ray, G. Christie, M. Bansal, D. Batra, and D. Parikh. Question relevance in VQA: Identifying non-visual and false-premise questions. arXiv preprint arXiv:1606.06622, 2016. 3, 6
  35. 35.M. Ren, R. Kiros, and R. S. Zemel. Exploring models and data for image question answering. In Advances in Neural Information Processing Systems (NIPS), pages 2935–2943, 2015. 1, 3, 4
  36. 36.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al. Imagenet large scale visual recognition challenge. International Journal on Computer Vision (IJCV), 115(3):211–252, 2015. 1, 2
  37. 37.N. Silberman, D. Hoiem, P. Kohli, and R. Fergus. Indoor segmentation and support inference from RGBD images. Computer Vision–ECCV 2012, pages 746–760, 2012. 4
  38. 38.N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. Journal of machine learning research, 15(1):1929–1958, 2014. 12
  39. 39.B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L. Li. YFCC100M: The new data in multimedia research. Communications of the ACM, 59(2):64–73, 2016. 4
  40. 40.A. S. Toor, H. Wechsler, and M. Nappi. Question part relevance and editing for cooperative and context-aware vqa (c2vqa). In International Workshop on Content-Based Multimedia Indexing, page 4. ACM, 2017. 3, 6
  41. 41.M. Vazquez and A. Steinfeld. An assisted photography framework to help visually impaired users properly aim a camera. In ACM Transactions on Computer-Human Interaction (TOCHI), volume 21, page 25, 2014. 3
  42. 42.R. Vedantam, L. C. Zitnick, and D. Parikh. CIDEr: Consensus-based image description evaluation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4566–4575, 2015. 7
  43. 43.P. Wang, Q. Wu, C. Shen, A. Dick, and A. Hengel. FVQA: fact-based visual question answering. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017. 1, 3, 4
  44. 44.P. Wang, Q. Wu, C. Shen, A. Hengel, and A. Dick. Explicit knowledge-based reasoning for visual question answering. arXiv preprint arXiv:1511.02570, 2015. 1, 3, 4
  45. 45.S. Wu, J. Wieland, O. Farivar, and J. Schiller. Automatic alttext: Computer-generated image descriptions for blind users on a social network service. In CSCW, pages 1180–1992, 2017. 1
  46. 46.J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba. Sun database: Large-scale scene recognition from abbey to zoo. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 3485–3492. IEEE, 2010. 1, 2
  47. 47.L. Yu, E. Park, A. C. Berg, and T. L. Berg. Visual madlibs: Fill in the blank image generation and question answering. In IEEE International Conference on Computer Vision (ICCV), pages 2461–2469, 2015. 1, 3, 4
  48. 48.Y. Zhong, P. J. Garrigues, and J. P. Bigham. Real time object scanning using a mobile phone and cloud-based visual search engine. In SIGACCESS Conference on Computers and Accessibility, page 20, 2013. 3
  49. 49.Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei. Visual7W: Grounded question answering in images. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 4995–5004, 2016. 1, 3, 4

Citation

MLA
Gurari, D., et al. “VizWiz Grand Challenge: Answering Visual Questions from Blind People”. arXiv, 2018, http://arxiv.org/abs/1802.08218v4.
APA
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., & Bigham, J. P. (2018). VizWiz Grand Challenge: Answering Visual Questions from Blind People. arXiv. http://arxiv.org/abs/1802.08218v4
Chicago
Gurari, D., Q. Li, A. J. Stangl, et al. 2018. “VizWiz Grand Challenge: Answering Visual Questions from Blind People”. arXiv. http://arxiv.org/abs/1802.08218v4.
Harvard
Gurari, D. et al. (2018) “VizWiz Grand Challenge: Answering Visual Questions from Blind People”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1802.08218v4.
Vancouver
1. Gurari D, Li Q, Stangl AJ, Guo A, Lin C, Grauman K, Luo J, Bigham JP (2018) VizWiz Grand Challenge: Answering Visual Questions from Blind People. arXiv

BibTeX

@article{gurari2018vizwiz,
  title = {VizWiz Grand Challenge: Answering Visual Questions from Blind People},
  author = {Gurari, Danna and Li, Qing and Stangl, Abigale J. and Guo, Anhong and Lin, Chi and Grauman, Kristen and Luo, Jiebo and Bigham, Jeffrey P.},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1802.08218v4},
  eprint = {1802.08218}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE