Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Amazon Mechanical Turk

Amazon Mechanical Turk is a crowdsourcing marketplace operated by Amazon that connects individuals and businesses with a global, remote workforce to perform discrete tasks requiring human intelligence. Requesters post microtasks, known as Human Intelligence Tasks, which workers complete online in exchange for monetary compensation. These assignments typically involve activities that computers struggle to perform autonomously, such as data classification, image and video annotation, audio transcription, survey participation, and content moderation. Consequently, the platform is widely utilized by researchers and developers to efficiently collect and label large volumes of human-annotated data for scientific studies and machine learning model development.

5 items

Hate Speech and Counter Speech Detection: Conversational Context Does Matter

Hate Speech and Counter Speech Detection: Conversational Context Does Matter

Xinchen Yu, Eduardo Blanco, Lingzi Hong

OrganizationsArizona State UniversityUniversity of North Texas

Why you should read this

Presents a context-aware dataset of Reddit comments to demonstrate that incorporating conversational history substantially alters human annotations and significantly boosts neural network performance when detecting hate speech and counter speech.

Hate speech is plaguing the cyberspace along with user-generated content. This paper investigates the role of conversational context in the annotation and detection of online hate and counter speech, where context is defined as the preceding comment in a conversation thread. We created a context-aware dataset for a 3-way classification task on Reddit comments: hate speech, counter speech, or neutral. Our analyses indicate that context is critical to identify hate and counter speech: human judgments change for most comments depending on whether we show annotators the context. A linguistic analysis draws insights into the language people use to express hate and counter speech. Experimental results show that neural networks obtain significantly better results if context is taken into account. We also present qualitative error analyses shedding light into (a) when and why context is beneficial and (b) the remaining errors made by our best model when context is taken into account.

Added

2026-09-26

VizWiz Grand Challenge: Answering Visual Questions from Blind People

VizWiz Grand Challenge: Answering Visual Questions from Blind People

Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, Jeffrey P. Bigham

OrganizationsCarnegie Mellon UniversityUniversity of Colorado BoulderUniversity of RochesterUniversity of Science and Technology of ChinaUniversity of Texas at Austin

Why you should read this

Introduces VizWiz, a dataset of over 31,000 real-world visual questions from blind users that challenges visual question answering models to handle conversational queries, imperfect mobile photos, and unanswerable prompts in genuine assistive settings.

The study of algorithms to automatically answer visual questions currently is motivated by visual question answering (VQA) datasets constructed in artificial VQA settings. We propose VizWiz, the first goal-oriented VQA dataset arising from a natural VQA setting. VizWiz consists of over 31,000 visual questions originating from blind people who each took a picture using a mobile phone and recorded a spoken question about it, together with 10 crowdsourced answers per visual question. VizWiz differs from the many existing VQA datasets because (1) images are captured by blind photographers and so are often poor quality, (2) questions are spoken and so are more conversational, and (3) often visual questions cannot be answered. Evaluation of modern algorithms for answering visual questions and deciding if a visual question is answerable reveals that VizWiz is a challenging dataset. We introduce this dataset to encourage a larger community to develop more generalized algorithms that can assist blind people.

Added

2026-09-25

Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding

Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding

Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, Abhinav Gupta

OrganizationsAllen Institute for AICarnegie Mellon UniversityINRIAUniversity of Washington

Why you should read this

Introduces a crowdsourced video collection methodology that captures realistic home activities across hundreds of participants, establishing the Charades dataset to advance human action recognition and video description generation.

Computer vision has a great potential to help our daily lives by searching for lost keys, watering flowers or reminding us to take a pill. To succeed with such tasks, computer vision methods need to be trained from real and diverse examples of our daily dynamic scenes. While most of such scenes are not particularly exciting, they typically do not appear on YouTube, in movies or TV broadcasts. So how do we collect sufficiently many diverse but boring samples representing our lives? We propose a novel Hollywood in Homes approach to collect such data. Instead of shooting videos in the lab, we ensure diversity by distributing and crowdsourcing the whole process of video creation from script writing to video recording and annotation. Following this procedure we collect a new dataset, Charades, with hundreds of people recording videos in their own homes, acting out casual everyday activities. The dataset is composed of 9,848 annotated videos with an average length of 30 seconds, showing activities of 267 people from three continents. Each video is annotated by multiple free-text descriptions, action labels, action intervals and classes of interacted objects. In total, Charades provides 27,847 video descriptions, 66,500 temporally localized intervals for 157 action classes and 41,104 labels for 46 object classes. Using this rich data, we evaluate and provide baseline results for several tasks including action recognition and automatic description generation. We believe that the realism, diversity, and casual nature of this dataset will present unique challenges and new opportunities for computer vision community.

Added

2026-09-25

LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop

LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop

Fisher Yu, Ari Seff, Yinda Zhang, Shuran Song, Thomas Funkhouser, Jianxiong Xiao

OrganizationsPrinceton University

Why you should read this

Presents an efficient human-in-the-loop labeling framework that scales dataset construction using deep model predictions, producing the massive LSUN visual recognition benchmark spanning millions of scene and object images.

While there has been remarkable progress in the performance of visual recognition algorithms, the state-of-the-art models tend to be exceptionally data-hungry. Large labeled training datasets, expensive and tedious to produce, are required to optimize millions of parameters in deep network models. Lagging behind the growth in model capacity, the available datasets are quickly becoming outdated in terms of size and density. To circumvent this bottleneck, we propose to amplify human effort through a partially automated labeling scheme, leveraging deep learning with humans in the loop. Starting from a large set of candidate images for each category, we iteratively sample a subset, ask people to label them, classify the others with a trained model, split the set into positives, negatives, and unlabeled based on the classification confidence, and then iterate with the unlabeled set. To assess the effectiveness of this cascading procedure and enable further progress in visual recognition research, we construct a new image dataset, LSUN. It contains around one million labeled images for each of 10 scene categories and 20 object categories. We experiment with training popular convolutional networks and find that they achieve substantial performance gains when trained on this dataset.

Added

2026-09-14