keyword
situation recognition
Situation recognition is a computer vision and multimodal artificial intelligence task that involves generating a structured semantic summary of an activity depicted in an image or video. Rather than merely detecting isolated objects or classifying an overall scene, situation recognition identifies the primary action or event and determines the specific semantic roles played by participating entities, such as the agent performing the action, the object being acted upon, the instrument used, and the location. Drawing upon the principles of semantic role labeling from linguistics, this framework organizes visual understanding into a standardized representation of who is doing what, with what, and where, often extending to grounding each participating entity to its corresponding visual region.
2 items

VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
Letitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank, Iacer Calixto, Albert Gatt
Why you should read this
Proposes a diagnostic benchmark designed to evaluate how effectively pretrained vision-and-language models visually ground specific linguistic constructs, revealing critical failure modes across state-of-the-art architectures when tested against carefully controlled adversarial foils.
We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for their visio-linguistic grounding capabilities on specific linguistic phenomena. VALSE offers a suite of six tests covering various linguistic constructs. Solving these requires models to ground linguistic phenomena in the visual modality, allowing more fine-grained evaluations than hitherto possible. We build VALSE using methods that support the construction of valid foils, and report results from evaluating five widely-used V&L models. Our experiments suggest that current models have considerable difficulty addressing most phenomena. Hence, we expect VALSE to serve as an important benchmark to measure future progress of pretrained V&L models from a linguistic perspective, complementing the canonical task-centred V&L evaluations.
Added
2026-10-01

Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, Aniruddha Kembhavi
Why you should read this
Presents Unified-IO, a groundbreaking model that unifies over 90 diverse vision, language, and multi-modal tasks into a single transformer-based architecture by homogenizing all inputs and outputs into discrete vocabulary tokens, achieving strong results across 16 benchmarks without task-specific fine-tuning.
We propose Unified-IO, a model that performs a large variety of AI tasks spanning classical computer vision tasks, including pose estimation, object detection, depth estimation and image generation, vision-and-language tasks such as region captioning and referring expression, to natural language processing tasks such as question answering and paraphrasing. Developing a single unified model for such a large variety of tasks poses unique challenges due to the heterogeneous inputs and outputs pertaining to each task, including RGB images, per-pixel maps, binary masks, bounding boxes, and language. We achieve this unification by homogenizing every supported input and output into a sequence of discrete vocabulary tokens. This common representation across all tasks allows us to train a single transformer-based architecture, jointly on over 90 diverse datasets in the vision and language fields. Unified-IO is the first model capable of performing all 7 tasks on the GRIT benchmark and produces strong results across 16 diverse benchmarks like NYUv2-Depth, ImageNet, VQA2.0, OK-VQA, Swig, VizWizGround, BoolQ, and SciTail, with no task-specific fine-tuning. Code and demos for Unified-IO are available at: this https URL.
Added
2026-01-28

