Built independently by an author, for readers. Read the story and support ChapterPal

keyword

image descriptors

Image descriptors are structured representations that encode the visual characteristics, features, or semantic content of an image or image region into a format suitable for computational processing and analysis. These descriptors can be global, summarizing the overall properties of an entire image, or local, capturing specific interest points, textures, shapes, or patches. They are generated using traditional handcrafted algorithms or learned through deep neural networks and vision models, producing compact numerical vectors, embeddings, or semantic attributes. By transforming raw pixel data into standardized and informative abstractions, image descriptors enable computer vision systems to efficiently perform tasks such as image retrieval, feature matching, visual geo-localization, object recognition, and multimodal reasoning.

3 items

Deep Visual Geo-localization Benchmark

Deep Visual Geo-localization Benchmark

Gabriele Moreno Berton, Riccardo Mereu, Gabriele Trivigno, Carlo Masone, Gabriela Csurka, Torsten Sattler, Barbara Caputo

OrganizationsConsorzio Interuniversitario Nazionale per l'InformaticaCzech Technical University in PragueNaver Labs EuropePolitecnico di Torino

Why you should read this

Establishes a modular visual geo-localization benchmarking framework and conducts extensive evaluations across pipeline components, offering practical guidelines that balance retrieval accuracy with real-world computational and memory costs.

In this paper, we propose a new open-source benchmark-ing framework for Visual Geo-localization (VG) that allows to build, train, and test a wide range of commonly used ar-chitectures, with the flexibility to change individual compo-nents of a geo-localization pipeline. The purpose of this framework is twofold: i) gaining insights into how differ-ent components and design choices in a VG pipeline im-pact the final results, both in terms of performance (re-call@N metric) and system requirements (such as execu-tion time and memory consumption); ii) establish a system-atic evaluation protocol for comparing different methods. Using the proposed framework, we perform a large suite of experiments which provide criteria for choosing back-bone, aggregation and negative mining depending on the use-case and requirements. We also assess the impact of engineering techniques like pre/post-processing, data aug-mentation and image resizing, showing that better perfor-mance can be obtained through somewhat simple proce-dures: for example, downscaling the images' resolution to 80% can lead to similar results with a 36% savings in ex-traction time and dataset storage requirement. Code and trained models are available at dataset storage require-ment. https://deep-vg-bench.herokuapp.com/.

Added

2026-09-26

Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners

Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners

Zhenhailong Wang, Manling Li, Ruochen Xu, Luowei Zhou, Jie Lei, Xudong Lin, Shuohang Wang, Ziyi Yang, Chenguang Zhu, Derek Hoiem, Shih-Fu Chang, Mohit Bansal, Heng Ji

OrganizationsColumbia UniversityGoogleMicrosoftUniversity of Illinois Urbana-ChampaignUniversity of North Carolina at Chapel Hill

Why you should read this

Proposes VidIL, a framework that decomposes video content into multi-level textual descriptions via frozen image-language models, enabling large language models to perform generative video tasks with few-shot prompting without requiring any video pretraining or finetuning.

The goal of this work is to build flexible video-language models that can generalize to various video-to-text tasks from few examples. Existing few-shot video-language learners focus exclusively on the encoder, resulting in the absence of a video-to-text decoder to handle generative tasks. Video captioners have been pretrained on large-scale video-language datasets, but they rely heavily on finetuning and lack the ability to generate text for unseen tasks in a few-shot setting. We propose VidIL, a few-shot Video-language Learner via Image and Language models, which demonstrates strong performance on few-shot video-to-text tasks without the necessity of pretraining or finetuning on any video datasets. We use image-language models to translate the video content into frame captions, object, attribute, and event phrases, and compose them into a temporal-aware template. We then instruct a language model, with a prompt containing a few in-context examples, to generate a target output from the composed content. The flexibility of prompting allows the model to capture any form of text input, such as automatic speech recognition (ASR) transcripts. Our experiments demonstrate the power of language models in understanding videos on a wide variety of video-language tasks, including video captioning, video question answering, video caption retrieval, and video future event prediction. Especially, on video future event prediction, our few-shot model significantly outperforms state-of-the-art supervised models trained on large-scale video datasets. Code and processed data are publicly available for research purposes at https://github.com/MikeWangWZHL/VidIL.

Added

2026-09-26