keyword
attribute prediction
Attribute prediction is a computer vision and machine learning task that involves automatically identifying and estimating the presence or degree of descriptive properties, semantic traits, or visual characteristics associated with an input entity, such as an image or object. Unlike standard categorical classification, which maps an input to a single high-level identity label, attribute prediction identifies multiple mid-level properties—such as color, texture, material, shape, or parts—that can span across diverse object classes. By decoupling visual descriptions from rigid category boundaries, this approach facilitates a more detailed understanding of entities and supports downstream applications such as zero-shot learning, fine-grained object recognition, image retrieval, and cross-category knowledge transfer.
4 items

An embarrassingly simple approach to zero-shot learning
Bernardino Romera-Paredes, Philip H. S. Torr
Why you should read this
Proposes a two-layer linear framework for zero-shot learning that is implementable in a single line of code, establishes theoretical generalization bounds through domain adaptation, and surpasses complex state-of-the-art methods across standard benchmarks by up to 17%.
Zero-shot learning consists in learning how to recognise new concepts by just having a description of them. Many sophisticated approaches have been proposed to address the challenges this problem comprises. In this paper we describe a zero-shot learning approach that can be implemented in just one line of code, yet it is able to outperform state of the art approaches on standard datasets. The approach is based on a more general framework which models the relationships between features, attributes, and classes as a two linear layers network, where the weights of the top layer are not learned but are given by the environment. We further provide a learning bound on the generalisation error of this kind of approaches, by casting them as domain adaptation methods. In experiments carried out on three standard real datasets, we found that our approach is able to perform significantly better than the state of art on all of them, obtaining a ratio of improvement up to 17%.
Added
2026-09-25

Cross-Stitch Networks for Multi-task Learning
Ishan Misra, Abhinav Shrivastava, Abhinav Gupta, Martial Hebert
Why you should read this
Proposes cross-stitch units, an end-to-end trainable mechanism that automatically learns the optimal balance of shared and task-specific representations across multiple convolutional networks, eliminating manual architecture search while boosting performance on data-scarce tasks.
Multi-task learning in Convolutional Networks has displayed remarkable success in the field of recognition. This success can be largely attributed to learning shared representations from multiple supervisory tasks. However, existing multi-task approaches rely on enumerating multiple network architectures specific to the tasks at hand, that do not generalize. In this paper, we propose a principled approach to learn shared representations in ConvNets using multi-task learning. Specifically, we propose a new sharing unit: "cross-stitch" unit. These units combine the activations from multiple networks and can be trained end-to-end. A network with cross-stitch units can learn an optimal combination of shared and task-specific representations. Our proposed method generalizes across multiple tasks and shows dramatically improved performance over baseline methods for categories with few training examples.
Added
2026-09-24

DeepFashion: Powering Robust Clothes Recognition and Retrieval with Rich Annotations
Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, Xiaoou Tang
Why you should read this
Introduces DeepFashion, a large-scale clothing dataset containing over 800,000 images with rich attribute, landmark, and cross-domain annotations, alongside FashionNet, a deep architecture that jointly predicts landmarks and attributes to advance clothing recognition and retrieval.
Recent advances in clothes recognition have been driven by the construction of clothes datasets. Existing datasets are limited in the amount of annotations and are difficult to cope with the various challenges in real-world applications. In this work, we introduce DeepFashion¹, a large-scale clothes dataset with comprehensive annotations. It contains over 800,000 images, which are richly annotated with massive attributes, clothing landmarks, and correspondence of images taken under different scenarios including store, street snapshot, and consumer. Such rich annotations enable the development of powerful algorithms in clothes recognition and facilitating future researches. To demonstrate the advantages of DeepFashion, we propose a new deep model, namely FashionNet, which learns clothing features by jointly predicting clothing attributes and landmarks. The estimated landmarks are then employed to pool or gate the learned features. It is optimized in an iterative manner. Extensive experiments demonstrate the effectiveness of FashionNet and the usefulness of DeepFashion.
Added
2026-09-18

Describing Objects by their Attributes
Ali Farhadi, Ian Endres, Derek Hoiem, David Forsyth
Why you should read this
Demonstrates how describing objects through learned visual attributes enables computers to recognize unfamiliar objects, learn new categories from text descriptions alone, and identify unusual variations—capabilities far beyond traditional naming-based recognition.
We propose to shift the goal of recognition from naming to describing. Doing so allows us not only to name familiar objects, but also: to report unusual aspects of a familiar object (“spotty dog”, not just “dog”); to say something about unfamiliar objects (“hairy and four-legged”, not just “unknown”); and to learn how to recognize new objects with few or no visual examples. Rather than focusing on identity assignment, we make inferring attributes the core problem of recognition. These attributes can be semantic (“spotty”) or discriminative (“dogs have it but sheep do not”). Learning attributes presents a major new challenge: generalization across object categories, not just across instances within a category. In this paper, we also introduce a novel feature selection method for learning attributes that generalize well across categories. We support our claims by thorough evaluation that provides insights into the limitations of the standard recognition paradigm of naming and demonstrates the new abilities provided by our attribute-based framework.
Added
2026-02-21
