Challenges in representation learning: A report on three machine learning contests
Ian J. GoodfellowDumitru ErhanPierre Luc CarrierAaron CourvilleMehdi MirzaBen HamnerWill CukierskiYichuan TangDavid ThalerDong-Hyun Lee
Examines the top-performing representation learning methods across three machine learning competitions—covering black-box, facial expression, and multimodal data—while providing actionable insights on algorithm performance and effective benchmark design.
- Paper: Representation Learning: A Review and New Perspectives, Yoshua Bengio et al. (2012). This seminal survey establishes the foundational principles, priors, and objectives of representation learning that motivate the workshop's three benchmark challenges.
- Paper: Multimodal Deep Learning, Jiquan Ngiam et al. (2011). It provides the foundational deep autoencoder architectures for learning joint cross-modal representations that directly inform the workshop's multimodal learning challenge.
- Paper: Why Does Unsupervised Pre-training Help Deep Learning?, D. Erhan et al. (2010). It explores the mechanisms behind unsupervised pre-training and feature discovery that underpin the black box and representation learning problems tackled in the contests.
- Paper: An Analysis of Single-Layer Networks in Unsupervised Feature Learning, Adam Coates et al. (2011). It analyzes key baseline methodologies in unsupervised feature learning that contextualize the competitive feature-extraction pipelines explored in the workshop.
- Paper: Building high-level features using large scale unsupervised learning, Quoc V. Le et al. (2011). It demonstrates how high-level facial and object representations emerge from large-scale unsupervised training, setting a precedent for the workshop's facial expression and feature learning tasks.
- Paper: Multimodal Machine Learning: A Survey and Taxonomy, Tadas Baltrušaitis et al. (2017). This comprehensive taxonomy systematizes multimodal representation, translation, and fusion, substantially generalizing beyond the early multimodal contest paradigms.
- Paper: Deep Learning Face Representation by Joint Identification-Verification, Yi Sun et al. (2014). It advances deep facial representation learning by formulating joint identification-verification objectives, extending the facial analysis benchmarks introduced in the contest.
- Paper: Deep Face Recognition, Omkar M. Parkhi et al. (2015). It scales up deep facial representation and metric learning architectures, building upon the baseline visual feature representations evaluated in the workshop.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). It presents modern crossmodal attention mechanisms for unaligned multimodal sequences, significantly advancing the fusion approaches benchmarked in early multimodal challenges.
- Paper: Contrastive Multiview Coding, Yonglong Tian et al. (2019). It extends multimodal and multi-view representation learning into modern self-supervised contrastive frameworks across diverse visual and sensory channels.
- Paper: ImageBind One Embedding Space to Bind Them All, Rohit Girdhar et al. (2023). It scales the multimodal representation challenge to unify six disparate sensory modalities within a single joint embedding space.
