keyword
semantic navigation
Semantic navigation is the capability of an embodied artificial intelligence agent or autonomous robot to navigate through an environment to find and reach a target specified by high-level semantic criteria, such as an object category, room type, or natural language description, rather than explicit geometric coordinates. Unlike traditional navigation methods that rely strictly on pre-mapped metric layouts and spatial coordinates, semantic navigation requires an agent to perceive, interpret, and reason about the visual and contextual meaning of its surroundings. To accomplish this, the system combines visual perception, semantic scene understanding, and spatial or commonsense reasoning to make sequential movement decisions, effectively explore unfamiliar or dynamic environments, and identify relevant goal objects or locations based on semantic relationships.
2 items

ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object Navigation
Kaiwen Zhou, Kaizhi Zheng, Connor Pryor, Yilin Shen, Hongxia Jin, Lise Getoor, Xin Eric Wang
Why you should read this
Proposes a training-free zero-shot object navigation framework that converts commonsense reasoning from pre-trained vision-language and large language models into soft logic exploration constraints, achieving state-of-the-art performance across multiple indoor benchmarks.
The ability to accurately locate and navigate to a specific object is a crucial capability for embodied agents that operate in the real world and interact with objects to complete tasks. Such object navigation tasks usually require large-scale training in visual environments with labeled objects, which generalizes poorly to novel objects in unknown environments. In this work, we present a novel zero-shot object navigation method, Exploration with Soft Commonsense constraints (ESC), that transfers commonsense knowledge in pre-trained models to open-world object navigation without any navigation experience nor any other training on the visual environments. First, ESC leverages a pre-trained vision and language model for open-world prompt-based grounding and a pre-trained commonsense language model for room and object reasoning. Then ESC converts commonsense knowledge into navigation actions by modeling it as soft logic predicates for efficient exploration. Extensive experiments on MP3D (Chang et al., 2017), HM3D (Ramakrishnan et al., 2021), and RoboTHOR (Deitke et al., 2020) benchmarks show that our ESC method improves significantly over baselines, and achieves new state-of-the-art results for zero-shot object navigation (e.g., 288% relative Success Rate improvement than CoW (Gadre et al., 2022) on MP3D).
Added
2026-10-05

AI2-THOR: An Interactive 3D Environment for Visual AI
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Matt Deitke, Kiana Ehsani, Daniel Gordon, Yuke Zhu, Aniruddha Kembhavi, Abhinav Gupta, Ali Farhadi
Why you should read this
Introduces a near-photorealistic 3D indoor simulation platform that enables embodied AI agents to physically interact with objects and explore complex environments for visual reinforcement learning and task planning.
We introduce The House Of inteRactions (THOR), a framework for visual AI research, available at this http URL. AI2-THOR consists of near photo-realistic 3D indoor scenes, where AI agents can navigate in the scenes and interact with objects to perform tasks. AI2-THOR enables research in many different domains including but not limited to deep reinforcement learning, imitation learning, learning by interaction, planning, visual question answering, unsupervised representation learning, object detection and segmentation, and learning models of cognition. The goal of AI2-THOR is to facilitate building visually intelligent models and push the research forward in this domain.
Added
2026-09-24
