Built independently by an author, for readers. Read the story and support ChapterPal

keyword

OMG-LLaVA

OMG-LLaVA is a multimodal artificial intelligence framework designed to bridge image-level conversational reasoning, object-level identification, and pixel-level visual understanding within a single system. Unlike traditional multimodal large language models that primarily handle high-level visual descriptions or specialized computer vision tools limited to segmentation tasks, OMG-LLaVA integrates a universal segmentation encoder and decoder with a large language model. This unified architecture processes user text instructions alongside diverse visual prompts by converting perception priors and visual features into tokens that the language model can interpret. Consequently, the system can engage in complex visual reasoning and concurrently produce conversational text responses and precise pixel-level segmentation masks.

1 item

OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Tao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan, Shengqiong Wu, Shunping Ji, Chen Change Loy, Shuicheng Yan

OrganizationsByteDanceNanyang Technological UniversitySkywork AIWuhan University

Why you should read this

Unifies image-level conversation, object-level visual prompting, and pixel-level segmentation into a single multimodal framework powered by one visual encoder, one decoder, and one LLM trained end-to-end.

Current universal segmentation methods demonstrate strong capabilities in pixel-level image and video understanding. However, they lack reasoning abilities and cannot be controlled via text instructions. In contrast, large vision-language multimodal models exhibit powerful vision-based conversation and reasoning capabilities but lack pixel-level understanding and have difficulty accepting visual prompts for flexible user interaction. This paper proposes OMG-LLaVA, a new and elegant framework combining powerful pixel-level vision understanding with reasoning abilities. It can accept various visual and text prompts for flexible user interaction. Specifically, we use a universal segmentation method as the visual encoder, integrating image information, perception priors, and visual prompts into visual tokens provided to the LLM. The LLM is responsible for understanding the user’s text instructions and providing text responses and pixel-level segmentation results based on the visual information. We propose perception prior embedding to better integrate perception priors with image features. OMG-LLaVA achieves image-level, object-level, and pixel-level reasoning and understanding in a single model, matching or surpassing the performance of specialized methods on multiple benchmarks. Rather than using LLM to connect each specialist, our work aims at end-to-end training on one encoder, one decoder, and one LLM. The code and model have been released for further research.

Added

2026-09-26