GSVA: Generalized Segmentation via Multimodal Large Language Models
Zhuofan XiaDongchen HanYizeng HanXuran PanShiji SongGao Huang
Presents a multimodal framework that enables language models to segment multiple target objects simultaneously and explicitly reject non-existent targets using dedicated tokens, setting a new state of the art on the gRefCOCO benchmark.
Real-world artificial intelligence applications, such as robotic navigation and embodied visual assistants, require computer vision systems to interpret complex natural language instructions and accurately locate referenced objects within images. Traditional referring expression segmentation methods operate under the strict assumption that every language prompt maps to exactly one target present in the visual scene. In practice, however, users frequently refer to multiple objects simultaneously or describe items that do not exist in the scene. When existing models encounter non-existent objects, they often force a segmentation mask onto unrelated visual areas, creating substantial operational and safety risks for autonomous systems.
The main objective of the article is to demonstrate a new framework, named Generalized Segmentation Vision Assistant (GSVA), designed to resolve Generalized Referring Expression Segmentation (GRES). The article evaluates whether combining multimodal large language models with specialized segmentation decoders can accurately segment multiple targets from a single prompt and explicitly reject non-existent targets.
To achieve this, the authors designed a unified vision-language architecture that pairs a multimodal large language model (based on Vicuna or LLaMA-2 backbones with visual encoders) with a segmentation foundation model (the Segment Anything Model). The approach introduces two primary mechanisms: generating multiple shared-weight segmentation tokens—each preceded by its corresponding descriptive text prompt to maintain context—and outputting a dedicated rejection token when a referent is absent from an image. These rejection tokens assign empty masks immediately without forcing the visual decoder to segment nonexistent areas. The authors evaluated GSVA against prior baselines, such as LISA and specialized non-LLM architectures, across standard benchmarks including the gRefCOCO dataset (comprising 278,232 expressions across 19,994 images) and classic datasets such as RefCOCO, RefCOCO+, and RefCOCOg.
The experimental findings show significant performance gains. First, GSVA established new state-of-the-art results on the gRefCOCO generalized segmentation benchmark; the fine-tuned 13-billion parameter version achieved an average mask accuracy of 70.04% on the validation set, outperforming specialized baselines and prior multimodal language models. Second, GSVA demonstrated a dramatic improvement in identifying absent targets, achieving null-target classification accuracies between 60% and 67%, whereas standard models without rejection capabilities often scored below 10% before task-specific tuning. Third, ablation experiments revealed that removing the rejection token reduced null-target identification accuracy by more than 25% and overall mask accuracy by about 10%. Finally, the system transferred effectively to classic single-target segmentation and bounding-box comprehension tasks, consistently outperforming comparable baseline models across various evaluation splits.
These findings imply that large multimodal models can successfully navigate complex spatial relationships and visual reasoning without requiring rigid, one-to-one prompt constraints. Explicit rejection tokens significantly reduce hallucination and false-positive mask generations, providing a safer, more reliable foundation for robotic manipulation and vision-guided automation where incorrect actions on non-existent objects carry high operational risk.
Stakeholders and developers deploying vision-language systems for autonomous platforms should integrate explicit rejection tokens and multi-entity prompting mechanisms into their multimodal architectures to mitigate hallucination risks. While the findings provide high confidence on benchmark datasets, future initiatives should evaluate the framework's real-time computational latency on embedded hardware and test its robustness across complex, real-world robotic deployments.
- Paper: Segment Anything, Alexander M. Kirillov et al. (2023). Segment Anything Model (SAM) introduces the foundation promptable mask decoder that GSVA and LISA prompt using special tokens for generalized segmentation.
- Paper: Visual Instruction Tuning, Haotian Liu et al. (2023). LLaVA establishes the core visual instruction tuning architecture that connects vision encoders to large language models for general multimodal understanding.
- Paper: Evaluating Object Hallucination in Large Vision-Language Models, Yifan Li et al. (2023). This paper investigates object hallucination in vision-language models, providing key context for why GSVA needs explicit target rejection mechanisms.
- Paper: Kosmos-2: Grounding Multimodal Large Language Models to the World, Zhiliang Peng et al. (2023). KOSMOS-2 details the integration of spatial location tokens into multimodal LLMs to ground natural language directly in image regions.
- Paper: Grounded Language-Image Pre-training, Liunian Harold Li et al. (2022). GLIP redefines object detection as visual phrase grounding, establishing fundamental techniques for aligning language tokens to spatial instance regions.
- Paper: Modeling Context in Referring Expressions, Licheng Yu et al. (2016). This work establishes the RefCOCO benchmark and context modeling for referring expressions, which GRES directly builds upon and generalizes.
- Paper: Generation and Comprehension of Unambiguous Object Descriptions, Junhua Mao et al. (2015). This foundational paper formalizes the task and evaluation of unambiguous referring expression comprehension and generation in natural scenes.
- Paper: Masked-attention Mask Transformer for Universal Image Segmentation, Bowen Cheng et al. (2022). Mask2Former provides the transformer-based mask attention framework widely leveraged in modern instance and generalized segmentation decoders.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). This survey systematically outlines the architectural components and instruction-tuning paradigms of modern multimodal large language models.
- Paper: SAM 3: Segment Anything with Concepts, Nicolas Carion et al. (2025). SAM 3 advances promptable concept segmentation to detect and track all instances described by text prompts or examples across images and video.
- Paper: Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models, Siddharth Karamcheti et al. (2024). Prismatic VLMs explores the broader design space of visually conditioned language models, systematically analyzing visual representations and localization performance.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). LLaVA-OneVision scales multimodal instruction following across single images, multi-image contexts, and video streams.
- Paper: Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces, Jihan Yang et al. (2025). This work evaluates the spatial reasoning, layout recall, and 3D perception capabilities of advanced multimodal large language models.
