LISA: Reasoning Segmentation via Large Language Model
Xin LaiZhuotao TianYukang ChenYanwei LiYuhui YuanShu LiuJiaya Jia
Introduces the task of reasoning segmentation alongside LISA, an end-to-end multimodal architecture that decodes special text tokens into binary masks to segment targets specified by implicit, knowledge-intensive user instructions.
Existing visual perception systems typically require explicit instructions or pre-defined labels to identify target objects, making them incapable of understanding implicit user intent or applying broader world knowledge. As automated systems and robotics encounter more complex, natural human commands, this limitation hinders their practical utility. The article addresses this gap by defining a new task called reasoning segmentation, where a system must interpret complex, indirect text queries and isolate the referenced objects within an image.
The main objective of the article is to introduce and evaluate LISA (Large Language Instructed Segmentation Assistant), a system that enables multimodal language models to generate precise segmentation masks from implicit instructions. To assess this capability, the article establishes ReasonSeg, a benchmark dataset consisting of 1,218 image-instruction-mask samples annotated with short and long implicit queries requiring visual and world knowledge reasoning. LISA integrates a vision backbone with an autoregressive language model via a newly introduced segmentation token, decoding its hidden embedding directly into visual masks in an end-to-end framework while preserving underlying dialogue capabilities.
The experimental findings show substantial improvements over existing segmentation and open-vocabulary systems. First, LISA achieves superior performance on the reasoning segmentation benchmark, outperforming established baseline models by more than 20 percentage points in global intersection-over-union metrics. Second, LISA demonstrates strong zero-shot reasoning abilities when trained purely on explicit, standard segmentation datasets, which further improves with minimal targeted fine-tuning on just 239 reasoning samples. Third, the unified end-to-end architecture significantly outperforms two-stage decoupled pipelines that use text descriptions as intermediaries. Fourth, scaling the underlying language model from a 7-billion to a 13-billion parameter variant yields noticeable accuracy gains, particularly on complex, long-sentence instructions, while also achieving state-of-the-art results on standard referring segmentation tasks.
These results indicate that embedding fine-grained visual mask generation directly into large multimodal models creates an effective path toward intuitive, instruction-following visual assistants. By reducing the need to engineer rigid explicit prompts, this approach lowers development complexity and improves interaction reliability for downstream robotics and automated inspection systems. Organizations developing perception or robotic systems should consider adopting end-to-end embedding-as-mask architectures and can fine-tune these models efficiently with lightweight parameter tuning and minimal task-specific reasoning annotations.
The primary limitations identified include the model's reliance on high-capacity language models to properly parse long or complex queries, which may create computational bottlenecks during deployment. While confidence in the benchmark performance is high, practical implementation across highly specialized operational environments will require further domain-specific validation and safety testing.
- Paper: Visual Instruction Tuning, Haotian Liu et al. (2023). LISA adapts LLaVA’s vision-encoder-to-language-model instruction-tuning framework, so this paper clarifies the multimodal assistant architecture it extends with mask generation.
- Paper: GSVA: Generalized Segmentation via Multimodal Large Language Models, Zhuofan Xia et al. (2024). GSVA extends LISA-style language-guided segmentation to prompts with multiple targets and explicit rejection of objects absent from the image.
