LISA: Reasoning Segmentation via Large Language Model

Xin LaiZhuotao TianYukang ChenYanwei LiYuhui YuanShu LiuJiaya Jia

article2024CVPR1,037 citations

Introduces the task of reasoning segmentation alongside LISA, an end-to-end multimodal architecture that decodes special text tokens into binary masks to segment targets specified by implicit, knowledge-intensive user instructions.

Listen

Existing visual perception systems typically require explicit instructions or pre-defined labels to identify target objects, making them incapable of understanding implicit user intent or applying broader world knowledge. As automated systems and robotics encounter more complex, natural human commands, this limitation hinders their practical utility. The article addresses this gap by defining a new task called reasoning segmentation, where a system must interpret complex, indirect text queries and isolate the referenced objects within an image.

The main objective of the article is to introduce and evaluate LISA (Large Language Instructed Segmentation Assistant), a system that enables multimodal language models to generate precise segmentation masks from implicit instructions. To assess this capability, the article establishes ReasonSeg, a benchmark dataset consisting of 1,218 image-instruction-mask samples annotated with short and long implicit queries requiring visual and world knowledge reasoning. LISA integrates a vision backbone with an autoregressive language model via a newly introduced segmentation token, decoding its hidden embedding directly into visual masks in an end-to-end framework while preserving underlying dialogue capabilities.

The experimental findings show substantial improvements over existing segmentation and open-vocabulary systems. First, LISA achieves superior performance on the reasoning segmentation benchmark, outperforming established baseline models by more than 20 percentage points in global intersection-over-union metrics. Second, LISA demonstrates strong zero-shot reasoning abilities when trained purely on explicit, standard segmentation datasets, which further improves with minimal targeted fine-tuning on just 239 reasoning samples. Third, the unified end-to-end architecture significantly outperforms two-stage decoupled pipelines that use text descriptions as intermediaries. Fourth, scaling the underlying language model from a 7-billion to a 13-billion parameter variant yields noticeable accuracy gains, particularly on complex, long-sentence instructions, while also achieving state-of-the-art results on standard referring segmentation tasks.

These results indicate that embedding fine-grained visual mask generation directly into large multimodal models creates an effective path toward intuitive, instruction-following visual assistants. By reducing the need to engineer rigid explicit prompts, this approach lowers development complexity and improves interaction reliability for downstream robotics and automated inspection systems. Organizations developing perception or robotic systems should consider adopting end-to-end embedding-as-mask architectures and can fine-tune these models efficiently with lightweight parameter tuning and minimal task-specific reasoning annotations.

The primary limitations identified include the model's reliance on high-capacity language models to properly parse long or complex queries, which may create computational bottlenecks during deployment. While confidence in the benchmark performance is high, practical implementation across highly specialized operational environments will require further domain-specific validation and safety testing.

  • Paper: Visual Instruction Tuning, Haotian Liu et al. (2023). LISA adapts LLaVA’s vision-encoder-to-language-model instruction-tuning framework, so this paper clarifies the multimodal assistant architecture it extends with mask generation.
Cover for LISA: Reasoning Segmentation via Large Language Model

Abstract

Although perception systems have made remarkable advancements in recent years, they still rely on explicit human instruction or pre-defined categories to identify the target objects before executing visual recognition tasks. Such systems cannot actively reason and comprehend implicit user intention. In this work, we propose a new segmentation task — reasoning segmentation. The task is designed to output a segmentation mask given a complex and implicit query text. Furthermore, we establish a benchmark comprising over one thousand image-instruction-mask data samples, incorporating intricate reasoning and world knowledge for evaluation purposes. Finally, we present LISA: large Language Instructed Segmentation Assistant, which inherits the language generation capabilities of multimodal Large Language Models (LLMs) while also possessing the ability to produce segmentation masks. We expand the original vocabulary with a <SEG> token and propose the embedding-as-mask paradigm to unlock the segmentation capability. Remarkably, LISA can handle cases involving complex reasoning and world knowledge. Also, it demonstrates robust zero-shot capability when trained exclusively on reasoning-free datasets. In addition, fine-tuning the model with merely 239 reasoning segmentation data samples results in further performance enhancement. Both quantitative and qualitative experiments show our method effectively unlocks new reasoning segmentation capabilities for multimodal LLMs. Code, models, and data are available at github.com/dvlab-research/LISA.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Image Segmentation
  • 2.2 Multimodal Large Language Model
  • 3 Reasoning Segmentation
  • 3.1 Problem Definition
  • 3.2 Benchmark
  • 4 Our Method
  • 4.1 Architecture
  • 4.2 Training
  • 5 Experiment
  • 5.1 Experimental Setting
  • 5.2 Reasoning Segmentation Results
  • 5.3 Vanilla Referring Segmentation Results
  • 5.4 Ablation Study
  • 5.5 Qualitative Results
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — LISA turns a language-model token embedding into a segmentation mask

    model/method

    LISA adds a vocabulary token, <SEG>, to a multimodal large language model (LLM). Given an image ximgx_{img} and text instruction xtxtx_{txt}, the LLM F\mathcal{F} generates a text response y^txt\hat{y}_{txt}. When the response calls for a mask, LISA takes the final-layer hidden embedding h~seg\tilde{h}_{seg} of the <SEG> token and projects it with a multilayer perceptron γ\gamma to obtain hsegh_{seg}. In parallel, a vision encoder Fenc\mathcal{F}_{enc} extracts dense image features ff. A mask decoder Fdec\mathcal{F}_{dec} combines the projected token embedding and image features to predict the binary mask M^\hat{M}:

    y^txt=F(ximg,xtxt),hseg=γ(h~seg),f=Fenc(ximg),M^=Fdec(hseg,f).\hat{y}_{txt}=\mathcal{F}(x_{img},x_{txt}),\qquad h_{seg}=\gamma(\tilde{h}_{seg}),\qquad f=\mathcal{F}_{enc}(x_{img}),\qquad \hat{M}=\mathcal{F}_{dec}(h_{seg},f).

    Here, ximgx_{img} is the input image, xtxtx_{txt} is the instruction, y^txt\hat{y}_{txt} is the generated response, and M^\hat{M} is the predicted mask. The embedding-as-mask design lets the model generate text and fine-grained masks in one end-to-end system; placing multiple <SEG> tokens in a response also allows it to produce multiple masks.

  2. Knowl 2 — Reasoning segmentation targets implicit, reasoning-intensive image queries

    definition

    Reasoning segmentation takes an image and an implicit text instruction and returns a binary mask for the object or region that satisfies the instruction. Unlike conventional referring segmentation, where a query may explicitly name or describe its target, a reasoning-segmentation query can require inference or world knowledge—for example, identifying a container from what it is used for, or selecting food by a nutritional property. The task therefore requires interpreting the query jointly with the image before locating the target pixels.

  3. Knowl 3 — ReasonSeg provides 1,218 annotated image-instruction-mask examples

    data/table

    ReasonSeg is a benchmark for evaluating reasoning segmentation. Its images come from OpenImages and ScanNetv2, and each example pairs an image and an implicit instruction with a high-quality target mask. Instructions include both short phrases and longer sentences, with queries intended to involve complex reasoning or world knowledge. The dataset has 1,218 examples: 239 in the training split, 200 in validation, and 779 in test. The larger validation and test splits reflect the benchmark's primary evaluation purpose.

  4. Knowl 4 — LISA jointly trains language generation and mask prediction

    equation

    LISA is trained with a text-generation loss and a segmentation-mask loss. Its total objective is

    L=λtxtLtxt+λmaskLmask,Ltxt=CE(y^txt,ytxt),Lmask=λbceBCE(M^,M)+λdiceDICE(M^,M).\mathcal{L}=\lambda_{txt}\mathcal{L}_{txt}+\lambda_{mask}\mathcal{L}_{mask},\qquad \mathcal{L}_{txt}=\mathrm{CE}(\hat{y}_{txt},y_{txt}),\qquad \mathcal{L}_{mask}=\lambda_{bce}\mathrm{BCE}(\hat{M},M)+\lambda_{dice}\mathrm{DICE}(\hat{M},M).

    Here, y^txt\hat{y}_{txt} and ytxty_{txt} are the generated and target text sequences; M^\hat{M} and MM are the predicted and ground-truth binary masks. CE\mathrm{CE} is autoregressive cross-entropy, while BCE\mathrm{BCE} and DICE\mathrm{DICE} are the per-pixel binary cross-entropy and DICE mask losses. In the reported training setup, λtxt=1.0\lambda_{txt}=1.0, λmask=1.0\lambda_{mask}=1.0, λbce=2.0\lambda_{bce}=2.0, and λdice=0.5\lambda_{dice}=0.5.

  5. Knowl 5 — Training combines segmentation, referring-expression, and VQA data

    experimental setup

    LISA's training mixture uses three kinds of public data, none of which contains reasoning-segmentation examples in the standard training setup. For semantic segmentation, the experiments use ADE20K, COCO-Stuff, PACO-LVIS, PartImageNet, and PASCAL-Part. The training pipeline selects up to three categories from an image and converts each selected category's ground-truth labels into a binary mask and a question-answer example. For referring segmentation, it uses refCLEF, refCOCO, refCOCO+, and refCOCOg, converting the explicit target description into an instruction and supervising the corresponding mask. For text-only visual question answering, it uses LLaVA-Instruct-150k with LLaVA v1 or LLaVA-v1.5-mix665k with LLaVA v1.5, to retain general VQA ability.

    For efficient fine-tuning, the LLM is adapted with LoRA; the vision backbone is frozen, while the mask decoder is fully fine-tuned. The LLM token embeddings, language-model head, and projection layer γ\gamma are also trainable. The reported setup uses eight NVIDIA 24 GB 3090 GPUs, AdamW with learning rate 0.00030.0003 and weight decay 00, a WarmupDecayLR scheduler with 100 warmup iterations, batch size 2 per device, and 10 gradient-accumulation steps. To avoid data leakage, COCO images that occur in refCOCO-family validation sets are excluded from training.

  6. Knowl 6 — LISA substantially outperforms prior methods on ReasonSeg

    empirical result

    The ReasonSeg evaluation reports generalized IoU (gIoU), the mean of per-image intersection-over-union scores, and cumulative IoU (cIoU), the cumulative intersection divided by cumulative union. The table gives validation results overall and by query length, plus overall test results. Scores are reported as in the paper. “ft” means fine-tuning on the 239 ReasonSeg training examples; unqualified LISA models use LLaVA v1 unless the model name specifies LLaVA1.5.

    Method Val gIoU Val cIoU Short gIoU Short cIoU Long gIoU Long cIoU Test gIoU Test cIoU
    OVSeg 28.5 18.6 18.0 15.5 28.7 22.5 26.1 20.8
    GRES 22.4 19.9 17.6 15.0 22.6 23.8 21.3 22.0
    X-Decoder 22.6 17.9 20.4 11.6 22.2 17.5 21.7 16.3
    SEEM 25.5 21.2 20.1 11.5 25.6 20.8 24.3 18.7
    Grounded-SAM 26.0 14.5 17.8 10.8 22.4 18.6 21.3 16.4
    LISA-7B 44.4 46.0 37.6 34.4 36.6 34.7 36.8 34.1
    LISA-7B (ft) 52.9 54.0 40.6 40.6 49.4 51.0 47.3 48.4
    LISA-13B 48.9 46.9 39.9 43.3 46.4 46.5 44.8 45.8
    LISA-13B (ft) 56.2 62.9 44.3 42.0 54.0 54.3 51.7 51.1
    LLaVA1.5-7B + OVSeg 38.2 23.5 24.2 18.7 44.6 37.1 39.7 31.8
    LISA-7B-LLaVA1.5 53.6 52.3 47.1 48.5 49.2 48.9 48.7 48.8
    LISA-7B-LLaVA1.5 (ft) 61.3 62.9 48.3 46.3 57.9 59.7 55.6 56.9
    LLaVA1.5-13B + OVSeg 37.9 26.4 27.1 19.4 46.1 40.6 41.5 34.1
    LISA-13B-LLaVA1.5 57.7 60.3 50.8 50.0 54.7 50.9 53.8 50.8
    LISA-13B-LLaVA1.5 (ft) 65.0 72.9 55.4 50.6 63.2 65.3 61.3 62.2

    Even without reasoning-segmentation training data, LISA-7B reaches test gIoU 36.8, compared with 26.1 for the strongest listed prior baseline, OVSeg. Fine-tuned LISA-13B-LLaVA1.5 reaches test gIoU 61.3 and cIoU 62.2. The table also shows that the LLaVA1.5-based LISA models outperform the corresponding LLaVA1.5-plus-OVSeg two-stage systems.

  7. Knowl 7 — Fine-tuning on 239 reasoning examples improves results, and rephrasing adds gains

    empirical result

    For LISA-7B, fine-tuning on the 239 ReasonSeg training examples raises validation gIoU/cIoU from 44.4/46.0 to 52.9/54.0. On the ReasonSeg test set, training on those 239 examples yields 51.7 gIoU and 51.1 cIoU; adding the 200 validation examples to the fine-tuning data, for 439 examples total, yields 54.0 gIoU and 54.9 cIoU. In a separate validation ablation, rephrasing reasoning instructions with GPT-3.5 and randomly choosing a rephrasing increases performance from 50.7 gIoU and 51.1 cIoU to 52.9 gIoU and 54.0 cIoU, gains of 2.2 and 2.9 points respectively.

  8. Knowl 8 — The vision backbone affects ReasonSeg performance

    empirical result

    On the ReasonSeg validation set, the frozen SAM vision backbone performs better than the tested alternatives in the reported LISA-7B experiments. The scores below are gIoU/cIoU; “ft” denotes fine-tuning on the ReasonSeg training split.

    Vision backbone gIoU cIoU
    Mask2Former-Swin-L 42.4 38.8
    SAM (with LoRA) 41.5 37.3
    SAM 44.4 46.0
    Mask2Former-Swin-L (ft) 50.7 52.3
    SAM with LoRA (ft) 51.8 51.9
    SAM (ft) 52.9 54.0

    The comparison shows that LISA can use a Mask2Former backbone, but SAM gives the highest score among these choices both before and after reasoning-data fine-tuning. Applying LoRA to SAM is below using the unadapted SAM backbone in both settings; the authors suggest that fine-tuning may impair SAM's generalization.

  9. Knowl 9 — Semantic-segmentation data is important to LISA's training mixture

    empirical result

    A training-data ablation on the ReasonSeg validation set reports 52.9 gIoU and 54.0 cIoU when the full mixture—including semantic segmentation, referring segmentation, VQA, and reasoning-segmentation data—is used. With the standard semantic, referring, and VQA data but no reasoning-segmentation examples, the scores are 44.4 and 46.0. In the ablation that removes semantic-segmentation datasets, the scores fall to 30.4 gIoU and 20.4 cIoU. The authors attribute the importance of semantic-segmentation data in part to its many labeled masks: a multiclass segmentation label map can provide multiple binary-mask training targets.

  10. Knowl 10 — LISA also performs strongly on conventional referring segmentation

    empirical result

    The authors evaluate LISA-7B on conventional referring segmentation using cIoU. Without additional referring-segmentation fine-tuning, its scores are 74.1/76.5/71.1 on refCOCO val/testA/testB, 62.4/67.4/56.5 on refCOCO+ val/testA/testB, and 66.4/68.5 on refCOCOg val(U)/test(U). After fine-tuning on referring-segmentation data, the corresponding scores are 74.9/79.1/72.3, 65.1/70.8/58.1, and 67.9/70.6. These results show that the same model framework supports explicit referring queries as well as reasoning-intensive queries.

Coverage note — The qualitative example gallery is omitted because it illustrates behaviors already captured by the task definition and quantitative results, without adding a separate measured finding.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. NeurIPS, 2022. 1, 3
  2. 2.Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. TPAMI, 2017. 2
  3. 3.Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. In CVPR, 2018. 6
  4. 4.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020. 2
  5. 5.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. TPAMI, 2018. 2
  6. 6.Xianjie Chen, Roozbeh Mottaghi, Xiaobai Liu, Sanja Fidler, Raquel Urtasun, and Alan Yuille. Detect what you can: Detecting and representing objects using holistic models and body parts. In CVPR, 2014. 6
  7. 7.Bowen Cheng, Maxwell D Collins, Yukun Zhu, Ting Liu, Thomas S Huang, Hartwig Adam, and Liang-Chieh Chen. Panoptic-deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation. In CVPR, 2020. 2
  8. 8.Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation. NeurIPS, 2021. 2
  9. 9.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In CVPR, 2022. 2, 4
  10. 10.Angela Dai, Angel X. Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In CVPR, 2017. 2, 3
  11. 11.Henghui Ding, Chang Liu, Suchen Wang, and Xudong Jiang. Vision-language transformer and query generation for referring segmentation. In ICCV, 2021. 7
  12. 12.Jun Fu, Jing Liu, Haijie Tian, Yong Li, Yongjun Bao, Zhiwei Fang, and Hanqing Lu. Dual attention network for scene segmentation. In CVPR, 2019. 2
  13. 13.Ju He, Shuo Yang, Shaokang Yang, Adam Kortylewski, Xiaoding Yuan, Jie-Neng Chen, Shuai Liu, Cheng Yang, Qihang Yu, and Alan Yuille. Partimagenet: A large, high-quality dataset of parts. In ECCV, 2022. 6
  14. 14.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In ICCV, 2017. 2
  15. 15.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv:2106.09685, 2021. 4, 5
  16. 16.Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. Ccnet: Criss-cross attention for semantic segmentation. In ICCV, 2019. 2
  17. 17.Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. Referitgame: Referring to objects in photographs of natural scenes. In EMNLP, 2014. 2, 3, 6
  18. 18.Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Dollar. Panoptic segmentation. In CVPR, 2019. 2
  19. 19.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv:2304.02643, 2023. 2, 4, 6
  20. 20.Jing Yu Koh, Ruslan Salakhutdinov, and Daniel Fried. Grounding language models to images for multimodal inputs and outputs. 2023. 3
  21. 21.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, Tom Duerig, and Vittorio Ferrari. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. IJCV, 2020. 2, 3
  22. 22.Xin Lai, Zhuotao Tian, Li Jiang, Shu Liu, Hengshuang Zhao, Liwei Wang, and Jiaya Jia. Semi-supervised semantic segmentation with directional context-aware consistency. In CVPR, 2021. 2
  23. 23.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, and Ziwei Liu. Otter: A multi-modal model with in-context instruction tuning. arXiv:2305.03726, 2023. 1, 3
  24. 24.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv:2301.12597, 2023. 1, 3
  25. 25.Yanwei Li, Hengshuang Zhao, Xiaojuan Qi, Liwei Wang, Zeming Li, Jian Sun, and Jiaya Jia. Fully convolutional networks for panoptic segmentation. In CVPR, 2021. 2
  26. 26.Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu. Open-vocabulary semantic segmentation with mask-adapted clip. In CVPR, 2023. 6
  27. 27.Chang Liu, Henghui Ding, and Xudong Jiang. Gres: Generalized referring expression segmentation. In CVPR, 2023. 6, 7
  28. 28.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. arXiv preprint, 2023. 1, 5, 6, 7
  29. 29.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. arXiv:2304.08485, 2023. 1, 3, 4, 5, 6
  30. 30.Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint, 2023. 6
  31. 31.Wei Liu, Andrew Rabinovich, and Alexander C. Berg. Parsenet: Looking wider to see better. arXiv preprint, 2015. 2
  32. 32.Zhaoyang Liu, Yinan He, Wenhai Wang, Weiyun Wang, Yi Wang, Shoufa Chen, Qinglong Zhang, Yang Yang, Qingyun Li, Jiashuo Yu, et al. Internchat: Solving vision-centric tasks by interacting with chatbots beyond language. arXiv:2305.05662, 2023. 3
  33. 33.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv:1711.05101, 2017. 6
  34. 34.Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Liujuan Cao, Chenglin Wu, Cheng Deng, and Rongrong Ji. Multi-task collaborative network for joint referring expression comprehension and segmentation. In CVPR, 2020. 7
  35. 35.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In CVPR, 2016. 6
  36. 36.Varun K Nagaraja, Vlad I Morariu, and Larry S Davis. Modeling context between objects for referring expression understanding. In ECCV, 2016. 2
  37. 37.Hyeonwoo Noh, Seunghoon Hong, and Bohyung Han. Learning deconvolution network for semantic segmentation. In ICCV, 2015. 2
  38. 38.Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. Kosmos-2: Grounding multimodal large language models to the world. arXiv:2306.14824, 2023. 3
  39. 39.Renjie Pi, Jiahui Gao, Shizhe Diao, Rui Pan, Hanze Dong, Jipeng Zhang, Lewei Yao, Jianhua Han, Hang Xu, Lingpeng Kong, and Tong Zhang. Detgpt: Detect what you need via reasoning. arXiv:2305.14167, 2023. 3
  40. 40.Vignesh Ramanathan, Anmol Kalia, Vladan Petrovic, Yi Wen, Baixue Zheng, Baishan Guo, Rui Wang, Aaron Marquez, Rama Kovvuri, Abhishek Kadian, et al. Paco: Parts and attributes of common objects. In CVPR, 2023. 6
  41. 41.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In SIGKDD, 2020. 6
  42. 42.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In MICCAI, 2015. 2
  43. 43.Evan Shelhamer, Jonathan Long, and Trevor Darrell. Fully convolutional networks for semantic segmentation. TPAMI, 2017. 2
  44. 44.Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface. arXiv:2303.17580, 2023. 3
  45. 45.Zhuotao Tian, Pengguang Chen, Xin Lai, Li Jiang, Shu Liu, Hengshuang Zhao, Bei Yu, Ming-Chang Yang, and Jiaya Jia. Adaptive perspective distillation for semantic segmentation. TPAMI, 2022. 2
  46. 46.Zhuotao Tian, Jiequan Cui, Li Jiang, Xiaojuan Qi, Xin Lai, Yixin Chen, Shu Liu, and Jiaya Jia. Learning context-aware classifier for semantic segmentation. AAAI, 2023. 2
  47. 47.Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, et al. Visionllm: Large language model is also an open-ended decoder for vision-centric tasks. arXiv:2305.11175, 2023. 3
  48. 48.Zhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao, Yandong Guo, Mingming Gong, and Tongliang Liu. Cris: Clip-driven referring image segmentation. In CVPR, 2022. 7
  49. 49.Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. Visual chatgpt: Talking, drawing and editing with visual foundation models. arXiv:2303.04671, 2023. 3
  50. 50.Yuwen Xiong, Renjie Liao, Hengshuang Zhao, Rui Hu, Min Bai, Ersin Yumer, and Raquel Urtasun. Upsnet: A unified panoptic segmentation network. In CVPR, 2019. 2
  51. 51.Maoke Yang, Kun Yu, Chi Zhang, Zhiwei Li, and Kuiyuan Yang. Denseaspp for semantic segmentation in street scenes. In CVPR, 2018. 2
  52. 52.Rui Yang, Lin Song, Yanwei Li, Sijie Zhao, Yixiao Ge, Xiu Li, and Ying Shan. Gpt4tools: Teaching large language model to use tools via self-instruction. arXiv:2305.18752, 2023. 3
  53. 53.Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Hengshuang Zhao, and Philip HS Torr. Lavt: Language-aware vision transformer for referring image segmentation. In CVPR, 2022. 7
  54. 54.Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Ehsan Azarnasab, Faisal Ahmed, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang. Mm-react: Prompting chatgpt for multimodal reasoning and action. arXiv:2303.11381, 2023. 3
  55. 55.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. mplug-owl: Modularization empowers large language models with multimodality. arXiv:2304.14178, 2023. 1, 3
  56. 56.Fisher Yu and Vladlen Koltun. Multi-scale context aggregation by dilated convolutions. In ICLR, 2016. 2
  57. 57.Shilong Zhang, Peize Sun, Shoufa Chen, Min Xiao, Wenqi Shao, Wenwei Zhang, Kai Chen, and Ping Luo. Gpt4roi: Instruction tuning large language model on region-of-interest. arXiv:2307.03601, 2023. 3
  58. 58.Wenwei Zhang, Jiangmiao Pang, Kai Chen, and Chen Change Loy. K-net: Towards unified image segmentation. NeurIPS, 2021. 2
  59. 59.Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In CVPR, 2017. 2
  60. 60.Hengshuang Zhao, Xiaojuan Qi, Xiaoyong Shen, Jianping Shi, and Jiaya Jia. Icnet for real-time semantic segmentation on high-resolution images. In ECCV, 2018.
  61. 61.Hengshuang Zhao, Yi Zhang, Shu Liu, Jianping Shi, Chen Change Loy, Dahua Lin, and Jiaya Jia. Psanet: Point-wise spatial attention network for scene parsing. In ECCV, 2018. 2
  62. 62.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In CVPR, 2017. 6
  63. 63.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv:2304.10592, 2023. 1, 3
  64. 64.Zhen Zhu, Mengde Xu, Song Bai, Tengteng Huang, and Xiang Bai. Asymmetric non-local neural networks for semantic segmentation. In ICCV, 2019. 2
  65. 65.Xueyan Zou, Zi-Yi Dou, Jianwei Yang, Zhe Gan, Linjie Li, Chunyuan Li, Xiyang Dai, Harkirat Behl, Jianfeng Wang, Lu Yuan, et al. Generalized decoding for pixel, image, and language. In CVPR, 2023. 2, 6, 7
  66. 66.Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Gao, and Yong Jae Lee. Segment everything everywhere all at once. arXiv:2304.06718, 2023. 2, 4, 6, 7

Citation

MLA
Lai, X., et al. “LISA: Reasoning Segmentation via Large Language Model”. arXiv, 2023, http://arxiv.org/abs/2308.00692v3.
APA
Lai, X., Tian, Z., Chen, Y., Li, Y., Yuan, Y., Liu, S., & Jia, J. (2023). LISA: Reasoning Segmentation via Large Language Model. arXiv. http://arxiv.org/abs/2308.00692v3
Chicago
Lai, X., Z. Tian, Y. Chen, et al. 2023. “LISA: Reasoning Segmentation via Large Language Model”. arXiv. http://arxiv.org/abs/2308.00692v3.
Harvard
Lai, X. et al. (2023) “LISA: Reasoning Segmentation via Large Language Model”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2308.00692v3.
Vancouver
1. Lai X, Tian Z, Chen Y, Li Y, Yuan Y, Liu S, Jia J (2023) LISA: Reasoning Segmentation via Large Language Model. arXiv

BibTeX

@article{lai2023lisa,
  title = {LISA: Reasoning Segmentation via Large Language Model},
  author = {Lai, Xin and Tian, Zhuotao and Chen, Yukang and Li, Yanwei and Yuan, Yuhui and Liu, Shu and Jia, Jiaya},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2308.00692v3},
  eprint = {2308.00692}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE