OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Tao ZhangXiangtai LiHao FeiHaobo YuanShengqiong WuShunping JiChen Change LoyShuicheng Yan

article2024NeurIPS186 citations

Unifies image-level conversation, object-level visual prompting, and pixel-level segmentation into a single multimodal framework powered by one visual encoder, one decoder, and one LLM trained end-to-end.

Listen

Current multimodal artificial intelligence systems typically specialize in either high-level visual reasoning, such as conversational description, or low-level visual perception, such as pixel-level segmentation. Systems that attempt to combine these capabilities often rely on complex multi-model pipelines or specialized tools, resulting in heavy computational overhead, severe task interference, and an inability to accept flexible visual prompts. The article demonstrates and evaluates a unified architecture called OMG-LLaVA, which integrates image-level, object-level, and pixel-level reasoning and understanding into a single model powered by one visual encoder, one decoder, and one large language model.

To bridge these tasks, the method models perception and reasoning as a unified token-to-token generation process. A frozen universal perception module extracts dense image features and object queries from user prompts such as points, boxes, and masks, and a perception prior embedding strategy merges these features into visual tokens fed to the language model. The language model then generates textual responses alongside special segmentation tokens that are decoded into precise masks. The framework was evaluated across standard benchmarks, including image-level conversation, referring expression segmentation, grounded conversation generation, and standard panoptic segmentation.

The findings show that OMG-LLaVA matches or outperforms specialized architectures across key perception and language metrics while maintaining a lightweight design. On referring expression segmentation, it achieved up to 78.0 cumulative intersection-over-union, outperforming comparable systems such as LISA and PixelLM. In grounded conversation tasks, it achieved stronger description and segmentation metrics (29.9 average precision at 50% overlap and 65.5 mean intersection-over-union) despite using significantly less pretraining data than competing models. Crucially, ablation experiments confirmed that the perception prior embedding is vital, providing an improvement of more than 10 to 13 points in segmentation accuracy compared to a naive model baseline, while preserving general conversational capabilities.

These results demonstrate that organizations can deploy versatile visual reasoning and fine-grained segmentation within a streamlined, single-backbone architecture, drastically reducing system complexity, training cost, and inference latency. Leaders should consider this unified approach when building interactive computer vision tools for complex visual question answering, object targeting, and automated scene parsing. Moving forward, engineering teams should explore expanding instruction-tuning datasets and extending the architecture to spatial-temporal reasoning for video feeds.

Confidence in these findings is supported by consistent cross-benchmark performance and thorough ablation studies. However, practical deployment considerations remain: joint training with dense segmentation data still induces a moderate drop in broader image-level reasoning compared to models trained purely on descriptive tasks, and the system cannot currently perform fine-grained, part-level object segmentation due to the underlying perception module's design constraints.

arXiv: 2406.19389
Cover for OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Abstract

Current universal segmentation methods demonstrate strong capabilities in pixel-level image and video understanding. However, they lack reasoning abilities and cannot be controlled via text instructions. In contrast, large vision-language multimodal models exhibit powerful vision-based conversation and reasoning capabilities but lack pixel-level understanding and have difficulty accepting visual prompts for flexible user interaction. This paper proposes OMG-LLaVA, a new and elegant framework combining powerful pixel-level vision understanding with reasoning abilities. It can accept various visual and text prompts for flexible user interaction. Specifically, we use a universal segmentation method as the visual encoder, integrating image information, perception priors, and visual prompts into visual tokens provided to the LLM. The LLM is responsible for understanding the user’s text instructions and providing text responses and pixel-level segmentation results based on the visual information. We propose perception prior embedding to better integrate perception priors with image features. OMG-LLaVA achieves image-level, object-level, and pixel-level reasoning and understanding in a single model, matching or surpassing the performance of specialized methods on multiple benchmarks. Rather than using LLM to connect each specialist, our work aims at end-to-end training on one encoder, one decoder, and one LLM. The code and model have been released for further research.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Task Unification
  • 3.2 OMG-LLaVA Framework
  • 3.3 Training and Testing Setup
  • 4 Experiment
  • 4.1 Main Results
  • 4.2 Ablation and Analysis
  • 5 Conclusion
  • References
  • A Appendix
  • A.1 More Implementation Details
  • A.2 More Experiment Results.
  • A.3 More Detailed Ablation Studies.
  • A.4 More Visualization Results
  • A.5 Limitation and Future Work Discussion

Knowls

  1. Knowl 1 — OMG-LLaVA Unified Token-to-Token Architecture

    model/method

    OMG-LLaVA unifies image-level understanding (e.g., captioning, dialogue), object-level reasoning (e.g., region captioning, visual prompt conversations), and pixel-level reasoning (e.g., universal segmentation, referring expression segmentation, grounded conversation generation) into a single autoregressive token generation framework. The system consists of three primary components: a single visual encoder (ConvNeXt-L CLIP), a universal perception decoder (frozen OMG-Seg decoder), and a Large Language Model (LLM).

    All tasks are expressed as mapping input tokens to output tokens across three distinct token types:

    • Text tokens TtT_t, representing linguistic instructions and generated text.
    • Pixel-centric visual tokens TpvT_{pv}, representing dense spatial features across the entire image.
    • Object-centric visual tokens TovT_{ov}, representing individual detected objects or user-prompted regions that can be decoded into segmentation masks.

    The unified model generation is defined as:

    Ttout,Tovout=LLM(Tpvin,Tovin,Ttin)T_{t}^{out}, T_{ov}^{out} = \text{LLM}(T_{pv}^{in}, T_{ov}^{in}, T_{t}^{in})

    Three special tokens structure multimodal interaction:

    • <Image>: replaced by the combined visual tokens Tv=(Tpv,Tov)T_v = (T_{pv}, T_{ov}).
    • <Region>: replaced by object-centric visual tokens corresponding to specific prompted regions.
    • [SEG]: generated by the LLM when mask output is requested and passed to the frozen OMG decoder to decode binary segmentation masks.

    For high-resolution inputs (1024×10241024 \times 1024), the visual encoder extracts features at 32×32\times downsampling, which are further downsampled to 64×64\times via a pixel shuffle operator, yielding exactly 256 pixel-centric visual tokens for the LLM.

  2. Knowl 2 — Perception Prior Embedding Formulation

    model/method

    To transfer segmentation priors from a frozen perception module into the LLM without retraining the perception backbone, OMG-LLaVA integrates object queries into dense image features via a non-trainable Perception Prior Embedding module.

    Let F∈RHW×CF \in \mathbb{R}^{HW \times C} denote the dense image feature map produced by the image encoder, where HWHW is the spatial token count and CC is the feature dimension. Let Q∈RNq×CQ \in \mathbb{R}^{N_q \times C} denote NqN_q object queries produced by the OMG decoder, accompanied by predicted segmentation masks M∈RNq×HWM \in \mathbb{R}^{N_q \times HW} and confidence score vector S∈R1×NqS \in \mathbb{R}^{1 \times N_q}.

    The per-pixel query mask score matrix MS∈RHW×NqMS \in \mathbb{R}^{HW \times N_q} is computed by soft-weighting masks by confidence scores across all object queries:

    MS=Softmax(M⊙S,dim=−1)MS = \text{Softmax}(M \odot S, \text{dim} = -1)

    where ⊙\odot represents element-wise broadcast multiplication, and the softmax normalization is evaluated over the query dimension.

    The pixel-centric visual tokens Tpv∈RHW×CT_{pv} \in \mathbb{R}^{HW \times C} are formed by adding the mask-score weighted object queries to the image features:

    Tpv=MS⋅Q+FT_{pv} = MS \cdot Q + F

    Foreground object queries from QQ are retained as object-centric visual tokens TovT_{ov}. The final visual token sequence TvT_v passed to the LLM visual projector is the concatenation:

    Tv=(Tpv,Tov)T_v = (T_{pv}, T_{ov})

  3. Knowl 3 — Visual Prompt Encoding via Attention Mask Constraints

    model/method

    OMG-LLaVA supports interactive user prompting via points, bounding boxes, and segmentation masks without converting all prompts into lossy point approximations or requiring separate vision prompt backbones.

    The OMG decoder employs alternating masked cross-attention and self-attention layers operating over learnable object queries and prompt queries:

    • Point Prompts: Encoded directly into visual prompt queries via prompt position embeddings and fed into the decoder.
    • Bounding Box Prompts: Spatial bounding box coordinates are converted into a binary spatial mask where all pixels outside the box are masked out (set to −∞-\infty in cross-attention). This restricts the cross-attention layers to query image features located strictly within the box.
    • Mask Prompts: The user-provided binary mask is applied directly as the attention mask within the masked cross-attention layers.

    This mechanism constrains the receptive field of masked cross-attention layers to the exact geometric prompt boundary, allowing the decoder to tokenize object-centric representations for downstream text generation.

  4. Knowl 4 — Two-Stage Training Objective and Regularization

    model/method

    OMG-LLaVA is trained using a two-stage optimization pipeline:

    1. Pre-training Stage: The image encoder, OMG decoder, and LLM are frozen. Only the visual projector PvP_v (mapping visual tokens to the LLM embedding space) and the text projector PtP_t (mapping LLM hidden states to the visual embedding space) are updated. In addition to standard autoregressive next-token language modeling loss Ltext\mathcal{L}_{text}, a cycle-consistency reconstruction penalty Lreg\mathcal{L}_{reg} preserves object-centric semantics in the visual projection space:

    Lpretrain=Ltext+Lreg\mathcal{L}_{pretrain} = \mathcal{L}_{text} + \mathcal{L}_{reg}

    Lreg=∥Tov−Pt(Pv(Tov))∥22\mathcal{L}_{reg} = \|T_{ov} - P_t(P_v(T_{ov}))\|_2^2

    where TovT_{ov} are the object-centric visual tokens.

    1. Instruction Tuning Stage: The image encoder and OMG decoder remain frozen. The visual and text projectors are fully fine-tuned, and the LLM is adapted using Low-Rank Adaptation (LoRA). The objective optimizes language generation and mask prediction decoded from the output [SEG] token hidden states:

    Linstruction=Ltext+Lmask\mathcal{L}_{instruction} = \mathcal{L}_{text} + \mathcal{L}_{mask}

    Lmask=αLCE+βLDICE\mathcal{L}_{mask} = \alpha \mathcal{L}_{CE} + \beta \mathcal{L}_{DICE}

    where LCE\mathcal{L}_{CE} is cross-entropy loss, LDICE\mathcal{L}_{DICE} is Dice loss, α=5\alpha = 5, and β=2\beta = 2.

  5. Knowl 5 — Ablation of Perception Prior Embedding and Object Query Input

    empirical result

    Ablation experiments quantify the necessity of the Perception Prior Embedding (M1) and direct object query token inputs (M2) on Referring Expression Segmentation (RES) and Grounded Conversation Generation (GCG). The baseline model (M0) directly couples frozen OMG-Seg and LLaVA without prior embeddings.

    Methods refCOCO refCOCO+ refCOCOg GCG
    cIoU gIoU cIoU gIoU cIoU gIoU METEOR mIoU
    Baseline (M0) 58.7 61.0 52.6 55.0 55.8 58.1 13.2 51.0
    + Perception prior embedding (M1) 72.5 74.3 63.2 65.4 67.8 70.6 13.6 62.1
    + Object query input (M2) 74.4 75.9 64.4 66.2 68.5 71.5 13.8 63.6

    Directly connecting a frozen perception module to an LLM without prior embedding (M0) results in poor segmentation (58.758.7 cIoU on refCOCO, 51.051.0 mIoU on GCG) because the LLM lacks perceptual priors to generate suitable decoding queries. Adding Perception Prior Embedding (M1) provides gains of +13.8+13.8 cIoU on refCOCO, +10.6+10.6 on refCOCO+, +11.7+11.7 on refCOCOg, and +11.1+11.1 mIoU on GCG. Providing foreground object queries directly as input tokens (M2) contributes an additional +1.9+1.9 cIoU on refCOCO and +1.5+1.5 mIoU on GCG.

  6. Knowl 6 — Referring Expression Segmentation Performance

    empirical result

    OMG-LLaVA was evaluated on the refCOCO, refCOCO+, and refCOCOg benchmarks using cumulative Intersection-over-Union (cIoU). Evaluated variants include OMG-LLaVA with a frozen OMG decoder and fine-tuned on referring expression datasets ("ft") with frozen or unfrozen decoders.

    Method Freeze Visual refCOCO refCOCO+ refCOCOg
    Decoder Encoder Val TestA TestB Val TestA TestB Val Test
    LISA ×\times 2 74.1 76.5 71.1 62.4 67.4 56.5 66.4 68.5
    LISA (ft) ×\times 2 74.9 79.1 72.3 65.1 70.8 58.1 67.9 70.6
    PixelLM ×\times 1 73.0 76.5 68.2 66.3 71.7 58.3 69.3 70.5
    GSVA (ft) ×\times 2 77.2 78.9 73.5 65.9 69.6 59.8 72.7 73.3
    OMG-LLaVA ✓\checkmark 1 75.6 77.7 71.2 65.6 69.7 58.9 70.7 70.2
    OMG-LLaVA (ft) ×\times 1 78.0 80.3 74.1 69.1 73.1 63.0 72.9 72.9
    OMG-LLaVA (ft) ✓\checkmark 1 77.2 79.8 74.1 68.7 73.0 61.6 71.7 71.9

    With a single visual encoder and a frozen segmentation decoder, OMG-LLaVA scores 75.675.6, 65.665.6, and 70.770.7 cIoU on the validation sets of refCOCO, refCOCO+, and refCOCOg, outperforming LISA (which uses 2 visual encoders and an unfrozen decoder) by 1.51.5, 3.23.2, and 4.34.3 cIoU. When fine-tuned with an unfrozen decoder, OMG-LLaVA reaches 78.078.0, 69.169.1, and 72.972.9 cIoU across the three validation sets.

  7. Knowl 7 — Grounded Conversation Generation Performance on GranDf

    empirical result

    Grounded Conversation Generation (GCG) requires producing complete scene descriptions interleaving text and localized instance masks. Evaluated on the GranDf dataset using METEOR and CIDEr for text quality and AP50 and mIoU for mask grounding quality ("ft" indicates fine-tuning on GranDf; †\dagger denotes pre-training on the full GranD dataset):

    Methods ft Visual Val Test
    Encoder METEOR CIDEr AP50 mIoU METEOR CIDEr AP50 mIoU
    Kosmos-2 ✓\checkmark 1 16.1 27.6 17.1 55.6 15.8 27.2 17.2 56.8
    LISA ✓\checkmark 2 13.0 33.9 25.2 62.0 12.9 32.2 24.8 61.7
    GLaMM†^\dagger ✓\checkmark 2 15.2 43.1 28.9 65.8 14.6 37.9 27.2 64.6
    OMG-LLaVA ×\times 1 13.8 36.2 26.9 64.6 13.5 33.1 26.1 62.8
    OMG-LLaVA ✓\checkmark 1 14.9 41.2 29.9 65.5 14.5 38.5 28.6 64.7

    OMG-LLaVA (ft) achieves 14.514.5 METEOR, 38.538.5 CIDEr, 28.628.6 AP50, and 64.764.7 mIoU on the GranDf test set using a single visual encoder. It outperforms LISA (using 2 encoders) by 1.61.6 METEOR, 6.36.3 CIDEr, 3.83.8 AP50, and 3.03.0 mIoU. Despite not pretraining on GranD, OMG-LLaVA outperforms GLaMM on test CIDEr (38.538.5 vs. 37.937.9), AP50 (28.628.6 vs. 27.227.2), and mIoU (64.764.7 vs. 64.664.6).

  8. Knowl 8 — Preservation of Image-Level Multimodal Capabilities

    empirical result

    Instruction tuning multimodal models on dense pixel-level tasks often induces catastrophic forgetting on standard image-level conversation and visual question answering benchmarks. OMG-LLaVA was benchmarked against existing grounding MLLMs across MME, MMBench, SEED-Bench, POPE, and AI2D:

    Method MME MMBench SEED-Bench POPE AI2D
    Training only with LLaVA dataset
    LLaVA 1.5 1422/267 68.5 65.9 86.7 56.6
    OMG-LLaVA 1448/282 67.5 68.9 89.7 61.7
    Co-training with LLaVA dataset and segmentation datasets
    LISA 1/1 0.4 - 0.0 0.0
    PixelLM 309/135 17.4 - 0.0 0.0
    LaSagnA 0/0 0.0 - 0.0 0.0
    GLaMM 14/9 36.8 - 0.94 28.2
    OMG-LLaVA 1177/235 47.9 56.5 80.0 42.9

    While co-trained models such as LISA, PixelLM, LaSagnA, and GLaMM experience catastrophic degradation on POPE (≤0.94\le 0.94), AI2D (≤28.2\le 28.2), and MME, OMG-LLaVA preserves balanced multimodal performance (1177/2351177/235 MME, 47.947.9 MMBench, 56.556.5 SEED-Bench, 80.080.0 POPE, and 42.942.9 AI2D).

  9. Knowl 9 — Ablations on Projector Architecture, Output Format, and Mask Hidden States

    empirical result

    Ablation studies isolate structural design decisions for the token projector, answer formatting, and segmentation embedding extraction:

    1. Object-Centric Projector Architecture: Using a shared linear MLP projector achieves 74.574.5 cIoU on refCOCO and 13.613.6 METEOR on refCOCOg(C). Introducing a cross-attention layer to the projector degrades refCOCO to 72.372.3 cIoU and refCOCOg(C) to 13.213.2 METEOR because cross-attention causes object-centric tokens to absorb excessive pixel-centric features. Using separate (unshared) MLPs for visual prompts and object queries drops performance to 72.372.3 cIoU and 13.113.1 METEOR.

    2. Segmentation Answer Formulation: Unifying segmentation responses under the flexible template "<p> Expression </p> [SEG]" yields 75.675.6 cIoU on refCOCO and 26.926.9 AP50 on GCG, whereas fixed templates like "Sure, it is [SEG]" (75.575.5 cIoU on refCOCO) impair the LLM's general instruction-following flexibility.

    3. Segmentation Hidden State Layer Selection: Extracting the [SEG] embedding strictly from the LLM's final layer hidden state attains 70.070.0 cIoU on refCOCOg. In contrast, averaging hidden states across all transformer layers drops refCOCOg performance to 68.768.7 cIoU, and concatenating all layer hidden states degrades performance to 62.362.3 cIoU.

  10. Knowl 10 — Limitations of OMG-LLaVA

    limitation

    OMG-LLaVA exhibits three specific structural limitations:

    1. Image-Level Benchmark Degradation under Co-training: Joint training on pixel-level segmentation data degrades general image-level conversational benchmark scores (e.g., MME performance decreases from 1448/2821448/282 under image-only training to 1177/2351177/235 under joint co-training).
    2. Inability to Perform Part-Level Segmentation: Constrained by the pre-trained OMG-Seg perception backbone, OMG-LLaVA is limited to object-level and stuff-level segmentation and cannot segment fine-grained sub-object parts.
    3. Absence of Spatio-Temporal Video Reasoning: Although the OMG-Seg perception backbone supports sequential video inputs, OMG-LLaVA is not trained on pixel-level spatio-temporal video instruction datasets, precluding video-grounded spatio-temporal reasoning.

Coverage note — Omitted qualitative visual response dialogues and descriptions of third-party datasets that followed established standard procedures without introducing new methodological designs.

References

  1. 1.Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219, 2024.
  2. 2.Ali Athar, Alexander Hermans, Jonathon Luiten, Deva Ramanan, and Bastian Leibe. Tarvis: A unified architecture for target-based video segmentation. In CVPR, 2023.
  3. 3.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966, 2023.
  4. 4.Amir Bar, Yossi Gandelsman, Trevor Darrell, Amir Globerson, and Alexei Efros. Visual prompting via image inpainting. In NeurIPS, 2022.
  5. 5.Adam Botach, Evgenii Zheltonozhskii, and Chaim Baskin. End-to-end referring video object segmentation with multimodal transformers. In CVPR, 2022.
  6. 6.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. In NeurIPS, 2020.
  7. 7.Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1209–1218, 2018.
  8. 8.Mu Cai, Haotian Liu, Siva Karthik Mustikovela, Gregory P. Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee. Making large multimodal models understand arbitrary visual prompts. In CVPR, 2024.
  9. 9.Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, et al. Internlm2 technical report. arXiv preprint arXiv:2403.17297, 2024.
  10. 10.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020.
  11. 11.Chi Chen, Ruoyu Qin, Fuwen Luo, Xiaoyue Mi, Peng Li, Maosong Sun, and Yang Liu. Position-enhanced visual instruction tuning for multimodal large language models. arXiv preprint arXiv:2308.13437, 2023.
  12. 12.Kai Chen, Jiangmiao Pang, Jiaqi Wang, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jianping Shi, Wanli Ouyang, Chen Change Loy, and Dahua Lin. Hybrid task cascade for instance segmentation. In CVPR, 2019.
  13. 13.Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. Shikra: Unleashing multimodal llm’s referential dialogue magic. arXiv preprint arXiv:2306.15195, 2023.
  14. 14.Lin Chen, Jisong Li, Xiaoyi Dong, Pan Zhang, Conghui He, Jiaqi Wang, Feng Zhao, and Dahua Lin. Sharegpt4v: Improving large multi-modal models with better captions. arXiv preprint arXiv:2311.12793, 2023.
  15. 15.Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, et al. Are we on the right way for evaluating large vision-language models? arXiv preprint arXiv:2403.20330, 2024.
  16. 16.Ting Chen, Saurabh Saxena, Lala Li, David J Fleet, and Geoffrey Hinton. Pix2seq: A language modeling framework for object detection. ICLR, 2022.
  17. 17.Wei-Ge Chen, Irina Spiridonova, Jianwei Yang, Jianfeng Gao, and Chunyuan Li. Llava-interactive: An all-in-one demo for image chat, segmentation, generation and editing. arXiv preprint arXiv:2311.00571, 2023.
  18. 18.Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. arXiv preprint arXiv:2404.16821, 2024.
  19. 19.Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In CVPR, 2022.
  20. 20.Bowen Cheng, Alexander G. Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation. In NeurIPS, 2021.
  21. 21.Yangming Cheng, Liulei Li, Yuanyou Xu, Xiaodi Li, Zongxin Yang, Wenguan Wang, and Yi Yang. Segment and track anything. arXiv preprint arXiv:2305.06558, 2023.
  22. 22.XTuner Contributors. Xtuner: A toolkit for efficiently fine-tuning llm. https://github.com/InternLM/xtuner, 2023.
  23. 23.Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, and Chen Change Loy. Mevis: A large-scale benchmark for video segmentation with motion expressions. In ICCV, 2023.
  24. 24.Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Xilin Wei, Songyang Zhang, Haodong Duan, Maosong Cao, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Xinyue Zhang, Wei Li, Jingwen Li, Kai Chen, Conghui He, Xingcheng Zhang, Yu Qiao, Dahua Lin, and Jiaqi Wang. Internlm-xcomposer2: Mastering free-form text-image composition and comprehension in vision-language large model. arXiv preprint arXiv:2401.16420, 2024.
  25. 25.Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang, Haodong Duan, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Zhe Chen, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Kai Chen, Conghui He, Xingcheng Zhang, Jifeng Dai, Yu Qiao, Dahua Lin, and Jiaqi Wang. Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd. arXiv preprint arXiv:2404.06512, 2024.
  26. 26.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. ICLR, 2021.
  27. 27.Zhongbin Fang, Xiangtai Li, Xia Li, Joachim M Buhmann, Chen Change Loy, and Mengyuan Liu. Explore in-context learning for 3d point cloud understanding. NeurIPS, 2024.
  28. 28.Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang, Meishan Zhang, Mong-Li Lee, and Wynne Hsu. Video-of-thought: Step-by-step video reasoning from perception to cognition. In ICML, 2024.
  29. 29.Hao Fei, Shengqiong Wu, Hanwang Zhang, Tat-Seng Chua, and Shuicheng Yan. Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing. arxiv preprint, 2024.
  30. 30.Hao Fei, Shengqiong Wu, Meishan Zhang, Min Zhang, Tat-Seng Chua, and Shuicheng Yan. Enhancing video-language representations with structural spatio-temporal alignment. PAMI, 2024.
  31. 31.Guang Feng, Zhiwei Hu, Lihe Zhang, and Huchuan Lu. Encoder fusion network with co-attention embedding for referring image segmentation. In CVPR, 2021.
  32. 32.Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, Mengdan Zhang, Xu Lin, Jinrui Yang, Xiawu Zheng, Ke Li, Xing Sun, Yunsheng Wu, and Rongrong Ji. Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394, 2023.
  33. 33.Xiuye Gu, Yin Cui, Jonathan Huang, Abdullah Rashwan, Xuan Yang, Xingyi Zhou, Golnaz Ghiasi, Weicheng Kuo, Huizhong Chen, Liang-Chieh Chen, et al. Dataseg: Taming a universal multi-dataset multi-task segmentation model. arXiv preprint arXiv:2306.01736, 2023.
  34. 34.Meng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu, Ming-Ming Cheng, and Shi-Min Hu. Segnext: Rethinking convolutional attention design for semantic segmentation. NeurIPS, 2022.
  35. 35.Junwen He, Yifan Wang, Lijun Wang, Huchuan Lu, Jun-Yan He, Jin-Peng Lan, Bin Luo, and Xuansong Xie. Multi-modal instruction tuned llms with fine-grained visual perception. arXiv preprint arXiv:2403.02969, 2024.
  36. 36.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In ICCV, 2017.
  37. 37.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  38. 38.Kuan-Chih Huang, Xiangtai Li, Lu Qi, Shuicheng Yan, and Ming-Hsuan Yang. Reason3d: Searching and reasoning 3d segmentation via large language model. arXiv preprint arXiv:2405.17427, 2024.
  39. 39.Drew A Hudson and Christopher D Manning. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In CVPR, 2019.
  40. 40.Louis Martin Hugo Touvron, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, and et al. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288, 2023.
  41. 41.Jitesh Jain, Jiachen Li, MangTik Chiu, Ali Hassani, Nikita Orlov, and Humphrey Shi. OneFormer: One Transformer to Rule Universal Image Segmentation. CVPR, 2023.
  42. 42.Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. Referitgame: Referring to objects in photographs of natural scenes. In EMNLP, pages 787–798, 2014.
  43. 43.Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is worth a dozen images. In ECCV, pages 235–251. Springer, 2016.
  44. 44.Anna Khoreva, Anna Rohrbach, and Bernt Schiele. Video object segmentation with language referring expressions. In ACCV, 2018.
  45. 45.Dahun Kim, Jun Xie, Huiyu Wang, Siyuan Qiao, Qihang Yu, Hong-Seok Kim, Hartwig Adam, In So Kweon, and Liang-Chieh Chen. Tubeformer-deeplab: Video mask transformer. In CVPR, 2022.
  46. 46.Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Dollár. Panoptic segmentation. In CVPR, 2019.
  47. 47.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. ICCV, 2023.
  48. 48.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. IJCV, 2020.
  49. 49.Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. Lisa: Reasoning segmentation via large language model. arXiv preprint arXiv:2308.00692, 2023.
  50. 50.Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yixiao Ge, and Ying Shan. Seed-bench: Benchmarking multimodal llms with generative comprehension. arXiv preprint arXiv:2307.16125, 2023.
  51. 51.Feng Li, Qing Jiang, Hao Zhang, Tianhe Ren, Shilong Liu, Xueyan Zou, Huaizhe Xu, Hongyang Li, Chunyuan Li, Jianwei Yang, et al. Visual in-context prompting. arXiv preprint arXiv:2311.13601, 2023.
  52. 52.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023.
  53. 53.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML, 2022.
  54. 54.Xiangtai Li, Henghui Ding, Wenwei Zhang, Haobo Yuan, Guangliang Cheng, Pang Jiangmiao, Kai Chen, Ziwei Liu, and Chen Change Loy. Transformer-based visual segmentation: A survey. arXiv pre-print, 2023.
  55. 55.Xiang Li, Jinglu Wang, Xiaohao Xu, Xiao Li, Bhiksha Raj, and Yan Lu. Robust referring video object segmentation with cyclic structural consensus. In ICCV, 2023.
  56. 56.Xiangtai Li, Shilin Xu, Yibo Yang, Guangliang Cheng, Yunhai Tong, and Dacheng Tao. Panoptic-partformer: Learning a unified model for panoptic part segmentation. In ECCV, 2022.
  57. 57.Xiangtai Li, Ansheng You, Zeping Zhu, Houlong Zhao, Maoke Yang, Kuiyuan Yang, and Yunhai Tong. Semantic flow for fast and accurate scene parsing. In ECCV, 2020.
  58. 58.Xiangtai Li, Haobo Yuan, Wei Li, Henghui Ding, Size Wu, Wenwei Zhang, Yining Li, Kai Chen, and Chen Change Loy. Omg-seg: Is one model good enough for all segmentation? CVPR, 2024.
  59. 59.Xiangtai Li, Haobo Yuan, Wenwei Zhang, Guangliang Cheng, Jiangmiao Pang, and Chen Change Loy. Tube-link: A flexible cross tube baseline for universal video segmentation. In ICCV, 2023.
  60. 60.Xiangtai Li, Li Zhang, Guangliang Cheng, Kuiyuan Yang, Yunhai Tong, Xiatian Zhu, and Tao Xiang. Global aggregation then local distribution for scene parsing. IEEE TIP, 2021.
  61. 61.Xiangtai Li, Wenwei Zhang, Jiangmiao Pang, Kai Chen, Guangliang Cheng, Yunhai Tong, and Chen Change Loy. Video k-net: A simple, strong, and unified baseline for video segmentation. In CVPR, 2022.
  62. 62.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. EMNLP, 2023.
  63. 63.Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. 2023.
  64. 64.Yanwei Li, Chengyao Wang, and Jiaya Jia. Llama-vid: An image is worth 2 tokens in large language models. arXiv preprint arXiv:2311.17043, 2023.
  65. 65.Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Yixin Chen, Ruihang Chu, Shaoteng Liu, and Jiaya Jia. Mini-gemini: Mining the potential of multi-modality vision language models. arXiv preprint arXiv:2403.18814, 2024.
  66. 66.Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan. Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122, 2023.
  67. 67.Weifeng Lin, Xinyu Wei, Ruichuan An, Peng Gao, Bocheng Zou, Yulin Luo, Siyuan Huang, Shanghang Zhang, and Hongsheng Li. Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want. arXiv preprint arXiv:2403.20271, 2024.
  68. 68.Chang Liu, Henghui Ding, and Xudong Jiang. GRES: Generalized referring expression segmentation. In CVPR, 2023.
  69. 69.Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning, 2023.
  70. 70.Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.
  71. 71.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In NeurIPS, 2023.
  72. 72.Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, et al. Llava-plus: Learning to use tools for creating multimodal agents. arXiv preprint arXiv:2311.05437, 2023.
  73. 73.Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023.
  74. 74.Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281, 2023.
  75. 75.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. CVPR, 2022.
  76. 76.Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. Learn to explain: Multimodal reasoning via thought chains for science question answering. In The 36th Conference on Neural Information Processing Systems (NeurIPS), 2022.
  77. 77.Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Liujuan Cao, Chenglin Wu, Cheng Deng, and Rongrong Ji. Multi-task collaborative network for joint referring expression comprehension and segmentation. In CVPR, 2020.
  78. 78.Chuofan Ma, Yi Jiang, Jiannan Wu, Zehuan Yuan, and Xiaojuan Qi. Groma: Localized visual tokenization for grounding multimodal large language models. arXiv preprint arXiv:2404.13013, 2024.
  79. 79.Sachin Mehta and Mohammad Rastegari. Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer. In ICLR, 2022.
  80. 80.Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 3DV, 2016.
  81. 81.Ting Pan, Lulu Tang, Xinlong Wang, and Shiguang Shan. Tokenize anything via prompting. arXiv preprint arXiv:2312.09128, 2023.
  82. 82.Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023.
  83. 83.Lu Qi, Yi-Wen Chen, Lehan Yang, Tiancheng Shen, Xiangtai Li, Weidong Guo, Yu Xu, and Ming-Hsuan Yang. Generalizable entity grounding via assistance of large language model. arXiv preprint arXiv:2402.02555, 2024.
  84. 84.Lu Qi, Jason Kuen, Tiancheng Shen, Jiuxiang Gu, Weidong Guo, Jiaya Jia, Zhe Lin, and Ming-Hsuan Yang. High-quality entity segmentation. In ICCV, 2023.
  85. 85.Lu Qi, Jason Kuen, Yi Wang, Jiuxiang Gu, Hengshuang Zhao, Philip Torr, Zhe Lin, and Jiaya Jia. Open world entity segmentation. TPAMI, 2022.
  86. 86.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021.
  87. 87.Hanoona Rasheed, Muhammad Maaz, Sahal Shaji, Abdelrahman Shaker, Salman Khan, Hisham Cholakkal, Rao M. Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S. Khan. Glamm: Pixel grounding large multimodal model. CVPR, 2024.
  88. 88.Zhongwei Ren, Zhicheng Huang, Yunchao Wei, Yao Zhao, Dongmei Fu, Jiashi Feng, and Xiaojie Jin. Pixellm: Pixel reasoning with large multimodal model. arXiv preprint arXiv:2312.02228, 2023.
  89. 89.Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. In CVPR, 2019.
  90. 90.Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiying Yu, Zhengxiong Luo, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang, et al. Generative multimodal models are in-context learners. arXiv:2312.13286, 2023.
  91. 91.Shuyang Sun, Weijun Wang, Andrew Howard, Qihang Yu, Philip Torr, and Liang-Chieh Chen. Remax: Relaxing for better training on efficient panoptic segmentation. NeurIPS, 2023.
  92. 92.Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023.
  93. 93.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. arXiv:2302.13971, 2023.
  94. 94.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, 2017.
  95. 95.Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Max-deeplab: End-to-end panoptic segmentation with mask transformers. In CVPR, 2021.
  96. 96.Junchi Wang and Lei Ke. Llm-seg: Bridging image segmentation and large language model reasoning. arXiv preprint arXiv:2404.08767, 2024.
  97. 97.Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, and Jifeng Dai. Visionllm: Large language model is also an open-ended decoder for vision-centric tasks. arXiv preprint arXiv:2305.11175, 2023.
  98. 98.Xinshun Wang, Zhongbin Fang, Xia Li, Xiangtai Li, and Mengyuan Liu. Skeleton-in-context: Unified skeleton sequence modeling with in-context learning. CVPR, 2024.
  99. 99.Xinlong Wang, Wen Wang, Yue Cao, Chunhua Shen, and Tiejun Huang. Images speak in images: A generalist painter for in-context visual learning. In CVPR, 2023.
  100. 100.Xinlong Wang, Xiaosong Zhang, Yue Cao, Wen Wang, Chunhua Shen, and Tiejun Huang. Seggpt: Segmenting everything in context. In ICCV, 2023.
  101. 101.Zhendong Wang, Yifan Jiang, Yadong Lu, Yelong Shen, Pengcheng He, Weizhu Chen, Zhangyang Wang, and Mingyuan Zhou. In-context learning unlocked for diffusion models. arXiv preprint arXiv:2305.01115, 2023.
  102. 102.Cong Wei, Haoxian Tan, Yujie Zhong, Yujiu Yang, and Lin Ma. Lasagna: Language-based segmentation assistant for complex queries. arXiv preprint arXiv:2404.08506, 2024.
  103. 103.Jianzong Wu, Xiangtai Li, Xia Li, Henghui Ding, Yunhai Tong, and Dacheng Tao. Towards robust referring image segmentation. IEEE-TIP, 2024.
  104. 104.Jianzong Wu, Xiangtai Li, Chenyang Si, Shangchen Zhou, Jingkang Yang, Jiangning Zhang, Yining Li, Kai Chen, Yunhai Tong, Ziwei Liu, et al. Towards language-driven video inpainting via multimodal large language models. CVPR, 2024.
  105. 105.Jianzong Wu, Xiangtai Li, Shilin Xu, Haobo Yuan, Henghui Ding, Yibo Yang, Xia Li, Jiangning Zhang, Yunhai Tong, Xudong Jiang, Bernard Ghanem, and Dacheng Tao. Towards open vocabulary learning: A survey. arXiv pre-print, 2023.
  106. 106.Shengqiong Wu, Hao Fei, Xiangtai Li, Jiayi Ji, Hanwang Zhang, Tat-Seng Chua, and Shuicheng Yan. Towards semantic equivalence of tokenization in multimodal llm. arXiv preprint arXiv:2406.05127, 2024.
  107. 107.Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. Next-gpt: Any-to-any multimodal llm. ICML, 2024.
  108. 108.Tsung-Han Wu, Giscard Biamby, David Chan, Lisa Dunlap, Ritwik Gupta, Xudong Wang, Joseph E Gonzalez, and Trevor Darrell. See, say, and segment: Teaching lmms to overcome false premises. arXiv preprint arXiv:2312.08366, 2023.
  109. 109.Zhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan, Shiji Song, and Gao Huang. Gsva: Generalized segmentation via multimodal large language models. arXiv preprint arXiv:2312.10103, 2023.
  110. 110.Ruyi Xu, Yuan Yao, Zonghao Guo, Junbo Cui, Zanlin Ni, Chunjiang Ge, Tat-Seng Chua, Zhiyuan Liu, Maosong Sun, and Gao Huang. Llava-uhd: an lmm perceiving any aspect ratio and high-resolution images. arXiv preprint arXiv:2403.11703, 2024.
  111. 111.Shilin Xu, Haobo Yuan, Qingyu Shi, Lu Qi, Jingbo Wang, Yibo Yang, Yining Li, Kai Chen, Yunhai Tong, Bernard Ghanem, Xiangtai Li, and Ming-Hsuan Yang. Rap-sam: Towards real-time all-purpose segment anything. arXiv preprint, 2024.
  112. 112.Bin Yan, Yi Jiang, Jiannan Wu, Dong Wang, Zehuan Yuan, Ping Luo, and Huchuan Lu. Universal instance perception as object discovery and retrieval. In CVPR, 2023.
  113. 113.An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfeng Xue, Na Ni, Pei Zhang, Peng Wang, Ru Peng, Rui Men, Ruize Gao, Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Yang Fan, Yang Yao, Yichang Zhang, Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zhihao Fan. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2024.
  114. 114.Senqiao Yang, Tianyuan Qu, Xin Lai, Zhuotao Tian, Bohao Peng, Shu Liu, and Jiaya Jia. An improved baseline for reasoning segmentation with large language model. arXiv preprint arXiv:2312.17240, 2023.
  115. 115.Zongxin Yang, Jiaxu Miao, Yunchao Wei, Wenguan Wang, Xiaohan Wang, and Yi Yang. Scalable video object segmentation with identification mechanism. TPAMI, 2024.
  116. 116.Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Hengshuang Zhao, and Philip HS Torr. Lavt: Language-aware vision transformer for referring image segmentation. In CVPR, 2022.
  117. 117.Zongxin Yang, Yunchao Wei, and Yi Yang. Collaborative video object segmentation by multi-scale foreground-background integration. TPAMI, 2021.
  118. 118.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. mplug-owl: Modularization empowers large language models with multimodality, 2023.
  119. 119.Haoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du, Bowen Zhang, Zirui Wang, Liangliang Cao, Shih-Fu Chang, and Yinfei Yang. Ferret: Refer and ground anything anywhere at any granularity. ICLR, 2024.
  120. 120.Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg. Mattnet: Modular attention network for referring expression comprehension. In CVPR, 2018.
  121. 121.Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. Modeling context in referring expressions. In ECCV, pages 69–85. Springer, 2016.
  122. 122.Qihang Yu, Huiyu Wang, Dahun Kim, Siyuan Qiao, Maxwell Collins, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Cmt-deeplab: Clustering mask transformers for panoptic segmentation. In CVPR, 2022.
  123. 123.Qihang Yu, Huiyu Wang, Siyuan Qiao, Maxwell Collins, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. k-means mask transformer. In ECCV, 2022.
  124. 124.Qihang Yu, Huiyu Wang, Siyuan Qiao, Maxwell Collins, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. k-means Mask Transformer. In ECCV, 2022.
  125. 125.Haobo Yuan, Xiangtai Li, Chong Zhou, Yining Li, Kai Chen, and Chen Change Loy. Open-vocabulary sam: Segment and recognize twenty-thousand classes interactively. arXiv preprint, 2024.
  126. 126.Yuqian Yuan, Wentong Li, Jian Liu, Dongqi Tang, Xinjie Luo, Chi Qin, Lei Zhang, and Jianke Zhu. Osprey: Pixel understanding with visual instruction tuning. arXiv preprint arXiv:2312.10032, 2023.
  127. 127.Yuhang Zang, Wei Li, Jun Han, Kaiyang Zhou, and Chen Change Loy. Contextual object detection with multimodal large language models. arXiv preprint arXiv:2305.18279, 2023.
  128. 128.Ao Zhang, Liming Zhao, Chen-Wei Xie, Yun Zheng, Wei Ji, and Tat-Seng Chua. Next-chat: An lmm for chat, detection and segmentation. arXiv preprint arXiv:2311.04498, 2023.
  129. 129.Hao Zhang, Hongyang Li, Feng Li, Tianhe Ren, Xueyan Zou, Shilong Liu, Shijia Huang, Jianfeng Gao, Lei Zhang, Chunyuan Li, and Jianwei Yang. Llava-grounding: Grounded visual chat with large multimodal models, 2023.
  130. 130.Hang Zhang, Xin Li, and Lidong Bing. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858, 2023.
  131. 131.Pan Zhang, Xiaoyi Dong, Bin Wang, Yuhang Cao, Chao Xu, Linke Ouyang, Zhiyuan Zhao, Shuangrui Ding, Songyang Zhang, Haodong Duan, Wenwei Zhang, Hang Yan, Xinyue Zhang, Wei Li, Jingwen Li, Kai Chen, Conghui He, Xingcheng Zhang, Yu Qiao, Dahua Lin, and Jiaqi Wang. Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition. arXiv preprint arXiv:2309.15112, 2023.
  132. 132.Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, and Yu Qiao. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199, 2023.
  133. 133.Shilong Zhang, Peize Sun, Shoufa Chen, Min Xiao, Wenqi Shao, Wenwei Zhang, Kai Chen, and Ping Luo. Gpt4roi: Instruction tuning large language model on region-of-interest. arXiv preprint arXiv:2307.03601, 2023.
  134. 134.Yichi Zhang, Ziqiao Ma, Xiaofeng Gao, Suhaila Shakiah, Qiaozi Gao, and Joyce Chai. Groundhog: Grounding large language models to holistic segmentation. In CVPR, pages 14227–14238, 2024.
  135. 135.Zheng Zhang, Yeyao Ma, Enming Zhang, and Xiang Bai. Psalm: Pixelwise segmentation with large multi-modal model. arXiv preprint arXiv:2403.14598, 2024.
  136. 136.Xiangyu Zhao, Xiangtai Li, Haodong Duan, Haian Huang, Yining Li, Kai Chen, and Hua Yang. Towards semantic equivalence of tokenization in multimodal llm. arXiv preprint, 2024.
  137. 137.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Semantic understanding of scenes through the ADE20K dataset. CVPR, 2017.
  138. 138.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. In IJCV, 2022.
  139. 139.Qianyu Zhou, Zhengyang Feng, Qiqi Gu, Jiangmiao Pang, Guangliang Cheng, Xuequan Lu, Jianping Shi, and Lizhuang Ma. Context-aware mixup for domain adaptive semantic segmentation. IEEE TCSVT, 2023.
  140. 140.Qianyu Zhou, Xiangtai Li, Lu He, Yibo Yang, Guangliang Cheng, Yunhai Tong, Lizhuang Ma, and Dacheng Tao. Transvod: End-to-end video object detection with spatial-temporal transformers. TPAMI, 2022.
  141. 141.Yikang Zhou, Tao Zhang, Shunping Ji, Shuicheng Yan, and Xiangtai Li. Dvis-daq: Improving video segmentation via dynamic anchor queries. arXiv preprint arXiv:2404.00086, 2024.
  142. 142.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023.

Citation

MLA
Zhang, T., et al. “OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 71737–67, https://proceedings.neurips.cc/paper_files/paper/2024/file/83eb86be3e2f9fd66c44d9073c51ba4d-Paper-Conference.pdf.
APA
Zhang, T., Li, X., Fei, H., Yuan, H., Wu, S., Ji, S., Loy, C. C., & Yan, S. (2024). OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding. Advances in Neural Information Processing Systems, 37, 71737–71767. https://proceedings.neurips.cc/paper_files/paper/2024/file/83eb86be3e2f9fd66c44d9073c51ba4d-Paper-Conference.pdf
Chicago
Zhang, T., X. Li, H. Fei, et al. 2024. “OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding”. Advances in Neural Information Processing Systems 37: 71737–67. https://proceedings.neurips.cc/paper_files/paper/2024/file/83eb86be3e2f9fd66c44d9073c51ba4d-Paper-Conference.pdf.
Harvard
Zhang, T. et al. (2024) “OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 71737–71767. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/83eb86be3e2f9fd66c44d9073c51ba4d-Paper-Conference.pdf.
Vancouver
1. Zhang T, Li X, Fei H, Yuan H, Wu S, Ji S, Loy CC, Yan S (2024) OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 71737–71767

BibTeX

@inproceedings{zhang2024omg,
  title = {OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding},
  author = {Zhang, Tao and Li, Xiangtai and Fei, Hao and Yuan, Haobo and Wu, Shengqiong and Ji, Shunping and Loy, Chen Change and Yan, Shuicheng},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {71737-71767},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/83eb86be3e2f9fd66c44d9073c51ba4d-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors