GSVA: Generalized Segmentation via Multimodal Large Language Models

Zhuofan XiaDongchen HanYizeng HanXuran PanShiji SongGao Huang

article2024CVPR209 citations

Presents a multimodal framework that enables language models to segment multiple target objects simultaneously and explicitly reject non-existent targets using dedicated tokens, setting a new state of the art on the gRefCOCO benchmark.

Listen

Real-world artificial intelligence applications, such as robotic navigation and embodied visual assistants, require computer vision systems to interpret complex natural language instructions and accurately locate referenced objects within images. Traditional referring expression segmentation methods operate under the strict assumption that every language prompt maps to exactly one target present in the visual scene. In practice, however, users frequently refer to multiple objects simultaneously or describe items that do not exist in the scene. When existing models encounter non-existent objects, they often force a segmentation mask onto unrelated visual areas, creating substantial operational and safety risks for autonomous systems.

The main objective of the article is to demonstrate a new framework, named Generalized Segmentation Vision Assistant (GSVA), designed to resolve Generalized Referring Expression Segmentation (GRES). The article evaluates whether combining multimodal large language models with specialized segmentation decoders can accurately segment multiple targets from a single prompt and explicitly reject non-existent targets.

To achieve this, the authors designed a unified vision-language architecture that pairs a multimodal large language model (based on Vicuna or LLaMA-2 backbones with visual encoders) with a segmentation foundation model (the Segment Anything Model). The approach introduces two primary mechanisms: generating multiple shared-weight segmentation tokens—each preceded by its corresponding descriptive text prompt to maintain context—and outputting a dedicated rejection token when a referent is absent from an image. These rejection tokens assign empty masks immediately without forcing the visual decoder to segment nonexistent areas. The authors evaluated GSVA against prior baselines, such as LISA and specialized non-LLM architectures, across standard benchmarks including the gRefCOCO dataset (comprising 278,232 expressions across 19,994 images) and classic datasets such as RefCOCO, RefCOCO+, and RefCOCOg.

The experimental findings show significant performance gains. First, GSVA established new state-of-the-art results on the gRefCOCO generalized segmentation benchmark; the fine-tuned 13-billion parameter version achieved an average mask accuracy of 70.04% on the validation set, outperforming specialized baselines and prior multimodal language models. Second, GSVA demonstrated a dramatic improvement in identifying absent targets, achieving null-target classification accuracies between 60% and 67%, whereas standard models without rejection capabilities often scored below 10% before task-specific tuning. Third, ablation experiments revealed that removing the rejection token reduced null-target identification accuracy by more than 25% and overall mask accuracy by about 10%. Finally, the system transferred effectively to classic single-target segmentation and bounding-box comprehension tasks, consistently outperforming comparable baseline models across various evaluation splits.

These findings imply that large multimodal models can successfully navigate complex spatial relationships and visual reasoning without requiring rigid, one-to-one prompt constraints. Explicit rejection tokens significantly reduce hallucination and false-positive mask generations, providing a safer, more reliable foundation for robotic manipulation and vision-guided automation where incorrect actions on non-existent objects carry high operational risk.

Stakeholders and developers deploying vision-language systems for autonomous platforms should integrate explicit rejection tokens and multi-entity prompting mechanisms into their multimodal architectures to mitigate hallucination risks. While the findings provide high confidence on benchmark datasets, future initiatives should evaluate the framework's real-time computational latency on embedded hardware and test its robustness across complex, real-world robotic deployments.

Cover for GSVA: Generalized Segmentation via Multimodal Large Language Models

Abstract

Generalized Referring Expression Segmentation (GRES) extends the scope of classic RES to refer to multiple objects in one expression or identify the empty targets absent in the image. GRES poses challenges in modeling the complex spatial relationships of the instances in the image and identifying non-existing referents. Multimodal Large Language Models (MLLMs) have recently shown tremendous progress in these complicated vision-language tasks. Connecting Large Language Models (LLMs) and vision models, MLLMs are proficient in understanding contexts with visual inputs. Among them, LISA, as a representative, adopts a special [SEG] token to prompt a segmentation mask decoder, e.g., SAM, to enable MLLMs in the RES task. However, existing solutions to GRES remain unsatisfactory since current segmentation MLLMs cannot correctly handle the cases where users might reference multiple subjects in a singular prompt or provide descriptions incongruent with any image target. In this paper, we propose Generalized Segmentation Vision Assistant (GSVA) to address this gap. Specifically, GSVA reuses the [SEG] token to prompt the segmentation model towards supporting multiple mask references simultaneously and innovatively learns to generate a [REJ] token to reject the null targets explicitly. Experiments validate GSVA’s efficacy in resolving the GRES issue, marking a notable enhancement and setting a new record on the GRES benchmark gRefCOCO dataset. GSVA also proves effective across various classic referring segmentation and comprehension tasks. Code is available at https://github.com/LeapLabTHU/GSVA.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Generalized Segmentation Vision Assistant
  • 3.1. Model Architecture
  • 3.2. GRES: Task and Challenges
  • 3.3. Multiple [SEG] Tokens for Multiple Targets
  • 3.4. Rejecting Empty Targets via [REJ] Tokens
  • 4. Experiments
  • 4.1. GRES
  • 4.2. Referring Expression Segmentation
  • 4.3. Referring Expression Comprehension
  • 4.4. Ablation Study
  • 4.5. Visualization
  • 5. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Architecture of Generalized Segmentation Vision Assistant

    model/method

    Generalized Segmentation Vision Assistant (GSVA) integrates a Multimodal Large Language Model (MLLM) with a query-based Segmentation Foundation Model (SFM) to perform Generalized Referring Expression Segmentation (GRES).

    The MLLM comprises a vision encoder FV1F_{V1} (CLIP-ViT-L/14 operating on images ximg∈Rh×w×3x_{\text{img}} \in \mathbb{R}^{h \times w \times 3} resized to h=w=224h = w = 224), a linear projection layer ϕ\phi, and a decoder-based language model FLLMF_{\text{LLM}} (Vicuna-7B, Vicuna-13B, or LLaMA-2-13B). The vision encoder and projection layer produce visual token embeddings:

    himg=ϕ(FV1(ximg))∈Rnimg×dh_{\text{img}} = \phi(F_{V1}(x_{\text{img}})) \in \mathbb{R}^{n_{\text{img}} \times d}

    where the visual token sequence length is nimg=hw142=256n_{\text{img}} = \frac{hw}{14^2} = 256, and dd is the language model hidden dimension (d=4096d = 4096 for 7B models and d=5120d = 5120 for 13B models). Input text prompts xtxtx_{\text{txt}} are tokenized by tokenizer T\mathcal{T} into htxt=T(xtxt)h_{\text{txt}} = \mathcal{T}(x_{\text{txt}}). Prepending fixed prompt tokens hprompth_{\text{prompt}}, the concatenated sequence [hprompt∥himg∥htxt][h_{\text{prompt}} \parallel h_{\text{img}} \parallel h_{\text{txt}}] is processed autoregressively by FLLMF_{\text{LLM}} to generate output token embeddings y~txt\tilde{y}_{\text{txt}}:

    y~txt=FLLM([hprompt∥himg∥htxt])\tilde{y}_{\text{txt}} = F_{\text{LLM}}([h_{\text{prompt}} \parallel h_{\text{img}} \parallel h_{\text{txt}}])

    The SFM is instantiated using the Segment Anything Model (SAM) with a ViT-H backbone. A frozen vision encoder FV2F_{V2} encodes high-resolution input images (H=W=1024H = W = 1024) into dense features:

    fseg=FV2(ximg)∈RH16×W16×Cf_{\text{seg}} = F_{V2}(x_{\text{img}}) \in \mathbb{R}^{\frac{H}{16} \times \frac{W}{16} \times C}

    with channel dimension C=256C = 256. For each valid referent, the LLM emits a [SEG][\text{SEG}] token. The hidden state corresponding to each [SEG][\text{SEG}] token, y~txt[SEG]\tilde{y}_{\text{txt}}[\text{SEG}], is projected into the SFM prompt space via a multi-layer perceptron (MLP) projector ψ\psi:

    hseg=ψ(y~txt[SEG])∈RNseg×Ch_{\text{seg}} = \psi(\tilde{y}_{\text{txt}}[\text{SEG}]) \in \mathbb{R}^{N_{\text{seg}} \times C}

    A trainable mask decoder FmaskF_{\text{mask}} decodes segmentation masks y~mask\tilde{y}_{\text{mask}} from the query embeddings hsegh_{\text{seg}} conditioned on dense features fsegf_{\text{seg}}:

    y~mask=Fmask(hseg∣fseg)\tilde{y}_{\text{mask}} = F_{\text{mask}}(h_{\text{seg}} \mid f_{\text{seg}})

    where each query in hsegh_{\text{seg}} decodes into an individual object mask in y~mask\tilde{y}_{\text{mask}}.

  2. Knowl 2 — Referent-Guided Shared-Weight [SEG] Tokens for Multi-Target Segmentation

    model/method

    In conventional segmentation MLLMs (such as LISA), the model relies on a single [SEG][\text{SEG}] token. When an instruction references multiple objects simultaneously, forcing a single [SEG][\text{SEG}] token requires one query embedding to correspond to multiple distinct entities, causing mask distortion and localization errors.

    GSVA supports multi-target segmentation by generating multiple weight-sharing [SEG][\text{SEG}] tokens within a single autoregressive response sequence. To prevent semantic ambiguity and assign each [SEG][\text{SEG}] token to its corresponding visual instance, each [SEG][\text{SEG}] token is explicitly prefixed with the textual description of the individual referent. For a prompt requesting kk targets {obj1,…,objk}\{\text{obj}_1, \dots, \text{obj}_k\}, the conversation prompt is structured as:

    • User: What are {obj_1}, {obj_2}, ..., {obj_k} in this image? Please output the segmentation masks.
    • Assistant: Sure, {obj_1}:[SEG], {obj_2}:[SEG], ..., {obj_k}:[SEG].

    Prepending {obji}\{\text{obj}_i\} functions as an implicit multimodal In-Context Learning (ICL) cue and dynamic conditioning mechanism. It guides the autoregressive LLM to ground the specific entity in the visual scene and inject its location-specific representation into the succeeding [SEG][\text{SEG}] token's hidden state y~txt[SEGi]\tilde{y}_{\text{txt}}[\text{SEG}_i].

  3. Knowl 3 — Explicit Empty-Target Rejection via the [REJ] Token

    model/method

    In Generalized Referring Expression Segmentation (GRES), instructions frequently describe objects that do not exist in the visual scene due to absent entities, invalid attributes, or inaccurate spatial locations. Standard segmentation MLLMs invariably emit a [SEG][\text{SEG}] token and invoke the mask decoder, leading to false-positive hallucinations on irrelevant image regions.

    GSVA incorporates a dedicated rejection token, [REJ][\text{REJ}], into the vocabulary of the Multimodal Large Language Model. When an entity requested in the user prompt is absent from the image, the MLLM generates the rejection token formatted as {obji}:[REJ]\{\text{obj}_i\}:[\text{REJ}]. For example, when {obj1}\{\text{obj}_1\} and {obj2}\{\text{obj}_2\} are absent and {obj3}\{\text{obj}_3\} is present:

    • User: What are {obj_1} (absent), {obj_2} (absent), {obj_3} in this image? Please output the segmentation masks.
    • Assistant: Sure, {obj_1}:[REJ], {obj_2}:[REJ], {obj_3}:[SEG].

    During inference, all [REJ][\text{REJ}] tokens are discarded from the query sequence passed to the segmentation foundation model and directly assigned all-zero negative masks. This unburdens the segmentation foundation model's mask decoder from learning empty-target classification and delegates spatial non-existence verification entirely to the multimodal reasoning capabilities of the language model.

  4. Knowl 4 — Generalized Referring Expression Segmentation Metrics: gIoU, cIoU, and N-acc.

    definition

    Generalized Referring Expression Segmentation (GRES) relaxes classical referring segmentation assumptions by supporting multiple targets and empty targets per expression. Performance on GRES benchmarks is evaluated using three metrics:

    1. Generalized Intersection-over-Union (gIoUgIoU): The per-expression Intersection-over-Union averaged across all expressions in the dataset. For non-empty expressions, IoU=∣Y∩Y^∣∣Y∪Y^∣IoU = \frac{|Y \cap \hat{Y}|}{|Y \cup \hat{Y}|} where YY is the ground-truth mask and Y^\hat{Y} is the predicted mask. For an empty-target expression where Y=∅Y = \emptyset, IoUIoU is set to 1.01.0 if the model correctly predicts an all-negative mask (Y^=∅\hat{Y} = \emptyset), and 0.00.0 if the model predicts any positive pixels (Y^≠∅\hat{Y} \neq \emptyset).

    2. Cumulative Intersection-over-Union (cIoUcIoU): The ratio of total intersection area to total union area accumulated globally across the dataset:

    cIoU=∑i∣Yi∩Y^i∣∑i∣Yi∪Y^i∣cIoU = \frac{\sum_{i} |Y_i \cap \hat{Y}_i|}{\sum_{i} |Y_i \cup \hat{Y}_i|}

    Correctly classified empty-target samples do not contribute to either the numerator or the denominator. Misclassified empty-target samples (false positives) add their predicted mask area ∣Y^i∣|\hat{Y}_i| to the denominator.

    1. No-Target Accuracy (N-acc.N\text{-acc.}): The classification accuracy on the subset of empty-target referring expressions:

    N-acc.=Number of correctly identified empty targetsTotal number of ground-truth empty targetsN\text{-acc.} = \frac{\text{Number of correctly identified empty targets}}{\text{Total number of ground-truth empty targets}}

  5. Knowl 5 — GRES Benchmark Evaluation on gRefCOCO

    data/table

    Performance of GSVA and baseline models on the gRefCOCO dataset across the Validation Set, Test Set A, and Test Set B. Models are evaluated with zero-shot pretraining (pretrained for 50,000 steps on a mixed dataset including gRefCOCO) and after fine-tuning (denoted by (ft)) for 10 epochs on the gRefCOCO training split.

    Method Validation Set Test Set A Test Set B
    gIoU cIoU N-acc. gIoU cIoU N-acc. gIoU cIoU N-acc.
    MattNet 48.24 47.51 41.15 59.30 58.66 44.04 46.14 45.33 41.32
    LTS 52.70 52.30 - 62.64 61.87 - 50.42 49.96 -
    VLT 52.00 52.51 47.17 63.20 62.19 48.74 50.88 50.52 47.82
    CRIS 56.27 55.34 - 63.42 63.82 - 51.79 51.04 -
    LAVT 58.40 57.64 49.32 65.90 65.32 49.25 55.83 55.04 48.46
    ReLA 63.60 62.42 56.37 70.03 69.26 59.02 61.02 59.88 58.40
    LISA-Vicuna-7B 32.21 38.72 2.71 48.54 52.55 6.37 39.65 44.79 5.00
    GSVA-Vicuna-7B 63.32 61.70 56.45 70.11 69.23 63.50 61.34 60.26 58.42
    LISA-Vicuna-7B (ft) 61.63 61.76 54.67 66.27 68.50 50.01 58.84 60.63 51.91
    GSVA-Vicuna-7B (ft) 66.47 63.29 62.43 71.08 69.93 65.31 62.23 60.47 60.56
    LISA-Vicuna-13B 32.73 39.85 3.66 48.76 53.62 4.89 39.49 45.35 4.41
    GSVA-Vicuna-13B 62.97 60.18 58.44 67.17 67.59 54.60 58.06 57.28 52.22
    LISA-Vicuna-13B (ft) 63.45 62.99 55.25 68.18 69.65 52.16 61.84 62.24 56.15
    GSVA-Vicuna-13B (ft) 68.01 64.05 65.36 71.75 70.51 67.25 63.83 61.28 63.11
    LISA-Llama2-13B 33.26 39.64 3.27 49.76 53.80 7.28 40.49 45.41 5.73
    GSVA-Llama2-13B 63.20 62.38 54.51 69.52 69.86 57.84 62.06 60.77 58.30
    LISA-Llama2-13B (ft) 65.24 63.96 57.49 69.99 71.00 55.43 62.11 62.29 56.34
    GSVA-Llama2-13B (ft) 70.04 66.38 66.02 73.29 72.79 64.72 65.45 63.20 62.47

    Without fine-tuning, LISA fails on GRES due to near-zero empty target accuracy (N-acc.≤7.28%N\text{-acc.} \le 7.28\%), achieving only 32.21%32.21\% to 33.26%33.26\% gIoUgIoU on the validation set. In contrast, GSVA achieves 63.32%63.32\% gIoUgIoU zero-shot (comparable to the fully-supervised non-LLM model ReLA at 63.60%63.60\%). After fine-tuning on gRefCOCO, GSVA-Llama2-13B achieves 70.04%70.04\% gIoUgIoU and 66.02%66.02\% N-acc.N\text{-acc.} on the validation set, outperforming fine-tuned LISA-Llama2-13B by 4.80%4.80\% in gIoUgIoU and 8.53%8.53\% in N-acc.N\text{-acc.}.

  6. Knowl 6 — Ablation of GSVA Architectural Components

    empirical result

    The contributions of referent description prefixes, multiple [SEG][\text{SEG}] tokens, and the [REJ][\text{REJ}] token are ablated on the gRefCOCO validation set using the GSVA-Vicuna-7B model.

    RefExp.+[SEG] Multiple [SEG] [REJ] Token gIoU cIoU N-acc.
    ✓ ✓ ✓ 63.32 61.70 56.45
    ✓ ✓ 51.57 60.95 30.32
    ✓ 44.86 59.37 11.96
    32.21 38.72 2.71
    ✓ ✓ 21.83 27.22 0.00
    1. Impact of the [REJ][\text{REJ}] Token: Removing the [REJ][\text{REJ}] token while retaining multiple [SEG][\text{SEG}] tokens with referent expressions leads to a 26.13%26.13\% absolute drop in N-acc.N\text{-acc.} (from 56.45%56.45\% to 30.32%30.32\%) and an 11.75%11.75\% absolute decline in gIoUgIoU (from 63.32%63.32\% to 51.57%51.57\%).
    2. Impact of Multiple [SEG][\text{SEG}] Tokens: Restricting the model to a single [SEG][\text{SEG}] token further decreases gIoUgIoU by 6.71%6.71\% (to 44.86%44.86\%) and N-acc.N\text{-acc.} by 18.36%18.36\% (to 11.96%11.96\%).
    3. Impact of Referent Description Hints: Removing the referent expressions before output tokens (i.e., generating plain [SEG], [SEG] or [REJ] sequences) causes null-target identification to completely fail (N-acc.=0.00%N\text{-acc.} = 0.00\%) and severely degrades segmentation quality (gIoU=21.83%gIoU = 21.83\%, cIoU=27.22%cIoU = 27.22\%). This demonstrates that explicit entity description prefixes are critical for binding individual [SEG][\text{SEG}] tokens to their corresponding visual referents.
  7. Knowl 7 — Performance on Classical Referring Expression Segmentation Benchmarks

    data/table

    Performance comparison on classical single-target Referring Expression Segmentation (RES) datasets: RefCOCO (UNC split), RefCOCO+ (UNC split), and RefCOCOg (UMD split). Performance is reported using cumulative Intersection-over-Union (cIoU,%cIoU, \%) across validation and test splits before and after joint fine-tuning ((ft)) on the three RES training sets for 10 epochs.

    Method RefCOCO (UNC) RefCOCO+ (UNC) RefCOCOg (UMD)
    Val Test-A Test-B Val Test-A Test-B Val Test
    MCN 62.4 64.2 59.7 50.6 55.0 44.7 49.2 49.4
    VLT 67.5 70.5 65.2 56.3 61.0 50.1 55.0 57.7
    CRIS 70.5 73.2 66.1 62.3 68.1 53.7 59.9 60.4
    LAVT 72.7 75.8 68.8 62.1 68.4 55.1 61.2 62.1
    ReLA 73.8 76.5 70.2 66.0 71.0 57.7 65.0 66.0
    PolyFormer-L 76.0 78.3 73.3 69.3 74.6 61.9 69.2 70.2
    LISA-Vicuna-7B 74.1 76.5 71.1 62.4 67.4 56.5 66.4 68.5
    GSVA-Vicuna-7B 76.4 77.4 72.8 64.5 67.7 58.6 71.1 72.0
    LISA-Vicuna-7B (ft) 74.9 79.1 72.3 65.1 70.8 58.1 67.9 70.6
    GSVA-Vicuna-7B (ft) 77.2 78.9 73.5 65.9 69.6 59.8 72.7 73.3
    LISA-Vicuna-13B 71.7 74.7 68.1 59.4 64.2 52.9 65.2 66.1
    GSVA-Vicuna-13B 74.6 77.5 70.5 62.5 66.5 55.5 69.6 71.2
    LISA-Vicuna-13B (ft) 76.0 78.8 72.9 65.0 70.2 58.1 69.5 70.5
    GSVA-Vicuna-13B (ft) 78.2 80.4 74.2 67.4 71.5 60.9 74.2 75.6
    LISA-Llama2-13B 73.4 76.2 69.5 62.3 66.6 56.3 68.2 68.5
    GSVA-Llama2-13B 77.7 79.9 74.2 68.0 71.5 61.5 73.2 73.9
    LISA-Llama2-13B (ft) 76.3 78.7 72.4 66.2 71.0 59.3 70.1 71.1
    GSVA-Llama2-13B (ft) 79.2 81.7 77.1 70.3 73.8 63.6 75.7 77.0

    GSVA consistently outperforms corresponding LISA baselines across all datasets and splits. Fine-tuned GSVA-Llama2-13B achieves state-of-the-art cIoUcIoU scores: 79.2%79.2\% on RefCOCO Val, 70.3%70.3\% on RefCOCO+ Val, and 75.7%75.7\% on RefCOCOg Val (outperforming LISA-Llama2-13B (ft) by 2.9%2.9\%, 4.1%4.1\%, and 5.6%5.6\%, respectively).

  8. Knowl 8 — Referring Expression Comprehension via Mask-Derived Bounding Boxes

    data/table

    Evaluation of GSVA transferred zero-shot to Referring Expression Comprehension (REC) on RefCOCO, RefCOCO+, and RefCOCOg. Bounding boxes are directly extracted from the predicted segmentation masks without any bounding box supervision during training. The evaluation metric is [email protected]\text{[email protected]} (the percentage of predictions where the bounding box IoU with ground truth is ≥0.5\ge 0.5).

    Method RefCOCO RefCOCO+ RefCOCOg
    Val Test-A Test-B Val Test-A Test-B Val Test
    u-LLaVA-7B (LoRA) 83.47 87.13 80.21 68.74 76.32 60.98 76.19 78.24
    u-LLaVA-7B (full-ft) 86.04 89.47 82.26 74.09 81.16 66.61 79.87 81.68
    LISA-Vicuna-7B 78.68 81.72 75.74 62.92 68.93 56.49 70.10 72.47
    GSVA-Vicuna-7B 85.50 88.01 82.49 70.21 75.62 65.11 79.00 79.21
    LISA-Vicuna-7B (ft) 85.39 88.84 82.59 74.23 79.46 68.40 79.34 80.42
    GSVA-Vicuna-7B (ft) 86.27 89.22 83.77 72.81 78.78 68.01 81.58 81.83
    LISA-Vicuna-13B 80.01 83.26 76.26 63.77 70.24 57.42 71.79 73.34
    GSVA-Vicuna-13B 83.12 87.01 80.54 68.14 73.90 62.00 77.08 78.89
    LISA-Vicuna-13B (ft) 85.92 89.05 83.16 74.86 81.08 68.87 80.09 81.48
    GSVA-Vicuna-13B (ft) 87.71 90.49 84.57 76.52 81.69 70.35 83.90 84.85
    LISA-Llama2-13B 82.52 85.56 78.82 67.91 73.77 62.25 75.37 76.83
    GSVA-Llama2-13B 86.99 89.54 84.08 73.89 79.10 69.38 80.68 82.07
    LISA-Llama2-13B (ft) 85.91 88.84 81.73 74.46 80.56 68.26 80.09 81.27
    GSVA-Llama2-13B (ft) 89.16 92.08 87.17 79.74 84.45 73.41 85.47 86.18

    Without fine-tuning, GSVA-Vicuna-7B surpasses LISA-Vicuna-7B by over 6.8%6.8\% on RefCOCO Val and over 8.9%8.9\% on RefCOCOg Val. Fine-tuned GSVA-Llama2-13B sets the highest performance across all splits (89.16%89.16\% on RefCOCO Val, 79.74%79.74\% on RefCOCO+ Val, and 85.47%85.47\% on RefCOCOg Val), outperforming the fully fine-tuned u-LLaVA-7B baseline.

Coverage note — None was omitted; all contributed models, prompting mechanisms, GRES metrics, experimental evaluations across GRES, RES, and REC, and ablation studies are fully preserved.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. In NeurIPS, 2022. 2
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. In NeurIPS, 2020. 2, 5
  3. 3.Ding-Jie Chen, Songhao Jia, Yi-Chen Lo, Hwann-Tzong Chen, and Tyng-Luh Liu. See-through-text grouping for referring image segmentation. In IEEE ICCV, 2019. 2
  4. 4.Yixin Chen, Qing Li, Deqian Kong, Yik Lun Kei, Song-Chun Zhu, Tao Gao, Yixin Zhu, and Siyuan Huang. Yourefit: Embodied reference understanding with language and gesture. In IEEE ICCV, 2021. 2, 4
  5. 5.Ming-Ming Cheng, Shuai Zheng, Wen-Yan Lin, Vibhav Vineet, Paul Sturgess, Nigel Crook, Niloy J Mitra, and Philip Torr. Imagespirit: Verbal guided image parsing. ACM ToG, 2014. 1, 2
  6. 6.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023. 2, 3
  7. 7.Yong Xien Chng, Henry Zheng, Yizeng Han, Xuchong Qiu, and Gao Huang. Mask grounding for referring image segmentation. In IEEE CVPR, 2024. 2
  8. 8.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022. 2
  9. 9.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards generalpurpose vision-language models with instruction tuning. In NeurIPS, 2023. 2, 5
  10. 10.Henghui Ding, Chang Liu, Suchen Wang, and Xudong Jiang. Vision-language transformer and query generation for referring segmentation. In IEEE ICCV, 2021. 6, 7
  11. 11.Xiaohan Ding, Xiangyu Zhang, Ningning Ma, Jungong Han, Guiguang Ding, and Jian Sun. Repvgg: Making vgg-style convnets great again. In IEEE ICCV, 2021. 2
  12. 12.Xiaohan Ding, Xiangyu Zhang, Jungong Han, and Guiguang Ding. Scaling up your kernels to 31x31: Revisiting large kernel design in cnns. In IEEE CVPR, 2022. 2
  13. 13.Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. A survey for in-context learning. arXiv preprint arXiv:2301.00234, 2022. 5
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. 2
  15. 15.Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. Eva: Exploring the limits of masked visual representation learning at scale. In IEEE CVPR, 2023. 2
  16. 16.Guang Feng, Zhiwei Hu, Lihe Zhang, and Huchuan Lu. Encoder fusion network with co-attention embedding for referring image segmentation. In IEEE CVPR, 2021. 2
  17. 17.Chen Gao, Jinyu Chen, Si Liu, Luting Wang, Qiong Zhang, and Qi Wu. Room-and-object aware knowledge reasoning for remote embodied referring expression. In IEEE CVPR, 2021. 1, 4
  18. 18.Dongchen Han, Tianzhu Ye, Yizeng Han, Zhuofan Xia, Shiji Song, and Gao Huang. Agent attention: On the integration of softmax and linear attention. arXiv preprint arXiv:2312.08874, 2023. 2
  19. 19.Yizeng Han, Gao Huang, Shiji Song, Le Yang, Honghui Wang, and Yulin Wang. Dynamic neural networks: A survey. IEEE TPAMI, 2021. 2
  20. 20.Yizeng Han, Gao Huang, Shiji Song, Le Yang, Yitian Zhang, and Haojun Jiang. Spatially adaptive feature refinement for efficient inference. IEEE TIP, 2021. 2
  21. 21.Yizeng Han, Yifan Pu, Zihang Lai, Chaofei Wang, Shiji Song, Junfeng Cao, Wenhui Huang, Chao Deng, and Gao Huang. Learning to weight samples for dynamic earlyexiting networks. In ECCV, 2022.
  22. 22.Yizeng Han, Zhihang Yuan, Yifan Pu, Chenhao Xue, Shiji Song, Guangyu Sun, and Gao Huang. Latency-aware spatialwise dynamic networks. In NeurIPS, 2022. 2
  23. 23.Yizeng Han, Dongchen Han, Zeyu Liu, Yulin Wang, Xuran Pan, Yifan Pu, Chao Deng, Junlan Feng, Shiji Song, and Gao Huang. Dynamic perceiver for efficient visual recognition. In IEEE ICCV, 2023. 2
  24. 24.Yizeng Han, Zeyu Liu, Zhihang Yuan, Yifan Pu, Chaofei Wang, Shiji Song, and Gao Huang. Latency-aware unified dynamic networks for efficient image recognition. arXiv preprint arXiv:2308.15949, 2023. 2
  25. 25.Ronghang Hu, Marcus Rohrbach, and Trevor Darrell. Segmentation from natural language expressions. In ECCV, 2016. 1, 2
  26. 26.Yutao Hu, Qixiong Wang, Wenqi Shao, Enze Xie, Zhenguo Li, Jungong Han, and Ping Luo. Beyond one-to-one: Rethinking the referring image segmentation. In IEEE ICCV, 2023. 2
  27. 27.Gao Huang, Yulin Wang, Kangchen Lv, Haojun Jiang, Wenhui Huang, Pengfei Qi, and Shiji Song. Glance and focus networks for dynamic visual recognition. IEEE TPAMI, 2022. 2
  28. 28.Rui Huang, Xuran Pan, Henry Zheng, Haojun Jiang, Zhifeng Xie, Cheng Wu, Shiji Song, and Gao Huang. Joint representation learning for text and 3d point cloud. Pattern Recognition, 2024. 1
  29. 29.Ya Jing, Tao Kong, Wei Wang, Liang Wang, Lei Li, and Tieniu Tan. Locate then segment: A strong pipeline for referring image segmentation. In IEEE CVPR, 2021. 6
  30. 30.Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. Referitgame: Referring to objects in photographs of natural scenes. In EMNLP, 2014. 1, 4, 6, 7
  31. 31.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In IEEE ICCV, 2023. 2, 3
  32. 32.Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. Lisa: Reasoning segmentation via large language model. arXiv preprint arXiv:2308.00692, 2023. 1, 2, 3, 4, 5, 6, 7, 8
  33. 33.Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Jingkang Yang, Chunyuan Li, and Ziwei Liu. Mimicit: Multi-modal in-context instruction tuning. arXiv preprint arXiv:2306.05425, 2023. 5
  34. 34.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, 2023. 2, 5
  35. 35.Ruiyu Li, Kaican Li, Yi-Chun Kuo, Michelle Shu, Xiaojuan Qi, Xiaoyong Shen, and Jiaya Jia. Referring image segmentation via recurrent refinement networks. In IEEE CVPR, 2018. 2
  36. 36.Ziyi Lin, Chris Liu, Renrui Zhang, Peng Gao, Longtian Qiu, Han Xiao, Han Qiu, Chen Lin, Wenqi Shao, Keqin Chen, et al. Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models. arXiv preprint arXiv:2311.07575, 2023. 2
  37. 37.Chenxi Liu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, and Alan Yuille. Recurrent multimodal interaction for referring image segmentation. In IEEE CVPR, 2017. 2
  38. 38.Chang Liu, Henghui Ding, and Xudong Jiang. Gres: Generalized referring expression segmentation. In IEEE CVPR, 2023. 1, 2, 4, 5, 6, 7, 8
  39. 39.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In NeurIPS, 2023. 2, 3, 5
  40. 40.Jiang Liu, Hui Ding, Zhaowei Cai, Yuting Zhang, Ravi Kumar Satzoda, Vijay Mahadevan, and R Manmatha. Polyformer: Referring image segmentation as sequential polygon generation. In IEEE CVPR, 2023. 7
  41. 41.Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023. 2
  42. 42.Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Liujuan Cao, Chenglin Wu, Cheng Deng, and Rongrong Ji. Multi-task collaborative network for joint referring expression comprehension and segmentation. In IEEE CVPR, 2020. 7
  43. 43.Tengchao Lv, Yupan Huang, Jingye Chen, Lei Cui, Shuming Ma, Yaoyao Chang, Shaohan Huang, Wenhui Wang, Li Dong, Weiyao Luo, et al. Kosmos-2.5: A multimodal literate model. arXiv preprint arXiv:2309.11419, 2023. 2
  44. 44.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In IEEE CVPR, 2016. 2, 4, 6, 7
  45. 45.Khanh Nguyen, Debadeepta Dey, Chris Brockett, and Bill Dolan. Vision-based navigation with language-based assistance via imitation learning with indirect intervention. In IEEE CVPR, 2019. 4
  46. 46.Zanlin Ni, Yulin Wang, Jiangwei Yu, Haojun Jiang, Yue Cao, and Gao Huang. Deep incubation: Training large models by divide-and-conquering. In IEEE CVPR, 2023. 2
  47. 47.Zanlin Ni, Yulin Wang, Renping Zhou, Jiayi Guo, Jinyi Hu, Zhiyuan Liu, Shiji Song, Yuan Yao, and Gao Huang. Revisiting non-autoregressive transformers for efficient image synthesis. In CVPR, 2024. 2
  48. 48.OpenAI. Gpt-4 technical report, 2023. 2
  49. 49.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. In NeurIPS, 2022. 2
  50. 50.Xuran Pan, Chunjiang Ge, Rui Lu, Shiji Song, Guanfu Chen, Zeyi Huang, and Gao Huang. On the integration of selfattention and convolution. In IEEE CVPR, 2022. 2
  51. 51.Xuran Pan, Tianzhu Ye, Zhuofan Xia, Shiji Song, and Gao Huang. Slide-transformer: Hierarchical vision transformer with local self-attention. In IEEE CVPR, 2023. 2
  52. 52.Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023. 2, 5
  53. 53.Yifan Pu, Yiru Wang, Zhuofan Xia, Yizeng Han, Yulin Wang, Weihao Gan, Zidong Wang, Shiji Song, and Gao Huang. Adaptive rotated convolution for rotated object detection. In IEEE ICCV, 2023. 2
  54. 54.Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. Reverie: Remote embodied visual referring expression in real indoor environments. In IEEE CVPR, 2020. 1, 4
  55. 55.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021. 2, 3
  56. 56.Hanoona Rasheed, Muhammad Maaz, Sahal Shaji, Abdelrahman Shaker, Salman Khan, Hisham Cholakkal, Rao M Anwer, Erix Xing, Ming-Hsuan Yang, and Fahad S Khan. Glamm: Pixel grounding large multimodal model. arXiv preprint arXiv:2311.03356, 2023. 2
  57. 57.Hengcan Shi, Hongliang Li, Fanman Meng, and Qingbo Wu. Key-word-aware network for referring expression image segmentation. In ECCV, 2018. 2
  58. 58.Qie Sima, Sinan Tan, Huaping Liu, Fuchun Sun, Weifeng Xu, and Ling Fu. Embodied referring expression for manipulation question answering in interactive environment. In IEEE ICRA, 2023. 1, 4
  59. 59.Yan Tai, Weichen Fan, Zhao Zhang, Feng Zhu, Rui Zhao, and Ziwei Liu. Link-context learning for multimodal llms. arXiv preprint arXiv:2308.07891, 2023. 5
  60. 60.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model, 2023. 2
  61. 61.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  62. 62.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023. 2
  63. 63.Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al. Image as a foreign language: Beit pretraining for vision and visionlanguage tasks. In IEEE CVPR, 2023. 2
  64. 64.Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, et al. Visionllm: Large language model is also an openended decoder for vision-centric tasks. In NeurIPS, 2023. 2
  65. 65.Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, et al. Cogvlm: Visual expert for pretrained language models. arXiv preprint arXiv:2311.03079, 2023. 2
  66. 66.Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. Reinforced cross-modal matching and selfsupervised imitation learning for vision-language navigation. In IEEE CVPR, 2019. 1, 4
  67. 67.Yulin Wang, Zhaoxi Chen, Haojun Jiang, Shiji Song, Yizeng Han, and Gao Huang. Adaptive focus for efficient video recognition. In IEEE ICCV, 2021. 2
  68. 68.Yulin Wang, Rui Huang, Shiji Song, Zeyi Huang, and Gao Huang. Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition. In NeurIPS, 2021.
  69. 69.Yulin Wang, Yang Yue, Yuanze Lin, Haojun Jiang, Zihang Lai, Victor Kulikov, Nikita Orlov, Humphrey Shi, and Gao Huang. Adafocus v2: End-to-end training of spatial dynamic networks for video recognition. In IEEE CVPR, 2022.
  70. 70.Yulin Wang, Yang Yue, Xinhong Xu, Ali Hassani, Victor Kulikov, Nikita Orlov, Shiji Song, Humphrey Shi, and Gao Huang. Adafocusv3: On unified spatial-temporal dynamic video recognition. In ECCV, 2022. 2
  71. 71.Yulin Wang, Yang Yue, Rui Lu, Tianjiao Liu, Zhao Zhong, Shiji Song, and Gao Huang. Efficienttrain: Exploring generalized curriculum learning for training visual backbones. In IEEE ICCV, 2023. 2
  72. 72.Zhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao, Yandong Guo, Mingming Gong, and Tongliang Liu. Cris: Clip-driven referring image segmentation. In IEEE CVPR, 2022. 6, 7
  73. 73.Zhichao Wei, Xiaohao Chen, Mingqiang Chen, and Siyu Zhu. Learning aligned cross-modal representations for referring image segmentation. arXiv preprint arXiv:2301.06429, 2023. 2
  74. 74.Penghao Wu and Saining Xie. V*: Guided visual search as a core mechanism in multimodal llms. arXiv preprint arXiv:2312.14135, 2023. 2
  75. 75.Tsung-Han Wu, Giscard Biamby, David Chan, Lisa Dunlap, Ritwik Gupta, Xudong Wang, Joseph E Gonzalez, and Trevor Darrell. See, say, and segment: Teaching lmms to overcome false premises. arXiv preprint arXiv:2312.08366, 2023. 2
  76. 76.Zhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li, and Gao Huang. Vision transformer with deformable attention. In IEEE CVPR, 2022. 2
  77. 77.Zhuofan Xia, Xuran Pan, Xuan Jin, Yuan He, Hui Xue, Shiji Song, and Gao Huang. Budgeted training for vision transformer. In ICLR, 2023. 2
  78. 78.Zhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li, and Gao Huang. Dat++: Spatially dynamic vision transformer with deformable attention. arXiv preprint arXiv:2309.01430, 2023. 2
  79. 79.Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, and Lu Yuan. Florence-2: Advancing a unified representation for a variety of vision tasks. arXiv preprint arXiv:2311.06242, 2023. 2
  80. 80.Jinjin Xu, Liwu Xu, Yuzhe Yang, Xiang Li, Yanchun Xie, Yi-Jie Huang, and Yaqian Li. u-llava: Unifying multimodal tasks via large language model. arXiv preprint arXiv:2311.05348, 2023. 2, 6, 7
  81. 81.Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Hengshuang Zhao, and Philip HS Torr. Lavt: Language-aware vision transformer for referring image segmentation. In IEEE CVPR, 2022. 2, 6, 7
  82. 82.Linwei Ye, Mrigank Rochan, Zhi Liu, and Yang Wang. Cross-modal self-attention network for referring image segmentation. In IEEE CVPR, 2019. 2
  83. 83.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178, 2023. 2
  84. 84.Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. Modeling context in referring expressions. In ECCV, 2016. 4, 5
  85. 85.Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg. Mattnet: Modular attention network for referring expression comprehension. In IEEE CVPR, 2018. 6
  86. 86.Ao Zhang, Liming Zhao, Chen-Wei Xie, Yun Zheng, Wei Ji, and Tat-Seng Chua. Next-chat: An lmm for chat, detection and segmentation. arXiv preprint arXiv:2311.04498, 2023. 2
  87. 87.Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199, 2023. 2
  88. 88.Shilong Zhang, Peize Sun, Shoufa Chen, Min Xiao, Wenqi Shao, Wenwei Zhang, Kai Chen, and Ping Luo. Gpt4roi: Instruction tuning large language model on region-of-interest. arXiv preprint arXiv:2307.03601, 2023. 5
  89. 89.Zicheng Zhang, Yi Zhu, Jianzhuang Liu, Xiaodan Liang, and Wei Ke. Coupalign: Coupling word-pixel with sentencemask alignments for referring image segmentation. In NeurIPS, 2022. 2
  90. 90.Haozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma, Kaikai An, Liang Chen, Zixuan Liu, Sheng Wang, Wenjuan Han, and Baobao Chang. Mmicl: Empowering vision-language model with multi-modal in-context learning. arXiv preprint arXiv:2309.07915, 2023. 5
  91. 91.Yang Zhao, Zhijie Lin, Daquan Zhou, Zilong Huang, Jiashi Feng, and Bingyi Kang. Bubogpt: Enabling visual grounding in multi-modal llms. arXiv preprint arXiv:2307.08581, 2023. 5
  92. 92.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023. 5
  93. 93.Xueyan Zou, Zi-Yi Dou, Jianwei Yang, Zhe Gan, Linjie Li, Chunyuan Li, Xiyang Dai, Harkirat Behl, Jianfeng Wang, Lu Yuan, et al. Generalized decoding for pixel, image, and language. In IEEE CVPR, 2023. 2, 7
  94. 94.Xueyan Zou, Jianwei Yang, Hao Zhang, Feng Li, Linjie Li, Jianfeng Gao, and Yong Jae Lee. Segment everything everywhere all at once. In NeurIPS, 2023. 2, 7

Citation

MLA
Xia, Z., et al. “GSVA: Generalized Segmentation via Multimodal Large Language Models”. arXiv, 2023, http://arxiv.org/abs/2312.10103v3.
APA
Xia, Z., Han, D., Han, Y., Pan, X., Song, S., & Huang, G. (2023). GSVA: Generalized Segmentation via Multimodal Large Language Models. arXiv. http://arxiv.org/abs/2312.10103v3
Chicago
Xia, Z., D. Han, Y. Han, X. Pan, S. Song, and G. Huang. 2023. “GSVA: Generalized Segmentation via Multimodal Large Language Models”. arXiv. http://arxiv.org/abs/2312.10103v3.
Harvard
Xia, Z. et al. (2023) “GSVA: Generalized Segmentation via Multimodal Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.10103v3.
Vancouver
1. Xia Z, Han D, Han Y, Pan X, Song S, Huang G (2023) GSVA: Generalized Segmentation via Multimodal Large Language Models. arXiv

BibTeX

@article{xia2023gsva,
  title = {GSVA: Generalized Segmentation via Multimodal Large Language Models},
  author = {Xia, Zhuofan and Han, Dongchen and Han, Yizeng and Pan, Xuran and Song, Shiji and Huang, Gao},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.10103v3},
  eprint = {2312.10103}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE