NExT-Chat: An LMM for Chat, Detection and Segmentation

Ao ZhangYuan YaoWei JiZhiyuan LiuTat-Seng Chua

article2024ICML96 citations

Proposes a pixel-to-embedding framework that unifies conversational understanding, bounding-box detection, and pixel-level segmentation within a single multimodal language model by decoding learned location representations into multiple spatial formats.

Listen

Large multimodal models have shown rapid progress in describing images and answering visual queries. However, most existing frameworks understand images only at a holistic level and struggle to accurately pinpoint, segment, or describe specific regions. Prior efforts to enable region-level comprehension typically convert bounding box coordinates into text tokens. This text-based method introduces significant computational overhead—requiring dozens of tokens per bounding box—and cannot naturally output pixel-level segmentation masks needed for advanced visual applications.

The article demonstrates a new paradigm called "pixel-to-embedding" (pix2emb) and introduces NExT-Chat, a unified multimodal model capable of conversational interaction, object detection, and image segmentation. The core objective is to evaluate whether representing spatial locations as continuous numerical representations (embeddings) rather than discrete text tokens enables an AI system to efficiently process location inputs and generate both bounding boxes and segmentation masks within a single conversational architecture.

To achieve this, the authors built NExT-Chat using a standard visual encoder paired with a language model, augmented with lightweight location encoders and decoders. Instead of generating word tokens for coordinates, the model outputs a single trigger token whose internal hidden state is fed directly into a box decoder or a mask decoder based on the Segment Anything Model. The framework utilizes a three-stage training process: pre-training on bounding box conversations, instruction fine-tuning, and a fast three-hour stage adapting the model to generate segmentation masks. To ensure that input locations and output predictions align seamlessly, the authors introduced a cycle-consistency training loss that ties the location encoder and decoder together. Experiments were conducted across established benchmarks for referring expression segmentation, referring expression comprehension, region captioning, and image hallucination diagnosis.

The evaluation produced several key findings. First, in referring expression segmentation, NExT-Chat achieved an average intersection-over-union score of 71.3 (reaching up to 80.3 when fine-tuned), outperforming comparable systems like LISA (67.9) despite using an order of magnitude fewer mask annotations—only 127,000 masks compared to millions in baseline training sets. Second, in region captioning, NExT-Chat achieved a CIDEr score of 79.6 (increasing to 114.0 upon fine-tuning), significantly outperforming models like Kosmos-2 (62.3). Third, the approach achieved massive computational efficiency: representing a bounding box requires only two tokens and a single added vocabulary item, making it up to 169 times more computationally efficient than standard coordinate text representations. Finally, the model demonstrated strong reliability against generating false information, scoring 87.7% accuracy on the random split of the POPE hallucination benchmark.

These findings indicate that treating location as continuous embeddings rather than text sequences dramatically reduces computational cost, lowers training data requirements, and unifies diverse visual tasks into one framework. By cutting down the sequence lengths needed to ground objects, organizations can deploy conversational models with spatial awareness at lower inference costs and lower latency. Furthermore, the ability to train segmentation modules in just three hours allows rapid task adaptation without catastrophic forgetting of core conversational skills.

Based on these results, organizations deploying visual AI assistants should adopt embedding-based spatial modeling over text-coordinate methods to improve throughput and support fine-grained mask outputs. Before deploying into specialized operational settings, teams should conduct domain-specific fine-tuning and implement content moderation filters, as the model was trained primarily on open-domain data and can occasionally hallucinate ungrounded facts. For future technical development, the authors note that dynamic weighting between text and detection loss functions should be explored to optimize spatial regression. Confidence in the results is high across general benchmark settings, though caution is warranted when applying the current system to multi-image inputs or specialized domains such as satellite and medical imaging.

Cover for NExT-Chat: An LMM for Chat, Detection and Segmentation

Abstract

The development of large language models (LLMs) has greatly advanced the field of multimodal understanding, leading to the emergence of large multimodal models (LMMs). In order to enhance visual comprehension, recent studies have equipped LMMs with region-level understanding capabilities by representing object bounding box coordinates as a series of text sequences (pix2seq). In this paper, we introduce a novel paradigm for object location modeling called the pix2emb method, where we ask the LMM to output the location embeddings and then decode them with different decoders. This paradigm allows us to use different location formats (such as bounding boxes and masks) in multimodal conversations. Leveraging the proposed pix2emb method, we train an LMM named NExT-Chat and demonstrate its capability of handling multiple tasks like visual grounding, region captioning, and grounded reasoning. Comprehensive experiments show the effectiveness of our NExT-Chat on various tasks, e.g., NExT-Chat (87.7) vs. Shikra (86.9) on POPE-Random, NExT-Chat (71.3) vs. LISA (67.9) on referring expression segmentation task, and NExT-Chat (79.6) vs. Kosmos-2 (62.3) on region caption task.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Method
  • 3.1. LMM Architecture
  • 3.2. Pix2Emb Method
  • 3.3. Training Process
  • 4. Qualitative Results
  • 5. Experiment
  • 5.1. Hallucination
  • 5.2. Referring Expression Segmentation
  • 5.3. Referring Expression Comprehension
  • 5.4. Region Captioning
  • 6. Ablation Study
  • 7. Conclusion
  • Impact Statement
  • References
  • A. Pix2Emb v.s. Pix2Seq
  • B. Additional Ablation Studies
  • D. Limitations
  • E. Additional Qualitative Results
  • C. Detection v.s. Segmentation

Knowls

  1. Knowl 1 — Pix2emb represents locations as embeddings rather than coordinate text

    model/method

    Pix2emb is a location-modeling paradigm in which object locations are represented by embeddings that can be encoded from location inputs and decoded into different output formats. For output, a new <trigger> token marks a localization request; its hidden state supplies an embedding to a location decoder, rather than asking the language model to generate coordinate strings as ordinary text. For input, a location encoder maps a bounding box into an embedding that can be supplied to the language model. This allows the same representation to support bounding-box regression and mask prediction. The paper does not train a mask encoder for user-provided mask inputs; such inputs are converted to boxes.

    The paper compares pix2emb with three pix2seq variants. Pix2emb uses regression, two sequence positions per box (the <trigger> and the location embedding), and one added vocabulary token. The 4-bin classification variant uses 6 tokens and adds 224 vocabulary tokens; the 2-bin variant uses 4 tokens and adds 1,024; the textual-number variant uses 26 tokens and adds none. The authors note that the lower token count can reduce the self-attention cost of processing boxes.

  2. Knowl 2 — Cycle consistency aligns location input and output embeddings

    model/method

    NExT-Chat uses a location encoder GG to map an input bounding box bb to an embedding, and a location decoder FF to map an embedding to a box. The paper adds a cycle-consistency loss so the input encoder and output decoder learn compatible representations. In the second cycle, tt is the hidden-state embedding of a generated <trigger>; F(t)F(t) predicts a box, which GG then re-encodes.

    Lcyc=L1(b,F(G(b)))+L2(t,G(F(t))),\mathcal{L}_{\mathrm{cyc}}=\mathcal{L}_1\bigl(b,F(G(b))\bigr)+\mathcal{L}_2\bigl(t,G(F(t))\bigr),

    Here, bb is a provided bounding box, FF is the box decoder, GG is the box encoder, and L1\mathcal{L}_1 and L2\mathcal{L}_2 denote the paper's coordinate-space L1L_1 loss and embedding-space L2L_2 loss, respectively. The authors report that this loss improves both region-captioning and referring-expression-comprehension results. In the RefCOCOg region-captioning evaluation, CIDEr/METEOR rises from 65.1/10.9 without the loss to 68.7/11.3 with it. In referring-expression comprehension, scores rise from 59.8 to 61.9 on RefCOCO, 45.3 to 48.9 on RefCOCO+, and 49.6 to 52.2 on RefCOCOg.

  3. Knowl 3 — NExT-Chat decodes a shared trigger embedding into boxes or masks

    model/method

    NExT-Chat uses a LLaVA-like vision-language architecture: CLIP ViT-L/14 at 336-pixel resolution produces a 24×2424\times24 grid of image patch embeddings, which are projected to the Vicuna-1.5 word-embedding dimension and passed to a decoder-only language model for conditional text generation.

    For detection, the hidden state t∈Rnt\in\mathbb{R}^n of the generated <trigger> is sent to a two-layer MLP box decoder FF. It predicts b=F(t)b=F(t), where b∈R4b\in\mathbb{R}^4 contains the box coordinates [x0,y0,x1,y1][x_0,y_0,x_1,y_1]. Detection training uses

    Ldet=2 L1(b,bgt)+0.8 GIoU(b,bgt),\mathcal{L}_{\mathrm{det}}=2\,\mathcal{L}_1(b,b_{\mathrm{gt}})+0.8\,\mathrm{GIoU}(b,b_{\mathrm{gt}}),

    where bgtb_{\mathrm{gt}} is the ground-truth box, and the terms are coordinate L1L_1 loss and generalized-IoU loss.

    For segmentation, a linear projector maps the same trigger hidden state to SAM's prompt-embedding dimension. SAM receives this prompt and the original image, and outputs a mask. The segmentation loss is Lseg=2 BCE(m,mgt)+0.5 Dice(m,mgt)\mathcal{L}_{\mathrm{seg}}=2\,\mathrm{BCE}(m,m_{\mathrm{gt}})+0.5\,\mathrm{Dice}(m,m_{\mathrm{gt}}), where mm is the predicted mask and mgtm_{\mathrm{gt}} is the ground-truth mask. In the authors' ablation, using the trigger embedding as the SAM prompt outperforms using only the predicted box, while combining box and embedding prompts gives slightly lower scores than embedding alone: cIoU for embedding versus box versus both is 76.6/75.3/76.1 on RefCOCO, 66.9/65.5/66.5 on RefCOCO+, and 69.9/68.2/69.5 on RefCOCOg. Stage-3 adaptation also improves over using unadapted SAM with the predicted box: cIoU rises from 69.6 to 76.6 on RefCOCO, 60.8 to 66.9 on RefCOCO+, and 63.2 to 69.9 on RefCOCOg.

  4. Knowl 4 — NExT-Chat is trained in three stages

    experimental setup

    NExT-Chat training first develops conversation and box-location abilities, then strengthens conversation, and finally adapts the model for mask output.

    Stage 1 pretrains on Flickr30K Entities, Visual Genome, RefCOCO, RefCOCO+, RefCOCOg, VQAv2, PointQA, Visual7W, and VCR. It uses batch size 64, learning rate 2×10−52\times10^{-5}, and 65,000 steps; the language model and box encoder/decoder are trained while the image encoder is frozen. The loss is Ltext+Ldet+Lcyc\mathcal{L}_{\mathrm{text}}+\mathcal{L}_{\mathrm{det}}+\mathcal{L}_{\mathrm{cyc}}. For the 7B model, this stage took about 59 hours on eight 80-GB A100 GPUs.

    Stage 2 fine-tunes on VQAv2, RefCOCO, Flickr30K Entities, LLaVA-instruct, VCR, and Shikra-RD, using batch size 64, learning rate 2×10−52\times10^{-5}, and the same losses as stage 1. It took about 10 hours on eight 80-GB A100 GPUs for the 7B model.

    Stage 3 trains the linear projector between the language model and SAM, along with SAM's decoder, on referring-segmentation splits of the RefCOCO datasets. It uses only segmentation loss and freezes the other parameters to limit catastrophic forgetting. This stage took about 3 hours on eight 80-GB A100 GPUs.

  5. Knowl 5 — NExT-Chat improves referring-expression segmentation, especially after task-specific fine-tuning

    data/table

    The paper evaluates referring-expression segmentation using cIoU on RefCOCO, RefCOCO+, and RefCOCOg. The table compares NExT-Chat before and after task-specific fine-tuning with fine-tuned LMM baselines. Fine-tuned NExT-Chat has the highest score in every listed split among these methods. The authors also report that NExT-Chat's segmentation training uses 127,000 object masks, whereas LISA uses more than an order of magnitude more mask annotations.

    Method RefCOCO RefCOCO+ RefCOCOg
    val testA testB val testA testB val test
    LISA-7B (ft) 74.9 79.1 72.3 65.1 70.8 58.1 67.9 70.6
    GLaMM (ft) 78.3 81.5 74.4 68.0 75.7 61.8 72.5 72.0
    NExT-Chat 76.9 80.5 72.4 67.6 73.7 59.4 69.5 70.3
    NExT-Chat (ft) 80.3 82.4 76.1 73.5 78.5 66.0 74.8 75.3
  6. Knowl 6 — NExT-Chat is competitive on referring-expression comprehension

    data/table

    Referring-expression comprehension is evaluated by [email protected] on the validation and test splits of RefCOCO, RefCOCO+, and RefCOCOg. NExT-Chat-7B is competitive with Shikra-7B but does not exceed it on these reported splits. It scores 3.3 points above VisionLLM-H on RefCOCO testA (90.0 versus 86.7); the paper also reports that it outperforms Kosmos-2 and VisionLLM-H overall in this comparison.

    Method RefCOCO RefCOCO+ RefCOCOg
    val testA testB val testA testB val test
    Shikra-7B 87.0 90.6 80.2 81.6 87.4 72.1 82.3 82.2
    NExT-Chat-7B 85.5 90.0 77.9 77.2 84.5 68.0 80.1 79.8

    The authors hypothesize that NExT-Chat's shortfall relative to Shikra may reflect the fixed detection-loss weight and the difficulty of training a language model on regression outputs; these are proposed explanations, not experimentally established causes.

  7. Knowl 7 — Cycle consistency and SAM prompt choice have measurable effects

    empirical result

    Ablations show that cycle consistency improves region-captioning and referring-expression-comprehension scores, with the values reported in the cycle-consistency knowl. A separate segmentation ablation finds that stage-3 adaptation is beneficial and that the trigger embedding is a better SAM prompt than the predicted box alone. On the RefCOCO, RefCOCO+, and RefCOCOg evaluations, embedding-only prompting yields cIoU scores of 76.6, 66.9, and 69.9; box-only prompting yields 75.3, 65.5, and 68.2; and using both yields 76.1, 66.5, and 69.5. Without stage-3 adaptation, the box-prompt configuration scores 69.6, 60.8, and 63.2, respectively. The authors suggest that the trigger embedding already carries the location information, so adding the box provides no new information.

  8. Knowl 8 — NExT-Chat improves region-captioning scores after fine-tuning

    data/table

    Region captioning is evaluated on RefCOCOg (google) using CIDEr and METEOR. Without task-specific fine-tuning, NExT-Chat scores higher in CIDEr than the listed baselines except GLaMM; after fine-tuning, it achieves the highest score on both metrics. The paper notes that GLaMM includes this dataset in its training, motivating the task-specific fine-tuning comparison.

    Method CIDEr METEOR
    GRIT 71.6 15.2
    Kosmos-2 (0-shot) 60.3 12.2
    Kosmos-2 (2-shot) 62.2 13.8
    Kosmos-2 (4-shot) 62.3 14.1
    ASM 41.9 13.6
    GLaMM 104.0 15.7
    GLaMM (ft) 105.0 16.2
    NExT-Chat 79.6 12.0
    NExT-Chat (ft) 114.0 17.4
  9. Knowl 9 — POPE results show strong but split-dependent hallucination performance

    data/table

    On the POPE image-hallucination benchmark, NExT-Chat achieves the highest accuracy among the compared methods on the Random and Popular splits, and the second-highest accuracy on the Adversarial split. The table reports NExT-Chat's other metrics alongside Shikra accuracy for comparison. The “Yes” values are also reproduced as reported by the paper.

    Split Accuracy Precision Recall F1-Score Yes Shikra accuracy
    Random 87.70 93.46 81.87 87.28 45.15 86.90
    Popular 84.57 86.54 81.87 84.14 47.30 83.97
    Adversarial 81.93 82.02 81.80 81.91 49.87 83.10

    The benchmark comparison includes Shikra, InstructBLIP, MiniGPT-4, LLaVA, MM-GPT, and mPLUG-Owl; NExT-Chat's adversarial accuracy trails Shikra's 83.10.

  10. Knowl 10 — Qualitative examples demonstrate grounding, captioning, and region-aware reasoning

    empirical result

    The paper's qualitative examples show NExT-Chat locating objects from text descriptions and from spatial relations to an input box, including a skateboard and a boat to the left of a marked boat. It generates grounded captions that associate multiple mentioned objects with their locations, such as four teddy bears, and describes a region supplied as input, including a small background light switch. In region-aware reasoning examples, it links an answer to localized evidence—for example, it suggests a mounted-police role based on a man's uniform and horse. These examples illustrate task breadth; they are not quantitative measurements of accuracy.

  11. Knowl 11 — The authors identify limits in multi-image, domain, and safety coverage

    limitation

    NExT-Chat is trained primarily on individual-image inputs, which limits its ability to handle multiple images in one interaction. The authors also identify insufficient training data from diverse domains as a barrier to reliable performance on medical and satellite imagery. They acknowledge possible visual hallucinations and offensive generated content. They suggest more high-quality and domain-specific training data, further alignment, and content-filtering algorithms as possible mitigations; the paper does not establish that these measures fully resolve the limitations.

Coverage note — No substantial contributed material was omitted; repetitive supplementary qualitative examples were consolidated into the capability knowl.

References

  1. 1.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35:23716–23736, 2022.
  2. 2.Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pp. 2425–2433, 2015.
  3. 3.Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. End-to-end object detection with transformers. In European conference on computer vision, pp. 213–229. Springer, 2020.
  4. 4.Chen, C., Qin, R., Luo, F., Mi, X., Li, P., Sun, M., and Liu, Y. Position-enhanced visual instruction tuning for multimodal large language models. arXiv preprint arXiv:2308.13437, 2023a.
  5. 5.Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., and Zhao, R. Shikra: Unleashing multimodal llm’s referential dialogue magic. arXiv preprint arXiv:2306.15195, 2023b.
  6. 6.Chen, T., Saxena, S., Li, L., Fleet, D. J., and Hinton, G. Pix2seq: A language modeling framework for object detection. arXiv preprint arXiv:2109.10852, 2021.
  7. 7.Chen, Y.-C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J. Uniter: Universal image-text representation learning. In European conference on computer vision, pp. 104–120. Springer, 2020.
  8. 8.Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S. Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.
  9. 9.Deng, J., Yang, Z., Chen, T., Zhou, W., and Li, H. Transvg: End-to-end visual grounding with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1769–1779, 2021.
  10. 10.Ding, H., Liu, C., Wang, S., and Jiang, X. Vision-language transformer and query generation for referring segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 16321–16330, 2021.
  11. 11.Gan, Z., Chen, Y.-C., Li, L., Zhu, C., Cheng, Y., and Liu, J. Large-scale adversarial training for vision-and-language representation learning. Advances in Neural Information Processing Systems, 33:6616–6628, 2020.
  12. 12.Gong, T., Lyu, C., Zhang, S., Wang, Y., Zheng, M., Zhao, Q., Liu, K., Zhang, W., Luo, P., and Chen, K. Multimodal-gpt: A vision and language model for dialogue with humans. arXiv preprint arXiv:2305.04790, 2023.
  13. 13.Huang, S., Dong, L., Wang, W., Hao, Y., Singhal, S., Ma, S., Lv, T., Cui, L., Mohammed, O. K., Liu, Q., et al. Language is not all you need: Aligning perception with language models. arXiv preprint arXiv:2302.14045, 2023.
  14. 14.Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Dollar, P., and Girshick, R. Segment anything. arXiv:2304.02643, 2023.
  15. 15.Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123:32–73, 2017.
  16. 16.Lai, X., Tian, Z., Chen, Y., Li, Y., Yuan, Y., Liu, S., and Jia, J. Lisa: Reasoning segmentation via large language model. arXiv preprint arXiv:2308.00692, 2023.
  17. 17.Li, B., Zhang, Y., Chen, L., Wang, J., Pu, F., Yang, J., Li, C., and Liu, Z. Mimic-it: Multi-modal in-context instruction tuning. 2023a.
  18. 18.Li, B., Zhang, Y., Chen, L., Wang, J., Yang, J., and Liu, Z. Otter: A multi-modal model with in-context instruction tuning. arXiv preprint arXiv:2305.03726, 2023b.
  19. 19.Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023c.
  20. 20.Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W. X., and Wen, J.-R. Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:2305.10355, 2023d.
  21. 21.Liu, C., Ding, H., and Jiang, X. Gres: Generalized referring expression segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 23592–23601, 2023a.
  22. 22.Liu, F., Lin, K., Li, L., Wang, J., Yacoob, Y., and Wang, L. Aligning large multi-modal model with robust instruction tuning. arXiv preprint arXiv:2306.14565, 2023b.
  23. 23.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning, 2023c.
  24. 24.Liu, J., Ding, H., Cai, Z., Zhang, Y., Satzoda, R. K., Mahadevan, V., and Manmatha, R. Polyformer: Referring image segmentation as sequential polygon generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18653–18663, 2023d.
  25. 25.Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Li, C., Yang, J., Su, H., Zhu, J., et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023e.
  26. 26.Luo, G., Zhou, Y., Sun, X., Cao, L., Wu, C., Deng, C., and Ji, R. Multi-task collaborative network for joint referring expression comprehension and segmentation. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pp. 10034–10043, 2020.
  27. 27.Mani, A., Yoo, N., Hinthorn, W., and Russakovsky, O. Point and ask: Incorporating pointing into visual question answering. 2020.
  28. 28.Mao, J., Huang, J., Toshev, A., Camburu, O., Yuille, A. L., and Murphy, K. Generation and comprehension of unambiguous object descriptions. In Proceedings of CVPR, pp. 11–20, 2016.
  29. 29.Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., and Wei, F. Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023.
  30. 30.Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In Proceedings of the IEEE international conference on computer vision, pp. 2641–2649, 2015.
  31. 31.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  32. 32.Rasheed, H., Maaz, M., Shaji, S., Shaker, A., Khan, S., Cholakkal, H., Anwer, R. M., Xing, E., Yang, M.-H., and Khan, F. S. Glamm: Pixel grounding large multimodal model. ArXiv 2311.03356, 2023.
  33. 33.Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., and Savarese, S. Generalized intersection over union: A metric and a loss for bounding box regression. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 658–666, 2019.
  34. 34.Wang, P., Yang, A., Men, R., Lin, J., Bai, S., Li, Z., Ma, J., Zhou, C., Zhou, J., and Yang, H. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International Conference on Machine Learning, pp. 23318–23340. PMLR, 2022a.
  35. 35.Wang, W., Chen, Z., Chen, X., Wu, J., Zhu, X., Zeng, G., Luo, P., Lu, T., Zhou, J., Qiao, Y., and Dai, J. Visionllm: Large language model is also an open-ended decoder for vision-centric tasks. arXiv preprint arXiv:2305.11175, 2023a.
  36. 36.Wang, W., Shi, M., Li, Q., Wang, W., Huang, Z., Xing, L., Chen, Z., Li, H., Zhu, X., Cao, Z., et al. The all-seeing project: Towards panoptic visual recognition and understanding of the open world. arXiv preprint arXiv:2308.01907, 2023b.
  37. 37.Wang, Z., Lu, Y., Li, Q., Tao, X., Guo, Y., Gong, M., and Liu, T. Cris: Clip-driven referring image segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11686–11695, 2022b.
  38. 38.Wu, J., Wang, J., Yang, Z., Gan, Z., Liu, Z., Yuan, J., and Wang, L. Grit: A generative region-to-text transformer for object understanding. arXiv preprint arXiv:2212.00280, 2022.
  39. 39.Yang, Z., Gan, Z., Wang, J., Hu, X., Ahmed, F., Liu, Z., Lu, Y., and Wang, L. Unitab: Unifying text and box outputs for grounded vision-language modeling. In European Conference on Computer Vision, pp. 521–539. Springer, 2022a.
  40. 40.Yang, Z., Wang, J., Tang, Y., Chen, K., Zhao, H., and Torr, P. H. Lavt: Language-aware vision transformer for referring image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18155–18165, 2022b.
  41. 41.Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., et al. mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178, 2023.
  42. 42.Yu, L., Poirson, P., Yang, S., Berg, A. C., and Berg, T. L. Modeling context in referring expressions. In Proceedings of ECCV, pp. 69–85. Springer, 2016.
  43. 43.Yu, L., Lin, Z., Shen, X., Yang, J., Lu, X., Bansal, M., and Berg, T. L. Mattnet: Modular attention network for referring expression comprehension. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1307–1315, 2018.
  44. 44.Zellers, R., Bisk, Y., Farhadi, A., and Choi, Y. From recognition to cognition: Visual commonsense reasoning. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  45. 45.Zhang, A., Fei, H., Yao, Y., Ji, W., Li, L., Liu, Z., and Chua, T.-S. Transfer visual prompt generator across llms. arXiv preprint arXiv:2305.01278, 2023a.
  46. 46.Zhang, S., Sun, P., Chen, S., Xiao, M., Shao, W., Zhang, W., Chen, K., and Luo, P. Gpt4roi: Instruction tuning large language model on region-of-interest. arXiv preprint arXiv:2307.03601, 2023b.
  47. 47.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023.
  48. 48.Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023.
  49. 49.Zhu, Y., Groth, O., Bernstein, M., and Fei-Fei, L. Visual7w: Grounded question answering in images, 2016.
  50. 50.Zou, X., Dou, Z.-Y., Yang, J., Gan, Z., Li, L., Li, C., Dai, X., Behl, H., Wang, J., Yuan, L., et al. Generalized decoding for pixel, image, and language. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15116–15127, 2023a.
  51. 51.Zou, X., Yang, J., Zhang, H., Li, F., Li, L., Gao, J., and Lee, Y. J. Segment everything everywhere all at once. arXiv preprint arXiv:2304.06718, 2023b.

Citation

MLA
Zhang, A., et al. “NExT-Chat: An LMM for Chat, Detection and Segmentation”. arXiv, 2023, http://arxiv.org/abs/2311.04498v4.
APA
Zhang, A., Yao, Y., Ji, W., Liu, Z., & Chua, T.-S. (2023). NExT-Chat: An LMM for Chat, Detection and Segmentation. arXiv. http://arxiv.org/abs/2311.04498v4
Chicago
Zhang, A., Y. Yao, W. Ji, Z. Liu, and T.-S. Chua. 2023. “NExT-Chat: An LMM for Chat, Detection and Segmentation”. arXiv. http://arxiv.org/abs/2311.04498v4.
Harvard
Zhang, A. et al. (2023) “NExT-Chat: An LMM for Chat, Detection and Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.04498v4.
Vancouver
1. Zhang A, Yao Y, Ji W, Liu Z, Chua T-S (2023) NExT-Chat: An LMM for Chat, Detection and Segmentation. arXiv

BibTeX

@article{zhang2023next,
  title = {NExT-Chat: An LMM for Chat, Detection and Segmentation},
  author = {Zhang, Ao and Yao, Yuan and Ji, Wei and Liu, Zhiyuan and Chua, Tat-Seng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.04498v4},
  eprint = {2311.04498}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/