NExT-Chat: An LMM for Chat, Detection and Segmentation
Ao ZhangYuan YaoWei JiZhiyuan LiuTat-Seng Chua
Proposes a pixel-to-embedding framework that unifies conversational understanding, bounding-box detection, and pixel-level segmentation within a single multimodal language model by decoding learned location representations into multiple spatial formats.
Large multimodal models have shown rapid progress in describing images and answering visual queries. However, most existing frameworks understand images only at a holistic level and struggle to accurately pinpoint, segment, or describe specific regions. Prior efforts to enable region-level comprehension typically convert bounding box coordinates into text tokens. This text-based method introduces significant computational overhead—requiring dozens of tokens per bounding box—and cannot naturally output pixel-level segmentation masks needed for advanced visual applications.
The article demonstrates a new paradigm called "pixel-to-embedding" (pix2emb) and introduces NExT-Chat, a unified multimodal model capable of conversational interaction, object detection, and image segmentation. The core objective is to evaluate whether representing spatial locations as continuous numerical representations (embeddings) rather than discrete text tokens enables an AI system to efficiently process location inputs and generate both bounding boxes and segmentation masks within a single conversational architecture.
To achieve this, the authors built NExT-Chat using a standard visual encoder paired with a language model, augmented with lightweight location encoders and decoders. Instead of generating word tokens for coordinates, the model outputs a single trigger token whose internal hidden state is fed directly into a box decoder or a mask decoder based on the Segment Anything Model. The framework utilizes a three-stage training process: pre-training on bounding box conversations, instruction fine-tuning, and a fast three-hour stage adapting the model to generate segmentation masks. To ensure that input locations and output predictions align seamlessly, the authors introduced a cycle-consistency training loss that ties the location encoder and decoder together. Experiments were conducted across established benchmarks for referring expression segmentation, referring expression comprehension, region captioning, and image hallucination diagnosis.
The evaluation produced several key findings. First, in referring expression segmentation, NExT-Chat achieved an average intersection-over-union score of 71.3 (reaching up to 80.3 when fine-tuned), outperforming comparable systems like LISA (67.9) despite using an order of magnitude fewer mask annotations—only 127,000 masks compared to millions in baseline training sets. Second, in region captioning, NExT-Chat achieved a CIDEr score of 79.6 (increasing to 114.0 upon fine-tuning), significantly outperforming models like Kosmos-2 (62.3). Third, the approach achieved massive computational efficiency: representing a bounding box requires only two tokens and a single added vocabulary item, making it up to 169 times more computationally efficient than standard coordinate text representations. Finally, the model demonstrated strong reliability against generating false information, scoring 87.7% accuracy on the random split of the POPE hallucination benchmark.
These findings indicate that treating location as continuous embeddings rather than text sequences dramatically reduces computational cost, lowers training data requirements, and unifies diverse visual tasks into one framework. By cutting down the sequence lengths needed to ground objects, organizations can deploy conversational models with spatial awareness at lower inference costs and lower latency. Furthermore, the ability to train segmentation modules in just three hours allows rapid task adaptation without catastrophic forgetting of core conversational skills.
Based on these results, organizations deploying visual AI assistants should adopt embedding-based spatial modeling over text-coordinate methods to improve throughput and support fine-grained mask outputs. Before deploying into specialized operational settings, teams should conduct domain-specific fine-tuning and implement content moderation filters, as the model was trained primarily on open-domain data and can occasionally hallucinate ungrounded facts. For future technical development, the authors note that dynamic weighting between text and detection loss functions should be explored to optimize spatial regression. Confidence in the results is high across general benchmark settings, though caution is warranted when applying the current system to multi-image inputs or specialized domains such as satellite and medical imaging.
- Paper: Kosmos-2: Grounding Multimodal Large Language Models to the World, Zhiliang Peng et al. (2023). Kosmos-2 grounds multimodal language in image regions with discrete location tokens, the earlier region-level approach that NExT-Chat contrasts with its location-embedding method.
- Paper: GSVA: Generalized Segmentation via Multimodal Large Language Models, Zhuofan Xia et al. (2024). GSVA carries language-guided segmentation into generalized settings, extending grounded multimodal assistants to segment multiple referents and reject objects absent from the image.
- Paper: OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding, Tao Zhang et al. (2024). OMG-LLaVA continues unified grounded conversation by joining image-level reasoning with object- and pixel-level outputs, including segmentation masks.
