Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Tianhe RenShilong LiuAiling ZengJing LinKunchang LiHe CaoJia-Yu ChenXin-Yu HuangYukang ChenFeng Yan

article2024arXiv1,315 citations

Introduces Grounded SAM, a framework that unifies Grounding DINO and the Segment Anything Model to achieve text-prompted zero-shot detection and segmentation for downstream tasks like automated data labeling and controllable image editing.

Listen

The rapid advancement of artificial intelligence requires proven research leadership across large language model agents, multimodal learning, and visual recognition. Organizations seeking to deploy cutting-edge foundation models face the challenge of identifying researchers capable of delivering both foundational theoretical innovations and widely adopted open-source software tools. The article evaluates and outlines the professional profile, academic achievements, and research trajectory of Shilong Liu, a Postdoctoral Research Fellow at the Princeton AI Lab.

The article summarizes Liu's research career by analyzing his education, industry and academic appointments, honors, and a portfolio of more than forty peer-reviewed and preprint publications. The scope covers his doctoral training in computer science at Tsinghua University through his industry research appointments at leading institutions, including Microsoft Research, NVIDIA Research, IDEA Research, Shengshu-Tech, and ByteDance Seed.

The findings establish Liu as a high-impact contributor across modern artificial intelligence. First, his scholarship has received more than 13,000 academic citations and over 30,000 repository stars on GitHub, with the majority earned as a lead or co-first author. Second, four of his papers achieved recognition among the top fifteen most influential publications at their respective academic conferences. Third, his open-source visual models exceed two million monthly downloads, highlighted by Grounding DINO, which stands as the most downloaded zero-shot object detection model on the Hugging Face platform. Finally, his research spans key breakthroughs, including tool-augmented vision models like LLaVA-Plus, foundational detection architectures, generative video systems, and self-evolving autonomous agents.

These results demonstrate significant practical influence on performance, software standardization, and model deployment across the artificial intelligence ecosystem. By bridging core academic discovery with scalable open-source tools, Liu's work helps lower implementation barriers and accelerates the development cycle for commercial and scientific vision-language applications.

Stakeholders and research leaders evaluating advanced multimodal strategies should consider integrating or benchmarking against Liu's open-source frameworks, particularly Grounding DINO and related transformer architectures. Continued tracking of his emerging work in autonomous agents and spatial reasoning is recommended to identify future technological opportunities.

The primary limitation of the article is that it reflects an individual curriculum vitae and publication record without providing detailed comparative performance metrics across all models discussed. Nevertheless, the provided metrics regarding citations, downloads, and peer recognition offer high confidence in the candidate's established technical contributions and ongoing leadership in artificial intelligence.

arXiv: 2401.14159
  • Paper: SAM 2: Segment Anything in Images and Videos, Nikhila Ravi et al. (2025). SAM 2 extends the promptable segmentation foundation of SAM into the video domain, enabling memory-conditioned spatio-temporal mask tracking.
  • Paper: SAM 3: Segment Anything with Concepts, Nicolas Carion et al. (2025). SAM 3 advances promptable segmentation by natively integrating open-vocabulary concept detection and tracking directly into the Segment Anything framework.

Abstract

We introduce Grounded SAM, which uses Grounding DINO as an open-set object detector to combine with the segment anything model (SAM). This integration enables the detection and segmentation of any regions based on arbitrary text inputs and opens a door to connecting various vision models. As shown in Fig.1, a wide range of vision tasks can be achieved by using the versatile Grounded SAM pipeline. For example, an automatic annotation pipeline based solely on input images can be realized by incorporating models such as BLIP and Recognize Anything. Additionally, incorporating Stable-Diffusion allows for controllable image editing, while the integration of OSX facilitates promptable 3D human motion analysis. Grounded SAM also shows superior performance on open-vocabulary benchmarks, achieving 48.7 mean AP on SegInW (Segmentation in the wild) zero-shot benchmark with the combination of Grounding DINO-Base and SAM-Huge models.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Task-specific Vision Models
  • 2.2 Unified Models
  • 2.3 Model Assembly with a Controller System
  • 3 Grounded SAM Playground
  • 3.1 Preliminary
  • 3.2 Grounded SAM: Open-Vocabulary Detection and Segmentation
  • 3.3 RAM-Grounded-SAM: Automatic Dense Image Annotation
  • 3.4 Grounded-SAM-SD: Highly Accurate and Controllable Image Editing
  • 3.5 Grounded-SAM-OSX: Promptable Human Motion Analysis
  • 3.6 More Extensions for Grounded SAM
  • 4 Effectiveness of Grounded SAM
  • 5 Conclusion and Prospects
  • 6 Contributions and Acknowledgments
  • References

Knowls

  1. Knowl 1 — Research Profile and Scientific Impact of Shilong Liu

    empirical result

    Shilong Liu is a Postdoctoral Research Fellow at the Princeton AI Lab specializing in LLM agents, multimodal learning, and physical intelligence. His research portfolio includes over 13,000 Google Scholar citations and over 30,000 GitHub stars, predominantly as first or leading author. Four of his works have been recognized among the "Top 15 Most Influential Papers" at their respective conferences by Paper Digest, and his open-sourced models achieve over 2 million monthly downloads, with Grounding DINO being the most downloaded zero-shot object detection model on Hugging Face.

  2. Knowl 2 — Grounding DINO Open-Set Object Detection Framework

    model/method

    Grounding DINO is an open-set object detection framework that couples the transformer-based DINO detector architecture with grounded pre-training on paired image-text data. By fusing language feature embeddings into multi-scale image feature maps via cross-modality attention at multiple architectural stages, the framework enables zero-shot object detection, open-vocabulary phrase localization, and text-guided visual grounding.

  3. Knowl 3 — DAB-DETR Dynamic Anchor Box Positional Query Architecture

    model/method

    DAB-DETR (Dynamic Anchor Boxes for DETR) reformulates the positional queries in Detection Transformers by explicitly defining queries as 4-dimensional bounding box coordinates (x,y,w,h)(x, y, w, h). This geometric formulation allows the positional queries to be updated layer-by-layer dynamically across the transformer decoder, providing explicit spatial priors that accelerate training convergence and improve object localization accuracy.

  4. Knowl 4 — DINO End-to-End Object Detection with Contrastive Denoising

    model/method

    DINO (DETR with Improved Denoising Anchor Boxes) is an end-to-end transformer-based object detector that integrates contrastive query denoising, dynamic 4D anchor box queries, and mixed query selection. The contrastive denoising training strategy supplies the decoder with both positive (perturbed ground-truth) and negative noisy boxes, stabilizing bipartite matching optimization and mitigating ambiguities in Hungarian loss assignment.

  5. Knowl 5 — LLaVA-Plus Tool-Augmented Multimodal Agent Framework

    model/method

    LLaVA-Plus is a multimodal agent architecture that endows large vision-language models (VLMs) with tool-use capabilities across a diverse suite of specialized vision models. By maintaining an interactive execution loop where the central VLM plans tool invocations, formats inputs for vision experts (such as detectors, segmenters, and keypoint estimators), and integrates their outputs into language reasoning, the system executes complex visual and interactive workflows.

  6. Knowl 6 — Alita and Alita-g Self-Evolving Autonomous Agent Architectures

    model/method

    Alita and its generative counterpart Alita-g are autonomous agent architectures designed for scalable agentic reasoning. They minimize manually engineered predefined scaffolding by relying on self-evolution mechanisms, enabling the agent to autonomously generate, evaluate, and iteratively refine new sub-agents and tool execution graphs to solve complex reasoning tasks.

  7. Knowl 7 — CubeBench Interactive Spatial Reasoning Benchmark

    experimental setup

    CubeBench is an evaluation benchmark formulated to diagnose interactive, long-horizon 3D spatial reasoning in multimodal models and embodied agents under partial observability. The benchmark requires models to systematically explore, manipulate, and infer global spatial configurations despite receiving only limited, view-dependent visual observations at each step.

  8. Knowl 8 — Grounded-SAM Open-World Perception Pipeline

    model/method

    Grounded-SAM is a modular open-world visual perception pipeline combining open-set object detection from Grounding DINO with zero-shot instance segmentation from the Segment Anything Model (SAM). The system first detects bounding boxes based on free-form arbitrary text prompts and subsequently forwards those predicted boxes as spatial prompts to SAM to generate high-quality instance masks for arbitrary object categories.

Coverage note — Specific employment dates and standard bibliographic listing details were omitted in favor of substantive descriptions of the candidate's core research systems and empirical milestones.

References

  1. 1.H.-a. Gao, Z. Zhang, T. Luo, K. Yang, X. Juan, J. Qiu, T. Chen, B. He, H. Zhao, H. Zhou, S. Liu, and M. Wang, Cubebench: Diagnosing interactive, long-horizon spatial reasoning under partial observations, 2026. arXiv: 2512.23328 [cs.AI]. URL: https://arxiv.org/abs/2512.23328.
  2. 2.J. Feng, Y. Zhang, C. Zhang, Y. Lu, S. Liu, and M. Wang, Web world models, 2025. arXiv: 2512.23676 [cs.AI]. URL: https://arxiv.org/abs/2512.23676.
  3. 3.S. Liu, H. Cheng, H. Liu, H. Zhang, F. Li, T. Ren, X. Zou, J. Yang, H. Su, J. Zhu, L. Zhang, J. Gao, and C. Li, “Llava-plus: Learning to use tools for creating multimodal agents,” in Computer Vision – ECCV 2024, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds., Cham: Springer Nature Switzerland, 2025, pp. 126–142, isbn: 978-3-031-72970-6.
  4. 4.S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su, J. Zhu, and L. Zhang, “Grounding dino: Marrying dino with grounded pre-training for open-set object detection,” in Computer Vision – ECCV 2024, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds., Cham: Springer Nature Switzerland, 2025, pp. 38–55, isbn: 978-3-031-72970-6.
  5. 5.J. Qiu, X. Qi, H. Wang, X. Juan, Y. Wang, Z. Zhao, J. Geng, J. Guo, P. Li, J. Shi, S. Liu, and M. Wang, Alita-g: Self-evolving generative agent for agent generation, 2025. arXiv: 2510.23601 [cs.AI]. URL: https://arxiv.org/abs/2510.23601.
  6. 6.J. Qiu, X. Qi, T. Zhang, X. Juan, J. Guo, Y. Lu, Y. Wang, Z. Yao, Q. Ren, X. Jiang, X. Zhou, D. Liu, L. Yang, Y. Wu, K. Huang, S. Liu, H. Wang, and M. Wang, Alita: Generalist agent enabling scalable agentic reasoning with minimal predefinition and maximal self-evolution, 2025. arXiv: 2505.20286 [cs.AI]. URL: https://arxiv.org/abs/2505.20286.
  7. 7.H. Zhang, F. Li, S. Liu, L. Zhang, H. Su, J. Zhu, L. M. Ni, and H.-Y. Shum, “Dino: Detr with improved denoising anchor boxes for end-to-end object detection,” in ICLR 2023, 2023.
  8. 8.S. Liu, F. Li, H. Zhang, X. Yang, X. Qi, H. Su, J. Zhu, and L. Zhang, “DAB-DETR: Dynamic anchor boxes are better queries for DETR,” in International Conference on Learning Representations, 2022. URL: https://openreview.net/forum?id=oMI9PjOb9Jl.
  9. 9.H.-a. Gao, Z. Zhang, T. Luo, K. Yang, X. Juan, J. Qiu, T. Chen, B. He, H. Zhao, H. Zhou, S. Liu, and M. Wang, Cubebench: Diagnosing interactive, long-horizon spatial reasoning under partial observations, 2026. arXiv: 2512.23328 [cs.AI]. URL: https://arxiv.org/abs/2512.23328.
  10. 10.J. Feng, Y. Zhang, C. Zhang, Y. Lu, S. Liu, and M. Wang, Web world models, 2025. arXiv: 2512.23676 [cs.AI]. URL: https://arxiv.org/abs/2512.23676.
  11. 11.H.-a. Gao, J. Geng, W. Hua, M. Hu, X. Juan, H. Liu, S. Liu, J. Qiu, X. Qi, Y. Wu, H. Wang, H. Xiao, Y. Zhou, S. Zhang, J. Zhang, J. Xiang, Y. Fang, Q. Zhao, D. Liu, Q. Ren, C. Qian, Z. Wang, M. Hu, H. Wang, Q. Wu, H. Ji, and M. Wang, A survey of self-evolving agents: On path to artificial super intelligence, 2025. arXiv: 2507.21046 [cs.AI]. URL: https://arxiv.org/abs/2507.21046.
  12. 12.Q. Jiang, F. Li, Z. Zeng, T. Ren, S. Liu, and L. Zhang, “T-rex2: Towards generic object detection via text-visual prompt synergy,” in Computer Vision – ECCV 2024, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds., Cham: Springer Nature Switzerland, 2025, pp. 38–57, isbn: 978-3-031-73414-4.
  13. 13.Z. Jiang, S. Xie, W. Li, W. Zu, P. Li, J. Qiu, S. Pei, L. Ma, T. Huang, M. Wang, and S. Liu, Zoom in, click out: Unlocking and evaluating the potential of zooming for gui grounding, 2025. arXiv: 2512.05941 [cs.CV]. URL: https://arxiv.org/abs/2512.05941.
  14. 14.A. Y. Li, B. Yu, D. Lei, T. Ren, and S. Liu, Chain-of-ground: Improving gui grounding via iterative reasoning and reference feedback, 2025. arXiv: 2512.01979 [cs.AI]. URL: https://arxiv.org/abs/2512.01979.
  15. 15.F. Li, H. Zhang, P. Sun, X. Zou, S. Liu, C. Li, J. Yang, L. Zhang, and J. Gao, “Segment and recognize anything at any granularity,” in Computer Vision – ECCV 2024, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds., Cham: Springer Nature Switzerland, 2025, pp. 467–484, isbn: 978-3-031-73195-2.
  16. 16.H. Li, H. Zhang, S. Liu, Z. Zeng, T. Ren, F. Li, and L. Zhang, “Taptr: Tracking any point with transformers as detection,” in Computer Vision – ECCV 2024, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds., Cham: Springer Nature Switzerland, 2025, pp. 57–75, isbn: 978-3-031-72640-8.
  17. 17.Z. Li, G. Chen, S. Liu, S. Wang, V. VS, Y. Ji, S. Lan, H. Zhang, Y. Zhao, S. Radhakrishnan, N. Chang, K. Sapra, A. S. Deshmukh, T. Rintamaki, M. Le, I. Karmanov, L. Voegtle, P. Fischer, D.-A. Huang, T. Roman, T. Lu, J. M. Alvarez, B. Catanzaro, J. Kautz, A. Tao, G. Liu, and Z. Yu, Eagle 2: Building post-training data strategies from scratch for frontier vision-language models, 2025. arXiv: 2501.14818 [cs.CV]. URL: https://arxiv.org/abs/2501.14818.
  18. 18.S. Liu, H. Cheng, H. Liu, H. Zhang, F. Li, T. Ren, X. Zou, J. Yang, H. Su, J. Zhu, L. Zhang, J. Gao, and C. Li, “Llava-plus: Learning to use tools for creating multimodal agents,” in Computer Vision – ECCV 2024, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds., Cham: Springer Nature Switzerland, 2025, pp. 126–142, isbn: 978-3-031-72970-6.
  19. 19.S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su, J. Zhu, and L. Zhang, “Grounding dino: Marrying dino with grounded pre-training for open-set object detection,” in Computer Vision – ECCV 2024, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds., Cham: Springer Nature Switzerland, 2025, pp. 38–55, isbn: 978-3-031-72970-6.
  20. 20.J. Qiu, X. Qi, H. Wang, X. Juan, Y. Wang, Z. Zhao, J. Geng, J. Guo, P. Li, J. Shi, S. Liu, and M. Wang, Alita-g: Self-evolving generative agent for agent generation, 2025. arXiv: 2510.23601 [cs.AI]. URL: https://arxiv.org/abs/2510.23601.
  21. 21.J. Qiu, X. Qi, T. Zhang, X. Juan, J. Guo, Y. Lu, Y. Wang, Z. Yao, Q. Ren, X. Jiang, X. Zhou, D. Liu, L. Yang, Y. Wu, K. Huang, S. Liu, H. Wang, and M. Wang, Alita: Generalist agent enabling scalable agentic reasoning with minimal predefinition and maximal self-evolution, 2025. arXiv: 2505.20286 [cs.AI]. URL: https://arxiv.org/abs/2505.20286.
  22. 22.J. Qiu, F. Xiao, Y. Wang, Y. Mao, Y. Chen, X. Juan, S. Zhang, S. Wang, X. Qi, T. Zhang, Z. Yao, J. Guo, Y. Lu, C. Argon, J. Cui, D. Chen, J. Zhou, S. Zhou, Z. Zhou, L. Yang, S. Liu, H. Wang, K. Huang, X. Jiang, Y. Cao, Y. Chen, Y. Chen, Z. Chen, R. Dai, M. Deng, J. Fu, Y. Gu, Z. Guan, Z. Huang, X. Ji, Y. Jiang, D. Kong, H. Li, J. Li, R. Li, T. Li, Z. Li, H. Lian, M. Lin, X. Liu, J. Lu, J. Lu, W. Luo, Z. Luo, Z. Pu, Z. Qiao, R. Ren, L. Wan, R. Wang, T. Wang, Y. Wang, Z. Wang, Z. Wang, Y. Wu, Z. Wu, H. Xin, W. Xing, R. Xiong, W. Xu, Y. Shu, Y. Xiao, X. Yang, Y. Yang, N. Yi, J. Yu, Y. Yu, H. Zeng, D. Zhang, Y. Zhang, Z. Zhang, Z. Zhang, X. Zheng, P. Zhou, L. Zhong, X. Zong, Y. Zhao, Z. Chen, L. Ding, X. Gao, B. Gong, Y. Li, Y. Liao, G. Ma, T. Ma, X. Sun, T. Wang, H. Xia, R. Xian, G. Ye, T. Yu, W. Zhang, Y. Wang, X. Gao, and M. Wang, On path to multimodal historical reasoning: Histbench and histagent, 2025. arXiv: 2505.20246 [cs.AI]. URL: https://arxiv.org/abs/2505.20246.
  23. 23.J. Yang, A. Zeng, T. Ren, S. Liu, F. Li, R. Zhang, and L. Zhang, “Ed-pose++: Enhanced explicit box detection for conventional and interactive multi-object keypoint detection,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 7, pp. 5636–5654, 2025. DOI: 10.1109/TPAMI.2025.3555527.
  24. 24.H. Zhang, H. Li, F. Li, T. Ren, X. Zou, S. Liu, S. Huang, J. Gao, Leizhang, C. Li, and J. Yang, “Llava-grounding: Grounded visual chat with large multimodal models,” in Computer Vision – ECCV 2024, A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, and G. Varol, Eds., Cham: Springer Nature Switzerland, 2025, pp. 19–35, isbn: 978-3-031-72775-7.
  25. 25.F. Bao, C. Xiang, G. Yue, G. He, H. Zhu, K. Zheng, M. Zhao, S. Liu, Y. Wang, and J. Zhu, Vidu: A highly consistent, dynamic and skilled text-to-video generator with diffusion models, 2024. arXiv: 2405.04233 [cs.CV]. URL: https://arxiv.org/abs/2405.04233.
  26. 26.F. Bao, C. Xiang, G. Yue, G. He, H. Zhu, K. Zheng, M. Zhao, S. Liu, Y. Wang, and J. Zhu, “Vidu: A highly consistent, dynamic and skilled text-to-video generator with diffusion models,” arXiv preprint arXiv:2405.04233, 2024.
  27. 27.C. Chen, Y. Guo, F. Tian, S. Liu, W. Yang, Z. Wang, J. Wu, H. Su, H. Pfister, and S. Liu, “A unified interactive model evaluation for classification, object detection, and instance segmentation in computer vision,” IEEE Transactions on Visualization and Computer Graphics, vol. 30, no. 1, pp. 76–86, 2024. DOI: 10.1109/TVCG.2023.3326588.
  28. 28.F. Li, Q. Jiang, H. Zhang, T. Ren, S. Liu, X. Zou, H. Xu, H. Li, J. Yang, C. Li, L. Zhang, and J. Gao, “Visual in-context prompting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2024, pp. 12 861–12 871.
  29. 29.J. Li, S. Liu, Z. Liu, Y. Wang, K. Zheng, J. Xu, J. Li, and J. Zhu, “Instructpix2nerf: Instructed 3d portrait editing from a single image,” ICLR, 2024.
  30. 30.Y. Li, F. Wei, C. Zhang, and H. Zhang, Eagle-2: Faster inference of language models with dynamic draft trees, 2024. arXiv: 2406.16858 [cs.CL]. URL: https://arxiv.org/abs/2406.16858.
  31. 31.T. Ren, Q. Jiang, S. Liu, Z. Zeng, W. Liu, H. Gao, H. Huang, Z. Ma, X. Jiang, Y. Chen, Y. Xiong, H. Zhang, F. Li, P. Tang, K. Yu, and L. Zhang, Grounding dino 1.5: Advance the "edge" of open-set object detection, 2024. arXiv: 2405.10300 [cs.CV]. URL: https://arxiv.org/abs/2405.10300.
  32. 32.Y. Zhang, X. Huang, J. Ma, Z. Li, Z. Luo, Y. Xie, Y. Qin, T. Luo, Y. Li, S. Liu, Y. Guo, and L. Zhang, “Recognize anything: A strong image tagging model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, Jun. 2024, pp. 1724–1732.
  33. 33.X. Zou, L. Li, J. Wang, J. Yang, M. Ding, Z. Yang, F. Li, H. Zhang, S. Liu, A. Aravinthan, Y. J. Lee†, and L. Wang†, Interfacing foundation models’ embeddings, 2024.
  34. 34.F. Li, A. Zeng, S. Liu, H. Zhang, H. Li, L. Zhang, and L. M. Ni, “Lite detr: An interleaved multi-scale encoder for efficient detr,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2023, pp. 18 558–18 567.
  35. 35.F. Li, H. Zhang, H. Xu, S. Liu, L. Zhang, L. M. Ni, and H.-Y. Shum, “Mask dino: Towards a unified transformer-based framework for object detection and segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2023, pp. 3041–3050.
  36. 36.H. Li, H. Zhang, Z. Zeng, S. Liu, F. Li, T. Ren, and L. Zhang, “Dfa3d: 3d deformable attention for 2d-to-3d feature lifting,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2023, pp. 6684–6693.
  37. 37.S. Liu, S. Huang, F. Li, H. Zhang, Y. Liang, H. Su, J. Zhu, and L. Zhang, “Dq-detr: Dual query detection transformer for phrase extraction and grounding,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 2, pp. 1728–1736, Jun. 2023. DOI: 10.1609/aaai.v37i2.25261.
  38. 38.S. Liu, T. Ren, J. Chen, Z. Zeng, H. Zhang, F. Li, H. Li, J. Huang, H. Su, J. Zhu, and L. Zhang, “Detection transformer with stable matching,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2023, pp. 6491–6500.
  39. 39.T. Ren, S. Liu, F. Li, H. Zhang, A. Zeng, J. Yang, X. Liao, D. Jia, H. Li, H. Cao, J. Wang, Z. Zeng, X. Qi, Y. Yuan, J. Yang, and L. Zhang, Detrex: Benchmarking detection transformers, 2023. arXiv: 2306.07265 [cs.CV]. URL: https://arxiv.org/abs/2306.07265.
  40. 40.T. Ren, S. Liu, A. Zeng, J. Lin, K. Li, H. Cao, J. Chen, X. Huang, Y. Chen, F. Yan, Z. Zeng, H. Zhang, F. Li, J. Yang, H. Li, Q. Jiang, and L. Zhang, “Grounded sam: Assembling open-world models for diverse visual tasks,” in ICCV 2023 Demo, 2023. arXiv: 2401.14159 [cs.CV].
  41. 41.Y. Shi, J. Wang, H. Cao, B. Tang, X. Qi, T. Yang, Y. Huang, S. Liu, L. Zhang, and H.-Y. Shum, “Toss: High-quality text-guided novel view synthesis from a single image,” ICLR, 2023.
  42. 42.J. Yang, A. Zeng, F. Li, S. Liu, R. Zhang, and L. Zhang, “Neural interactive keypoint detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2023, pp. 15 122–15 132.
  43. 43.J. Yang, A. Zeng, S. Liu, F. Li, R. Zhang, and L. Zhang, “Explicit box detection unifies end-to-end multi-person pose estimation,” in International Conference on Learning Representations, 2023. URL: https://openreview.net/forum?id=s4WVupnJjmX.
  44. 44.H. Zhang, F. Li, S. Liu, L. Zhang, H. Su, J. Zhu, L. M. Ni, and H.-Y. Shum, “Dino: Detr with improved denoising anchor boxes for end-to-end object detection,” in ICLR 2023, 2023.
  45. 45.H. Zhang, F. Li, H. Xu, S. Huang, S. Liu, L. M. Ni, and L. Zhang, “Mp-former: Mask-piloted transformer for image segmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2023, pp. 18 074–18 083.
  46. 46.H. Zhang, F. Li, X. Zou, S. Liu, C. Li, J. Yang, and L. Zhang, “A simple framework for open-vocabulary segmentation and detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2023, pp. 1020–1031.
  47. 47.F. Li, H. Zhang, S. Liu, J. Guo, L. M. Ni, and L. Zhang, “Dn-detr: Accelerate detr training by introducing query denoising,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 13 619–13 627.
  48. 48.F. Li, H. Zhang, Y.-F. Zhang, S. Liu, J. Guo, L. M. Ni, P. Zhang, and L. Zhang, Vision-language intelligence: Tasks, representation learning, and large models, 2022. arXiv: 2203.01922 [cs.CV]. URL: https://arxiv.org/abs/2203.01922.
  49. 49.S. Liu, F. Li, H. Zhang, X. Yang, X. Qi, H. Su, J. Zhu, and L. Zhang, “DAB-DETR: Dynamic anchor boxes are better queries for DETR,” in International Conference on Learning Representations, 2022. URL: https://openreview.net/forum?id=oMI9PjOb9Jl.
  50. 50.X. Yang, S. Liu, Y. Dong, H. Su, L. Zhang, and J. Zhu, “Towards generalizable detection of face forgery via self-guided model-agnostic learning,” Pattern Recognition Letters, vol. 160, pp. 98–104, 2022, issn: 0167-8655. DOI: https://doi.org/10.1016/j.patrec.2022.06.007.
  51. 51.S. Liu, L. Zhang, X. Yang, H. Su, and J. Zhu, Query2label: A simple transformer way to multi-label classification, 2021. arXiv: 2107.10834 [cs.CV]. URL: https://arxiv.org/abs/2107.10834.
  52. 52.S. Liu, L. Zhang, X. Yang, H. Su, and J. Zhu, “Unsupervised part segmentation through disentangling appearance and shape,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2021, pp. 8355–8364.

Citation

MLA
Ren, T., et al. “Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks”. arXiv, 2024, https://doi.org/10.48550/arXiv.2401.14159.
APA
Ren, T., Liu, S., Zeng, A., Lin, J., Li, K., Cao, H., Chen, J., Huang, X., Chen, Y., Yan, F., Zeng, Z., Zhang, H., Li, F., Yang, J., Li, H., Jiang, Q., & Zhang, L. (2024). Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks. arXiv. https://doi.org/10.48550/arXiv.2401.14159
Chicago
Ren, T., S. Liu, A. Zeng, et al. 2024. “Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks”. Preprint, ArXiv. https://doi.org/10.48550/arXiv.2401.14159.
Harvard
Ren, T. et al. (2024) “Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks”. arXiv. Available at: https://doi.org/10.48550/arXiv.2401.14159.
Vancouver
1. Ren T, Liu S, Zeng A, et al (2024) Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks. https://doi.org/10.48550/arXiv.2401.14159

BibTeX

@misc{https://doi.org/10.48550/arxiv.2401.14159,
  doi = {10.48550/ARXIV.2401.14159},
  url = {https://arxiv.org/abs/2401.14159},
  author = {Ren, Tianhe and Liu, Shilong and Zeng, Ailing and Lin, Jing and Li, Kunchang and Cao, He and Chen, Jiayu and Huang, Xinyu and Chen, Yukang and Yan, Feng and Zeng, Zhaoyang and Zhang, Hao and Li, Feng and Yang, Jie and Li, Hongyang and Jiang, Qing and Zhang, Lei},
  keywords = {Computer Vision and Pattern Recognition (cs.CV), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks},
  publisher = {arXiv},
  year = {2024},
  copyright = {arXiv.org perpetual, non-exclusive license}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF