The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

Weiyun WangMin ShiQingyun LiWenhai WangZhenhang HuangLinjie XingZhe ChenHao LiXizhou ZhuZhiguo Cao

article2024ICLR132 citations

Introduces a billion-region dataset covering 3.5 million concepts alongside a unified vision-language model that achieves strong zero-shot performance across region-level recognition, captioning, and question answering in the open world.

Listen

Current artificial intelligence systems struggle to perceive and understand specific objects within visual scenes in the same detailed manner that humans do. While modern large language models excel at text reasoning, and visual-language models can interpret whole images, existing models generally fail to comprehend individual visual regions. Progress has been blocked by a scarcity of large-scale, detailed region-level data and a lack of unified models capable of handling both discriminative tasks like object recognition and generative tasks like text captioning.

The article introduces the All-Seeing project to establish a comprehensive framework for panoptic visual recognition and understanding across the open world. It evaluates a semi-automated data generation engine alongside a location-aware multimodal foundation model designed to perceive, identify, and describe arbitrary regions within images.

To overcome the prohibitive expense of manual labeling, the authors built an automated data engine that pairs off-the-shelf vision and language models with human-in-the-loop verification. This engine produced the AS-1B dataset, containing 1.2 billion region annotations across 11 million images, covering 3.5 million concepts and 132.2 billion text tokens. Leveraging this resource, the authors developed the All-Seeing Model (ASM), which combines a location-aware image tokenizer with a large language model decoder to unify region-text alignment and text generation using shared parameters.

The evaluation produced several key findings. First, the resulting dataset provides 33 times more annotated regions and 159 times more open-world concepts than prior benchmarks. Second, in zero-shot object recognition, ASM outperformed standard contrastive models by 10.4 mean Average Precision on COCO and 14.3 on LVIS. Third, ASM achieved state-of-the-art results in region-level and image-level captioning, outperforming concurrent models on benchmarks such as RefCOCOg and Flickr30K. Finally, human evaluations confirmed that ASM generated more informative, accurate, and hallucination-free descriptions than leading competitors.

These results demonstrate that integrating region-level visual grounding with language decoders significantly improves perceptual precision while reducing costly factual errors. Organizations building downstream vision-language applications can leverage this unified approach to reduce training costs and streamline architectures across diverse visual tasks. Teams seeking to adopt this technology should implement iterative human-in-the-loop workflows, as fine-tuning on a small set of human-verified annotations yielded substantial performance gains.

Readers should note that automated pipelines generate noisy initial data, particularly for small or low-resolution regions where language models struggle. While human verification and closed-set detectors effectively mitigated these errors during testing, continuous oversight remains necessary when deploying such models to domain-specific or safety-critical visual environments.

  • Paper: Panoptic Segmentation, Alexander Kirillov et al. (2018). This paper defines panoptic segmentation—the unified recognition of amorphous regions and countable objects—that the All-Seeing model extends to open-world, language-guided visual understanding.
  • Paper: Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations, Ranjay Krishna et al. (2016). Visual Genome establishes the dense region descriptions, attributes, relationships, and visual question-answer pairs that directly anticipate the annotation structure of AS-1B.
  • Paper: Segment Anything, Alexander M. Kirillov et al. (2023). Segment Anything provides the promptable segmentation model and iterative human-in-the-loop data engine that inform the All-Seeing project's scalable region annotation pipeline.
  • Paper: LVIS: A Dataset for Large Vocabulary Instance Segmentation, Agrim Gupta et al. (2019). LVIS introduces large-vocabulary instance segmentation with explicit attention to rare categories, preparing the reader for All-Seeing's coverage of millions of common and long-tail concepts.
  • Paper: PACO: Parts and Attributes of Common Objects, Vignesh Ramanathan et al. (2023). PACO supplies the fine-grained parts-and-attributes perspective that helps explain All-Seeing's emphasis on detailed concept descriptions beyond whole-object labels.
  • Paper: VQA: Visual Question Answering, Stanislaw Antol et al. (2015). VQA establishes the visual question-answering task that becomes one of All-Seeing's central region-level supervision signals and zero-shot capabilities.
  • Paper: The Open Images Dataset V4, Alina Kuznetsova et al. (2018). Open Images demonstrates how broad, unified annotations across classification, localization, and relationships can support large-scale visual recognition, a foundation expanded by AS-1B.
  • Paper: SAM 3: Segment Anything with Concepts, Nicolas Carion et al. (2025). SAM 3 extends the All-Seeing direction from open-world region recognition toward detecting, segmenting, and tracking every instance described by a visual or linguistic concept.
Cover for The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

Abstract

We present the All-Seeing (AS) project: a large-scale data and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates human feedback and efficient models in the loop, we create a new dataset (AS-1B) with over 1 billion regions annotated with semantic tags, question-answering pairs, and detailed captions. It covers a wide range of 3.5 million common and rare concepts in the real world, and has 132.2 billion tokens that describe the concepts and their attributes. Leveraging this new dataset, we develop the All-Seeing model (ASM), a unified framework for panoptic visual recognition and understanding. The model is trained with open-ended language prompts and locations, which allows it to generalize to various vision and language tasks with remarkable zero-shot performance, including region-text retrieval, region recognition, captioning, and question-answering. We hope that this project can serve as a foundation for vision-language artificial general intelligence research. Models and the dataset shall be released at this https URL, and demo can be seen at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The All-Seeing Dataset (AS-1B)
  • 3.1 Data Annotation Engine
  • 3.2 Open-World Localization
  • 3.3 Open-World Semantic
  • 3.3.1 Semantic Tags
  • 3.3.2 Detailed Descriptions
  • 3.4 Matching Location and Semantic
  • 3.5 Human Verification
  • 3.6 Data Engine Iteration
  • 4 The All-Seeing Model (ASM)
  • 4.1 Overal Architecture
  • 4.2 Location-Aware Image Tokenizer
  • 4.3 LLM-Based Decoder
  • 5 Data Analysis
  • 5.1 Data Scale
  • 5.2 Data Diversity
  • 5.3 Data Quality
  • 6 Experiments
  • 6.1 Implementation Details
  • 6.2 Text Generation
  • 6.3 Zero-shot Region Recognition
  • 6.4 Data Engineering
  • 6.5 Human Evaluation
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — All-Seeing 1B (AS-1B) Dataset Structure and Statistics

    data/table

    The All-Seeing 1B (AS-1B) dataset is a large-scale multimodal dataset designed for open-world panoptic visual recognition and understanding. Built upon 11 million base images from SA-1B, it provides dense region annotations covering 3.5 million distinct open-world concepts with 132.2 billion total tokens.

    Dataset #Images #Regions #Concepts #Tokens Location Semantic
    ImageNet-22K 15M – 22,000 – – Closed-Set
    COCO Caption 0.1M – – 8.4M – Closed-Set
    SBU 0.8M – – 14.6M – Open-World
    CC12M 12.4M – – 250.9M – Open-World
    YFCC15M 15M – – 1.0B – Open-World
    COYO700M 700M – – 15.0B – Open-World
    Laion-5B 5B – – 135.0B – Open-World
    SA-1B 11M 1.1B – – Open-World –
    COCO 0.1M 0.9M 80 – Closed-Set Closed-Set
    LVIS 0.1M 1.5M 1,203 – Closed-Set Closed-Set
    Objects365 0.6M 10.1M 365 – Closed-Set Closed-Set
    Open Images 1.5M 14.8M 600 – Closed-Set Closed-Set
    BigDetection 3.5M 36.0M 600 – Closed-Set Closed-Set
    V3Det 0.2M 1.5M 13,029 – Closed-Set Closed-Set
    Visual Genome 0.1M 0.3M 18,136 51.2M Open-World Open-World
    AS-1B (ours) 11M 1.2B 3.5M 132.2B Open-World Open-World

    The annotations in AS-1B include 1.2 billion bounding boxes paired with semantic tags, 3.3 billion visual question-answering pairs (averaging 10.50 question tokens and 16.91 answer tokens per pair), and 1.2 billion detailed region captions (averaging 34.84 tokens). Region scales follow a normal distribution across tiny (<202<20^2 px, 4.2%), small (202∼40220^2 \sim 40^2 px, 8.7%), medium (402∼100240^2 \sim 100^2 px, 35.8%), large (1002∼2002100^2 \sim 200^2 px, 23.7%), xlarge (2002∼5002200^2 \sim 500^2 px, 18.3%), and huge (>5002>500^2 px, 9.5%).

  2. Knowl 2 — All-Seeing Model (ASM) Architecture and Location-Aware Tokenizer

    model/method

    The All-Seeing Model (ASM) is a unified vision-language foundation model designed to handle both discriminative (retrieval/matching) and generative (captioning/VQA) tasks at image and region levels. The architecture consists of two primary modules:

    1. Location-Aware Image Tokenizer: Given an input image, a vision transformer backbone (ViT-g/14) extracts spatial feature maps F∈RH×W×D\mathcal{F} \in \mathbb{R}^{H \times W \times D}, where H,WH, W are spatial dimensions and DD is feature channel dimension. For NN specified regions (given as bounding boxes, masks, or point sets), RoIAlign extracts region features R∈RHr×Wr×D\mathcal{R} \in \mathbb{R}^{H_r \times W_r \times D}. These features are flattened and mapped via two fully connected (FC) layers into region queries Qr∈RG×DqQ_r \in \mathbb{R}^{G \times D_q}, where GG is the query group size and DqD_q is the query token dimension. Randomly initialized query tokens Q′∈RG×DqQ' \in \mathbb{R}^{G \times D_q} represent the global image. Concatenating Q′Q' with the NN region queries produces location-aware query tokens Q∈R(N+1)G×DqQ \in \mathbb{R}^{(N+1)G \times D_q}. A 12-block Transformer decoder processes QQ alongside image features F\mathcal{F}, and the output is projected into the language model feature dimension DtD_t, yielding visual prompt tokens V∈R(N+1)G×Dt\mathcal{V} \in \mathbb{R}^{(N+1)G \times D_t}. If no region is specified, the bounding box defaults to the entire image.

    2. LLM-Based Decoder: The model employs Husky-7B as its core autoregressive language decoder, which receives visual tokens V\mathcal{V}, learnable soft task prompts, and text instruction tokens to execute both generative and discriminative objectives with shared parameters.

  3. Knowl 3 — Unified Task Prompting and Objective Formulation in ASM

    model/method

    To unify discriminative and generative vision-language tasks within a single autoregressive LLM decoder with shared parameters, ASM introduces soft task prompts and a special alignment token:

    1. Generative Tasks: The input token sequence is structured as: {Pg} ⟨bos⟩ Human: {V} Instruction Assistant:\{\mathcal{P}_g\} \ \langle\text{bos}\rangle \ \text{Human: } \{\mathcal{V}\} \ \text{Instruction} \ \text{Assistant:} where Pg∈RM×Dt\mathcal{P}_g \in \mathbb{R}^{M \times D_t} (M=5M=5, hidden dimension DtD_t) is a learnable generative task prompt, ⟨bos⟩\langle\text{bos}\rangle is the beginning-of-sequence token, and V\mathcal{V} are location-aware visual prompt tokens. The model generates response tokens autoregressively.

    2. Discriminative Tasks (Region-Text Matching): The vision and text sequences are embedded separately. The visual sequence prompt is: {Pd} ⟨bos⟩ Human: {V} Query ⟨align⟩\{\mathcal{P}_d\} \ \langle\text{bos}\rangle \ \text{Human: } \{\mathcal{V}\} \ \text{Query} \ \langle\text{align}\rangle and the text sequence prompt is: {Pd} ⟨bos⟩ Assistant: {Candidate Text} ⟨align⟩\{\mathcal{P}_d\} \ \langle\text{bos}\rangle \ \text{Assistant: } \{\text{Candidate Text}\} \ \langle\text{align}\rangle where Pd∈RM×Dt\mathcal{P}_d \in \mathbb{R}^{M \times D_t} is a learnable discriminative task prompt and ⟨align⟩\langle\text{align}\rangle is a dedicated token whose final output embedding gathers the holistic sequence representation. Similarity between region and text is computed directly on the normalized ⟨align⟩\langle\text{align}\rangle embeddings.

    3. Training Loss: The total loss is defined as: Ltotal=Lgen+Ldis\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{gen}} + \mathcal{L}_{\text{dis}} where Lgen\mathcal{L}_{\text{gen}} is the standard autoregressive cross-entropy language modeling loss over target tokens, and Ldis\mathcal{L}_{\text{dis}} is the symmetric InfoNCE contrastive loss over region and text embeddings computed at the ⟨align⟩\langle\text{align}\rangle token.

  4. Knowl 4 — Multi-Source Region Proposal Merging Algorithm

    algorithm

    Because different object detectors produce detection scores on incomparable scales, standard Non-Maximum Suppression (NMS) cannot merge region proposals from heterogeneous sources. The AS-1B localization pipeline merges proposals from class-agnostic models (SAM), closed-set detectors (InternImage-H, EVA-02), and visual grounding models (GLIP) using sequential IoU matching while preserving semantic tag metadata.

    Input: Existing region proposals R\mathcal{R}, new region proposals R′\mathcal{R}', IoU threshold TIoUT_{\text{IoU}}
    Output: Merged region proposals R\mathcal{R}
    for region r′∈R′r' \in \mathcal{R}' do
        Calculate IoU between r′r' and all proposals in R\mathcal{R}
        if max⁡r∈RIoU(r′,r)>TIoU\max_{r \in \mathcal{R}} \text{IoU}(r', r) > T_{\text{IoU}} then
            Merge semantic tags from r′r' into the tag list of the corresponding matching region r∈Rr \in \mathcal{R}
            Delete r′r'
        else
            Add r′r' into R\mathcal{R}
        end if
    end for

    In the data engine, R\mathcal{R} is initialized with SAM class-agnostic proposals (which account for 36.4% of final regions), and then sequentially merged with proposals from InternImage (20.5%), EVA-02 (22.5%), and GLIP (20.6%).

  5. Knowl 5 — Iterative Semi-Automatic Data Engine Loop

    algorithm

    To annotate over 1 billion regions cost-effectively, a 'data-human-model' loop progressively improves annotation quality by alternating between automated annotation generation, human verification, and model fine-tuning.

    Input: Iteration number nn, image collection I\mathcal{I}, model suite M\mathcal{M}, annotation pipeline P(M,I)P(\mathcal{M}, \mathcal{I})
    Output: Annotated dataset A\mathcal{A}, fine-tuned model suite M\mathcal{M}
    Generate initial coarse annotations A0←P(M,I)\mathcal{A}_0 \leftarrow P(\mathcal{M}, \mathcal{I}) using off-the-shelf foundation models
    Train baseline All-Seeing Model M0\mathcal{M}_0 using A0\mathcal{A}_0
    i←0i \leftarrow 0
    while i<ni < n do
        Sample annotations from Ai\mathcal{A}_i and perform human verification, yielding verified subset Ai′\mathcal{A}'_i
        Fine-tune model Mi\mathcal{M}_i on Ai′\mathcal{A}'_i to obtain updated model Mi+1\mathcal{M}_{i+1}
        Re-annotate and re-rank image dataset I\mathcal{I} via P(Mi+1,I)P(\mathcal{M}_{i+1}, \mathcal{I}) to produce Ai+1\mathcal{A}_{i+1}
        i←i+1i \leftarrow i + 1
    end while

    This loop reduces human verification time to approximately 10 seconds per region (for semantic tags) and 15 seconds per VQA pair, lowering human labor requirements by a factor of 1,000 relative to full manual annotation.

  6. Knowl 6 — Open-World Semantic and Detailed Description Generation Pipeline

    model/method

    The semantic generation module of the AS-1B data engine uses specialized LLM/VLLM agents to produce tags, attributes, and detailed captions:

    1. Semantic Tag Generation:

      • Spotter: MiniGPT-4 generates overall image captions from which noun phrases are extracted, supplemented by scene text from EasyOCR.
      • Imaginator: Vicuna infers plausible scene-specific co-occurring objects (e.g., classroom →\rightarrow chalkboard, teacher) expanded via web search.
      • Splitter: Vicuna decomposes physical object categories into component parts (e.g., airplane →\rightarrow wing, windshield), ignoring non-physical concepts.
      • Magnifier: BLIP generates captions for cropped region proposals to capture localized nouns missed in global descriptions.
    2. Semantic-Location Matching: Candidate tags are scored against image crops using CLIP (in round 1) or ASM (in subsequent rounds). To avoid mislabeling sub-objects (e.g., classifying a person wearing a backpack as 'backpack'), CLIPSeg generates segmentation masks for each candidate, modulating the alignment score by the predicted mask area.

    3. Detailed Region Descriptions:

      • Questioner: Vicuna generates 3 attribute/status questions per semantic tag.
      • Responder: Husky-7B answers each question based on the cropped visual region.
      • Writer: Vicuna synthesizes the three question-answer pairs into a coherent, single-sentence region caption.
  7. Knowl 7 — Zero-Shot Region Recognition Performance on COCO and LVIS

    empirical result

    When evaluated on zero-shot region recognition using ground-truth bounding boxes following the RegionCLIP evaluation protocol, the All-Seeing Model (ASM) and Region-Aware CLIP baseline (R-CLIP, OpenCLIP ViT-L/14 with an RoIAlign head trained on AS-1B) outperform standard vision-language models across object scales.

    Model COCO LVIS
    mAP APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L mAP APS\text{AP}_S APM\text{AP}_M APL\text{AP}_L
    CLIP 58.9 50.7 70.4 58.3 47.1 40.3 59.2 57.4
    OpenCLIP 63.3 47.8 75.6 60.9 49.1 37.4 62.8 66.5
    R-CLIP (baseline) 68.6 61.4 75.4 79.3 54.8 49.3 60.6 66.6
    ASM (ours) 69.3 64.3 78.0 71.0 61.4 56.7 67.9 69.2

    On the COCO benchmark (80 categories), ASM achieves 69.3 mAP (+10.4 over CLIP). On the 1,203-category LVIS benchmark, ASM achieves 61.4 mAP (+14.3 over CLIP and +6.6 over R-CLIP), showing substantial gains on small (APS=56.7\text{AP}_S = 56.7) and medium (APM=67.9\text{AP}_M = 67.9) regions.

  8. Knowl 8 — Region-Level and Image-Level Captioning Benchmark Results

    empirical result

    ASM evaluated on region-level captioning (Visual Genome, RefCOCOg) and zero-shot image-level captioning (Flickr30k, NoCaps) exhibits state-of-the-art text generation across scales.

    Region Captioning Performance:

    Model Zero-shot Visual Genome RefCOCOg
    Meteor CIDEr Meteor CIDEr
    GRiT No 17.1 142.0 15.2 71.6
    SLR No – – 15.4 59.2
    SLR+Rerank No – – 15.9 66.2
    Kosmos-2 Yes – – 12.2 60.3
    ASM Yes 12.6 44.2 13.6 41.9
    ASM-FT No 18.0 145.1 20.8 103.0

    Zero-Shot Image Captioning Performance:

    Model Zero-shot Flickr30k NoCaps
    CIDEr SPICE CIDEr SPICE
    Flamingo-9B Yes 61.5 – – –
    BLIP-2 Yes – – 121.6 15.8
    InstructBLIP Yes 82.8 – 123.1 –
    Shikra-13B Yes 73.9 – – –
    Kosmos-2 Yes 66.7 – – –
    ASM (ours) Yes 77.9 17.3 104.8 14.5
    ASM-FT (ours) Yes 87.7 18.7 117.2 15.6

    Under zero-shot evaluation on RefCOCOg, ASM achieves 13.6 Meteor (+1.4 over Kosmos-2). With second-stage fine-tuning (ASM-FT), it achieves 103.0 CIDEr on RefCOCOg and 87.7 CIDEr on Flickr30k.

  9. Knowl 9 — Ablation of Data Scale, Cleaning, and Human Verification

    empirical result

    Ablation experiments on the AS-1B data pipeline using the R-CLIP baseline demonstrate the separate quantitative benefits of training data scaling, automated candidate cleaning, and human-in-the-loop verification:

    1. Data Scaling: Increasing training semantic tags from 1M to 5M images improves COCO mAP from 67.8 to 68.6 and LVIS mAP from 54.0 to 54.8.
    2. Data Cleaning: Filtering the raw 2.14B region proposals down to 1.2B (by taking the top 100 CLIP-scoring regions per image and re-ranking tags with CLIPSeg mask areas) improves zero-shot recognition by +6.0 mAP on COCO (61.8 to 67.8) and +7.5 mAP on LVIS (46.5 to 54.0).
    3. Human Feedback Fine-Tuning: Fine-tuning on only 40k human-verified annotations for 4 epochs using human-marked positive tags and unselected candidates as hard negatives yields consistent improvements:
    Human Data Input Scale COCO mAP LVIS mAP
    No 224 67.8 54.8
    Yes 224 70.2 (+2.4) 55.0 (+0.2)
    No 896 76.7 65.7
    Yes 896 80.0 (+3.3) 68.4 (+2.7)
  10. Knowl 10 — Human Preference Evaluation on Generated Captions

    empirical result

    In a user study with 5 human evaluators assessing 20 random samples per dataset across Visual Genome, RefCOCOg, Flickr30K, and NoCaps, participants selected the most informative caption free of hallucinations and factual errors.

    Model / Source Visual Genome RefCOCOg Flickr30K NoCaps
    Win Rate (%) Length Win Rate (%) Length Win Rate (%) Length Win Rate (%) Length
    Human 47.8 13.6 10.3 6.3 30.0 16.0 27.3 15.1
    LLaVA 4.3 110.8 15.4 100.6 17.5 114.0 9.1 108.4
    MiniGPT4 8.7 110.9 15.4 113.5 14.2 114.6 13.6 101.0
    ASM (ours) 39.2 37.5 46.1 33.6 38.3 112.4 50.0 102.1

    While LLaVA and MiniGPT4 generate long descriptions (~100–114 words), they exhibit frequent hallucinations and over-association on region tasks. ASM produces region captions of moderate length (33.6–37.5 words) with fewer factual errors, winning 46.1% preference on RefCOCOg and 50.0% on NoCaps, outperforming ground-truth human annotations which were deemed overly brief.

Coverage note — None was omitted; all key contributed data specifications, architectural mechanisms, merging/iteration algorithms, quantitative benchmark results, data engineering ablations, and human evaluations are represented in the knowls.

References

  1. 1.Harsh Agrawal, Karan Desai, Yufei Wang, Xinlei Chen, Rishabh Jain, Mark Johnson, Dhruv Batra, Devi Parikh, Stefan Lee, and Peter Anderson. Nocaps: Novel object captioning at scale. In Int. Conf. Comput. Vis., 2019.
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al. Flamingo: a visual language model for few-shot learning. Adv. Neural Inform. Process. Syst., 2022.
  3. 3.Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. Spice: Semantic propositional image caption evaluation. In Eur. Conf. Comput. Vis., 2016.
  4. 4.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In Int. Conf. Comput. Vis., 2015.
  5. 5.Satanjeev Banerjee and Alon Lavie. METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. In Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare R. Voss, editors, Proceedings of the Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization@ACL 2005, Ann Arbor, Michigan, USA, June 29, 2005, pages 65–72. Association for Computational Linguistics, 2005.
  6. 6.Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Songhao Piao, and Furu Wei. Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. Adv. Neural Inform. Process. Syst., 2022.
  7. 7.Hangbo Bao, Wenhui Wang, Li Dong, and Furu Wei. Vl-beit: Generative vision-language pretraining. arXiv preprint arXiv:2206.01127, 2022.
  8. 8.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Adv. Neural Inform. Process. Syst., 2020.
  9. 9.Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim. Coyo-700m: Image-text pair dataset. https://github.com/kakaobrain/coyo-dataset, 2022.
  10. 10.Likun Cai, Zhi Zhang, Yi Zhu, Li Zhang, Mu Li, and Xiangyang Xue. Bigdetection: A large-scale benchmark for improved object detector pre-training. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  11. 11.Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In IEEE Conf. Comput. Vis. Pattern Recog., 2021.
  12. 12.Guo Chen, Yin-Dong Zheng, Jiahao Wang, Jilan Xu, Yifei Huang, Junting Pan, Yi Wang, Yali Wang, Yu Qiao, Tong Lu, et al. Videollm: Modeling video sequence with large language models. arXiv preprint arXiv:2305.13292, 2023.
  13. 13.Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao. Shikra: Unleashing multimodal llm’s referential dialogue magic. arXiv preprint arXiv:2306.15195, 2023.
  14. 14.Shoufa Chen, Peize Sun, Yibing Song, and Ping Luo. Diffusiondet: Diffusion model for object detection. arXiv preprint arXiv:2211.09788, 2022.
  15. 15.Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015.
  16. 16.Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Yu Qiao. Vision transformer adapter for dense predictions. In Int. Conf. Learn. Represent., 2023.
  17. 17.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023), 2023.
  18. 18.EasyOCR contributors. Easyocr. https://github.com/JaidedAI/EasyOCR, 2023.
  19. 19.MOSS contributors. Moss. https://github.com/OpenLMLab/MOSS, 2023.
  20. 20.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In IEEE Conf. Comput. Vis. Pattern Recog., 2016.
  21. 21.Yiming Cui, Ziqing Yang, and Xin Yao. Efficient and effective text encoding for chinese llama and alpaca. arXiv preprint arXiv:2304.08177, 2023.
  22. 22.Wenliang Dai, Junnan Li, Dongxu Li, AnthonyMeng Huat, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning. arXiv preprint arXiv:2305.06500, 2023.
  23. 23.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In IEEE Conf. Comput. Vis. Pattern Recog., 2009.
  24. 24.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. Int. J. Comput. Vis., 88:303–338, 2010.
  25. 25.Yuxin Fang, Quan Sun, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. Eva-02: A visual representation for neon genesis. arXiv preprint arXiv:2303.11331, 2023.
  26. 26.Yuxin Fang, Wen Wang, Binhui Xie, Quan Sun, Ledell Wu, Xinggang Wang, Tiejun Huang, Xinlong Wang, and Yue Cao. Eva: Exploring the limits of masked visual representation learning at scale. In IEEE Conf. Comput. Vis. Pattern Recog., 2023.
  27. 27.Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al. Datacomp: In search of the next generation of multimodal datasets. arXiv preprint arXiv:2304.14108, 2023.
  28. 28.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  29. 29.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. Pal: Program-aided language models. arXiv preprint arXiv:2211.10435, 2022.
  30. 30.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In IEEE Conf. Comput. Vis. Pattern Recog., 2017.
  31. 31.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2019.
  32. 32.Yaru Hao, Haoyu Song, Li Dong, Shaohan Huang, Zewen Chi, Wenhui Wang, Shuming Ma, and Furu Wei. Language models are general-purpose interfaces. arXiv preprint arXiv:2206.06336, 2022.
  33. 33.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. In Int. Conf. Comput. Vis., 2017.
  34. 34.Xiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, and Lijuan Wang. Scaling up vision-language pre-training for image captioning. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  35. 35.Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Qiang Liu, et al. Language is not all you need: Aligning perception with language models. arXiv preprint arXiv:2302.14045, 2023.
  36. 36.Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. Openclip. July 2021.
  37. 37.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning., 2021.
  38. 38.Sebastian Kalkowski, Christian Schulze, Andreas Dengel, and Damian Borth. Real-time analysis and visualization of the yfcc100m dataset. In Proceedings of the 2015 workshop on community-organized multimodal mining: opportunities for novel solutions, pages 25–30, 2015.
  39. 39.Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. Mdetr-modulated detection for end-to-end multi-modal understanding. In Int. Conf. Comput. Vis., 2021.
  40. 40.Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, et al. Icdar 2015 competition on robust reading. In 2015 13th International Conference on Document Analysis and Recognition (ICDAR), pages 1156–1160. IEEE, 2015.
  41. 41.Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Dollár. Panoptic segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2019.
  42. 42.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv preprint arXiv:2304.02643, 2023.
  43. 43.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. Int. J. Comput. Vis., 123:32–73, 2017.
  44. 44.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  45. 45.Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al. The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. Int. J. Comput. Vis., 128(7):1956–1981, 2020.
  46. 46.Feng Li, Hao Zhang, Huaizhe Xu, Shilong Liu, Lei Zhang, Lionel M Ni, and Heung-Yeung Shum. Mask dino: Towards a unified transformer-based framework for object detection and segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2023.
  47. 47.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597, 2023.
  48. 48.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning., 2022.
  49. 49.Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. Align before fuse: Vision and language representation learning with momentum distillation. Adv. Neural Inform. Process. Syst., 2021.
  50. 50.KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355, 2023.
  51. 51.Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al. Grounded language-image pre-training. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  52. 52.Yanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer, and Kaiming He. Scaling language-image pre-training via masking. In IEEE Conf. Comput. Vis. Pattern Recog., 2023.
  53. 53.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Eur. Conf. Comput. Vis., 2014.
  54. 54.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023.
  55. 55.Zhaoyang Liu, Yinan He, Wenhai Wang, Weiyun Wang, Yi Wang, Shoufa Chen, Qinglong Zhang, Yang Yang, Qingyun Li, Jiashuo Yu, et al. Interngpt: Solving vision-centric tasks by interacting with chatbots beyond language. arXiv preprint arXiv:2305.05662, 2023.
  56. 56.Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al. The flan collection: Designing data and methods for effective instruction tuning. arXiv preprint arXiv:2301.13688, 2023.
  57. 57.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In Int. Conf. Learn. Represent., 2019.
  58. 58.Timo Lüddecke and Alexander Ecker. Image segmentation using text and image prompts. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  59. 59.Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and comprehension of unambiguous object descriptions. In IEEE Conf. Comput. Vis. Pattern Recog., 2016.
  60. 60.Yao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang, Mingyu Ding, Jun Jin, Bin Wang, Jifeng Dai, Yu Qiao, and Ping Luo. Embodiedgpt: Vision-language pre-training via embodied chain of thought. arXiv preprint arXiv:2305.15021, 2023.
  61. 61.OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  62. 62.TB OpenAI. Chatgpt: Optimizing language models for dialogue. OpenAI, 2022.
  63. 63.Vicente Ordonez, Girish Kulkarni, and Tamara Berg. Im2text: Describing images using 1 million captioned photographs. Adv. Neural Inform. Process. Syst., 2011.
  64. 64.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Adv. Neural Inform. Process. Syst., 2022.
  65. 65.Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei. Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023.
  66. 66.Shi Qiu, Xiaogang Wang, and Xiaoou Tang. Visual semantic complex network for web images. In Int. Conf. Comput. Vis., 2013.
  67. 67.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning., 2021.
  68. 68.Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. OpenAI, 2018.
  69. 69.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI, 2019.
  70. 70.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
  71. 71.Christoph Schuhman, Andreas Köpf, Richard Vencu, Theo Coombes, and Romain Beaumont. Laion coco: 600m synthetic captions from laion2b-en. 2022.
  72. 72.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Adv. Neural Inform. Process. Syst., 2022.
  73. 73.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021.
  74. 74.Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. In Int. Conf. Comput. Vis., 2019.
  75. 75.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Ann. Meeting of the Assoc. for Comput. Linguistics, 2018.
  76. 76.Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface. arXiv preprint arXiv:2303.17580, 2023.
  77. 77.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. Alpaca: A strong, replicable instruction-following model. Stanford Center for Research on Foundation Models. https://crfm. stanford. edu/2023/03/13/alpaca. html, 2023.
  78. 78.InternLM Team. Internlm: A multilingual language model with progressively enhanced capabilities. 2023.
  79. 79.Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. Yfcc100m: The new data in multimedia research. Communications of the ACM, 59(2):64–73, 2016.
  80. 80.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  81. 81.Anne M Treisman and Garry Gelade. A feature-integration theory of attention. Cognitive psychology, 12(1):97–136, 1980.
  82. 82.Trieu H Trinh and Quoc V Le. A simple method for commonsense reasoning. arXiv preprint arXiv:1806.02847, 2018.
  83. 83.Daniel A Updegrove, Sheldon B Smith, and Wendy Rickard Bollentin. Ccnews: An online forum for newsletter editors. In Proceedings of the 16th annual ACM SIGUCCS conference on user services, 1988.
  84. 84.Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. The inaturalist species classification and detection dataset. In IEEE Conf. Comput. Vis. Pattern Recog., 2018.
  85. 85.Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evaluation. In IEEE Conf. Comput. Vis. Pattern Recog., 2015.
  86. 86.Jiaqi Wang, Pan Zhang, Tao Chu, Yuhang Cao, Yujie Zhou, Tong Wu, Bin Wang, Conghui He, and Dahua Lin. V3det: Vast vocabulary visual detection dataset. arXiv preprint arXiv:2304.03752, 2023.
  87. 87.Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, et al. Visionllm: Large language model is also an open-ended decoder for vision-centric tasks. arXiv preprint arXiv:2305.11175, 2023.
  88. 88.Wenhai Wang, Jifeng Dai, Zhe Chen, Zhenhang Huang, Zhiqi Li, Xizhou Zhu, Xiaowei Hu, Tong Lu, Lewei Lu, Hongsheng Li, et al. Internimage: Exploring large-scale vision foundation models with deformable convolutions. In Int. Conf. Comput. Vis., 2023.
  89. 89.Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv preprint arXiv:2208.10442, 2022.
  90. 90.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language model with self generated instructions. Ann. Meeting of the Assoc. for Comput. Linguistics, 2022.
  91. 91.Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao. Simvlm: Simple visual language model pretraining with weak supervision. Int. Conf. Learn. Represent., 2022.
  92. 92.Jialian Wu, Jianfeng Wang, Zhengyuan Yang, Zhe Gan, Zicheng Liu, Junsong Yuan, and Lijuan Wang. Grit: A generative region-to-text transformer for object understanding. arXiv preprint arXiv:2212.00280, 2022.
  93. 93.Enze Xie, Peize Sun, Xiaoge Song, Wenhai Wang, Xuebo Liu, Ding Liang, Chunhua Shen, and Ping Luo. Polarmask: Single shot instance segmentation with polar representation. In IEEE Conf. Comput. Vis. Pattern Recog., 2020.
  94. 94.Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Ehsan Azarnasab, Faisal Ahmed, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang. Mm-react: Prompting chatgpt for multimodal reasoning and action. arXiv preprint arXiv:2303.11381, 2023.
  95. 95.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. Int. Conf. Learn. Represent., 2022.
  96. 96.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178, 2023.
  97. 97.Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. TACL, 2:67–78, 2014.
  98. 98.Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. Trans. Mach. Learn. Res., 2022.
  99. 99.Licheng Yu, Hao Tan, Mohit Bansal, and Tamara L Berg. A joint speaker-listener-reinforcer model for referring expressions. In IEEE Conf. Comput. Vis. Pattern Recog., 2017.
  100. 100.Sha Yuan, Hanyu Zhao, Zhengxiao Du, Ming Ding, Xiao Liu, Yukuo Cen, Xu Zou, Zhilin Yang, and Jie Tang. Wudaocorpora: A super large-scale chinese corpora for pre-training language models. AI Open, 2:65–68, 2021.
  101. 101.Liu Yuliang, Jin Lianwen, Zhang Shuaitao, and Zhang Sheng. Detecting curve text in the wild: New dataset and new solution. arXiv preprint arXiv:1712.02170, 2017.
  102. 102.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al. Glm-130b: An open bilingual pre-trained model. Int. Conf. Learn. Represent., 2022.
  103. 103.Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. Vinvl: Making visual representations matter in vision-language models. arXiv preprint arXiv:2101.00529, 2021.
  104. 104.Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199, 2023.
  105. 105.Shilong Zhang, Peize Sun, Shoufa Chen, Min Xiao, Wenqi Shao, Wenwei Zhang, Kai Chen, and Ping Luo. Gpt4roi: Instruction tuning large language model on region-of-interest. arXiv preprint arXiv:2307.03601, 2023.
  106. 106.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  107. 107.Liang Zhao, En Yu, Zheng Ge, Jinrong Yang, Haoran Wei, Hongyu Zhou, Jianjian Sun, Yuang Peng, Runpei Dong, Chunrui Han, et al. Chatspot: Bootstrapping multimodal llms via precise referring instruction tuning. arXiv preprint arXiv:2307.09474, 2023.
  108. 108.Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, and Jianfeng Gao. Regionclip: Region-based language-image pretraining. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  109. 109.Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Semantic understanding of scenes through the ade20k dataset. Int. J. Comput. Vis., 127:302–321, 2019.
  110. 110.Deyao Zhu, Jun Chen, Kilichbek Haydarov, Xiaoqian Shen, Wenxuan Zhang, and Mohamed Elhoseiny. Chatgpt asks, blip-2 answers: Automatic questioning towards enriched visual descriptions. arXiv preprint arXiv:2303.06594, 2023.
  111. 111.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023.
  112. 112.Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, et al. Ghost in the minecraft: Generally capable agents for open-world enviroments via large language models with text-based knowledge and memory. arXiv preprint arXiv:2305.17144, 2023.
  113. 113.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. In Int. Conf. Learn. Represent., 2021.
  114. 114.Xizhou Zhu, Jinguo Zhu, Hao Li, Xiaoshi Wu, Hongsheng Li, Xiaohua Wang, and Jifeng Dai. Uni-perceiver: Pre-training unified architecture for generic perception for zero-shot and few-shot tasks. In IEEE Conf. Comput. Vis. Pattern Recog., 2022.
  115. 115.Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. Visual7w: Grounded question answering in images. In IEEE Conf. Comput. Vis. Pattern Recog., 2016.
  116. 116.Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Int. Conf. Comput. Vis., 2015.

Citation

MLA
Wang, W., et al. “The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World”. arXiv, 2023, http://arxiv.org/abs/2308.01907v1.
APA
Wang, W., Shi, M., Li, Q., Wang, W., Huang, Z., Xing, L., Chen, Z., Li, H., Zhu, X., Cao, Z., Chen, Y., Lu, T., Dai, J., & Qiao, Y. (2023). The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World. arXiv. http://arxiv.org/abs/2308.01907v1
Chicago
Wang, W., M. Shi, Q. Li, et al. 2023. “The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World”. arXiv. http://arxiv.org/abs/2308.01907v1.
Harvard
Wang, W. et al. (2023) “The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2308.01907v1.
Vancouver
1. Wang W, Shi M, Li Q, et al (2023) The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World. arXiv

BibTeX

@article{wang2023the,
  title = {The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World},
  author = {Wang, Weiyun and Shi, Min and Li, Qingyun and Wang, Wenhai and Huang, Zhenhang and Xing, Linjie and Chen, Zhe and Li, Hao and Zhu, Xizhou and Cao, Zhiguo and Chen, Yushi and Lu, Tong and Dai, Jifeng and Qiao, Yu},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2308.01907v1},
  eprint = {2308.01907}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors