MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Kaining YingFanqing MengJin WangZhiqian LiHan LinYue YangHao ZhangWenbo ZhangYuqi LinShuo Liu

article2024ICML215 citations

Presents MMT-Bench, an extensive evaluation benchmark spanning over 31,000 questions across 162 vision-language tasks, to expose key performance limits in advanced multimodal models and map their progress toward general visual intelligence.

Listen

Artificial intelligence systems that integrate vision and text are advancing rapidly, driving interest in multimodal artificial general intelligence capable of performing diverse human-level tasks across various domains. However, standard evaluation benchmarks remain constrained by narrow task coverage and simple visual questions, failing to effectively measure progress toward broad, expert-level visual intelligence. This creates a critical need for rigorous testing environments that assess how well models generalize across diverse image modalities and complex reasoning scenarios.

The article introduces MMT-Bench, a comprehensive benchmark designed to evaluate large vision-language models across massive multimodal tasks requiring expert knowledge, spatial localization, fine-grained perception, and deliberate reasoning. It aims to quantify current system performance, uncover task interdependencies, and identify specific domains where existing models succeed or struggle.

To establish this benchmark, the researchers curated 31,325 multiple-choice visual questions spanning 32 core meta-tasks, 162 subtasks, and 13 distinct visual input types such as natural scenes, medical images, depth maps, and graphical user interfaces. Using this dataset, the authors evaluated 32 leading closed-source and open-source models under standardized protocols. In addition, the study mapped the relationships among tasks by constructing task vectors through parameter-efficient fine-tuning, grouping tasks into clusters to analyze in-domain and out-of-domain performance trends.

The evaluation produced several critical findings regarding current model capabilities. First, the benchmark presents a significant challenge to state-of-the-art models: the top-performing closed-source system, GPT-4o, achieved an overall accuracy of only 65.5%, which fell to 59.5% when standard visual recognition tasks were excluded. Second, leading open-source models demonstrated strong competitiveness, with InternVL-Chat-v1.2 scoring 63.4% and outperforming several proprietary systems like GPT-4V (61.1%) and GeminiProVision (61.6%). Third, detailed error analyses showed that model failures are predominantly driven by perception errors (51% to 77% across top models) and complex reasoning deficits. Finally, task mapping revealed that while models excel at high-level recognition and captioning, they consistently underperform in out-of-domain areas requiring fine-grained spatial localization, coordinate detection, and user interface navigation.

These results demonstrate that high performance on conventional recognition tasks does not translate to robust spatial reasoning or operational execution in specialized environments. For organizations deploying these technologies, the findings indicate substantial operational risks when using current foundation models for tasks requiring precise pixel-level localization, document structuring, or graphical interface interaction. Notably, instruction tuning does not uniformly improve generalization, as un-tuned baseline models outperformed several instruction-tuned counterparts on structured closed-set evaluations.

To improve multimodal performance, developers should expand training regimens to include visual referring data, normalized coordinates, and multi-image sequence inputs. Organizations evaluating or deploying vision-language models should prioritize addressing core perception and spatial reasoning limitations rather than relying exclusively on high-level recognition benchmarks. Future work should focus on incorporating a broader range of multimodal tasks and developing training techniques that mitigate performance trade-offs across distinct task domains.

The benchmark's findings should be interpreted with awareness of its structured multiple-choice design and curated data sources, which may introduce domain weighting biases. While the evaluation demonstrates high internal consistency and reliability across tested models, stakeholders should exercise caution when extrapolating these scores to unconstrained real-world environments with unrepresented demographic contexts or visual modalities.

Cover for MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Abstract

Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of multimodal tasks testing rudimentary capabilities, falling short in tracking LVLM development. In this study, we present MMT-Bench, a comprehensive benchmark designed to assess LVLMs across massive multimodal tasks requiring expert knowledge and deliberate visual recognition, localization, and reasoning. MMT-Bench comprises 31,325 meticulously curated multi-choice visual questions from various multimodal scenarios such as vehicle driving and embodied navigation, covering 32 core meta-tasks and 162 subtasks in multimodal understanding. Due to its extensive task coverage, MMT-Bench enables the evaluation of LVLMs using a task map, facilitating the discovery of in- and out-of-domain tasks. Evaluation results involving 32 LVLMs such as the proprietary GPT-4o, GeminiProVision, and open-sourced InternVL-Chat, underscore the significant challenges posed by MMT-Bench. We anticipate that MMT-Bench will inspire the community to develop next-generation multimodal foundation models aimed at achieving general-purpose multimodal intelligence.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. MMT-Bench
  • 3.1. Hierarchical Task Structure
  • 3.2. Data Collection
  • 4. Experiments
  • 4.1. Evaluation Details
  • 4.2. Overall Evaluation
  • 4.3. Specific Task and Prompt Methods Analysis
  • 4.4. Error Analysis
  • 5. Taskonomy Analysis
  • 5.1. Analytical Tools
  • 5.2. Findings on Task Map
  • 6. Conclusion and Discussion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Task Map
  • B. Hierarchical Structure of MMT-Bench
  • C. Task Abbreviations
  • D. More Experimental Details
  • D.1. LVLMs Model Details
  • D.2. Multi-Images Prompt Experimental Details
  • D.3. Visual Referring Prompting Experimental Details
  • E. Pixel Coordinates vs. Normalized Coordinates
  • F. Analysis on Images Types and Capabilities
  • G. Case Study
  • H. Comparison of MMT-Bench with Other Benchmarks on OCR-Related Tasks
  • I. Some Details about the Benchmark Construction
  • I.1. Metadata
  • I.2. Prepare the Answer and Options
  • I.3. Statistics of Image and Video in MMT-Bench
  • J. OpenCompass' Protocol
  • K. Computaional Resources
  • L. Detailed Main Results

Knowls

  1. Knowl 1 — MMT-Bench Dataset and Benchmark Architecture

    definition

    MMT-Bench is a large-scale multimodal benchmark designed to evaluate Large Vision-Language Models (LVLMs) across massive multimodal tasks towards multitask Artificial General Intelligence (AGI).

    Benchmark Composition

    • Total Samples: 31,32531,325 curated multi-choice visual questions (maximum 8 choices per question).
    • Task Hierarchy: 3232 core meta-tasks decomposed into 162162 subtasks.
    • Input Modalities and Visual Data Types: Covers single images (25,73225,732), image pairs / multiple images (3,8003,800 sets comprising 14,80014,800 images), and video clips (1,7931,793 videos comprising 10,57210,572 frames) across 13 distinct visual types (natural scenes, synthetic images, depth maps, text-rich images, paintings, screenshots, point clouds, medical images, scientific diagrams, abstract drawings, remote sensing images, visual marks, and charts).
    • Evaluated Capabilities: Spans 14 multimodal capability dimensions: visual recognition, visual localization, visual reasoning, optical character recognition (OCR), counting, 3D perception, temporal understanding, pixel perception, expert knowledge utilization, multi-image analysis, visual description, retrieval, visual prompting understanding, and embodied planning.
    • Sampling Constraints: To ensure evaluation efficiency, each subtask caps test samples at a maximum of 200200 samples via uniform random sampling across source datasets.
  2. Knowl 2 — Task Vector and Task Map Construction via LoRA Probing

    model/method

    To evaluate structural relationships between diverse multimodal tasks, MMT-Bench formalizes each subtask tt as a continuous task vector VtV^t based on model parameter weight displacement after probing.

    Task Vector Formulation

    Let W0W_0 denote the initial pretrained weights of a probing model (specifically, pretrained QwenVL-Chat). For each subtask t∈{1,…,T}t \in \{1, \dots, T\} with instruction-formatted task data Dt\mathcal{D}^t, the task vector VtV^t is computed as the parameter variation after fine-tuning: Vt=arg⁡min⁡WL(W∣Dt)−W0V^t = \arg\min_{W} \mathcal{L}(W \mid \mathcal{D}^t) - W_0 where L\mathcal{L} represents the task loss. Rather than full fine-tuning, Low-Rank Adaptation (LoRA) fine-tuning is conducted for 3 epochs per subtask, reducing the parameter dimensionality of the task vector from 9.6B to 3.5M parameters.

    Task Map Distance Matrix

    For T=162T = 162 subtasks, the subtask relationships form a task map graph represented by the distance matrix G={Gst}s,t=1TG = \{G^{st}\}_{s,t=1}^T, where the pairwise distance Gst∈[0,2]G^{st} \in [0, 2] is defined by cosine distance: Gst=1−cos⁡(Vs,Vt)=1−⟨Vs,Vt⟩∥Vs∥2∥Vt∥2G^{st} = 1 - \cos(V^s, V^t) = 1 - \frac{\langle V^s, V^t \rangle}{\|V^s\|_2 \|V^t\|_2} where ⟨⋅,⋅⟩\langle \cdot, \cdot \rangle is the Euclidean inner product.

  3. Knowl 3 — Kendall's Tau Ranking Consistency Across Task Proximity

    theoretical result

    To measure the agreement of relative model performance between two subtasks ss and tt, the benchmark uses Kendall's rank correlation coefficient (Kendall's tau τst\tau^{st}) across a suite of MM evaluated LVLMs.

    Formulation

    Let PmsP_m^s denote the evaluation performance (accuracy) of model mm on subtask ss. The pairwise ranking correlation is defined as: τst=2M(M−1)∑1≤m<n≤Msign((Pms−Pns)(Pmt−Pnt))\tau^{st} = \frac{2}{M(M-1)} \sum_{1 \le m < n \le M} \text{sign}\left((P_m^s - P_n^s)(P_m^t - P_n^t)\right) where sign(x)=1\text{sign}(x) = 1 if x>0x > 0, −1-1 if x<0x < 0, and 00 if x=0x = 0. τst=1\tau^{st} = 1 indicates identical model rankings across both tasks.

    Distance-Thresholded Consistency

    For a normalized distance threshold δ\delta on the task map, let Δs={t:Gst≤δ}\Delta_s = \{t : G^{st} \le \delta\} be the neighborhood of subtask ss. The average ranking consistency τδ\tau_\delta is: τδ=1T∑s=1T1∣Δs∣∑t∈Δsτst\tau_\delta = \frac{1}{T} \sum_{s=1}^T \frac{1}{|\Delta_s|} \sum_{t \in \Delta_s} \tau^{st}

    Distance Threshold δ\delta 1 1/2 1/4 1/6 1/8
    Ranking Consistency τδ\tau_\delta 0.29 0.31 0.32 0.41 0.60

    As the task distance threshold δ\delta decreases, performance ranking consistency τδ\tau_\delta increases monotonically from 0.29 to 0.60, demonstrating that LVLMs exhibit significantly more consistent relative capabilities on structurally closer tasks.

  4. Knowl 4 — Task Clustering for In-Domain and Out-of-Domain Task Identification

    empirical result

    Applying hierarchical clustering with 12 clusters on the 162162-subtask task map enables systematic discovery of In-Domain (ID) and Out-of-Domain (OoD) capabilities across 32 LVLMs. A cluster is designated ID if it exhibits high mean accuracy and high rank correlation τ\tau with overall performance; it is designated OoD if it exhibits low accuracy and low correlation τ\tau.

    Cluster ID 1 2 3 4 5 6 7 8 9 10 11 12
    Number of Tasks 11 53 16 16 9 8 7 16 4 9 10 3
    Kendall's tau τ\tau 0.54 0.73 0.57 0.48 -0.05 0.62 0.63 0.34 0.12 0.57 0.38 0.59
    Accuracy (%) 40.4 64.7 61.9 39.9 55.9 30.0 33.1 40.2 31.4 61.2 33.2 50.7

    Primary Discoveries

    • In-Domain Clusters (Clusters 2, 3, 10): LVLMs perform robustly on general visual recognition (Cluster 2: 64.7% acc, τ=0.73\tau=0.73), specialized recognition such as medical modality and facial emotion (Cluster 3: 61.9% acc, τ=0.57\tau=0.57), and visual description/captioning (Cluster 10: 61.2% acc, τ=0.57\tau=0.57).
    • Out-of-Domain Clusters (Clusters 8, 9, 11): Current models fail at fine-grained spatial cognition, object detection, and tracking (Cluster 8: 40.2% acc, τ=0.34\tau=0.34), mobile GUI navigation (Cluster 9: 31.4% acc, τ=0.12\tau=0.12), and structural parsing tasks such as table structure recognition and code extraction (Cluster 11: 33.2% acc, τ=0.38\tau=0.38).
  5. Knowl 5 — Comparative Performance of Leading LVLMs on MMT-Bench

    empirical result

    Across 32 evaluated LVLMs on MMT-Bench, state-of-the-art models show severe performance limitations on multitask vision-language understanding.

    Key Evaluation Findings

    • Proprietary Top Model: GPT-4o achieves the highest overall accuracy across all 162 subtasks at 65.5%65.5\%. When excluding Visual Recognition (VR) tasks—where it scores 88.0%88.0\%—GPT-4o's overall accuracy falls to 59.5%59.5\%.
    • Leading Open-Source Model: InternVL-Chat-v1.2-34B achieves 63.4%63.4\% overall accuracy (58.2%58.2\% excluding VR), outperforming major proprietary systems including GPT-4V (61.1%61.1\% overall, 55.1%55.1\% non-VR) and GeminiProVision (61.6%61.6\% overall, 55.1%55.1\% non-VR).
    • Model Scaling and LLM Backbone: Increasing model parameters from 7B to 13B improves accuracy for LLaVA-v1.5 (from 49.5%49.5\% to 51.7%51.7\%) and LLaVA-v1.5-XTuner (from 50.2%50.2\% to 51.1%51.1\%). Upgrading the base LLM from InternLM to InternLM2 increases LLaVA-7B performance from 49.7%49.7\% to 50.8%50.8\%.
    • Unsupervised Pretraining vs. SFT: BLIP-2 (Flan-T5-XXL), which lacks visual instruction-following SFT, scores 54.8%54.8\% overall, outperforming multiple models fine-tuned on millions of instruction pairs (such as LLaVA-v1.5-13B at 51.7%51.7\% and mPLUG-Owl2 at 52.0%52.0\%).
  6. Knowl 6 — OpenCompass Multi-Choice Extraction Protocol

    algorithm

    To robustly parse multi-choice option predictions from arbitrary LVLM free-form text generations, MMT-Bench employs a cascading extraction pipeline based on OpenCompass.

    Input: Model response string R, set of valid option letters OptLetters = {A, B, C, ...}, option text mapping OptTexts = {A: text_A, B: text_B, ...}
    Output: Extracted option letter choice in OptLetters union {Z}
    // Step 1: Check for explicit option letter designation
    if R contains an unambiguous option letter indicator matched against OptLetters then
        return matched option letter
    end if
    // Step 2: Check for direct option content match
    for each letter L, text T in OptTexts do
        if T appears verbatim in R without conflicting options then
            return L
        end if
    end for
    // Step 3: Fallback extraction via LLM parser
    ExtractedLetter = QueryChatGPTForOptionExtraction(R, OptLetters, OptTexts)
    if ExtractedLetter is in OptLetters then
        return ExtractedLetter
    end if
    // Step 4: Refusal or complete extraction failure
    return 'Z'

    Reliability Statistics

    Step 1 succeeds on over 87%87\% of queries across evaluated models (e.g., LLaVA-1.5-7B reaches 99.994%99.994\%, BLIP-2 reaches 99.448%99.448\%, and InternVL-Chat-v1.2-34B reaches 99.211%99.211\%). Setting extraction failures and model refusals (e.g., GPT-4V at 10.136%10.136\%, Claude-3-Haiku at 4.023%4.023\%) to option letter 'Z' prevents random guessing artifacts from inflating benchmark scores.

  7. Knowl 7 — Empirical Verification of Visual Grounding in MMT-Bench

    empirical result

    To verify that MMT-Bench questions require genuine visual reasoning rather than exploiting linguistic artifacts or prior knowledge encoded in the language model backbone, LVLMs were evaluated in blind text-only conditions (input image replaced by a solid black image).

    Model Without Visual Input (%) With Visual Input (%) Delta (%)
    Random Guessing Baseline 28.5 - -
    Frequent Choice Baseline 31.7 - -
    ChatGPT-3.5 (Text-only) 33.2 - -
    LLaVA-1.5-7B 31.6 49.7 +18.1
    LLaVA-1.5-13B 33.3 51.7 +18.4
    QWen-VL-Chat 32.3 52.5 +20.2
    Claude-3-Haiku 33.1 52.2 +19.1

    Without visual inputs, all models perform near random chance (28.5%28.5\%) or frequency guessing (31.7%31.7\%), with accuracies between 31.6%31.6\% and 33.3%33.3\%. Providing visual inputs yields substantial accuracy deltas between +18.1%+18.1\% and +20.2%+20.2\%, confirming that benchmark questions are visually grounded.

  8. Knowl 8 — Error Taxonomy and Error Breakdown for Frontier LVLMs

    empirical result

    Manual expert error analysis on up to 5 incorrect predictions per subtask across 162 subtasks categorizes LVLM mistakes into six distinct failure modes: Perception Error, Reasoning Error, Lack of Knowledge, Lack of Capability, Refuse to Answer, and Fail to Follow Instruction.

    Error Distributions

    • GPT-4V: Perception Error (51.0%51.0\%), Lack of Capability (19.0%19.0\%), Reasoning Error (9.94%9.94\%), Lack of Knowledge (8.86%8.86\%), Refuse to Answer (6.11%6.11\%), Annotation Error (2.99%2.99\%), Fail to Follow Instruction (2.04%2.04\%).
    • GeminiProVision: Perception Error (76.9%76.9\%), Reasoning Error (10.4%10.4\%), Lack of Knowledge (6.99%6.99\%), Refuse to Answer (2.67%2.67\%), Lack of Capability (1.27%1.27\%), Fail to Follow Instruction (1.14%1.14\%), Annotation Error (0.635%0.635\%).
    • InternVL-Chat-v1.2: Perception Error (67.2%67.2\%), Reasoning Error (14.8%14.8\%), Lack of Knowledge (9.04%9.04\%), Fail to Follow Instruction (6.64%6.64\%), Lack of Capability (1.29%1.29\%), Annotation Error (1.11%1.11\%).

    Perception errors caused by visual encoder limitations represent the primary bottleneck (51.0%–76.9%51.0\%\text{--}76.9\%), followed by complex visual reasoning (9.94%–14.8%9.94\%\text{--}14.8\%).

  9. Knowl 9 — Visual Referring Prompting Underperforms Text Coordinate Formats

    empirical result

    Across 14 visual referring tasks (including human interaction understanding, keypoint detection, interactive segmentation, instance captioning, single object tracking, and referring detection), direct visual prompt editing (e.g., drawing bounding boxes or masks on the image) performs significantly worse than textual coordinate representations.

    Tested Prompt Formats

    1. Visual Prompting: Directly superimposing visual bounding boxes or masks on the input image.
    2. Pixel Coordinates: Supplying coordinates in [x,y,w,h][x, y, w, h] or [x1,y1,x2,y2][x_1, y_1, x_2, y_2] scaled to pixel dimensions [0,w][0, w] and [0,h][0, h].
    3. Normalized Coordinates: Supplying coordinates normalized within [0,1][0, 1].
    4. Combined Format: Both normalized coordinates and visual prompting.

    Across evaluated models, visual prompting consistently yields the lowest average accuracy compared to coordinate-based prompting. Normalized coordinates generally outperform pixel coordinates because instruction tuning pipelines predominantly train LVLMs with normalized bounding box templates.

  10. Knowl 10 — Multi-Image Prompting Dynamics in Multi-Frame Tasks

    empirical result

    Evaluating 28 multi-image subtasks (such as video captioning, face retrieval, and action quality assessment) reveals distinct behavioral differences between multi-image native and single-image LVLMs:

    • Native Multi-Image LVLMs: Models natively architected for multiple visual tokens (such as GeminiProVision, mPLUG-Owl2, and QWen-VL-Chat) achieve substantial gains when fed multiple distinct images rather than a single tiled composite image. For example, GeminiProVision accuracy on face retrieval (FR) jumps from 30.5%30.5\% under single-image composite prompting to 92.5%92.5\% under multi-image prompting.
    • Single-Image LVLMs with Feature Concatenation: Passing individual images sequentially through the vision encoder and concatenating visual token embeddings before the LLM backbone improves performance across most single-image models (including LLaVA-v1.5-7B, ShareGPT4V, and BLIP-2). BLIP-2 displays notable gains on General Action Recognition (GAR) and Video Captioning (VC) under concatenated multi-image prompting.

Coverage note — Omitted only the exhaustive per-subtask result tables (Tables A19 through A37) which expand the aggregated 32 meta-task benchmark results already synthesized into high-level empirical results, and individual qualitative visual examples from the appendix case studies.

References

  1. 1.Moeimouto face dataset. http://www.nurs.or.jp/~nagadomi/animeface-character-dataset/.
  2. 2.Wikiart. WikiArt.org, 2022.
  3. 3.Acharya, M., Kafle, K., and Kanan, C. Tallyqa: Answering complex counting questions. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pp. 8076–8084, 2019.
  4. 4.Achille, A., Lam, M., Tewari, R., Ravichandran, A., Maji, S., Fowlkes, C. C., Soatto, S., and Perona, P. Task2vec: Task embedding for meta-learning. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 6430–6439, 2019.
  5. 5.AI, ., :, Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., Yu, K., Liu, P., Liu, Q., Yue, S., Yang, S., Yang, S., Yu, T., Xie, W., Huang, W., Hu, X., Ren, X., Niu, X., Nie, P., Xu, Y., Liu, Y., Wang, Y., Cai, Y., Gu, Z., Liu, Z., and Dai, Z. Yi: Open foundation models by 01.ai, 2024.
  6. 6.Alam, F., Alam, T., Hasan, M. A., Hasnat, A., Imran, M., and Ofli, F. Medic: A multi-task learning dataset for disaster image classification. Neural Computing and Applications, 35:2609–2632, 2023.
  7. 7.ALESSIO, C. Animals-10. https://www.kaggle.com/datasets/alessiocorrado99/animals10?select=translate.py, 2020.
  8. 8.anchen li. Sketch2code. https://github.com/mzbac/sketch2code, 2018.
  9. 9.Andriluka, M., Pishchulin, L., Gehler, P., and Schiele, B.
  10. 10.Aneja, D., Colburn, A., Faigin, G., Shapiro, L., and Mones, B. Modeling stylized character expressions via deep learning. In Asian Conference on Computer Vision, pp. 136–153. Springer, 2016.
  11. 11.Anguita, D., Ghio, A., Oneto, L., Parra, X., Reyes-Ortiz, J. L., et al. A public domain dataset for human activity recognition using smartphones. In Esann, volume 3, pp. 3, 2013.
  12. 12.Anthropic. Claude, 2023. URL https://www.anthropic.com. Accessed: 2023-04-18.
  13. 13.Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pp. 2425–2433, 2015.
  14. 14.Artem. Image season recognition. https://www.kaggle.com/c/image-season-recognition/data, 2013.
  15. 15.Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv preprint arXiv:2308.12966, 2023.
  16. 16.BANERJEE, S. Animal image dataset (90 different animals). https://www.kaggle.com/datasets/iamsouravbanerjee/animal-image-dataset-90-different-animals, 2022.
  17. 17.Bell, S., Upchurch, P., Snavely, N., and Bala, K. Opensurfaces: A richly annotated catalog of surface appearance. ACM Transactions on graphics (TOG), 32(4):1–17, 2013.
  18. 18.Beltramelli, T. pix2code: Generating code from a graphical user interface screenshot. arXiv preprint arXiv:1705.07962, 2017.
  19. 19.Bergmann, P., Fauser, M., Sattlegger, D., and Steger, C. Mvtec ad–a comprehensive real-world dataset for unsupervised anomaly detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9592–9600, 2019.
  20. 20.Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3): 8, 2023.
  21. 21.BHATHENA, J. Weather image recognition. https://www.kaggle.com/datasets/jehanbhathena/weather-dataset, 2021.
  22. 22.Bitton-Guetta, N., Bitton, Y., Hessel, J., Schmidt, L., Elovici, Y., Stanovsky, G., and Schwartz, R. Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2616–2627, 2023.
  23. 23.Bossard, L., Guillaumin, M., and Van Gool, L. Food-101–mining discriminative components with random forests. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VI 13, pp. 446–461. Springer, 2014.
  24. 24.Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O. nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11621–11631, 2020.
  25. 25.Cai, M., Liu, H., Mustikovela, S. K., Meyer, G. P., Chai, Y., Park, D., and Lee, Y. J. Making large multimodal models understand arbitrary visual prompts. arXiv preprint arXiv:2312.00784, 2023.
  26. 26.Chao, Y.-W., Liu, Y., Liu, X., Zeng, H., and Deng, J. Learning to detect human-object interactions. In 2018 ieee winter conference on applications of computer vision (wacv), pp. 381–389. IEEE, 2018.
  27. 27.Chen, D. L. and Dolan, W. B. Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics (ACL-2011), Portland, OR, June 2011.
  28. 28.Chen, L., Li, J., Dong, X., Zhang, P., He, C., Wang, J., Zhao, F., and Lin, D. Sharegpt4v: Improving large multi-modal models with better captions. arXiv preprint arXiv:2311.12793, 2023a.
  29. 29.Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., Li, B., Luo, P., Lu, T., Qiao, Y., and Dai, J. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. arXiv preprint arXiv:2312.14238, 2023b.
  30. 30.Chen, Z., Wang, W., Tian, H., Ye, S., Gao, Z., Cui, E., Tong, W., Hu, K., Luo, J., Ma, Z., et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. arXiv preprint arXiv:2404.16821, 2024.
  31. 31.Cheng, M.-M., Mitra, N. J., Huang, X., Torr, P. H., and Hu, S.-M. Global contrast based salient region detection. IEEE transactions on pattern analysis and machine intelligence, 37(3):569–582, 2014a.
  32. 32.Cheng, Y., Fu, H., Wei, X., Xiao, J., and Cao, X. Depth enhanced saliency detection method. In Proceedings of international conference on internet multimedia computing and service, pp. 23–27, 2014b.
  33. 33.Chi, Z., Huang, H., Xu, H.-D., Yu, H., Yin, W., and Mao, X.-L. Complicated table structure recognition. arXiv preprint arXiv:1908.04729, 2019.
  34. 34.CHUKS, P. Fake/real logo detection dataset. https://www.kaggle.com/datasets/prosperchuks/fakereal-logo-detection-dataset, 2023.
  35. 35.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Narang, S., Mishra, G., Yu, A., Zhao, V., Huang, Y., Dai, A., Yu, H., Petrov, S., Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J. Scaling instruction-finetuned language models, 2022. URL https://arxiv.org/abs/2210.11416.
  36. 36.CIJOV, A. Images alike. https://www.kaggle.com/datasets/alincijov/images-alike, 2021.
  37. 37.Contributors, O. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass, 2023a.
  38. 38.Contributors, T.-M. Transcore-m. https://github.com/PCIResearch/TransCore-M, 2023b.
  39. 39.Contributors, X. Xtuner: A toolkit for efficiently fine-tuning llm. https://github.com/InternLM/xtuner, 2023c.
  40. 40.Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S. Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.
  41. 41.Danci, M.-D. Architecturalstyle recognition. https://github.com/dumitrux/architectural-style-recognition, 2019.
  42. 42.Davison, A. K., Lansley, C., Costen, N., Tan, K., and Yap, M. H. Samm: A spontaneous micro-facial movement dataset. IEEE transactions on affective computing, 9(1): 116–129, 2016.
  43. 43.Ding, H., Liu, C., He, S., Jiang, X., and Loy, C. C. Mevis: A large-scale benchmark for video segmentation with motion expressions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2694–2703, 2023.
  44. 44.Ding, J., Xue, N., Xia, G.-S., Bai, X., Yang, W., Yang, M., Belongie, S., Luo, J., Datcu, M., Pelillo, M., and Zhang, L. Object detection in aerial images: A large-scale benchmark and challenges. IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 1–1, 2021a. doi: 10.1109/TPAMI.2021.3117983.
  45. 45.Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., et al. Cogview: Mastering text-to-image generation via transformers. Advances in Neural Information Processing Systems, 34: 19822–19835, 2021b.
  46. 46.Doersch, C., Gupta, A., Markeeva, L., Recasens, A., Smaira, L., Aytar, Y., Carreira, J., Zisserman, A., and Yang, Y. Tap-vid: A benchmark for tracking any point in a video. Advances in Neural Information Processing Systems, 35: 13610–13626, 2022.
  47. 47.Dong, X., Zhang, P., Zang, Y., Cao, Y., Wang, B., Ouyang, L., Wei, X., Zhang, S., Duan, H., Cao, M., Zhang, W., Li, Y., Yan, H., Gao, Y., Zhang, X., Li, W., Li, J., Chen, K., He, C., Zhang, X., Qiao, Y., Lin, D., and Wang, J. Internlm-xcomposer2: Mastering free-form text-image composition and comprehension in vision-language large model. arXiv preprint arXiv:2401.16420, 2024.
  48. 48.Dr. Hamid Rahim Sheikh, Dr. Alan C. Bovik, D. L. C. D. Z. W. Live image quality assessment database. https://live.ece.utexas.edu/research/quality/subjective.htm, 2006.
  49. 49.El Korchi, A. and Ghanou, Y. 2d geometric shapes dataset–for machine learning and pattern recognition. Data in Brief, 32:106090, 2020.
  50. 50.Ertler, C., Mislej, J., Ollmann, T., Porzi, L., Neuhold, G., and Kuang, Y. The mapillary traffic sign dataset for detection and classification on a global scale. In European Conference on Computer Vision, pp. 68–84. Springer, 2020.
  51. 51.Everingham, M., Eslami, S. M. A., Van Gool, L., Williams, C. K. I., Winn, J., and Zisserman, A. The pascal visual object classes challenge: A retrospective. International Journal of Computer Vision, 111(1):98–136, January 2015.
  52. 52.Fan, D.-P., Ji, G.-P., Sun, G., Cheng, M.-M., Shen, J., and Shao, L. Camouflaged object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 2777–2787, 2020.
  53. 53.Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., et al. Mme: A comprehensive evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2306.13394, 2023.
  54. 54.GAJARE, N. Rock images. https://www.kaggle.com/datasets/neelgajare/rocks-dataset, 2022.
  55. 55.Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., Li, H., and Qiao, Y. Llama-adapter v2: Parameter-efficient visual instruction model. arXiv preprint arXiv:2304.15010, 2023.
  56. 56.Gbeminiyi, A. Multi-class weather dataset for image classification. Mendeley Data, 6:15–23, 2018.
  57. 57.Geiger, A., Lenz, P., and Urtasun, R. Are we ready for autonomous driving? the kitti vision benchmark suite. In 2012 IEEE conference on computer vision and pattern recognition, pp. 3354–3361. IEEE, 2012.
  58. 58.GERRY. 30 musical instruments -image classification. https://www.kaggle.com/datasets/gpiosenka/musical-instruments-image-classification/data, 2021.
  59. 59.GERRY. 100 sports image classification. https://www.kaggle.com/datasets/gpiosenka/sports-classification, 2023.
  60. 60.GHIMIRE, B. Landscape color and grayscale images. https://www.kaggle.com/datasets/theblackmamba31/landscape-image-colorization/code, 2021.
  61. 61.Hanhe Lin, F. W. S. Helmet dataset. https://osf.io/4pwj8/, 2020.
  62. 62.Haq, N. U., Fraz, M. M., Hashmi, T., and Shahzad, M. Orientation aware weapons detection in visual data: a benchmark dataset. Computing, 104(12):2581–2604, 2022.
  63. 63.HARI, P. Movie posters. https://www.kaggle.com/datasets/phiitm/movie-posters, 2020.
  64. 64.Hsieh, M.-R., Lin, Y.-L., and Hsu, W. H. Drone-based object counting by spatially regularized regional proposal network. In Proceedings of the IEEE international conference on computer vision, pp. 4145–4153, 2017.
  65. 65.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  66. 66.Hu, Y., Li, T., Lu, Q., Shao, W., He, J., Qiao, Y., and Luo, P. Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm. arXiv preprint arXiv:2402.09181, 2024.
  67. 67.Huang, G. B., Ramesh, M., Berg, T., and Learned-Miller, E. Labeled faces in the wild: A database for studying face recognition in unconstrained environments. Technical Report 07-49, University of Massachusetts, Amherst, October 2007.
  68. 68.Huang, M.-L., Liao, Y.-C., Shiau, K.-L., and Tseng, Y.-L. Traditional chinese god image dataset: A glimpse of chinese culture. Data in Brief, 46:108861, 2023.
  69. 69.Huang, T.-H., Ferraro, F., Mostafazadeh, N., Misra, I., Agrawal, A., Devlin, J., Girshick, R., He, X., Kohli, P., Batra, D., et al. Visual storytelling. In Proceedings of the 2016 conference of the North American chapter of the association for computational linguistics: Human language technologies, pp. 1233–1239, 2016.
  70. 70.Huang, Z., Chen, K., He, J., Bai, X., Karatzas, D., Lu, S., and Jawahar, C. Icdar2019 competition on scanned receipt ocr and information extraction. In 2019 International Conference on Document Analysis and Recognition (ICDAR), pp. 1516–1520. IEEE, 2019.
  71. 71.Hudson, D. A. and Manning, C. D. Gqa: A new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 6700–6709, 2019.
  72. 72.Huttunen, H. Tau vehicle type recognition competition. https://www.kaggle.com/c/vehicle, 2019.
  73. 73.Hwang, E. and Shwartz, V. Memecap: A dataset for captioning and interpreting memes. arXiv preprint arXiv:2305.13703, 2023.
  74. 74.ICARO. Best artworks of all time. https://www.kaggle.com/datasets/ikarus777/best-artworks-of-all-time, 2019.
  75. 75.Idrees, H., Zamir, A. R., Jiang, Y.-G., Gorban, A., Laptev, I., Sukthankar, R., and Shah, M. The thumos challenge on action recognition for videos –in the wild–. Computer Vision and Image Understanding, 155:1–23, 2017.
  76. 76.Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A. Editing models with task arithmetic. the 11th International Conference on Learning Representation (ICLR 2023), 2023.
  77. 77.INNAT. Van gogh paintings. https://www.kaggle.com/datasets/ipythonx/van-gogh-paintings, 2022.
  78. 78.Jain, V. and Learned-Miller, E. Fddb: A benchmark for face detection in unconstrained settings. Technical report, UMass Amherst technical report, 2010.
  79. 79.JENSEN, M. B. Lisa traffic light dataset. https://www.kaggle.com/datasets/mbornoe/lisa-traffic-light-dataset, 2018.
  80. 80.Jhamtani, H. and Berg-Kirkpatrick, T. Learning to describe differences between pairs of similar images. arXiv preprint arXiv:1808.10584, 2018.
  81. 81.JOHANN. Vehicle type recognition. https://www.kaggle.com/code/theoneandonlyp/vehicle-type-recognition, 2023.
  82. 82.Ju, R., Liu, Y., Ren, T., Ge, L., and Wu, G. Depth-aware salient object detection using anisotropic center-surround difference. Signal Processing: Image Communication, 38:115–126, 2015.
  83. 83.Karatzas, D., Shafait, F., Uchida, S., Iwamura, M., i Bigorda, L. G., Mestre, S. R., Mas, J., Mota, D. F., Almazan, J. A., and De Las Heras, L. P. Icdar 2013 robust reading competition. In 2013 12th international conference on document analysis and recognition, pp. 1484–1493. IEEE, 2013.
  84. 84.Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017.
  85. 85.Kiela, D., Firooz, H., Mohan, A., Goswami, V., Singh, A., Ringshia, P., and Testuggine, D. The hateful memes challenge: Detecting hate speech in multimodal memes. Advances in neural information processing systems, 33: 2611–2624, 2020.
  86. 86.Kondo, Y., Ukita, N., Yamaguchi, T., Hou, H.-Y., Shen, M.-Y., Hsu, C.-C., Huang, E.-M., Huang, Y.-C., Xia, Y.-C., Wang, C.-Y., Lee, C.-Y., Huo, D., Kastner, M. A., Liu, T., Kawanishi, Y., Hirayama, T., Komamizu, T., Ide, I., Shinya, Y., Liu, X., Liang, G., and Yasui, S. MVA2023 Small Object Detection Challenge for Spotting Birds: Dataset, Methods, and Results. In 2023 18th International Conference on Machine Vision and Applications (MVA), 2023. https://www.mva-org.jp/mva2023/challenge.
  87. 87.Kong, Y., Jia, Y., and Fu, Y. Learning human interaction by interactive phrases. In Computer Vision–ECCV 2012: 12th European Conference on Computer Vision, Florence, Italy, October 7-13, 2012, Proceedings, Part I 12, pp. 300–313. Springer, 2012.
  88. 88.Kosti, R., Alvarez, J. M., Recasens, A., and Lapedriza, A. Context based emotion recognition using emotic dataset. IEEE transactions on pattern analysis and machine intelligence, 42(11):2755–2766, 2019.
  89. 89.Krause, J., Johnson, J., Krishna, R., and Fei-Fei, L. A hierarchical approach for generating descriptive image paragraphs. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 317–325, 2017.
  90. 90.Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123:32–73, 2017a.
  91. 91.Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123:32–73, 2017b.
  92. 92.KUMAR, V. Face mask detection. https://www.kaggle.com/datasets/vijaykumar1799/face-mask-detection, 2021.
  93. 93.KUMARUJJAWAL1234. Religious symbols-image classification. https://www.kaggle.com/datasets/kumarujjawal123456/famous-religious-symbols, 2023.
  94. 94.Kylberg, G. The kylberg texture dataset v. 1.0. External report (Blue series) 35, Centre for Image Analysis, Swedish University of Agricultural Sciences and Uppsala University, Uppsala, Sweden, September 2011. URL http://www.cb.uu.se/~gustaf/texture/.
  95. 95.LABS, D. Electronics object image dataset — computer parts. https://www.kaggle.com/datasets/dataclusterlabs/electronics-mouse-keyboard-image-dataset, 2023a.
  96. 96.LABS, D. Transparent object images — indoor object dataset. https://www.kaggle.com/datasets/dataclusterlabs/transparent-object-detection, 2023b.
  97. 97.Lai, X., Tian, Z., Chen, Y., Li, Y., Yuan, Y., Liu, S., and Jia, J. Lisa: Reasoning segmentation via large language model. arXiv preprint arXiv:2308.00692, 2023.
  98. 98.Langley, P. Crafting papers on machine learning. In Langley, P. (ed.), Proceedings of the 17th International Conference on Machine Learning (ICML 2000), pp. 1207–1216, Stanford, CA, 2000. Morgan Kaufmann.
  99. 99.LARXEL. Helmet detection. https://www.kaggle.com/datasets/andrewmvd/helmet-detection, 2020a.
  100. 100.LARXEL. Road sign detection. https://www.kaggle.com/datasets/andrewmvd/road-sign-detection, 2020b.
  101. 101.Latif, E., Mai, G., Nyaaba, M., Wu, X., Liu, N., Lu, G., Li, S., Liu, T., and Zhai, X. Agi: Artificial general intelligence for education, 2023.
  102. 102.Lazebnik, S., Schmid, C., and Ponce, J. A sparse texture representation using local affine regions. IEEE transactions on pattern analysis and machine intelligence, 27(8): 1265–1278, 2005.
  103. 103.Le, Y. and Yang, X. Tiny imagenet visual recognition challenge. CS 231N, 7(7):3, 2015.
  104. 104.Lee, J., Kim, S., Kim, S., Park, J., and Sohn, K. Context-aware emotion recognition networks. In Proceedings of the IEEE/CVF international conference on computer vision, 2019.
  105. 105.Li, B., Wang, R., Wang, G., Ge, Y., Ge, Y., and Shan, Y. Seed-bench: Benchmarking multimodal llms with generative comprehension. arXiv preprint arXiv:2307.16125, 2023a.
  106. 106.Li, D., Rodriguez, C., Yu, X., and Li, H. Word-level deep sign language recognition from video: A new large-scale dataset and methods comparison. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp. 1459–1469, 2020a.
  107. 107.Li, J., Zhang, J., and Tao, D. Deep automatic natural image matting. arXiv preprint arXiv:2107.07235, 2021.
  108. 108.Li, J., Zhang, J., Maybank, S. J., and Tao, D. Bridging composite and real: towards end-to-end deep image matting. International Journal of Computer Vision, 130(2): 246–266, 2022.
  109. 109.Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023b.
  110. 110.Li, K., Wang, Y., He, Y., Li, Y., Wang, Y., Liu, Y., Wang, Z., Xu, J., Chen, G., Luo, P., et al. Mvbench: A comprehensive multi-modal video understanding benchmark. arXiv preprint arXiv:2311.17005, 2023c.
  111. 111.Li, X., Wei, T., Chen, Y. P., Tai, Y.-W., and Tang, C.-K. Fss-1000: A 1000-class dataset for few-shot segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 2869–2878, 2020b.
  112. 112.Li, Y., Yang, X., Sun, P., Qi, H., and Lyu, S. Celeb-df (v2): a new dataset for deepfake forensics [j]. arXiv preprint arXiv, 2019.
  113. 113.Li, Z., Yang, B., Liu, Q., Ma, Z., Zhang, S., Yang, J., Sun, Y., Liu, Y., and Bai, X. Monkey: Image resolution and text label are important things for large multi-modal models. arXiv preprint arXiv:2311.06607, 2023d.
  114. 114.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pp. 740–755. Springer, 2014a.
  115. 115.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pp. 740–755. Springer, 2014b.
  116. 116.Liu, B., Fu, J., Kato, M. P., and Yoshikawa, M. Beyond narrative description: Generating poetry from images by multi-adversarial training. In Proceedings of the 26th ACM international conference on Multimedia, pp. 783–791, 2018a.
  117. 117.Liu, F., Emerson, G. E. T., and Collier, N. Visual spatial reasoning. Transactions of the Association for Computational Linguistics, 2023a.
  118. 118.Liu, H., Li, C., Li, Y., and Lee, Y. J. Improved baselines with visual instruction tuning, 2023b.
  119. 119.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning, 2023c.
  120. 120.Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J. Llava-next: Improved reasoning, ocr, and world knowledge, 2024a.
  121. 121.Liu, S., Ying, K., Zhang, H., Yang, Y., Lin, Y., Zhang, T., Li, C., Qiao, Y., Luo, P., Shao, W., et al. Convbench: A multi-turn conversation evaluation benchmark with hierarchical capability for large vision-language models. arXiv preprint arXiv:2403.20194, 2024b.
  122. 122.Liu, W., W. Luo, D. L., and Gao, S. Future frame prediction for anomaly detection – a new baseline. In 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018b.
  123. 123.Liu, X., Liu, W., Mei, T., and Ma, H. A deep learning-based approach to progressive vehicle re-identification for urban surveillance. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14, pp. 869–884. Springer, 2016a.
  124. 124.Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., et al. Mmbench: Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281, 2023d.
  125. 125.Liu, Y., Zhu, M., Li, H., Chen, H., Wang, X., and Shen, C. Matcher: Segment anything with one shot using all-purpose feature matching. arXiv preprint arXiv:2305.13310, 2023e.
  126. 126.Liu, Z., Luo, P., Wang, X., and Tang, X. Deep learning face attributes in the wild. In Proceedings of International Conference on Computer Vision (ICCV), December 2015.
  127. 127.Liu, Z., Luo, P., Qiu, S., Wang, X., and Tang, X. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1096–1104, 2016b.
  128. 128.Lu, C., Shi, J., and Jia, J. Abnormal event detection at 150 fps in matlab. In Proceedings of the IEEE international conference on computer vision, pp. 2720–2727, 2013.
  129. 129.Lu, H., Liu, W., Zhang, B., Wang, B., Dong, K., Liu, B., Sun, J., Ren, T., Li, Z., Yang, H., Sun, Y., Deng, C., Xu, H., Xie, Z., and Ruan, C. Deepseek-vl: Towards real-world vision-language understanding, 2024.
  130. 130.Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255, 2023.
  131. 131.Lv, Y., Zhang, J., Dai, Y., Li, A., Liu, B., Barnes, N., and Fan, D.-P. Simultaneously localize, segment and rank the camouflaged objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11591–11601, 2021.
  132. 132.MA, A. T. Anime characters personality and facial images. https://www.kaggle.com/datasets/tianyimasf/anime-characters, 2023.
  133. 133.Ma, W. and Liang, S. Polar: Posture-level action recognition dataset. In 2019 6th International Conference on Systems and Informatics (ICSAI), pp. 427–433. IEEE, 2019.
  134. 134.Machajdik, J. and Hanbury, A. Affective image classification using features inspired by psychology and art theory. In Proceedings of the 18th ACM international conference on Multimedia, pp. 83–92, 2010.
  135. 135.Mallikarjuna, P., Targhi, A. T., Fritz, M., Hayman, E., Caputo, B., and Eklundh, J.-O. The kth-tips2 database. Computational Vision and Active Perception Laboratory, Stockholm, Sweden, 11:12, 2006.
  136. 136.Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, pp. 3195–3204, 2019.
  137. 137.Marti, U.-V. and Bunke, H. The iam-database: an english sentence database for offline handwriting recognition. International Journal on Document Analysis and Recognition, 5:39–46, 2002.
  138. 138.Martin, D., Fowlkes, C., Tal, D., and Malik, J. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics. In Proc. 8th Int’l Conf. Computer Vision, volume 2, pp. 416–423, July 2001.
  139. 139.Masry, A., Long, D. X., Tan, J. Q., Joty, S., and Hoque, E. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. arXiv preprint arXiv:2203.10244, 2022.
  140. 140.Mathew, M., Bagal, V., Tito, R., Karatzas, D., Valveny, E., and Jawahar, C. Infographicvqa. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 1697–1706, 2022.
  141. 141.MENON, S. S. Animals 151. https://www.kaggle.com/datasets/sharansmenon/animals141, 2022.
  142. 142.Mishra, A., Alahari, K., and Jawahar, C. V. Scene text recognition using higher order language priors. In BMVC, 2012.
  143. 143.MOHAMED, M. Garbage classification (12 classes). https://www.kaggle.com/datasets/mostafaabla/garbage-classification, 2021.
  144. 144.Mohamed, Y., Khan, F. F., Haydarov, K., and Elhoseiny, M. It is okay to not be okay: Overcoming emotional bias in affective image captioning by contrastive data collection. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), volume abs/2204.07660, 2022.
  145. 145.Morris, M. R., Sohl-dickstein, J., Fiedel, N., Warkentin, T., Dafoe, A., Faust, A., Farabet, C., and Legg, S. Levels of agi: Operationalizing progress on the path to agi. arXiv preprint arXiv:2311.02462, 2023.
  146. 146.Mouchere, H., Viard-Gaudin, C., Zanibbi, R., and Garain, U. Icfhr 2014 competition on recognition of on-line handwritten mathematical expressions (crohme 2014). In 2014 14th International Conference on Frontiers in Handwriting Recognition, pp. 791–796. IEEE, 2014.
  147. 147.mrayinteractive. The quick, draw! dataset. https://github.com/googlecreativelab/quickdraw-dataset, 2014.
  148. 148.Ng, X. L., Ong, K. E., Zheng, Q., Ni, Y., Yeo, S. Y., and Liu, J. Animal kingdom: A large and diverse dataset for animal behavior understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 19023–19034, June 2022.
  149. 149.Nilsback, M.-E. and Zisserman, A. Automated flower classification over a large number of classes. In 2008 Sixth Indian conference on computer vision, graphics & image processing, pp. 722–729. IEEE, 2008.
  150. 150.NOGRA, J. A. Face mask usage. https://www.kaggle.com/datasets/jamesnogra/face-mask-usage, 2022.
  151. 151.Obeid, J. and Hoque, E. Chart-to-text: Generating natural language descriptions for charts by adapting the transformer model. arXiv preprint arXiv:2010.09142, 2020.
  152. 152.OLAFENWA, M. Idenprof. https://github.com/OlafenwaMoses/IdenProf, 2018.
  153. 153.OpenAI. Gpt-4o. https://openai.com/index/hello-gpt-4o/, 2024.
  154. 154.Parmar, P. and Morris, B. Action quality assessment across multiple actions. In 2019 IEEE winter conference on applications of computer vision (WACV), pp. 1468–1476. IEEE, 2019.
  155. 155.Parmar, P. and Tran Morris, B. Learning to score olympic events. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pp. 20–28, 2017.
  156. 156.Peng, X., Bai, Q., Xia, X., Huang, Z., Saenko, K., and Wang, B. Moment matching for multi-source domain adaptation. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 1406–1415, 2019.
  157. 157.Philbin, J., Chum, O., Isard, M., Sivic, J., and Zisserman, A. Object retrieval with large vocabularies and fast spatial matching. In 2007 IEEE conference on computer vision and pattern recognition, pp. 1–8. IEEE, 2007.
  158. 158.Pont-Tuset, J., Perazzi, F., Caelles, S., Arbelaez, P., Sorkine-Hornung, A., and Van Gool, L. The 2017 davis challenge on video object segmentation. arXiv:1704.00675, 2017.
  159. 159.Qi, J., Gao, Y., Hu, Y., Wang, X., Liu, X., Bai, X., Belongie, S., Yuille, A., Torr, P. H., and Bai, S. Occluded video instance segmentation: A benchmark. International Journal of Computer Vision, 130(8):2022–2039, 2022.
  160. 160.Quattoni, A. and Torralba, A. Recognizing indoor scenes. In 2009 IEEE conference on computer vision and pattern recognition, pp. 413–420. IEEE, 2009.
  161. 161.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision, 2021.
  162. 162.Rahman, R., Hasan, R., Farhad, A. A., Laskar, M. T. R., Ashmafee, M. H., and Kamal, A. R. M. Chartsumm: A comprehensive benchmark for automatic chart summarization of long and short summaries. arXiv preprint arXiv:2304.13620, 2023.
  163. 163.Ranjan, V., Sharma, U., Nguyen, T., and Hoai, M. Learning to count everything. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3394–3403, 2021.
  164. 164.Rawles, C., Li, A., Rodriguez, D., Riva, O., and Lillicrap, T. Android in the wild: A large-scale dataset for android device control. arXiv preprint arXiv:2307.10088, 2023.
  165. 165.RBDash-Team. Rbdash. https://github.com/RBDash-Team/RBDash, 2023.
  166. 166.ROMAN, K. Facial emotion recognition dataset. https://www.kaggle.com/datasets/tapakah68/facial-emotion-recognition, 2023.
  167. 167.Rosenfeld, A., Solbach, M. D., and Tsotsos, J. K. Totally looks like-how humans compare, compared to machines. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pp. 1961–1964, 2018.
  168. 168.Rossler, A., Cozzolino, D., Verdoliva, L., Riess, C., Thies, J., and Nießner, M. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 1–11, 2019.
  169. 169.SAHA, B. cricket-football-baseball classification. https://www.kaggle.com/datasets/imbikramsaha/cricket-football-baseball, 2022.
  170. 170.Salman Ibne Eunus, JU Komol, R. A. N. S. H. A. M. A. Rock classification dataset. https://www.kaggle.com/datasets/salmaneunus/rock-classification, 2021.
  171. 171.SANYAL, S. Weapon detection dataset. https://www.kaggle.com/datasets/snehilsanyal/weapon-detection-test, 2023.
  172. 172.Sasaki, R., Fujinami, M., and Nakai, H. Comprehensive image dataset for enhancing object detection in chemical experiments. Data in Brief, 52:110054, 2024.
  173. 173.SEKAR, S. Waste classification data. https://www.kaggle.com/datasets/techsash/waste-classification-data, 2019.
  174. 174.Sermanet, P., Xu, K., and Levine, S. Unsupervised perceptual rewards for imitation learning. arXiv preprint arXiv:1612.06699, 2016.
  175. 175.SETH, K. Fruits and vegetables image recognition dataset. https://www.kaggle.com/datasets/kritikseth/fruit-and-vegetable-image-recognition, 2022.
  176. 176.Shao, W., Hu, Y., Gao, P., Lei, M., Zhang, K., Meng, F., Xu, P., Huang, S., Li, H., Qiao, Y., et al. Tiny lvlm-ehub: Early multimodal experiments with bard. arXiv preprint arXiv:2308.03729, 2023.
  177. 177.SHETTY, S. Image colorization. https://www.kaggle.com/datasets/shravankumar9892/image-colorization, 2018.
  178. 178.Silberman, N., Hoiem, D., Kohli, P., and Fergus, R. Indoor segmentation and support inference from rgbd images. In Computer Vision–ECCV 2012: 12th European Conference on Computer Vision, Florence, Italy, October 7-13, 2012, Proceedings, Part V 12, pp. 746–760. Springer, 2012.
  179. 179.Singh, A. asingh33/CNNGestureRecognizer: CNNGestureRecognizer, November 2017. URL https://doi.org/10.5281/zenodo.1064825.
  180. 180.Singh, S. S. Teaching machines to code: neural markup generation with visual attention. arXiv preprint arXiv:1802.05415, 2018.
  181. 181.Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al. Large language models encode clinical knowledge. Nature, 620(7972):172–180, 2023.
  182. 182.Smith, B., Yin, Q., Feiner, S., and Nayar, S. Gaze Locking: Passive Eye Contact Detection for Human?Object Interaction. In ACM Symposium on User Interface Software and Technology (UIST), pp. 271–280, Oct 2013.
  183. 183.Srivastava, N., Mansimov, E., and Salakhudinov, R. Unsupervised learning of video representations using lstms. In International conference on machine learning, pp. 843–852. PMLR, 2015.
  184. 184.stepanje. Metal parts defect detection dataset. https://github.com/stepanje/MPDD, 2021.
  185. 185.Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V., Tsui, P., Guo, J., Zhou, Y., Chai, Y., Caine, B., et al. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 2446–2454, 2020.
  186. 186.Sun, Q., Fang, Y., Wu, L., Wang, X., and Cao, Y. Eva-clip: Improved training techniques for clip at scale. arXiv preprint arXiv:2303.15389, 2023.
  187. 187.sunshine joe. Flickrsportlogos-10. https://github.com/sunshine-joe/FlickrSportLogos-10, 2018.
  188. 188.SUPERPOTATO9. Ai recognition dataset. https://www.kaggle.com/datasets/superpotato9/dalle-recognition-dataset, 2024.
  189. 189.TAMRAKAR, A. E waste image dataset. https://www.kaggle.com/datasets/akshat103/e-waste-image-dataset, 2023.
  190. 190.Team, G. Gemini: A family of highly capable multimodal models, 2023a.
  191. 191.Team, I. Internlm: A multilingual language model with progressively enhanced capabilities. https://github.com/InternLM/InternLM, 2023b.
  192. 192.Team, Q. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023c.
  193. 193.tensorflow. Flower photos. https://www.tensorflow.org/tutorials/load_data/images.
  194. 194.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. Llama: Open and efficient foundation language models, 2023a.
  195. 195.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open foundation and fine-tuned chat models, 2023b.
  196. 196.TYAGI, A. music instruments classification. https://www.kaggle.com/datasets/aayushme/music-instruments-classification/data, 2020.
  197. 197.tz28. Chinese-number-gestures-recognition. https://github.com/tz28/Chinese-number-gestures-recognition, 2018.
  198. 198.Uy, M. A., Pham, Q.-H., Hua, B.-S., Nguyen, D. T., and Yeung, S.-K. Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data. In International Conference on Computer Vision (ICCV), 2019.
  199. 199.VERMA, A. Disaster images dataset. https://www.kaggle.com/datasets/varpit94/disaster-images-dataset, 2021.
  200. 200.Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S. The caltech-ucsd birds-200-2011 dataset. 2011.
  201. 201.Wallace, B., Wu, Z., and Hariharan, B. Can we characterize tasks without labels or features? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1245–1254, 2021.
  202. 202.Wang, H., Ge, S., Lipton, Z., and Xing, E. P. Learning robust global representations by penalizing local predictive power. Advances in Neural Information Processing Systems, 32, 2019.
  203. 203.Wang, L., Lu, H., Wang, Y., Feng, M., Wang, D., Yin, B., and Ruan, X. Learning to detect salient objects with image-level supervision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 136–145, 2017.
  204. 204.Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., Ji, J., Yang, Z., Zhao, L., Song, X., Xu, J., Xu, B., Li, J., Dong, Y., Ding, M., and Tang, J. Cogvlm: Visual expert for pretrained language models. 2023.
  205. 205.Wang, Z., Yang, J., Jin, H., Shechtman, E., Agarwala, A., Brandt, J., and Huang, T. S. Deepfont: Identify your font from an image. In Proceedings of the 23rd ACM international conference on Multimedia, pp. 451–459, 2015.
  206. 206.Wu, B., Yu, S., Chen, Zhenfang, T. J. B., and Gan, C. Star: A benchmark for situated reasoning in real-world videos. In Thirty-fifth Conference on Neural Information Processing Systems (NeurIPS), 2021.
  207. 207.Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., and Xiao, J. 3d shapenets: A deep representation for volumetric shapes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1912–1920, 2015.
  208. 208.Xiao, H., Rasul, K., and Vollgraf, R. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017.
  209. 209.Xie, B., Zhang, S., Zhou, Z., Li, B., Zhang, Y., Hessel, J., Yang, J., and Liu, Z. Funqa: Towards surprising video comprehension. arXiv preprint arXiv:2306.14899, 2023.
  210. 210.Xie, E., Wang, W., Wang, W., Ding, M., Shen, C., and Luo, P. Segmenting transparent objects in the wild. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIII 16, pp. 696–711. Springer, 2020.
  211. 211.Xu, J., Mei, T., Yao, T., and Rui, Y. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5288–5296, 2016.
  212. 212.Xu, L., Jin, S., Zeng, W., Liu, W., Qian, C., Ouyang, W., Luo, P., and Wang, X. Pose for everything: Towards category-agnostic pose estimation. October 2022.
  213. 213.Xu, P., Shao, W., Zhang, K., Gao, P., Liu, S., Lei, M., Meng, F., Huang, S., Qiao, Y., and Luo, P. Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models. arXiv preprint arXiv:2306.09265, 2023.
  214. 214.Yan, W.-J., Wu, Q., Liu, Y.-J., Wang, S.-J., and Fu, X. Casme database: A dataset of spontaneous micro-expressions collected from neutralized faces. In 2013 10th IEEE international conference and workshops on automatic face and gesture recognition (FG), pp. 1–7. IEEE, 2013.
  215. 215.Yang, J., Zhang, H., Li, F., Zou, X., Li, C., and Gao, J. Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v. arXiv preprint arXiv:2310.11441, 2023a.
  216. 216.Yang, L., Fan, Y., and Xu, N. Video instance segmentation. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 5188–5197, 2019.
  217. 217.Yang, S., Luo, P., Loy, C. C., and Tang, X. Wider face: A face detection benchmark. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  218. 218.Yang, X., Yan, J., Liao, W., Yang, X., Tang, J., and He, T. Scrdet++: Detecting small, cluttered and rotated objects via instance-level feature denoising and rotation loss smoothing. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.
  219. 219.Yang, Z., Li, L., Lin, K., Wang, J., Lin, C.-C., Liu, Z., and Wang, L. The dawn of lmms: Preliminary explorations with gpt-4v (ision). arXiv preprint arXiv:2309.17421, 9 (1):1, 2023b.
  220. 220.Yang, Z., Li, L., Lin, K., Wang, J., Lin, C.-C., Liu, Z., and Wang, L. The dawn of lmms: Preliminary explorations with gpt-4v (ision). arXiv preprint arXiv:2309.17421, 9 (1):1, 2023c.
  221. 221.Yang, Z., Liu, J., Han, Y., Chen, X., Huang, Z., Fu, B., and Yu, G. Appagent: Multimodal agents as smartphone users. arXiv preprint arXiv:2312.13771, 2023d.
  222. 222.Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., et al. mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178, 2023a.
  223. 223.Ye, Q., Xu, H., Ye, J., Yan, M., Hu, A., Liu, H., Qian, Q., Zhang, J., Huang, F., and Zhou, J. mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration, 2023b.
  224. 224.Yin, Z., Wang, J., Cao, J., Shi, Z., Liu, D., Li, M., Sheng, L., Bai, L., Huang, X., Wang, Z., et al. Lamm: Language-assisted multi-modal instruction-tuning dataset, framework, and benchmark. arXiv preprint arXiv:2306.06687, 2023.
  225. 225.Young, P., Lai, A., Hodosh, M., and Hockenmaier, J. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2:67–78, 2014.
  226. 226.Yu, H., Xu, Y., Zhang, J., Zhao, W., Guan, Z., and Tao, D. Ap-10k: A benchmark for animal pose estimation in the wild. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021.
  227. 227.Yu, L., Poirson, P., Yang, S., Berg, A. C., and Berg, T. L. Modeling context in referring expressions. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14, pp. 69–85. Springer, 2016.
  228. 228.Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L. Mm-vet: Evaluating large multi-modal models for integrated capabilities. arXiv preprint arXiv:2308.02490, 2023.
  229. 229.Yu, X., Gong, Y., Jiang, N., Ye, Q., and Han, Z. Scale match for tiny person detection. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp. 1257–1265, 2020.
  230. 230.Yuan, Y., Liu, X., Dikubab, W., Liu, H., Ji, Z., Wu, Z., and Bai, X. Syntax-aware network for handwritten mathematical expression recognition. arXiv preprint arXiv:2203.01601, 2022.
  231. 231.Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., Wei, C., Yu, B., Yuan, R., Sun, R., Yin, M., Zheng, B., Yang, Z., Liu, Y., Huang, W., Sun, H., Su, Y., and Chen, W. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. arXiv preprint arXiv:2311.16502, 2023a.
  232. 232.Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., et al. Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. arXiv preprint arXiv:2311.16502, 2023b.
  233. 233.Zamir, A. R., Sax, A., Shen, W., Guibas, L. J., Malik, J., and Savarese, S. Taskonomy: Disentangling task transfer learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3712–3722, 2018.
  234. 234.Zhan, Y., Xiong, Z., and Yuan, Y. Rsvg: Exploring data and models for visual grounding on remote sensing data. IEEE Transactions on Geoscience and Remote Sensing, 61:1–13, 2023. doi: 10.1109/TGRS.2023.3250471.
  235. 235.Zhang, P., Dong, X., Wang, B., Cao, Y., Xu, C., Ouyang, L., Zhao, Z., Ding, S., Zhang, S., Duan, H., Zhang, W., Yan, H., Zhang, X., Li, W., Li, J., Chen, K., He, C., Zhang, X., Qiao, Y., Lin, D., and Wang, J. Internlm-xcomposer: A vision-language large model for advanced text-image comprehension and composition. arXiv preprint arXiv:2309.15112, 2023a.
  236. 236.Zhang, R., Han, J., Liu, C., Gao, P., Zhou, A., Hu, X., Yan, S., Lu, P., Li, H., and Qiao, Y. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199, 2023b.
  237. 237.Zhang, T., Zhang, X., Li, J., Xu, X., Wang, B., Zhan, X., Xu, Y., Ke, X., Zeng, T., Su, H., et al. Sar ship detection dataset (ssdd): Official release and comprehensive data analysis. Remote Sensing, 13(18):3690, 2021.
  238. 238.Zhang, W., Zhu, M., and Derpanis, K. G. From actemes to action: A strongly-supervised representation for detailed action understanding. In Proceedings of the IEEE international conference on computer vision, pp. 2248–2255, 2013.
  239. 239.Zhang, Y., Zhou, D., Chen, S., Gao, S., and Ma, Y. Single-image crowd counting via multi-column convolutional neural network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 589–597, 2016.
  240. 240.Zhang, Y., Pan, J., Zhou, Y., Pan, R., and Chai, J. Grounding visual illusions in language: Do vision-language models perceive illusions like humans? In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 5718–5728, 2023c.
  241. 241.Zhang, Z., Luo, P., Loy, C.-C., and Tang, X. Learning social relation traits from face images. In Proceedings of the IEEE international conference on computer vision, pp. 3631–3639, 2015.
  242. 242.Zhao, T., Zhang, T., Zhu, M., Shen, H., Lee, K., Lu, X., and Yin, J. Vl-checklist: Evaluating pre-trained vision-language models with objects, attributes and relations. arXiv preprint arXiv:2207.00221, 2022.
  243. 243.Zheng, L., Shen, L., Tian, L., Wang, S., Wang, J., and Tian, Q. Scalable person re-identification: A benchmark. In Proceedings of the IEEE international conference on computer vision, pp. 1116–1124, 2015.
  244. 244.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023.
  245. 245.Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., and Torralba, A. Places: A 10 million image database for scene recognition. IEEE transactions on pattern analysis and machine intelligence, 40(6):1452–1464, 2017.
  246. 246.Zhu, P., Wen, L., Du, D., Bian, X., Fan, H., Hu, Q., and Ling, H. Detection and tracking meet drones challenge. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7380–7399, 2021.
  247. 247.Zhutian Yang, Tomas Lozano-Perez, J. M. W. L. Kitchen worlds. https://github.com/Learning-and-Intelligent-Systems/kitchen-worlds, 2022.

Citation

MLA
Ying, K., et al. “MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI”. arXiv, 2024, http://arxiv.org/abs/2404.16006v1.
APA
Ying, K., Meng, F., Wang, J., Li, Z., Lin, H., Yang, Y., Zhang, H., Zhang, W., Lin, Y., Liu, S., Lei, J., Lu, Q., Chen, R., Xu, P., Zhang, R., Zhang, H., Gao, P., Wang, Y., Qiao, Y., … Shao, W. (2024). MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI. arXiv. http://arxiv.org/abs/2404.16006v1
Chicago
Ying, K., F. Meng, J. Wang, et al. 2024. “MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI”. arXiv. http://arxiv.org/abs/2404.16006v1.
Harvard
Ying, K. et al. (2024) “MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2404.16006v1.
Vancouver
1. Ying K, Meng F, Wang J, et al (2024) MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI. arXiv

BibTeX

@article{ying2024mmt,
  title = {MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI},
  author = {Ying, Kaining and Meng, Fanqing and Wang, Jin and Li, Zhiqian and Lin, Han and Yang, Yue and Zhang, Hao and Zhang, Wenbo and Lin, Yuqi and Liu, Shuo and Lei, Jiayi and Lu, Quanfeng and Chen, Runjian and Xu, Peng and Zhang, Renrui and Zhang, Haozhe and Gao, Peng and Wang, Yali and Qiao, Yu and Luo, Ping and Zhang, Kaipeng and Shao, Wenqi},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2404.16006v1},
  eprint = {2404.16006}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/