video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Guangzhi SunWenyi YuChangli TangXianzhao ChenTian TanWei LiLu LuZejun MaYuxuan WangChao Zhang

article2024ICML126 citations

Presents video-SALMONN, an end-to-end multimodal large language model that unifies speech, audio events, and visual perception using a multi-resolution causal Q-Former to enable fine-grained temporal understanding and audio-visual question answering.

Listen

Modern artificial intelligence has made substantial progress in interpreting text, still images, and environmental sounds. However, comprehending short video content remains a major bottleneck because existing audio-visual systems largely ignore spoken human language. Spoken communication conveys vital semantic context, speaker identity, and subtle emotional cues that traditional vision models or non-speech audio systems miss. Cascading multiple independent systems—such as separate speech transcribers and video analyzers—creates complex, fragile pipelines. As video consumption surges, there is an urgent practical need for a single, unified model capable of jointly understanding video frames, ambient sounds, music, and spoken language.

The article demonstrates and evaluates video-SALMONN, an end-to-end multimodal artificial intelligence system designed to process all primary elements of video within a single model. The primary objective is to prove that integrating fine-grained speech recognition directly into audio-visual language models significantly improves general video comprehension and cross-modal reasoning.

To achieve this, the authors designed a novel multi-resolution causal alignment module that bridges specialized audio and visual feature extractors with a large language model. This framework synchronizes video frames with audio inputs every half-second while operating across multiple time windows (spanning 0.5-second and 5.0-second scales) to capture both high-frequency speech nuances and broad video context. The model was trained efficiently on roughly one million open-source multimodal samples using parameter-efficient fine-tuning, alongside a diversity loss function to prevent information redundancy and an unpaired training strategy to prevent one modality from overpowering another. The authors evaluated the system across ten tasks within a new benchmark covering speech recognition, audio captioning, image understanding, and complex audio-visual question answering.

The findings establish that video-SALMONN sets a new state of the art for unified video and audio understanding. On video question answering focused on temporal causal reasoning, the model achieved nearly a 50% accuracy score, delivering an absolute gain of roughly 25% over a fine-tuned visual baseline. On audio-visual question-answering tasks containing human speech, the system achieved over 30% absolute accuracy improvements compared to leading audio-visual models that cannot process spoken words. It achieved a low 2.6% word error rate on clean speech transcription while demonstrating zero-shot emergent capabilities, such as accurately identifying which speaker in a video made a statement and determining why specific scenes are humorous or romantic by combining speech dialogue, background music, and visual actions.

These results show that unified speech-audio-visual models can eliminate the operational cost and latency of maintaining separate machine transcription, audio classification, and computer vision systems. Organizations deploying video analytics, automated content moderation, educational tooling, or interactive assistants can achieve significantly deeper semantic insight without building fragmented toolchains. The ablation studies confirm that multi-resolution processing is essential: high resolution is required for speech interpretation, while low resolution preserves the high-level semantic context required for overall video reasoning.

Decision-makers should consider piloting unified audio-visual models for complex video retrieval, accessibility tools, and automated multimedia analysis. Where high-resolution text or fine-grained visual details within still images are essential, practitioners should incorporate localized spatial scanning techniques, as the base model prioritizes temporal flow over ultra-dense spatial resolution. Future development should expand the training beyond short clips to handle long-form video archives and explore lightweight deployment configurations for real-time edge processing.

While the evaluation shows high confidence across diverse standard datasets, certain boundaries apply. The system relies on pre-trained foundation models, meaning it inherits their baseline demographic limitations and transcription biases. Furthermore, setting the diversity loss penalty too high during training can degrade speech recognition accuracy by introducing hallucinations. Overall, the evidence firmly supports video-SALMONN as a robust, highly capable foundation for integrated multimedia processing.

arXiv: 2406.15704

No sufficiently relevant recommendations were found.

Cover for video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. video-SALMONN
  • 3.1. Temporal Fine-grained Synchronisation
  • 3.2. MRC Q-Former
  • 3.2.1. CAUSAL STRUCTURE
  • 3.3. System Training
  • 4. Experimental Setup
  • 4.1. Speech-Audio-Visual Evaluation Benchmark
  • 4.2. Model Configurations
  • 4.3. Training Data and Specifications
  • 5. Results and Discussions
  • 5.1. Main Results
  • 5.2. Ablation Studies
  • 5.3. Analysis on Multi-resolution
  • 5.4. Analysis of the Diversity Loss
  • 5.5. Emergent Speech-Audio-Visual Co-reasoning
  • 6. Conclusions
  • Impact Statement
  • References
  • A. Training Set and Benchmark Details
  • B. Examples of the AVQA Dataset
  • C. Evaluation Details
  • D. GPT Scoring Prompt Design
  • E. Visualisation of Diversity Loss Effect
  • F. Additional Results on Lip Reading
  • G. Additional Results on MUSIC-AVQA
  • H. Comparison between Vicuna and Llama-2 as Backbone LLMs
  • I. Spotlight for Static Image

Knowls

  1. Knowl 1 — Multi-resolution causal Q-Former aligns video information at several time scales

    model/method

    video-SALMONN uses a multi-resolution causal Q-Former (MRC Q-Former) to convert synchronized audio-visual feature sequences into a fixed number of language-model input vectors while retaining information at different temporal scales. At resolution rr, it divides an input sequence of TT video-frame steps into windows of k(r)k^{(r)} steps. A set of N(r)N^{(r)} learned query vectors is applied to each window, and the Q-Former returns N(r)N^{(r)} joint audio-visual output vectors per window. If W(r)W^{(r)} is the number of windows, then

    W(r)=⌈Tk(r)⌉,W(r)N(r)=C,W^{(r)}=\left\lceil\frac{T}{k^{(r)}}\right\rceil,\qquad W^{(r)}N^{(r)}=C,

    where CC is the fixed output length supplied to the language model. Thus, fine-grained windows use fewer query vectors per window, while coarser windows use more; each resolution still produces CC vectors. The learned query vectors are distinct across resolutions, but the remaining MRC Q-Former parameters are shared. Outputs from the resolutions are projected and combined before being passed to the language model. The windowing permits variable-length video input without applying one Q-Former to the entire sequence, and the multiple scales are intended to preserve both fine temporal details and higher-level video content.

  2. Knowl 2 — End-to-end audio-visual input pipeline uses frame-level temporal synchronization

    model/method

    video-SALMONN combines features from visual, speech, and non-speech audio encoders before language generation. Its visual encoder is the InstructBLIP image encoder; for video, it encodes frames sampled at 2 Hz and concatenates their features over time. Whisper encodes speech and BEATs encodes non-speech audio from the same audio stream. When both audio and video are supplied, their feature sequences are aligned at video-frame steps (one step every 0.5 seconds), with zero padding used to make the sequences the same length; the aligned speech, audio-event, and visual features at each step are concatenated along the feature dimension. Missing modalities are represented by zero-padded sequences. An image without video is treated as one frame; when paired audio is supplied, the image is repeated to match the audio sequence length. The resulting representation is processed by the MRC Q-Former and then supplied, along with the text prompt, to a language-model backbone. The pretrained modality encoders are kept fixed during training.

  3. Knowl 3 — Causal self-attention carries information forward across video frames

    model/method

    The MRC Q-Former adds a causal self-attention module to the standard Q-Former structure. Its block-wise triangular causal mask prevents a frame's representation from attending to future frames, while allowing it to incorporate information from preceding frames. This is designed to represent temporal causal relations among independently encoded frames and to support questions about event order, including questions about what happens next. The paper motivates the structure as complementing positional embeddings, which alone may be insufficient for learning such relations.

  4. Knowl 4 — Diversity loss discourages redundant low-resolution query vectors

    equation

    video-SALMONN adds a query-diversity penalty at its low-resolution levels, where each window contains enough frames to support extracting different aspects of the input. For resolution rr, window ww, and output query vectors hw,i(r)h^{(r)}_{w,i}, the penalty sums pairwise cosine similarities within each window:

    Ldiverse=∑r=2R∑w=1W(r)∑i=1N(r)∑j=1j≠iN(r)sim⁡ ⁣(hw,i(r),hw,j(r)).\mathcal{L}_{\mathrm{diverse}}=\sum_{r=2}^{R}\sum_{w=1}^{W^{(r)}}\sum_{i=1}^{N^{(r)}}\sum_{\substack{j=1\\j\ne i}}^{N^{(r)}}\operatorname{sim}\!\left(h^{(r)}_{w,i},h^{(r)}_{w,j}\right).

    Here, RR is the number of resolution levels, W(r)W^{(r)} is the number of windows at resolution rr, N(r)N^{(r)} is the number of queries per window, and sim⁡\operatorname{sim} is cosine similarity between query vectors. The training objective is L=LCE+λLdiverse\mathcal{L}=\mathcal{L}_{\mathrm{CE}}+\lambda\mathcal{L}_{\mathrm{diverse}}, where LCE\mathcal{L}_{\mathrm{CE}} is the response-generation cross-entropy loss and λ\lambda sets the diversity penalty's weight. The penalty is intended to reduce repetitive query content; the paper also reports that an excessively large λ\lambda can confuse the language model and increase hallucination.

  5. Knowl 5 — Unpaired audio-visual mixed training reduces reliance on one modality

    model/method

    To counter modality dominance, video-SALMONN supplements the limited paired audio-visual training data with unpaired audio and visual data. For these mixed examples, the prompt combines the original audio-task and video-task instructions, requiring the model to process both inputs rather than solving the example from a single dominant modality. The authors report that this strategy improves balance between audio and visual features and is an important factor in the model's audio-visual understanding and co-reasoning. The paper does not specify a numerical proportion for the unpaired examples in the main description.

  6. Knowl 6 — Training configuration and supervision for video-SALMONN

    experimental setup

    The default system uses Vicuna-v1.5 13B as its language-model backbone (7B is also evaluated), Whisper large-v2 for speech, BEATs for non-speech audio, and InstructBLIP's ViT and Q-Former for visual encoding. The MRC Q-Former has two Transformer blocks with 768-dimensional hidden states. Its default resolutions are 0.5 seconds and 5 seconds, using 3 and 30 output queries per window, respectively; projected outputs have 5120 dimensions before entering the language model. LoRA with rank 32 adapts the language model's attention query, key, and value projections and feed-forward weights; these trainable LoRA parameters comprise 0.4% of the language-model parameters. Training updates the MRC Q-Former and LoRA parameters using multi-task instruction fine-tuning.

    Training supervision combines single-modality and audio-visual data. Examples include LibriSpeech for ASR, AudioCaps for audio captioning, LLAVA-150k and OCRVQA for visual question answering and text reading, TextCaps for image captioning, NExT-QA and VideoChat for video questions, and COCO images with spoken captions. Audio-visual supervision includes 600 hours of Ego4D video-caption data, 300 hours of How2 AVSR data, and AVSD dialogue data. The full training mixture contains about 1 million samples, fewer than 300,000 of which are video samples, and uses publicly available datasets.

  7. Knowl 7 — SAVE evaluates single-modality and joint speech-audio-visual tasks

    experimental setup

    The paper introduces the Speech-Audio-Visual Evaluation (SAVE) benchmark, comprising six single-modality tasks and four audio-visual tasks. The table gives the reported evaluation sets, sample counts, metrics, and whether the task is zero-shot; zero-shot indicates that the test instruction and audio-visual inputs were unseen during training. Ego4D-QA and Presentation-QA are the speech-focused audio-visual QA sets, with questions generated from video descriptions and ASR transcriptions. AVM tests whether spoken descriptions match images or whether audio matches video; AVSSD tests audio-visual sound-source understanding.

    Task Test set Samples Metric Zero-shot
    ASR LibriSpeech test-clean 2620 WER No
    AAC AudioCaps test 938 SPIDEr No
    IC Flickr30k test 1000 CIDEr Yes
    OCR TextVQA test 1000 Accuracy Yes
    VQA GQA test-dev balanced 1000 Accuracy Yes
    Video QA NExT-QA test 1000 Accuracy Yes
    AVSR How2 dev5 500 WER No
    AVQA Ego4D + Presentation-QA 2000 Accuracy Yes
    AVSSD VGGSS 850 Accuracy Yes
    AVM SpokenCOCO + VGGSS 1000 Accuracy Yes
  8. Knowl 8 — SAVE results show strong video QA and speech-dependent audio-visual performance

    empirical result

    On SAVE, video-SALMONN performs across the reported speech, audio, image, video, and audio-visual tasks. In the table, WER is lower-is-better; the other metrics are higher-is-better. The single-modality results show a 49.6% Video QA accuracy for the 13B model, compared with 24.7% for InstructBLIP fine-tuned on the same image and video training data. The audio-visual results show 49.8% on Ego4D-QA and 70.5% on Presentation-QA for video-SALMONN 13B, versus 18.2% and 21.3% for Video-LLaMA. Video-SALMONN 13B also records 79.7% AVM accuracy. These comparisons support the paper's emphasis on temporal video understanding and speech-inclusive audio-visual QA.

    System ASR WER AAC SPIDEr Video QA Acc. IC CIDEr OCR Acc. VQA Acc.
    InstructBLIP 13B – – 21.0% 84.5 36.5% 48.9%
    InstructBLIP 13B fine-tuned – – 24.7% 78.9 36.7% 45.6%
    Video-LLaMA 7B 100%+ 3.5 22.5% 22.0 16.4% 15.1%
    video-SALMONN 7B 4.1% 39.1 42.5% 78.1 34.6% 45.3%
    video-SALMONN 13B, visual-only – – 44.8% 74.0 34.2% 45.6%
    video-SALMONN 13B 2.6% 49.7 49.6% 89.6 37.8% 44.8%
    System AVSR WER Ego4D-QA Acc. Presentation-QA Acc. AVSSD Acc. AVM Acc.
    Whisper large-v2 8.3% – – – –
    InstructBLIP 13B – – – 1.1% –
    InstructBLIP 13B fine-tuned – – – 20.3% –
    Video-LLaMA 7B – 18.2% 21.3% 41.9% 52.3%
    video-SALMONN 13B, visual-only – 35.0% 46.5% 23.5% –
    video-SALMONN 7B 8.7% 36.2% 41.3% 50.5% 74.3%
    video-SALMONN 13B 7.7% 49.8% 70.5% 47.6% 79.7%
  9. Knowl 9 — Ablations support complementary temporal scales and both training strategies

    empirical result

    Ablations on SAVE show that the complete video-SALMONN configuration performs better overall than variants missing a resolution, the mixed training scheme, the diversity loss, or the MRC Q-Former. The 0.5-second-only model has better ASR and AVSR than the 5-second-only model, while the 5-second-only model has stronger Video QA; combining both resolutions gives the strongest balanced results. Removing mixed training reduces average AVQA accuracy from 60.2% to 54.0% and AVM from 79.7% to 75.3%. Removing diversity loss reduces AVQA to 53.5%. Removing the MRC Q-Former gives larger declines, including Video QA from 49.6% to 42.7% and AVQA from 60.2% to 45.3%. Further removing synchronization from that variant reduces AVSSD from 74.5% to 72.0% and AVSR from 8.5% to 8.9% WER.

    System ASR WER OCR Acc. Video QA Acc. AVSR WER AVQA Acc. AVM Acc.
    video-SALMONN 2.6% 37.8% 49.6% 7.7% 60.2% 79.7%
    Without 5s resolution 2.5% 35.4% 47.2% 7.7% 57.2% 77.5%
    Without 0.5s resolution 2.9% 37.1% 49.9% 8.3% 58.9% 80.6%
    Without mixed training scheme 2.6% 34.0% 46.9% 8.3% 54.0% 75.3%
    Without diversity loss 2.5% 36.8% 49.3% 7.7% 53.5% 78.6%
    Without MRC Q-Former 3.3% 34.6% 42.7% 8.5% 45.3% 74.5%
    Without MRC Q-Former, synchronization, and diversity 3.1% 34.7% 36.0% 8.9% 44.6% 72.0%

    The AVQA column is the mean over Ego4D-QA and Presentation-QA. In a separate resolution analysis, masking one resolution's output showed that the high-resolution stream primarily supports speech-content tasks while the low-resolution stream supports higher-level tasks such as Video QA. The authors also report that smaller windows improve speech recognition but can hurt video QA because each window has fewer output tokens available to represent its visual content.

  10. Knowl 10 — Qualitative examples demonstrate cross-modal reasoning beyond transcription

    empirical result

    The paper's qualitative examples show video-SALMONN using speech, other audio, and visual context jointly rather than treating speech only as text to transcribe. Examples include answering an image question spoken in audio; explaining an audio-image mismatch; composing a story that combines an audio event with a video even when the audio and video are unpaired; and explaining why a movie scene is romantic using its visuals, dialogue, and background music. Other examples show speech supplying facts needed to identify an object in a video, and visual context helping attribute an utterance to a particular character. These are demonstrations from selected cases, not a separate quantitative measurement of generalization.

  11. Knowl 11 — Image spotlight improves OCR but can slightly impair video-task performance

    limitation

    The paper identifies limited spatial resolution as a weakness for static-image tasks such as OCR. Its image-spotlight extension splits an image into sequential sub-images and sends their encodings through the MRC Q-Former as a scan from the top-left to the bottom-right. On SAVE, the spotlight version improves OCR accuracy from 37.8% to 56.1%, while Video QA changes from 49.6% to 49.1% and image captioning CIDEr from 89.6 to 87.3. The authors suggest that this small video-task degradation may arise because the image-scan input differs from the model's usual treatment of video frames and can confuse it. The results indicate a trade-off: spotlighting helps detailed image reading but does not improve every task.

Coverage note — Detailed dataset prompts, individual case-study transcripts, and the supplementary Llama-2 comparison are omitted because they do not add a comparably load-bearing method or result beyond the benchmark, main evaluations, and qualitative capability evidence included here.

References

  1. 1.Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., et al. Flamingo: A visual language model for few-shot learning. In Proc. NeurIPS, 2022.
  2. 2.Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., et al. PaLM 2 technical report. arXiv:2305.10403, 2023.
  3. 3.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., et al. Language models are few-shot learners. In Proc. NeurIPS, 2020.
  4. 4.Chen, F., Han, M., Zhao, H., Zhang, Q., Shi, J., Xu, S., and Xu, B. X-LLM: Bootstrapping advanced large language models by treating multi-modalities as foreign languages. arXiv:2305.04160, 2023a.
  5. 5.Chen, G., Zheng, Y.-D., Wang, J., Xu, J., Huang, Y., Pan, J., Wang, Y., Wang, Y., Qiao, Y., Lu, T., and Wang, L. VideoLLM: Modeling video sequence with large language models. arXiv:2305.13292, 2023b.
  6. 6.Chen, H., Xie, W., Vedaldi, A., and Zisserman, A. VGGSound: A large-scale audio-visual dataset. In Proc. ICASSP, 2020.
  7. 7.Chen, S., Li, H., Wang, Q., Zhao, Z., Sun, M., Zhu, X., and Liu, J. Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset. In Proc. NeurIPS, 2023c.
  8. 8.Chen, S., Wu, Y., Wang, C., Liu, S., Tompkins, D., Chen, Z., Che, W., Yu, X., and Wei, F. BEATs: Audio pre-training with acoustic tokenizers. In Proc. ICML, 2023d.
  9. 9.Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P. Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality, March 2023. URL https://lmsys.org/blog/2023-03-30-vicuna/.
  10. 10.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., et al. Scaling instruction-finetuned language models. arXiv:2210.11416, 2022.
  11. 11.Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S. InstructBLIP: Towards general-purpose vision-language models with instruction tuning. arXiv:2305.06500, 2023.
  12. 12.Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J. GLM: General language model pretraining with autoregressive blank infilling. In Proc. ACL, 2022.
  13. 13.Fiscus, J. G., Ajot, J., Michel, M., and Garofolo, J. S. The rich transcription 2006 spring meeting recognition evaluation. In Machine Learning for Multimodal Interaction: Third International Workshop, MLMI 2006, Bethesda, MD, USA, May 1-4, 2006, Revised Selected Papers 3, pp. 309–322. Springer, 2006a.
  14. 14.Fiscus, J. G., Radde, N., Garofolo, J. S., Le, A., Ajot, J., and Laprun, C. The rich transcription 2005 spring meeting recognition evaluation. In Machine Learning for Multimodal Interaction: Second International Workshop, MLMI 2005, Edinburgh, UK, July 11-13, 2005, Revised Selected Papers 2, pp. 369–389. Springer, 2006b.
  15. 15.Fiscus, J. G., Ajot, J., and Garofolo, J. S. The rich transcription 2007 meeting recognition evaluation. In CLEaR, 2007. URL https://api.semanticscholar.org/CorpusID:15113788.
  16. 16.Garofolo, J. S., Fiscus, J. G., and Laprun, C. D. The rich transcription 2004 spring meeting recognition evaluation. US Department of Commerce, National Institute of Standards and Technology, 2004.
  17. 17.Girdhar, R., El-Nouby, A., Liu, Z., Singh, M., Alwala, K. V., Joulin, A., and Misra, I. ImageBind: One embedding space to bind them all. arXiv:2305.05665, 2023.
  18. 18.Gong, Y., Luo, H., Liu, A. H., Karlinsky, L., and Glass, J. Listen, think, and understand. arXiv:2305.10790, 2023.
  19. 19.Grauman, K., Westbury, A., Byrne, E., et al. Ego4D: Around the world in 3,000 hours of egocentric video. In Proc. CVPR, 2022.
  20. 20.Hsu, W.-N., Harwath, D., Song, C., and Glass, J. Text-free image-to-speech synthesis using learned segmental units. In Proc. NeurIPS Workshop on Self-Supervised Learning for Speech and Audio Processing, 2020.
  21. 21.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. LoRA: Low-rank adaptation of large language models. In Proc. ICLR, 2022.
  22. 22.Hudson, D. A. and Manning, C. D. GQA: A new dataset for real-world visual reasoning and compositional question answering. In Proc. CVPR, 2019.
  23. 23.Kim, C. D., Kim, B., Lee, H., and Kim, G. AudioCaps: Generating captions for audios in the wild. In Proc. NAACL-HLT, 2019.
  24. 24.Li, J., Li, D., Savarese, S., and Hoi, S. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proc. ICML, 2023a.
  25. 25.Li, K., He, Y., Wang, Y., Li, Y., Wang, W., Luo, P., Wang, Y., Wang, L., and Qiao, Y. VideoChat: Chat-centric video understanding. arXiv:2305.06355, 2023b.
  26. 26.Lin, T.-Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Ramanan, D., Zitnick, C. L., and Dollar, P. Microsoft COCO: Common objects in context. In Proc. ECCV, 2014.
  27. 27.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. arXiv:2304.08485, 2023.
  28. 28.Liu, S., Zhu, Z., Ye, N., Guadarrama, S., and Murphy, K. Improved image captioning via policy gradient optimization of spider. In Proc. ICCV, Venice, 2017.
  29. 29.Luo, R., Zhao, Z., Yang, M., Dong, J., Qiu, M., Lu, P., Wang, T., and Wei, Z. Valley: Video assistant with large language model enhanced ability. arXiv: 2306.07207, 2023.
  30. 30.Lyu, C., Wu, M., Wang, L., Huang, X., Liu, B., Du, Z., Shi, S., and Tu, Z. Macaw-LLM: Multi-modal language modeling with image, audio, video, and text integration. arXiv:2306.09093, 2023.
  31. 31.Maaz, M., Rasheed, H., Khan, S., and Khan, F. S. VideoChatGPT: Towards detailed video understanding via large vision and language models. arXiv:2306.05424, 2023.
  32. 32.Mishra, A., Shekhar, S., Singh, A. K., and Chakraborty, A. OCR-VQA: Visual question answering by reading text in images. In Proc. ICDAR, 2019.
  33. 33.OpenAI. GPT-4 technical report. arXiv:2303.08774, 2023.
  34. 34.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., et al. Training language models to follow instructions with human feedback. In Proc. NeurIPS, 2022.
  35. 35.Panayotov, V., Chen, G., Povey, D., and Khudanpur, S. Librispeech: An ASR corpus based on public domain audio books. In Proc. ICASSP, 2015.
  36. 36.Peng, B., Li, C., He, P., Galley, M., and Gao, J. Instruction tuning with GPT-4. arXiv:2304.03277, 2023.
  37. 37.Piergiovanni, A., Noble, I., Kim, D., Ryoo, M. S., Gomes, V., and Angelova, A. Mirasol3b: A multimodal autoregressive model for time-aligned and contextual modalities. arXiv preprint arXiv:2311.05698, 2023.
  38. 38.Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I. Robust speech recognition via large-scale weak supervision. Proc. ICML, 2023.
  39. 39.Rubenstein, P. K., Asawaroengchai, C., Nguyen, D. D., et al. AudioPaLM: A large language model that can speak and listen. arXiv:2306.12925, 2023.
  40. 40.Sanabria, R., Caglayan, O., Palaskar, S., Elliott, D., Barrault, L., Specia, L., and Metze, F. How2: A large-scale dataset for multimodal language understanding. In Proc. ViGIL, 2018.
  41. 41.Shu, F., Zhang, L., Jiang, H., and Xie, C. Audio-visual llm for video understanding. arXiv: 2312.06720, 2023a.
  42. 42.Shu, F., Zhang, L., Jiang, H., and Xie, C. Audio-visual llm for video understanding. arXiv preprint arXiv:2312.06720, 2023b.
  43. 43.Sidorov, O., Hu, R., Rohrbach, M., and Singh, A. Textcaps: a dataset for image captioningwith reading comprehension. In Proc. European Conference on Computer Vision, 2020.
  44. 44.Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M. Towards vqa models that can read. In Proc. CVPR, 2019.
  45. 45.Su, Y., Lan, T., Li, H., Xu, J., Wang, Y., and Cai, D. PandaGPT: One model to instruction-follow them all. arXiv:2305.16355, 2023.
  46. 46.Sun, G., Yu, W., Tang, C., Chen, X., Tan, T., Li, W., Lu, L., Ma, Z., and Zhang, C. Fine-grained audio-visual joint representations for multimodal large language models. arXiv:2310.05863, 2023.
  47. 47.Tang, C., Yu, W., Sun, G., Chen, X., Tan, T., Li, W., Lu, L., Ma, Z., and Zhang, C. SALMONN: Towards generic hearing abilities for large language models. arXiv:2310.13289, 2023.
  48. 48.Tang, C., Yu, W., Sun, G., Chen, X., Tan, T., Li, W., Lu, L., Ma, Z., and Zhang, C. Extending large language models for speech and audio captioning. In To appear in Proc. ICASSP, 2024.
  49. 49.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. LLaMA: Open and efficient foundation language models. arXiv:2302.13971, 2023.
  50. 50.Xiao, J., Shang, X., Yao, A., and Chua, T.-S. NExT-QA: Next phase of question-answering to explaining temporal actions. In Proc. CVPR, 2021.
  51. 51.Young, P., Lai, A., Hodosh, M., and Hockenmaier, J. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2:67–78, 2014.
  52. 52.Yu, W., Tang, C., Sun, G., Chen, X., Tan, T., Li, W., Lu, L., Ma, Z., and Zhang, C. Connecting speech encoder and large language model for ASR. To appear in Proc. ICASSP, 2024.
  53. 53.Zeng, Y., Zhang, H., Zheng, J., Xia, J., Wei, G., Wei, Y., Zhang, Y., and Kong, T. What matters in training a gpt4-style language model with multimodal inputs? arXiv:2307.02469, 2023.
  54. 54.Zhang, D., Li, S., Zhang, X., Zhan, J., Wang, P., Zhou, Y., and Qiu, X. SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities. arXiv:2305.11000, 2023a.
  55. 55.Zhang, H., Li, X., and Bing, L. Video-LLaMA: An instruction-tuned audio-visual language model for video understanding. arXiv:2306.02858, 2023b.
  56. 56.Zhao, Y., Misra, I., Krahenbühl, P., and Girdhar, R. Learning video representations from large language models. In Proc. CVPR, 2022.
  57. 57.Zhao, Y., Lin, Z., Zhou, D., Huang, Z., Feng, J., and Kang, B. BuboGPT: Enabling visual grounding in multi-modal LLMs. arXiv:2307.08581, 2023.

Citation

MLA
Sun, G., et al. “video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models”. arXiv, 2024, http://arxiv.org/abs/2406.15704v1.
APA
Sun, G., Yu, W., Tang, C., Chen, X., Tan, T., Li, W., Lu, L., Ma, Z., Wang, Y., & Zhang, C. (2024). video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models. arXiv. http://arxiv.org/abs/2406.15704v1
Chicago
Sun, G., W. Yu, C. Tang, et al. 2024. “video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models”. arXiv. http://arxiv.org/abs/2406.15704v1.
Harvard
Sun, G. et al. (2024) “video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2406.15704v1.
Vancouver
1. Sun G, Yu W, Tang C, Chen X, Tan T, Li W, Lu L, Ma Z, Wang Y, Zhang C (2024) video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models. arXiv

BibTeX

@article{sun2024video,
  title = {video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models},
  author = {Sun, Guangzhi and Yu, Wenyi and Tang, Changli and Chen, Xianzhao and Tan, Tian and Li, Wei and Lu, Lu and Ma, Zejun and Wang, Yuxuan and Zhang, Chao},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2406.15704v1},
  eprint = {2406.15704}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/