Cross-Modal Fine-Tuning: Align then Refine
Junhong ShenLiam LiLucio M. DeryCorey StatenMikhail KhodakGraham NeubigAmeet Talwalkar
Proposes ORCA, a cross-modal fine-tuning framework that matches target-data distributions to pretraining modalities before model adaptation, achieving state-of-the-art results across 12 distinct modalities without requiring domain-specific pretrained architectures.
Large pretrained machine learning models have achieved strong results in fields with massive datasets, such as computer vision and natural language processing. However, domains with limited data—including the physical sciences, genomics, and specialized tabular applications—often struggle to benefit because training high-performing models from scratch requires extensive data, substantial computing power, and deep domain expertise. Existing attempts to reuse pretrained models across different data types have largely relied on rigid, ad-hoc techniques that do not generalize well.
The article introduces and evaluates ORCA, a general framework designed to adapt standard pretrained language and vision transformers to diverse, out-of-modality target tasks. The framework addresses the cross-modal transfer problem through an "align-then-refine" workflow: it first standardizes data shapes, trains an initial embedding network to align the target data distribution with the model's original pretraining data, and finally fine-tunes all model parameters on the target task.
To evaluate this approach, the authors tested ORCA across three major benchmarks spanning more than 60 datasets and 12 distinct data types, including partial differential equations, genomics, electrocardiograms, and tabular records. The evaluation compared ORCA against hand-designed models, automated machine learning (AutoML) systems, specialized cross-modal baselines, and general-purpose architectures, primarily utilizing pretrained RoBERTa and Swin Transformer backbones.
The findings show that cross-modal adaptation using data alignment consistently outperforms existing alternatives. On the multi-domain NAS-Bench-360 benchmark, ORCA achieved the lowest prediction error on 7 out of 10 tasks and placed in the top three across all 10, outperforming hand-designed and AutoML models. Crucially, the initial data-alignment stage proved essential: directly fine-tuning pretrained models without alignment resulted in worse performance and higher variance. On specialized benchmarks, ORCA matched or exceeded expert domain architectures on scientific differential equations and established AutoML systems on tabular classification. Furthermore, the framework delivered substantial performance advantages in low-data settings—matching the performance of standard fine-tuning while requiring only one-third of the target data—while adding minimal computing overhead, as the alignment phase accounted for only about 11% of the total fine-tuning time.
These results indicate that foundational knowledge captured in vision and language models can be transferred across unrelated fields without costly architecture redesigns. This approach significantly reduces the time, technical complexity, and data-gathering expenses required to deploy advanced machine learning in specialized industries. In particular, it offers an effective path for deploying high-performing models in domains where data collection is expensive or impractical.
Organizations developing machine learning for specialized or data-scarce domains should consider adapting existing pretrained models via distribution-aligned fine-tuning rather than building models from scratch. When implementing this workflow, practitioners should perform full parameter fine-tuning rather than restricting updates to specific layers, as full fine-tuning yields noticeably superior accuracy with minimal additional computational time. Additionally, teams should select pretrained backbones whose structural representations best align with the target task based on distribution-distance metrics.
While the empirical evidence across the tested benchmarks is strong, some limitations remain. The evaluations focused on one-dimensional and two-dimensional data, leaving higher-dimensional problems and complex decision-making setups for future validation. Furthermore, selecting the optimal pretrained model currently requires empirical comparison. Practitioners should view this framework as a proven, highly competitive baseline for cross-modal tasks, while validating performance on their specific domain data.
- Paper: Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, Junnan Li et al. (2021). It introduces the core 'align before fuse' paradigm that directly inspires and provides the conceptual blueprint for the align-then-refine cross-modal adaptation framework.
- Paper: Perceiver: General Perception with Iterative Attention, Andrew Jaegle et al. (2021). It establishes foundational principles for building general-purpose perceptual architectures capable of consuming arbitrary, unaligned sensory modalities.
- Paper: Cross-Modal Discrete Representation Learning, Alexander H. Liu et al. (2022). It provides foundational methods for bridging disparate modalities through shared embedding alignment before downstream task transfer.
- Paper: Towards a Unified View of Parameter-Efficient Transfer Learning, Junxian He et al. (2022). It provides the unified conceptual formulation of parameter-efficient transfer learning mechanisms used when adapting pretrained backbones to downstream tasks.
- Paper: CLIP-Adapter: Better Vision-Language Models with Feature Adapters, Peng Gao et al. (2021). It details how lightweight embedding and feature adapters can efficiently bridge new input domains into a frozen, pretrained foundation model.
- Paper: ImageBind One Embedding Space to Bind Them All, Rohit Girdhar et al. (2023). It scales cross-modal alignment across six diverse sensory modalities into a single joint embedding space without requiring all-to-all paired training.
- Paper: MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task Learning, Yi Xin et al. (2024). It extends cross-modal alignment and parameter-efficient adaptation strategies to multi-task learning settings across downstream domains.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). It advances cross-modal transfer learning by demonstrating unified alignment across diverse visual input formats including single images, multi-images, and video streams.
- Paper: Tuning LayerNorm in Attention: Towards Efficient Multi-Modal LLM Finetuning, Bingchen Zhao et al. (2024). It investigates highly parameter-efficient adaptation mechanisms for transitioning text-pretrained foundation models to multimodal domains.
