Exploring Data-Free LoRA Transferability for Video Diffusion Models
Yuchen Wang1^{1}1
Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)1^{1}1
Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)1^{1}1
Wenliang Zhong1^{1}1
Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)1^{1}1
Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)1^{1}1
Lichen Bai1^{1}1
Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)1^{1}1
Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)1^{1}1
Zikai Zhou1^{1}1
Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)1^{1}1
Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)1^{1}1
Shitong Shao1^{1}1
Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)1^{1}1
Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)1^{1}1
Bojun Cheng2^{2}2
Thrust of Microelectronics, The Hong Kong University of Science and Technology (Guangzhou)2^{2}2
Thrust of Microelectronics, The Hong Kong University of Science and Technology (Guangzhou)2^{2}2
Shuo Chen3^{3}3
School of Intelligence Science and Technology, Nanjing University3^{3}3
School of Intelligence Science and Technology, Nanjing University3^{3}3
Shuo Yang4^{4}4
Department of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)4^{4}4
Department of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)4^{4}4
Zeke Xie1^{1}1
Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)1^{1}1
Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)1^{1}1
1^{1}1 Thrust of Data Science and Analytics, The Hong Kong University of Science and Technology (Guangzhou)
2^{2}2 Thrust of Microelectronics, The Hong Kong University of Science and Technology (Guangzhou)
3^{3}3 School of Intelligence Science and Technology, Nanjing University
4^{4}4 Department of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)
Correspondence to: Zeke Xie [email protected]
2^{2}2 Thrust of Microelectronics, The Hong Kong University of Science and Technology (Guangzhou)
3^{3}3 School of Intelligence Science and Technology, Nanjing University
4^{4}4 Department of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)
Correspondence to: Zeke Xie [email protected]
Abstract
Video diffusion models leveraging step distillation or causal distillation have achieved remarkable performance. However, adapting existing LoRAs to these variants remains a critical challenge due to weight space mismatches. We observe that direct application leads to style degradation and structural collapse, yet the underlying mechanisms remain poorly understood. To fill this gap, we delve into the weight space and identify that the incompatibility stems from spectral interference within shared functional clusters defined over singular subspaces. Specifically, our analysis reveals that while both paradigms respect spectral rigidity, they establish conflicting routing pathways that clash through constructive overload or destructive cancellation. To address this issue, we propose Cluster-Aware Spectral Arbitration (CASA), a data-free framework that dynamically arbitrates between safeguarding the target’s manifold and restoring LoRA alignment based on spectral density. Extensive experiments demonstrate that CASA effectively mitigates artifacts and revives LoRA functionality. Our code is available at https://github.com/Noahwangyuchen/CASA.
1. Introduction
Video diffusion models (VDMs; [1, 2, 3, 4]) have recently emerged as a powerful paradigm for high-fidelity video generation, enabling coherent synthesis across both spatial and temporal dimensions. To reduce the substantial computational cost of diffusion-based video generation, a series of efficient variants have been proposed [5], including step distillation [6, 7, 8] and various causal distillation strategies [9, 10, 11]. These approaches are typically implemented via full fine-tuning (FFT) of a pretrained base model, resulting in distilled VDMs that preserve generation quality while significantly accelerating inference. As a consequence, modern VDM ecosystems increasingly consist of families of closely related models that share a common generative foundation but differ in their weight space due to fine-tuning.
In parallel, Low-Rank Adaptation (LoRA) has become a standard tool for efficient and controllable adaptation of large generative models [12]. However, when LoRAs trained on a base video diffusion model are applied directly to its distilled variants, they are prone to degrading or generating various artifacts, as shown in Figure 1. Re-training LoRAs on each distilled model is computationally expensive and requires access to user data, rendering it impractical in many real-world scenarios.
Resolving this incompatibility requires a mechanistic understanding of how FFT and LoRA modify VDMs. However, the effects of FFT and LoRA in VDMs have not been systematically studied. In this work, we fill this gap by conducting the first weight-space analysis of fine-tuning in video diffusion models. Our analysis reveals a consistent spectral rigidity property across layers and further shows that adaptation primarily manifests as structured rotations of singular subspaces that exhibit strong cluster-level coherence.
Building on these insights, we further utilize a routing-based perspective for characterizing the effect of FFT and LoRA, showing that the failure of LoRA reuse on distilled video diffusion models stems from spectral interference between incompatible routing patterns. While both adaptations respect spectral rigidity, they establish conflicting functional routes in the shared weight space, leading to style degradation and structural artifacts. To resolve this conflict, we propose Cluster-Aware Spectral Arbitration (CASA), a data-free framework that formulates LoRA transfer as a cluster-level arbitration problem in the spectral domain. CASA is designed to dynamically balance the preservation of generative pathways of the target model with the restoration of LoRA functionality. As a result, CASA enables effective LoRA reuse in distilled VDMs without additional training or access to user data. The contributions of this work can be summarized as follows:
- To the best of our knowledge, we provide the first weight-space analysis of full fine-tuning and LoRA in VDMs, and introduce a routing-based perspective for characterizing their effects in singular subspaces.
- We identify the root cause of LoRA failure on distilled video diffusion models as spectral interference between incompatible routing patterns introduced by full fine-tuning and LoRA.
- We propose Cluster-Aware Spectral Arbitration (CASA), a data-free method that enables effective reuse of LoRAs on distilled models without additional training or user data.
2. Related Works
Video Diffusion Models and Distillation. Video diffusion models (VDMs; [1, 13, 14, 15]) have recently demonstrated strong performance in high-fidelity video generation. However, these models largely rely on bidirectional attention, incurring high computational cost and preventing streaming generation. To alleviate this, step distillation methods [7, 8, 16] aim to reduce the number of denoising steps by training a student model to approximate the multi-step behavior of a pretrained teacher. In parallel, causal distillation methods [9, 17, 18, 19] reformulate bidirectional video diffusion into causal autoregressive processes, enabling streaming generation with significantly reduced latency. Despite their algorithmic differences, these approaches are typically implemented via full fine-tuning of a pretrained base model [9, 6], resulting in distilled variants that preserve generation quality while altering the underlying weight space. In this work, we analyze how such full fine-tuning reshapes the weight space of video diffusion models and how it interacts with parameter-efficient adaptations such as LoRA.
Weight Space Analysis of Model Adaptation. Understanding how fine-tuning modifies pretrained models in the weight space has attracted increasing attention. Prior works [20, 21, 22, 23, 24, 25, 26] have analyzed the effect of full fine-tuning and LoRAs from a weight space perspective, often by leveraging singular value decomposition (SVD) to study spectral properties induced by adaptation. However, existing analyses are largely limited to LLMs and do not directly extend to VDMs with fundamentally different generation dynamics. To the best of our knowledge, a systematic weight-space analysis of full fine-tuning and LoRA in video diffusion models remains largely unexplored.
LoRA Transferability. Low-Rank Adaptation (LoRA) [12] and its variants [27, 28, 29] have emerged as an efficient and widely adopted parameter-efficient fine-tuning technique. Recent work has explored LoRA transfer across models from different perspectives. X-Adapter [30] learns mapping layers between source and target models to align feature spaces for PEFT modules. Trans-LoRA [31] relies on synthetic data generated by large language models to facilitate cross-model transfer. Other methods aim to improve intrinsic transferability by constraining the update space, such as LoRA-X [32], which restricts updates within selected singular directions. More closely related to our setting, ProLoRA [33] enables data-free LoRA transfer by projecting source adaptations into the target weight space. Despite these advances, existing methods are primarily developed for large language models or image generation models. In contrast, we study LoRA transfer in VDMs from a weight-space perspective and propose a data-free transfer method grounded in the specific spectral structure of VDMs.
3. Analysis
In this section, we analyze how FFT and LoRA modify VDMs in the weight space and how their interaction leads to LoRA incompatibility on distilled models. We first reveal a consistent spectral rigidity across layers, then show that adaptation mainly manifests as structured, cluster-coherent perturbations of singular subspaces. Finally, we characterize FFT and LoRA routing patterns at the cluster level and identify spectral interference as the key failure mechanism. Together, our analysis establishes a mechanistic foundation for understanding LoRA incompatibility on distilled VDMs. We conduct the analysis using Wan2.1-T2V-1.3B [1] as the source model, FastWan2.1-T2V-1.3B [6] as the distilled target model, with Jinx-v2 and Steamboat-Willie-1.3B as LoRAs.
3.1 Spectral Rigidity
We first examine the global spectral effects of FFT and LoRA on video diffusion models. For each weight matrix, we compare the singular value spectrum of the base model with that of its FFT- and LoRA-adapted counterparts. The left part of Figure 2 shows singular value curves of a random layer, where we observe that the spectral shape is largely preserved under both adaptation strategies. To quantify this observation at scale, we measure the relative spectral change across all layers of the model using the ℓ2\ell_2ℓ2 norm,
where S\mathbf{S}S and S′\mathbf{S}'S′ denote the singular values before and after fine-tuning, respectively. As presented in the right part of Figure 2, both adaptations exhibit extremely small relative spectral changes not exceeding 0.3%0.3\%0.3%, indicating that fine-tuning does not substantially redistribute spectral energy.
These results reveal a pronounced spectral rigidity property in video diffusion models: despite significant differences in training objectives and optimization procedures, both FFT and LoRA preserve the singular value spectrum to a high degree. This suggests that fine-tuning does not alter the overall capacity allocation or destabilize the underlying generative manifold, but instead induces more subtle structural modifications. At the same time, rigidity in the singular values implies that the primary effects of fine-tuning must arise from changes in the associated singular subspaces. In the following subsection, we therefore turn to analyze how FFT and LoRA perturb these subspaces and show that such perturbations exhibit strong structured patterns.
3.2 Structured Perturbations of Singular Subspaces
To characterize how fine-tuning alters the internal representation of video diffusion models, we analyze the changes induced by FFT and LoRA at the level of singular subspaces. We quantify subspace alignment using the cosine similarity matrix ∣U⊤U′∣\lvert \mathbf{U}^\top \mathbf{U}' \rvert∣U⊤U′∣, where U\mathbf{U}U and U′\mathbf{U}'U′ denote the left singular bases before and after fine-tuning, respectively. Unless otherwise stated, we focus on U\mathbf{U}U, while observing qualitatively identical behavior for V\mathbf{V}V.
Figure 3 visualizes this similarity matrix together with the corresponding singular value spectrum of a random layer. We examine three representative regions along the spectrum, corresponding to head, middle, and tail, to illustrate the structure of subspace perturbations. In the head of the spectrum, where singular values are dominant and well separated, the similarity matrix is sharply concentrated along the diagonal. Each singular direction with top singular values remains closely aligned with its original counterpart, indicating that fine-tuning preserves the leading subspace structure almost identically. This behavior contrasts with observations in large language models, where LoRA has been shown to introduce intruder dimensions that exhibit extremely low cosine similarity to any pre-trained singular direction [24], or to significantly increase the leading singular values [23].
Moving to the middle of the spectrum, the similarity matrix exhibits clear block-wise patterns. Instead of a strict one-to-one correspondence, groups of neighboring singular vectors display mutual similarity, forming coherent clusters. Notably, these clusters align with step-like plateaus in the singular value curve, where local spectral gaps separate adjacent groups. Within each plateau, singular directions are more interchangeable under fine-tuning, while mixing across plateaus remains limited. In contrast, the tail of the spectrum shows a different behavior. Here, the singular values decay smoothly with much smaller local gaps, and the similarity matrix transitions into a diffuse, banded pattern. Many singular directions in this region are near-degenerate and therefore admit broader mixing under fine-tuning, resulting in the observed spread of similarity.
Importantly, the resulting structure is highly consistent across distinct LoRAs and when applying LoRA to the distilled model, indicating that it reflects a stable organization of the underlying model rather than an artifact of a specific adaptation. Overall, these observations are consistent with classical perturbation theory, which relates subspace stability to local spectral separation [34]. Regions with larger spectral gaps exhibit more localized and structured mixing, while near-degenerate regions allow more distributed perturbations.
Together, these results indicate that fine-tuning induces structured perturbations across the singular subspaces, giving rise to stable cluster-level organization in the leading and intermediate spectrum. In the following, we build upon this cluster-level view to analyze how FFT and LoRA establish routing patterns, and how their interaction within these clusters leads to incompatible spectral interference.
3.3 Cluster-level Routing
To analyze how FFT and LoRA modify the functional interactions within video diffusion models, we study the changes induced in the weight space through a routing perspective. Given a weight update Δ\DeltaΔ applied to a linear projection with singular value decomposition W=USV⊤\mathbf{W} = \mathbf{U} \mathbf{S} \mathbf{V}^\topW=USV⊤, we define the corresponding routing matrix as C=U⊤ΔV\mathbf{C} = \mathbf{U}^\top\Delta\mathbf{V}C=U⊤ΔV. Each entry Cij\mathbf{C}_{ij}Cij quantifies how the update introduces the interaction between the jjj-th right singular direction and the iii-th left singular direction, providing a natural view of fine-tuning as cross-subspace routing, where columns of C\mathbf{C}C correspond to senders in the input singular basis, while rows correspond to receivers in the output basis. The magnitude of each entry reflects the strength of information flow induced between the corresponding subspaces.
While routing can be examined at the level of individual singular directions, a meaningful characterization of global routing behavior requires assessing its consistency within the singular clusters identified above. We first group singular directions into clusters based on their mixing before and after fine-tuning1. For each cluster, we measure pattern coherence as the cosine similarity between routing directions of singular vectors within the same cluster, and quantify intensity stability as the coefficient of variation of their routing energy, both when acting as senders and receivers. As shown in Figure 4, both FFT and LoRA exhibit strong cluster-level pattern coherence with highly aligned routing directions. At the same time, routing intensity remains stable within clusters, as reflected by low intra-cluster variability in both sending and receiving energy. These results indicate that both methods maintain strong intra-cluster consistency, confirming that singular clusters serve as stable functional units for routing analysis.
1.
We construct a graph over singular vectors before and after fine-tuning, where two directions are connected if their cosine similarity exceeds a threshold of , and obtain connected components as clusters, considering only the top- singular vectors that cumulatively capture of the spectral energy.
We next examine the routing patterns at the cluster level. For each cluster, we measure the root mean square (RMS) energy density as senders and as receivers, normalized by the number of connections within the cluster. Figure 5a shows a typical cluster-wise energy profile. We observe that under FFT, routing energy is highly concentrated in a small number of clusters, which predominantly correspond to the head of the spectrum, indicating their dominant role in shaping the effective generative subspace2. These clusters exhibit significantly higher RMS energy as both senders and receivers, while energy density gradually decays toward later clusters. In contrast, LoRA does not exhibit such behavior. Its routing energy is distributed more uniformly across clusters, with no small subset of clusters dominating either outgoing or incoming signal intensity. This pattern is further reflected in the block-wise energy maps presented in Figure 5b, where LoRA induces broadly distributed connections across cluster pairs, while FFT concentrates energy in a limited set of high-intensity blocks.
2.
We further explore the generative subspace from the routing perspective in Appendix B.2.
Together, these observations indicate that FFT establishes a strongly centralized routing structure anchored at a few dominant clusters, while LoRA injects functional modifications in a more diffuse and globally distributed manner.
3.4 Cluster-level Routing Interference
We now analyze how LoRA and FFT interact at the cluster-routing level. We first quantify their routing overlap at the cluster level. For each input–output cluster pair, we compute the product of their RMS routing energy, which measures the intensity of co-activation on the same routing pathway. This quantity reflects the potential for interaction, but does not by itself determine whether the interaction is constructive or destructive. We therefore complement this analysis by examining the directional alignment between LoRA and model drift on the same cluster pairs. Interference arises when strong routing overlap coincides with incompatible directional alignment, indicating conflicting updates on the same functional pathway.
Figure 6 visualizes these two quantities. We observe that strong interactions are highly localized, concentrating on a small subset of cluster pairs. Notably, these high-interaction regions predominantly involve head clusters, consistent with their elevated routing energy under FFT. The corresponding direction map exhibits substantial misalignment. Across interacting cluster pairs, LoRA and FFT can be either strongly aligned or strongly opposed, with no global tendency toward constructive or destructive alignment. Importantly, interference in head clusters is particularly consequential. Since these clusters dominate the effective generative subspace, conflicting routing signals injected by LoRA into pathways already strongly modulated by distillation are more likely to perturb the generation behavior.
Together, these observations suggest that LoRA incompatibility on distilled video diffusion models stems from cluster-level interference. When LoRA injects signals into cluster pathways that are already strongly modulated by model distillation, their interaction can become spectrally incompatible, leading to unstable or distorted generation behavior. This insight suggests that a uniform transfer strategy is doomed to fail; instead, an arbitration mechanism is needed to distinguish between these spectrally conflicting regions.
4. Methodology
Overview and Motivation. Given a source VDM equipped with a LoRA and its distilled variant obtained by FFT, our goal is to reuse the original LoRA on the distilled model without additional training or user data. Motivated by our analysis, we operate in the singular subspace view, where adaptation is expressed as structured routing, and incompatibility arises from conflicting routing in the same functional subspaces. We therefore propose Cluster-Aware Spectral Arbitration (CASA), a principled routing arbitration framework that decomposes LoRA transfer into two complementary objectives: (i) preserving LoRA-induced functional routing in non-critical spectral regions of the distilled model, and (ii) preventing over-activation along dominant generative pathways shaped by distillation.
4.1 Cluster-Aware Routing Representation
Let Ws\mathbf{W}_sWs denote the weight matrix of a layer in the source model, with singular value decomposition Ws=UsSsVs⊤\mathbf{W}_s = \mathbf{U}_s \mathbf{S}_s \mathbf{V}_s^{\top}Ws=UsSsVs⊤. Let Δlora=BA\Delta_{\mathrm{lora}} = \mathbf{B}\mathbf{A}Δlora=BA be the LoRA update trained on the source model, and let Δfft\Delta_{\mathrm{fft}}Δfft denote the weight drift induced by FFT for the corresponding layer.3
3.
In practice, can be obtained by subtracting the source weights from the distilled weights at the same layer.
We project both updates into the singular bases of the source layer and obtain their routing representations by
Here, each entry C(i,j)\mathbf{C}(i,j)C(i,j) quantifies a routing connection from the jjj-th right singular direction (sender) to the iii-th left singular direction (receiver).
Following our analysis, we further construct cluster-level routing representations. We first restrict clustering to the leading singular subspace by selecting the smallest kkk that captures 90%90\%90% of the spectral energy, i.e.,
where σi\sigma_iσi represents the iii-th singular value. This focuses CASA on the spectrally stable area where structured cluster organization is most pronounced. We then build singular clusters in the top-kkk subspace using a gap-aware perturbation graph. For indices i,j≤ki,j\le ki,j≤k, we define a predicted rotation strength
and connect iii and jjj if R(i,j)\mathbf{R}(i,j)R(i,j) exceeds a threshold τ\tauτ. Connected components of the resulting undirected graph define clusters {Gm}m=1M\{\mathcal{G}_m\}_{m=1}^{M}{Gm}m=1M. This construction can be interpreted as grouping singular directions whose relative coupling strength exceeds their local spectral separation, consistent with our analysis in Section 3.2.
4.2 Cluster-Aware Spectral Arbitration
CASA produces an updated routing matrix Ccasa\mathbf{C}_{\mathrm{casa}}Ccasa by combining Clora\mathbf{C}_{\mathrm{lora}}Clora and Cfft\mathbf{C}_{\mathrm{fft}}Cfft with selective arbitration between reviving LoRA function and protecting the generative subspace of the target model.
Identifying Spectrally Dominant Routing Constraints. We characterize spectrally dominant routing regions as implicit constraints imposed by the distilled model, quantified via routing energy density in the singular routing space. Specifically, for each cluster Gm\mathcal{G}_mGm, we compute its sending and receiving density as
Clusters whose sending or receiving density exceeds a quantile threshold qdomq_\mathrm{dom}qdom are regarded as dominant sending or receiving clusters, denoted as Gdomsend\mathcal{G}^{\mathrm{send}}_{\mathrm{dom}}Gdomsend and Gdomrecv\mathcal{G}^{\mathrm{recv}}_{\mathrm{dom}}Gdomrecv, respectively. This criterion captures routing pathways that are heavily utilized by the distilled model and therefore play a central role in its effective generative subspace4.
4.
Please refer to Appendix B.2 for more analysis.
Formally, we define an indicator function D(i,j)\mathcal{D}(i,j)D(i,j) to determine whether a routing entry (i,j)(i,j)(i,j) lies in a dominant routing region:
where I[⋅]\mathbb{I}[\cdot]I[⋅] is the indicator function. A routing entry is considered dominant if D(i,j)=1\mathcal{D}(i,j)=1D(i,j)=1.
Restoring LoRA in Non-Dominant Regions. For routing entries outside dominant regions, we directly restore the LoRA update by compensating for the FFT drift to avoid destructive cancellation. For entries satisfying D(i,j)=0\mathcal{D}(i,j)=0D(i,j)=0:
Under the assumption that these regions are weakly coupled to the distilled model’s generative manifold, this operation corresponds to the minimum-interference solution that exactly restores LoRA-induced routing while leaving dominant pathways unchanged.
Arbitration in Dominant Routing Regions. Routing entries that fall within dominant regions, i.e., those satisfying D(i,j)=1\mathcal{D}(i,j)=1D(i,j)=1, require more cautious treatment due to their influence on the generative subspace. As analyzed in Section 3.4, directly restoring LoRA updates in these regions may lead to over-activation, disrupting critical generation pathways. Specifically, over-activation can be viewed as a form of constructive interference in routing space, where LoRA and FFT induce aligned updates along the same functional subspace, resulting in disproportionate energy amplification beyond the operating regime of the distilled model.
To capture this behavior, we model the risk of over-activation as a factorized score consisting of a local interaction term and a cluster-level contextual alignment term. We first define the local interaction matrix E\mathbf{E}E for dominant routing regions as
which assigns positive interaction energy only to routing entries where LoRA and full fine-tuning induce aligned routing directions. Entries with opposite directions yield zero interaction energy and are therefore excluded from further consideration.
To incorporate the cluster-level context, we further compute the directional alignment between Clora\mathbf{C}_{\mathrm{lora}}Clora and Cfft\mathbf{C}_{\mathrm{fft}}Cfft. For each pair of clusters (Gm,Gn)(\mathcal{G}_m,\mathcal{G}_n)(Gm,Gn), we define their routing blocks as submatrices
The directional alignment between the two routing blocks is then measured by cosine similarity:
where ⟨⋅,⋅⟩\langle \cdot,\cdot \rangle⟨⋅,⋅⟩ denotes the Frobenius inner product. This cluster-level alignment is then propagated back to individual routing entries as a contextual factor. Let g(⋅)g(\cdot)g(⋅) map a singular index to its cluster id in {Gm}m=1M\{\mathcal{G}_m\}_{m=1}^M{Gm}m=1M within the top-kkk subspace. We define
and obtain the score of over-activation risk as
For entries with S(i,j)\mathbf{S}(i,j)S(i,j) exceeding a quantile threshold qactq_\mathrm{act}qact, CASA applies a magnitude-based arbitration rule that restricts the recovered routing strength to a safe range by retaining only the stronger contribution between LoRA and FFT:
All remaining entries in dominant regions are kept as Clora(i,j)\mathbf{C}_{\mathrm{lora}}(i,j)Clora(i,j), preserving the FFT-induced generative structure while avoiding widespread suppression of LoRA routing. Finally, we obtain the updated LoRA parameters by projecting Ccasa\mathbf{C}_{\mathrm{casa}}Ccasa back to the source weight space as Δcasa=UsCcasaVs⊤\Delta_{\mathrm{casa}} = \mathbf{U}_s\mathbf{C}_{\mathrm{casa}}\mathbf{V}_s^{\top}Δcasa=UsCcasaVs⊤, and applying a low-rank factorization to recover (B,A)(\mathbf{B},\mathbf{A})(B,A).
From a unified perspective, CASA can be interpreted as a constrained routing recovery problem in the singular subspace. It constructs an update Ccasa\mathbf{C}_{\mathrm{casa}}Ccasa such that the effective routing Cfft+Ccasa\mathbf{C}_{\mathrm{fft}} + \mathbf{C}_{\mathrm{casa}}Cfft+Ccasa recovers Clora\mathbf{C}_{\mathrm{lora}}Clora as closely as possible, while respecting the stability constraints imposed by the distilled generative subspace. To this end, CASA enforces routing compensation in non-dominant regions, where interactions are weakly coupled to generation, and restricts the recovered routing in spectrally dominant regions to remain within a safe activation envelope. This yields a minimal-intervention solution that restores LoRA functionality while preserving the generative space of the distilled model.
5. Experiments
In this section, we will evaluate the effectiveness of CASA on the task of transferring LoRAs trained on a base video diffusion model to its distilled variants. Moreover, we conduct ablation study and analysis to validate the design choices. Additional analyses are provided in Appendix D.
5.1 Experimental Settings
Models. We conduct experiments on two scales of video diffusion models based on the Wan2.1 [1] text-to-video architecture. We use Wan2.1-T2V-1.3B and Wan2.1-T2V-14B as source models on which LoRAs are trained. For Wan2.1-T2V-1.3B, we evaluate LoRA transfer to two distilled variants, including FastWan2.1-T2V-1.3B [6] and Rolling Forcing [35]. For Wan2.1-T2V-14B, we consider FastWan2.1-T2V-14B [6] and Krea Realtime Video [36] as target models. The target models cover both step distillation and causal distillation strategies.
Datasets. We evaluate CASA using a diverse set of LoRAs. For experiments on Wan2.1-T2V-1.3B, we use Steamboat-Willie-1.3B and Jinx-v2. For Wan2.1-T2V-14B, we adopt Film-Noir and Steamboat-Willie-14B. All LoRAs are obtained from public Hugging Face repositories. Detailed model card links are provided in Appendix C.1.
Metrics. We evaluate both generation stability and LoRA style preservation using complementary metrics. For generation quality, we adopt VideoAlign [37] to obtain the visual quality and motion quality, reporting their average as a Quality Score. To evaluate style transfer fidelity, we adopt CSD [38] to measure the style similarity between videos generated by the target model and those generated by the source model with same LoRAs. Please refer to Appendix C.2 for more details.
5.2 Main Results
Table 1 reports the quantitative comparison between direct LoRA reuse and CASA-based transfer across different settings. Overall, CASA consistently improves or maintains generation quality while achieving higher style fidelity, as reflected by Quality Score and CSD. Notably, the gains in CSD are more pronounced across most settings, indicating that CASA more reliably preserves LoRA-induced style under cross-model transfer and better recovers the intended adaptation effect. At the same time, the improvement in Quality Scores suggests that CASA effectively mitigates structural collapse during generation, leading to more coherent and stable video outputs. This trend is consistent across different LoRAs, highlighting the robustness of the proposed method under varying adaptation patterns. We further observe that the magnitude of improvement on 14B models is generally smaller than that on 1.3B models, indicating that larger models may inherently possess a more stable generative space and are therefore less sensitive to routing interference, and thus suffer less structural degradation introduced by such conflicts. This observation is also consistent with our analysis that stronger base models tend to better absorb perturbations, reducing the relative impact of transfer-induced interference.
5.3 Ablation Study
We conduct ablation studies on FastWan2.1-T2V-1.3B to evaluate the necessity and individual roles of the two core components of CASA, restoration (R) and arbitration (A). As shown in Table 2, using either component alone leads to suboptimal behavior. Restoration without arbitration improves style similarity but may degrade generation stability, as it introduces LoRA signals into sensitive routing pathways without accounting for potential conflicts. Conversely, arbitration without restoration preserves stability at the cost of weaker LoRA recovery, since it primarily constrains dominant pathways but does not sufficiently reinstate LoRA functionality in under-activated regions. Combining both components achieves the best overall performance across different LoRAs, consistently yielding higher CSD while maintaining strong quality. These results support our analysis that effective LoRA transfer on distilled video diffusion models requires both functional restoration in non-dominant regions and interference-aware arbitration in dominant routing pathways.
5.4 More Analysis
Comparison with ProLoRA.
ProLoRA [33] is a representative data-free LoRA transfer method originally developed for image generative models, which decomposes LoRA updates into the subspace and null space of the source model and projects them onto the corresponding subspaces of the target model. We compare CASA with ProLoRA on FastWan2.1-T2V-1.3B, and report the result in Table 3. As presented, ProLoRA exhibits a substantial drop in CSD across both LoRAs, suggesting a pronounced loss of LoRA functionality and reduced effectiveness of the transferred adaptation. In contrast, CASA consistently improves CSD while maintaining generation quality. The degradation observed with ProLoRA aligns with our analysis that dominant routing pathways shaped by full fine-tuning play a critical role in VDMs, and that effective LoRA reuse requires interference-aware arbitration rather than uniform projection, particularly in regions with concentrated routing interactions.
Visualization of CASA Intervention Regions.
Figure 7 illustrates how CASA intervenes in different regions of the routing space. The left panel shows the cluster-level routing energy density induced by FFT, where a small number of clusters exhibit substantially higher energy, indicating concentrated generative pathways. The middle panel reports the cluster-level directional alignment between FFT-induced and LoRA-induced routing. We observe heterogeneous alignment patterns, ranging from strong agreement to strong opposition, with pronounced structure in regions associated with high routing density. The right panel overlays the regions where CASA applies restoration, preservation, and arbitration5. Restoration is primarily applied to low-density regions, where routing interactions are weak and LoRA-induced behavior can be safely recovered with minimal risk of interference. In contrast, arbitration is selectively triggered in high-density regions with strong directional alignment, where over-activation might harm the generative space and thus requires careful regulation.
5.
Since arbitration operates at pixel-level, we regard blocks with more than pixels within it detected with over-activation risk as arbitrated.
Sensitivity Analysis. CASA introduces two quantile-based hyperparameters: qdomq_{\mathrm{dom}}qdom for identifying spectrally dominant routing regions, and qactq_{\mathrm{act}}qact for selecting routing entries requiring arbitration due to potential over-activation. To evaluate the robustness of CASA with respect to these thresholds, we conduct a sensitivity analysis by varying one parameter while fixing the other, and report both generation quality and style fidelity. Figure 8 shows that CASA maintains stable performance across a broad range of values. For qdomq_{\mathrm{dom}}qdom, moderate values (around 0.450.450.45–0.500.500.50) achieve a good balance between generation quality and CSD. Smaller values lead to overly conservative intervention, limiting LoRA recovery, while larger values underestimate dominant pathways and allow residual interference. Similarly, CASA is robust to variations in qactq_{\mathrm{act}}qact, as generation quality remains stable and CSD varies only slightly, indicating that the arbitration mechanism is not sensitive to precise threshold tuning. Overall, these results show that CASA does not require fine-grained hyperparameter tuning and remains effective across a wide range of threshold settings.
Execution Time. We report the execution time of CASA and its preprocessing steps across different model scales, with all experiments conducted on a single NVIDIA RTX 4090 GPU and excluding model and data loading overhead. CASA involves two one-time preprocessing stages and a lightweight per-LoRA transfer. First, singular value decomposition (SVD) of the source model’s Transformer weights is computed once and reused across different LoRAs and target models, taking several tens of seconds for Wan2.1-T2V-1.3B and approximately 36 minutes for Wan2.1-T2V-14B. Second, we compute the FFT-induced weight drift Cfft\mathbf{C}_{\mathrm{fft}}Cfft by projecting the difference between target and source weights onto the source singular bases. This step also runs once per source–target pair and takes tens of seconds for the 1.3B model and about 12 minutes for the 14B model. Given these precomputed components, CASA performs LoRA transfer efficiently, requiring around 5 seconds per LoRA for the 1.3B model and about 1 minute for the 14B model. Overall, the computational cost is dominated by the one-time preprocessing, while the per-LoRA overhead remains minimal, making CASA practical for scenarios involving repeated LoRA reuse across shared source and distilled models.
Extension on other Model Family. To verify whether our findings are general behaviors of video diffusion models, we extend our analysis and evaluation to a different video diffusion model family. Specifically, we use HunyuanVideo-1.5-480P-T2V [39] and HunyuanVideo-1.5-480P-T2V-CFG-Distill [39] as the base and distilled models, together with Retro-Anime6 as the style LoRA. On this new model family, we conduct the same analysis as in Section 3 and observe consistent mechanisms. In particular, both the distilled and LoRA-adapted models exhibit spectral rigidity with extremely small changes in singular values, reinforcing the stability of the underlying spectral structure. We also observe structured subspace perturbations, with stable head, block-wise mixing in the middle spectrum, and diffuse behavior in the tail, mirroring the patterns identified in the Wan family. At the routing level, we again find that high-interaction regions concentrate in head clusters, and that the interaction between LoRA and FFT can be either strongly aligned or strongly opposed, without a global bias. Details of these analyses can be found in Appendix. We further evaluate our method on the new model family using models and LoRA mentioned above. As presented in Table 4, CASA outperforms direct reuse in both quality and style fidelity, indicating the effectiveness of our method on this model family. Overall, these analyses and results demonstrate that our findings generalize beyond the Wan family and capture common structural behaviors in video diffusion models.
6.
Please see Appendix C.1 for details.
6. Conclusion
We studied why LoRAs trained on base video diffusion models often fail when reused on distilled variants. Through a weight-space analysis, we uncovered a pronounced spectral rigidity in VDMs and showed that both full fine-tuning and LoRA primarily introduce structured routing patterns at the cluster level, with incompatibility arising from conflicting routing interactions within spectrally functional subspaces. Motivated by these insights, we proposed Cluster-Aware Spectral Arbitration (CASA), a data-free transfer method that restores LoRA routing in non-dominant regions while arbitrating updates on dominant pathways to prevent over-activation. Experiments across multiple distillation paradigms and model scales demonstrate that CASA improves both generation stability and style fidelity, enabling effective LoRA reuse on distilled VDMs without additional training or user data.
Impact Statement
This paper presents work whose goal is to advance the field of Machine Learning. There are many potential societal consequences of our work, none of which we feel must be specifically highlighted here.
Acknowledgement
This work was partially supported by The Hong Kong University of Science and Technology (Guangzhou) Kunpeng&Ascend Center of Cultivation.
Appendix
A. Additional Examples of Structured Perturbations of Singular Subspaces
In this appendix, we provide additional evidence to support the structured singular subspace behavior analyzed in Section 3.2. Specifically, we visualize the cosine similarity matrices between left singular bases before and after adaptation for a broader set of layers and model scales.
Figure 9 and Figure 10 show representative results of layers from both Wan2.1-T2V-1.3B and Wan2.1-T2V-14B. For each projection matrix, we display the subspace similarity ∣U⊤U′∣\lvert \mathbf{U}^\top \mathbf{U}' \rvert∣U⊤U′∣ over three spectral regions: the head of the spectrum, a middle region, and the tail. The curves indicate the corresponding singular value magnitudes, illustrating the relationship between spectral separation and subspace stability.
Across all examined layers and projection types, several consistent patterns emerge. First, singular directions associated with the largest singular values remain highly aligned after adaptation, resulting in sharply concentrated diagonal structures in the head region. Second, the middle spectrum consistently exhibits block-wise mixing patterns, where groups of neighboring singular vectors form coherent clusters with strong intra-cluster similarity. These blocks align closely with step-like plateaus in the singular value curves, indicating that local spectral gaps constrain subspace mixing. Third, the tail of the spectrum shows increasingly diffuse similarity patterns, reflecting near-degeneracy and weaker spectral separation in this regime. We observe that Wan2.1-T2V-14B shows less mixing in these regions, indicating a more robust generative space.
Importantly, these behaviors persist across self-attention (
q/k/v/o), cross-attention, and feed-forward projections. This consistency suggests that the observed cluster-level organization is not an artifact of a specific layer or adaptation instance, but rather reflects a stable structural property of video diffusion models under fine-tuning.Together with the results in Section 3.2, these additional visualizations further support our claim that fine-tuning primarily induces structured perturbations of singular subspaces with strong cluster-level coherence, rather than arbitrary or unstructured rotations. This cluster-level organization provides a natural foundation for the routing-based analysis and arbitration strategy developed in the main paper.
B. More analysis of Cluster-level Routing
B.1 Cluster-level Routing Behaviors
In this appendix, we provide more visualization of the cluster-level routing behaviors introduced in Section 3.3. Specifically, we contrast the routing patterns induced by full fine-tuning and LoRA using both aggregated density profiles and cluster-to-cluster routing maps, covering self-attention, cross-attention, and feed-forward modules.
Cluster-wise Routing Density Profiles. Figure 11 illustrates the routing energy density associated with each singular cluster when acting as input (sender) or output (receiver). For each projection matrix, we report the root-mean-square (RMS) routing energy aggregated over all connections originating from or terminating at a given cluster.
Distributions of FFT and LoRA differ markedly in how routing energy is allocated. FFT displays a heterogeneous routing profile, where several clusters (typically in head regions) exhibit substantially higher routing energy relative to the rest. At the same time, many other clusters retain moderate decaying routing contributions, indicating that FFT does not collapse routing to a small subset of clusters.
By contrast, LoRA induces more uniformly distributed cluster-to-cluster interactions. Routing energy is spread across a larger number of cluster pairs, resulting in denser but less sharply peaked heatmaps. This distributed structure is consistent across attention and feed-forward modules.
Cluster-to-Cluster Routing Structure. Figure 12 further visualizes the routing behavior at the level of cluster pairs, where each entry represents the RMS routing energy between an input cluster and an output cluster. These heatmaps reveal the internal organization of routing beyond marginal cluster densities.
Under FFT, routing energy is typically highly structured and localized. Strong interactions concentrate in a limited set of cluster blocks, forming pronounced high-intensity regions that correspond to dominant generative pathways. Outside these regions, routing energy is uniformly weak, indicating that FFT selectively reinforces specific cluster-level connections.
By contrast, LoRA usually produces dense and widespread cluster-to-cluster interactions. Routing energy spreads broadly across the heatmap, with numerous cluster pairs exhibiting moderate interaction strength. This distributed pattern aligns with the view that LoRA operates by introducing cross-subspace couplings across a wide range of functional clusters, rather than amplifying a small number of existing pathways.
Implications for Routing Interference. Together, the density profiles and routing heatmaps provide complementary perspectives on the cluster-level routing behavior of FFT and LoRA. FFT establishes a centralized routing structure dominated by a small number of spectrally important clusters, while LoRA induces diffuse and globally distributed routing. When these two forms of routing coexist on the same model, interference tends to concentrate on the dominant cluster pathways emphasized by FFT, while remaining regions are weakly coupled.
These observations provide empirical support for the cluster-level routing interference analyzed in Section 3.4, and motivate the need for selective, cluster-aware spectral arbitration when transferring LoRA to distilled video diffusion models.
B.2 Generative Space from the Routing Perspective
B.2.1 Relationship between Dominant Clusters and Generative Space
In Section 3.3, we show that the distilled-model drift Δfft\Delta_{\mathrm{fft}}Δfft induces highly structured cluster-level routing in the singular subspace. Here, we provide additional evidence that the high-energy routing clusters in Cfft\mathbf{C}_{\mathrm{fft}}Cfft are tightly connected to the model's effective generative space.
For a source weight matrix Ws=UsSsVs⊤\mathbf{W}_s=\mathbf{U}_s\mathbf{S}_s\mathbf{V}_s^{\top}Ws=UsSsVs⊤ and its distilled counterpart Wt\mathbf{W}_tWt, we define the distilled drift
and its routing representation in the source singular basis
Following Section 4.1, we partition the leading singular directions into clusters {Gm}m=1M\{\mathcal{G}_m\}_{m=1}^{M}{Gm}m=1M (constructed on the top-kkk subspace that captures 90%90\%90% spectral energy).
We quantify cluster-level routing density on the sender and receiver sides by
Given a quantile threshold q∈[0,1]q\in[0,1]q∈[0,1], we define dominant sender/receiver cluster sets as
where Qq(⋅)Q_q(\cdot)Qq(⋅) denotes the qqq-quantile of the given set. We then mark a routing entry (i,j)(i,j)(i,j) as dominant if it involves a dominant sender or receiver cluster:
Ablating Non-Dominant Routing and Constructing a Partial-Distilled Model. To test how much of the distilled generative behavior is preserved by dominant routing clusters alone, we ablate the non-dominant part of Cfft\mathbf{C}_{\mathrm{fft}}Cfft. Concretely, we form an ablation routing matrix
where ⊙\odot⊙ denotes element-wise product. Projecting back to weight space gives the corresponding ablation update
We then obtain an ablated weight matrix by applying this update to the distilled model:
By construction, Wt(q)\mathbf{W}_t^{(q)}Wt(q) removes the routing contributions of non-dominant clusters, while retaining routing pathways associated with high-energy sender/receiver clusters. The special case q=0q=0q=0 corresponds to the original distilled model.
Results and Interpretation. We evaluate videos generated by Wt(q)\mathbf{W}_t^{(q)}Wt(q) using VideoAlign [37] and VBench [40] Imaging Quality under different quantile thresholds. Table 5 and Table 6 report representative metrics, and Figure 13 visualizes sample frames. As qqq increases, the generative quality gradually degrades, which is expected since more routing pathways are removed. However, a key qualitative observation is that the generated videos remain structurally coherent (e.g., subject layout and global motion stay organized) rather than collapsing into chaotic artifacts. This suggests that the dominant routing clusters in Cfft\mathbf{C}_{\mathrm{fft}}Cfft capture a substantial portion of the distilled model's effective generative space.
Notably, using a moderate quantile (around the median, i.e., preserving roughly half of the highest-density clusters) already maintains most of the perceived generative structure, even though fine details and sharpness may decrease. Together, these findings support the connection between cluster-level routing density and the generative space, that high-energy routing clusters correspond to functionally critical pathways that anchor coherent generation, while the remaining lower-energy routing contributes more to refinement. A more detailed characterization of which generative functions map to specific dominant clusters is an interesting direction for future work.
B.2.2 Over-Activation of Dominant Routing Blocks and Generative Failure
The above results suggest that dominant routing clusters play a central role in anchoring the effective generative space of distilled VDMs. We further probe this connection by directly testing the effect of over-activation on these dominant routing pathways.
Specifically, we consider the cluster-level routing matrix Cfft\mathbf{C}_{\mathrm{fft}}Cfft and partition it into small sender–receiver blocks corresponding to cluster pairs (Gm,Gn)(\mathcal{G}_m,\mathcal{G}_n)(Gm,Gn). For each block, we compute its routing energy density as the Frobenius norm normalized by block size. We then identify the top 5%5\%5% of the blocks with the highest routing energy density, which corresponds to the functional pathways that are the most utilized in the distilled model.
To simulate over-activation, we perturb these high-energy routing blocks by adding Gaussian noise sampled from N(μ=2,σ2=5)\mathcal{N}(\mu=2,\sigma^2=5)N(μ=2,σ2=5), while leaving all other routing entries as zero. The perturbed routing matrix is then projected back to the weight space using the source singular bases and added to the distilled model, yielding a modified model that differs from the original distilled model only through amplified activation along a small subset of dominant routing pathways.
The effect of this targeted over-activation is immediate and severe. As illustrated in Figure 14, videos generated from the perturbed model exhibit pronounced structural failures, including but not limited to character duplication, limb hallucination, spatial clipping, and floating artifacts. Importantly, these failures arise despite the perturbation being confined to a very small fraction of routing blocks, highlighting the extreme sensitivity of the generative process to excessive activation along dominant pathways.
These observations provide direct causal evidence that dominant routing clusters are not only sufficient to sustain coherent generation (as shown in the previous subsection), but are also fragile to over-activation. Excessive amplification along these pathways destabilizes the generative space, leading to catastrophic structural artifacts. This behavior mirrors the failure modes observed when directly reusing LoRAs on distilled models, and strongly supports our interpretation of LoRA incompatibility as a form of spectral over-activation and interference within shared functional routing subspaces.
C. Experimental Details
C.1 LoRA Model Details
We provide the urls of the LoRAs we used in our experiments here.
- Steamboat-Willie-1.3B: https://huggingface.co/benjamin-paine/steamboat-willie-1.3b
- Wan-LoRA-Arcane-Jinx-v2: https://huggingface.co/Cseti/Wan-LoRA-Arcane-Jinx-v2
- Film-Noir: https://huggingface.co/Remade-AI/Film-Noir
- Steamboat-Willie-14B: https://huggingface.co/benjamin-paine/steamboat-willie-14b
- Retro-Anime: https://tensor.art/models/966330915603110253/HY1.5-Anime-Retro-Anime-Style-v1.0
C.2 Details of Evaluation Metrics
Generation Quality. We explain how we measure the quality of videos generated by directly reusing LoRA on the distilled VDM and by transferring it with CASA. We deploy VideoAlign [37], a VLM-based reward model designed for video generation assessment. VideoAlign provides fine-grained reward signals by jointly evaluating Visual Quality, which measures frame-level image reasonableness and clarity, and Motion Quality, which evaluates dynamic stability, motion clarity, and the absence of artifacts such as flickering.
For each generated video, we compute both Visual Quality and Motion Quality scores using the pretrained VideoAlign reward model, and report their average as a unified video quality score. We then average the video-level scores across all evaluation samples to obtain the final quality score reported in our experiments. Compared to traditional handcrafted metrics, VideoAlign offers a learned, perceptually aligned evaluation that better captures common failure modes in video generation, making it particularly suitable for assessing generation stability under our setting.
Style Fidelity. We explain how we measure the style similarity of videos generated by the source model and the target model with the same LoRA applied here. For each LoRA, we first generate 30 videos using a same group of prompts (generated by Gemini-3-Flash) using the source model and the target model. We next extract the first frame as the style proxy of the video, and use CSD [38] to generate the style embeddings of these frames. Finally, we compute cosine similarities for each pair of videos generated by the source model and target model using the same prompt, and obtain their average similarity as the final CSD-Score.
C.3 Hardware
CASA can be implemented on Ascend 910B or 910C with PyTorch 2.6.0, SciPy 1.16.3, Safetensors 0.6.2 and NumPy 1.26.4.
D. Additional Analyses
D.1 Discussion and Guideline for Hyperparameter Selection
top-kkk Subspace Selection
As analyzed in Section 3.2, the head and middle spectrum exhibit clear cluster-level organization, while the tail is near-degenerate and shows diffuse mixing. Including the tail would introduce many singular directions that cannot be reliably clustered, degrading routing analysis. In practice, selecting the top-kkk subspace that captures 90% of the spectral energy works consistently across our settings. For other models, this threshold can be determined by inspecting the singular value spectrum and selecting the range where clear cluster-level structure (e.g., spectral plateaus with noticeable gaps) is present.
Selection of τ\tauτ for Clustering
The threshold τ\tauτ separates singular direction pairs based on their normalized interaction relative to spectral gaps. As shown in Section 3.2, intra-cluster pairs exhibit strong coupling with small spectral gaps, while inter-cluster pairs are weakly connected, leading to a natural separation in the interaction score. As a result, τ\tauτ functions as a coarse structural separator rather than a sensitive hyperparameter. Empirically, we observe that clustering remains unchanged across a wide range of τ\tauτ values (e.g., 1–10). In practice, a fixed value works across models without tuning.
Quantile-based Thresholds CASA uses two quantile-based hyperparameters: qdomq_{\mathrm{dom}}qdom for identifying dominant routing regions and qactq_{\mathrm{act}}qact for detecting potential over-activation. For qdomq_{\mathrm{dom}}qdom, we select it based on the relationship between routing energy and generation quality as analyzed in Appendix B.2.1. We observe that a subset of high-energy clusters is sufficient to maintain coherent generation. We therefore choose qdomq_{\mathrm{dom}}qdom as a quantile threshold that clusters with routing energy above this threshold preserve satisfactory generation quality. In practice, this threshold typically falls in the range of 0.45–0.55. For qactq_{\mathrm{act}}qact, we base its selection on the observation that over-activation occurs in a small number of high-energy routing regions, as analyzed in Appendix B.2.2. We therefore use a high quantile to capture these sparse high-risk entries.
D.2 Weight Space analysis on Hunyuanvideo-1.5
Spectral Rigidity. As shown in Figure 15, on HunyuanVideo-1.5 family, we observe strong spectral rigidity showing as extremely small changes in singular values from both the distilled and LoRA-adapted models.
Structured Perturbation. As presented in Figure 16, we observe structured perturbation with stable head, block-wise mixing in the middle spectrum, and diffuse behavior in the tail on HunyuanVideo-1.5 family, mirroring the pattern we find on Wan family.
Routing Interference. As shown in Figure 17, at the routing level, we again observe that high-interaction regions concentrate in head clusters, and that the interaction between LoRA and FFT can be either strongly aligned or strongly opposed, without a global bias. We also observe some architecture-dependent differences. For example, in HunyuanVideo-1.5, routing interference in txt_attn layers is less concentrated than in img_attn, and img_mod/txt_mod exhibit much lower effective rank, consistent with their role as compact conditional scaling modules.
References
[1] Team Wan et al. (2025). Wan: Open and Advanced Large-Scale Video Generative Models. arXiv preprint arXiv:2503.20314.
[2] Kong et al. (2024). Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603.
[3] Tim Brooks et al. (2024). Video generation models as world simulators.
[4] Jiang et al. (2025). VACE: All-in-One Video Creation and Editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 17191-17202.
[5] Shao et al. (2026). Efficient Video Diffusion Models: Advancements and Challenges. arXiv preprint arXiv:2604.15911.
[6] Zhang et al. (2025). Vsa: Faster video diffusion with trainable sparse attention. arXiv preprint arXiv:2505.13389.
[7] Ding et al. (2025). DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). pp. 17961-17971.
[8] Shanchuan Lin et al. (2025). Diffusion Adversarial Post-Training for One-Step Video Generation. In Forty-second International Conference on Machine Learning.
[9] Xun Huang et al. (2025). Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion. In The Thirty-ninth Annual Conference on Neural Information Processing Systems.
[10] Kaifeng Gao et al. (2025). Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing. In Forty-second International Conference on Machine Learning.
[11] Yin et al. (2025). From Slow Bidirectional to Fast Autoregressive Video Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 22963-22974.
[12] Edward J Hu et al. (2022). LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations.
[13] Andreas Blattmann et al. (2023). Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets. arXiv:2311.15127.
[14] Zhuoyi Yang et al. (2025). CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer. In The Thirteenth International Conference on Learning Representations.
[15] Rombach et al. (2022). High-Resolution Image Synthesis With Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 10684-10695.
[16] Haocheng Xi et al. (2025). Sparse Video-Gen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity. In Forty-second International Conference on Machine Learning.
[17] Yunhong Lu et al. (2025). Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation. arXiv:2512.04678.
[18] Shuai Yang et al. (2025). LongLive: Real-time Interactive Long Video Generation. arXiv:2509.22622.
[19] Gu et al. (2025). Long-context autoregressive video modeling with next-frame prediction. arXiv preprint arXiv:2503.19325.
[20] Shih-Yang Liu et al. (2024). DoRA: Weight-Decomposed Low-Rank Adaptation. In Forty-first International Conference on Machine Learning.
[21] Chenghao Fan et al. (2025). Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment. In Forty-second International Conference on Machine Learning.
[22] Fanxu Meng et al. (2024). PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems.
[23] Chongjie Si et al. (2025). Weight Spectra Induced Efficient Model Adaptation. arXiv:2505.23099.
[24] Reece S Shuttleworth et al. (2025). LoRA vs Full Fine-tuning: An Illusion of Equivalence. In The Thirty-ninth Annual Conference on Neural Information Processing Systems.
[25] Chongjie Si et al. (2024). See Further for Parameter Efficient Fine-tuning by Standing on the Shoulders of Decomposition. arXiv:2407.05417.
[26] Chongjie Si et al. (2025). Unleashing the Power of Task-Specific Directions in Parameter Efficient Fine-tuning. In The Thirteenth International Conference on Learning Representations.
[27] Tim Dettmers et al. (2023). QLoRA: Efficient Finetuning of Quantized LLMs. In Thirty-seventh Conference on Neural Information Processing Systems.
[28] Dawid Jan Kopiczko et al. (2024). VeRA: Vector-based Random Matrix Adaptation. In The Twelfth International Conference on Learning Representations.
[29] Qingru Zhang et al. (2023). Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning. In The Eleventh International Conference on Learning Representations.
[30] Ran et al. (2024). X-adapter: Adding universal compatibility of plugins for upgraded diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 8775–8784.
[31] Wang et al. (2024). Trans-LoRA: towards data-free Transferable Parameter Efficient Finetuning. In Advances in Neural Information Processing Systems. pp. 61217–61237. doi:10.52202/079017-1957.
[32] Farzad Farhadzadeh et al. (2025). LoRA-X: Bridging Foundation Models with Training-Free Cross-Model Adaptation. In The Thirteenth International Conference on Learning Representations.
[33] Farzad Farhadzadeh et al. (2025). Zero-Shot Adaptation of Parameter-Efficient Fine-Tuning in Diffusion Models. In Forty-second International Conference on Machine Learning.
[34] Chandler Davis and W. M. Kahan (1970). The Rotation of Eigenvectors by a Perturbation. III. SIAM Journal on Numerical Analysis. 7(1). pp. 1–46.
[35] Kunhao Liu et al. (2025). Rolling Forcing: Autoregressive Long Video Diffusion in Real Time. arXiv:2509.25161.
[36] Erwann Millon (2025). Krea Realtime 14B: Real-time Video Generation.
[37] Jie Liu et al. (2025). Improving Video Generation with Human Feedback. In The Thirty-ninth Annual Conference on Neural Information Processing Systems.
[38] Somepalli et al. (2024). Investigating Style Similarity in Diffusion Models. In Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part LXVI. pp. 143–160. doi:10.1007/978-3-031-72848-8_9.
[39] Bing Wu et al. (2025). HunyuanVideo 1.5 Technical Report. arXiv:2511.18870.
[40] Huang et al. (2024). VBench: Comprehensive Benchmark Suite for Video Generative Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.



















