Decompose, Adjust, Compose: Effective Normalization by Playing with Frequency for Domain Generalization
Sangrok LeeJongseong BaeHa Young Kim
Proposes a frequency-domain normalization framework that separates and regulates amplitude and phase variations to eliminate domain-specific styles while preserving essential semantic content for domain generalization.
Deep learning models frequently suffer significant performance degradation when deployed on unseen data distributions, a major operational risk for computer vision applications. To build models that remain robust across different environments, researchers rely on domain generalization. Standard approaches often use feature normalization to strip away domain-specific style information, treating statistical measures as style and normalized features as content. However, this process unintentionally distorts the core content information, undermining the model's reliability.
The article aims to explain and resolve this content distortion by analyzing normalization through frequency analysis and introducing feature-level normalization techniques that systematically regulate content and style variations for robust computer vision models.
The researchers mathematically demonstrated using the Fourier transform that the statistical mean shift in traditional normalization directly causes phase distortion, which alters the underlying image content. Based on this finding, they developed Phase-Consistent Normalization (PCNorm), which preserves original content by recombining the phase of pre-normalized features with the amplitude (style) of post-normalized features. They further expanded this concept into Content-Controlling Normalization (CCNorm) and Style-Controlling Normalization (SCNorm), which use learnable parameters to adjust the exact balance between original and normalized content and style. These modules were integrated into standard ResNet architectures to create two specialized models, DAC-P and DAC-SC, and evaluated across five major domain generalization benchmark datasets spanning thousands of multi-domain images: VLCS, PACS, Office-Home, DomainNet, and TerraIncognita.
The evaluation produced several key findings: First, the primary model, DAC-SC, established a new benchmark by achieving an average accuracy of 65.6% across all five datasets, outperforming established baseline methods. Second, DAC-SC achieved top performances on individual benchmarks, including 87.5% on PACS, 70.3% on Office-Home, and 44.9% on DomainNet (a 3.7 percentage point increase over the standard baseline). Third, controlling content and style through learnable parameters (CCNorm and SCNorm) generally provided superior generalization over strict content preservation (PCNorm). Fourth, in datasets with minimal domain shift like TerraIncognita, preserving pure content via PCNorm proved more effective (49.8% accuracy) than active adjustment.
These findings indicate that completely eliminating style or rigidly freezing content is suboptimal for real-world robustness. Instead, allowing vision models to learn the appropriate degree of content preservation and style filtering at intermediate network stages significantly improves out-of-distribution accuracy. For organizations deploying vision systems in uncontrolled environments, adopting these normalization techniques reduces operational failure risks and improves model reliability without needing target domain retraining.
Organizations developing computer vision pipelines should consider replacing standard normalization in downsampling and transition layers with content- and style-controlling modules. For tasks with severe domain gaps, DAC-SC is recommended, whereas applications with minor domain variations should favor the simpler PCNorm approach. Future work should evaluate these frequency-based normalization layers across different neural architectures beyond ResNet and in broader vision tasks.
The results carry high confidence based on consistent multi-trial evaluations across standard industry benchmarks. However, stakeholders should note that the empirical validations were primarily conducted using ResNet backbones in image classification settings, and adjustments may be needed when transferring the approach to other model families or complex multimodal tasks.
- Paper: In Search of Lost Domain Generalization, Ishaan Gulrajani et al. (2020). Introduces the standardized DomainBed evaluation benchmark that the source adopts to rigorously measure domain generalization performance across unseen environments.
- Paper: Domain Generalization: A Survey, Kaiyang Zhou et al. (2021). Provides a comprehensive taxonomy of domain generalization paradigms, establishing the theoretical background for feature alignment and domain-invariant representation learning.
- Paper: Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift, Sergey Ioffe et al. (2015). Introduces foundational feature-level normalization in deep networks, whose statistical shifts the source analyzes in the frequency domain to address content distortion.
- Paper: Deeper, Broader and Artier Domain Generalization, Da Li et al. (2017). Establishes the widely used PACS benchmark across diverse visual styles, serving as a primary experimental testbed for the source paper.
- Paper: Moment Matching for Multi-Source Domain Adaptation, Xingchao Peng et al. (2018). Introduces the large-scale multi-domain DomainNet dataset utilized by the source to validate out-of-distribution accuracy under severe domain gaps.
No sufficiently relevant recommendations were found.
