Efficient Multi-Scale Attention Module with Cross-Spatial Learning

Daliang OuyangSu HeJian ZhanHuaiyong GuoZhijie HuangM.L. LuoGuo-Liang Zhang

article2023IEEE International Conference on Acoustics, Speech, and Signal Processing1,657 citations

Proposes an efficient multi-scale attention module that avoids channel dimensionality reduction and captures pixel-level interactions across parallel branches, boosting visual representation performance in image classification and object detection with minimal computational overhead.

Listen

Modern computer vision models rely heavily on deep neural networks to recognize and detect objects. To boost accuracy without making networks excessively deep and slow, engineers frequently use attention mechanisms, which help models focus on the most informative features. However, conventional attention mechanisms often reduce channel dimensions to save computational budget, which unintentionally degrades visual representation quality, or rely on complex sequential operations that increase latency.

The article evaluates a novel architectural component called the Efficient Multi-Scale Attention (EMA) module, demonstrating that retaining complete channel information across parallel processing branches substantially enhances accuracy without introducing significant computational overhead.

The researchers designed a modular architecture that divides feature channels into distinct groups and reshapes them to avoid dimensionality reduction. The module processes features across parallel multi-scale branches—using both 1x1 and 3x3 convolutions—and fuses the resulting representations through a cross-spatial matrix dot-product learning mechanism. To validate the design, extensive empirical evaluations were conducted across standard image classification benchmarks (CIFAR-100 and ImageNet-1k) and object detection datasets (MS COCO and VisDrone2019) integrated into standard backbones such as ResNet, MobileNetV2, and YOLOv5.

The key findings demonstrate consistent performance advantages across tasks. First, on CIFAR-100 classification using ResNet50, the module improved Top-1 accuracy by 3.43 percentage points over the baseline and outperformed established attention methods while maintaining a compact footprint. Second, when integrated with ResNet101, the module achieved 80.86% Top-1 accuracy with fewer parameters (42.96 million versus 46.22 million) and lower computational costs than Coordinate Attention. Third, on MobileNetV2 with ImageNet-1k, the approach achieved a state-of-the-art 74.32% Top-1 accuracy while requiring fewer parameters (3.55 million versus 3.95 million) than Coordinate Attention. Finally, on MS COCO object detection using YOLOv5s, the approach reached 57.8% mean average precision at IoU 0.5 with negligible parameter additions (0.01 million), outperforming baseline models and competing attention mechanisms.

These results indicate that computer vision models can achieve superior accuracy and spatial awareness without the performance trade-offs commonly imposed by channel reduction. By capturing both short- and long-range dependencies efficiently, organizations can deploy higher-performing vision models to edge devices, drones, and mobile terminals without requiring expanded computational budgets or costly hardware upgrades.

Decision-makers should consider integrating the module into existing vision pipelines and edge-deployed models where latency, memory footprint, and detection precision are critical constraints. The source code is publicly accessible for immediate testing and pilot integration. For subsequent development, research teams should evaluate the module across broader visual tasks, such as semantic segmentation, and test deployment across diverse edge hardware platforms.

The findings are supported by consistent, reproducible results across standard computer vision benchmarks. However, evaluation is currently confined to 2D image classification and object detection in standard experimental settings. Practical application in real-time embedded systems or distinct tasks such as video tracking will require further empirical validation.

Cover for Efficient Multi-Scale Attention Module with Cross-Spatial Learning

Abstract

Remarkable effectiveness of the channel or spatial attention mechanisms for producing more discernible feature representation are illustrated in various computer vision tasks. However, modeling the cross-channel relationships with channel dimensionality reduction may bring side effect in extracting deep visual representations. In this paper, a novel efficient multi-scale attention (EMA) module is proposed. Focusing on retaining the information on per channel and decreasing the computational overhead, we reshape the partly channels into the batch dimensions and group the channel dimensions into multiple sub-features which make the spatial semantic features well-distributed inside each feature group. Specifically, apart from encoding the global information to re-calibrate the channel-wise weight in each parallel branch, the output features of the two parallel branches are further aggregated by a cross-dimension interaction for capturing pixel-level pairwise relationship. We conduct extensive ablation studies and experiments on image classification and object detection tasks with popular benchmarks (e.g., CIFAR-100, ImageNet-1k, MS COCO and VisDrone2019) for evaluating its performance.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Efficient Multi-Scale Attention
  • 3.1. Revisit Coordinate Attention (CA)
  • 3.2. Multi-Scale Attention (EMA) Module
  • 4. Experiments
  • 4.1. Image Classification on CIFAR-100
  • 4.2. Image Classification on ImageNet-1k
  • 4.3. Object Detection on MS COCO
  • 4.4. Object Detection on VisDrone
  • 5. Ablation Study
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Efficient Multi-Scale Attention Module Architecture

    model/method

    The Efficient Multi-Scale Attention (EMA) module is an attention mechanism for deep convolutional neural networks designed to establish both short- and long-range dependencies across spatial and channel dimensions without applying channel dimensionality reduction.

    Given an intermediate feature tensor X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}, where CC is the number of channels, HH is feature height, and WW is feature width, the EMA module executes three stages:

    1. Feature Grouping: XX is partitioned along the channel dimension into GG sub-features: X=[X0,X1,…,XG−1]X = [X_0, X_1, \dots, X_{G-1}], where each sub-feature Xi∈R(C/G)×H×WX_i \in \mathbb{R}^{(C/G) \times H \times W} with G≪CG \ll C. The GG group dimensions are reshaped and permuted into the batch dimension (yielding effective batch size B×GB \times G), allowing standard convolution kernels to operate on each group independently without channel reduction or parameter overhead.

    2. Parallel Subnetworks: Within each group, features are passed in parallel into two branches:

      • 1×11 \times 1 Branch: Encodes directional global spatial context via two 1D global average pooling operations along the horizontal (XX) and vertical (YY) axes. The two 1D encoded feature vectors are concatenated along the spatial dimension, transformed via a shared 1×11 \times 1 convolution without channel dimensionality reduction, split back into two directional vectors, and passed through Sigmoid activations. The resulting channel-wise attention maps are multiplied together to model local cross-channel interaction.
      • 3×33 \times 3 Branch: Directly applies a single 3×33 \times 3 convolution kernel to the sub-feature tensor XiX_i to capture local multi-scale spatial structure and expand the receptive field.
    3. Cross-Spatial Learning: The output representations of the 1×11 \times 1 and 3×33 \times 3 branches are aggregated across spatial directions using 2D global average pooling, Softmax normalizations, and matrix dot-product operations to yield pixel-level pairwise attention maps that re-weight the grouped features.

  2. Knowl 2 — Cross-Spatial Learning Aggregation in EMA

    model/method

    The Cross-Spatial Learning mechanism in the Efficient Multi-Scale Attention (EMA) module aggregates spatial and channel contextual information across two parallel subnetwork branches (a 1×11 \times 1 branch and a 3×33 \times 3 branch) without channel reduction.

    Let F1×1∈R(C/G)×H×WF_{1\times 1} \in \mathbb{R}^{(C/G) \times H \times W} and F3×3∈R(C/G)×H×WF_{3\times 3} \in \mathbb{R}^{(C/G) \times H \times W} denote the output feature maps of the 1×11 \times 1 branch and 3×33 \times 3 branch for a given group of C/GC/G channels, where HH and WW are spatial height and width.

    1. First Spatial Attention Map: A 2D global average pooling operation is applied to F1×1F_{1\times 1} to extract global spatial context into a channel descriptor R(C/G)×1×1\mathbb{R}^{(C/G) \times 1 \times 1}, followed by a Softmax normalization to yield a representation S1∈R1×(C/G)S_1 \in \mathbb{R}^{1 \times (C/G)}. Concurrently, F3×3F_{3\times 3} is reshaped to R(C/G)×(HW)\mathbb{R}^{(C/G) \times (HW)}. The matrix dot-product of S1S_1 and the reshaped F3×3F_{3\times 3} produces an intermediate spatial attention map M1∈R1×H×WM_1 \in \mathbb{R}^{1 \times H \times W}.

    2. Second Spatial Attention Map: Symmetrically, 2D global average pooling and Softmax normalization are applied to F3×3F_{3\times 3} to yield S2∈R1×(C/G)S_2 \in \mathbb{R}^{1 \times (C/G)}, while F1×1F_{1\times 1} is reshaped to R(C/G)×(HW)\mathbb{R}^{(C/G) \times (HW)}. Their matrix dot-product generates a second spatial attention map M2∈R1×H×WM_2 \in \mathbb{R}^{1 \times H \times W}.

    3. Feature Aggregation: The two spatial attention maps M1M_1 and M2M_2 are summed and passed through a Sigmoid activation function σ(⋅)\sigma(\cdot) to produce the composite pixel-level spatial weight map. This weight map is multiplied with the original group feature map to generate the re-weighted output tensor of dimension (C/G)×H×W(C/G) \times H \times W, preserving global context and local multi-scale positional structure.

  3. Knowl 3 — Spatial Global Average Pooling Operations in EMA

    equation

    For an intermediate feature tensor X∈RC×H×WX \in \mathbb{R}^{C \times H \times W}, where c∈{1,…,C}c \in \{1, \dots, C\} is the channel index, HH is the spatial height, and WW is the spatial width, spatial pooling is formulated along 1D and 2D directions:

    1. 1D Horizontal Global Average Pooling: Encodes global positional information along the horizontal dimension at height h∈{0,…,H−1}h \in \{0, \dots, H-1\}: zcH(h)=1W∑i=0W−1xc(h,i)z_c^H(h) = \frac{1}{W} \sum_{i=0}^{W-1} x_c(h, i)

    2. 1D Vertical Global Average Pooling: Encodes global positional information along the vertical dimension at width w∈{0,…,W−1}w \in \{0, \dots, W-1\}: zcW(w)=1H∑j=0H−1xc(j,w)z_c^W(w) = \frac{1}{H} \sum_{j=0}^{H-1} x_c(j, w)

    3. 2D Global Average Pooling: Encodes spatial context across the entire 2D spatial area for channel cc: zc=1H×W∑i=1H∑j=1Wxc(i,j)z_c = \frac{1}{H \times W} \sum_{i=1}^{H} \sum_{j=1}^{W} x_c(i, j)

    where xc(i,j)x_c(i, j) denotes the feature value at spatial location (i,j)(i, j) in channel cc.

  4. Knowl 4 — CIFAR-100 Image Classification Benchmarks with EMA

    data/table

    Comparison of the Efficient Multi-Scale Attention (EMA) module against baseline CNN backbones (ResNet-50 and ResNet-101) and alternative attention mechanisms (CBAM, SA, ECA, NAM, CA) on the CIFAR-100 classification benchmark. All models were trained for 200 epochs using SGD (momentum 0.9, weight decay 4×10−54\times 10^{-5}, default batch size 128).

    Method Backbone #Param. FLOPs Top-1 (%) Top-5 (%)
    Baseline ResNet50 23.71M 1.30G 77.26 93.63
    + CBAM ResNet50 26.24M 1.31G 80.56 95.34
    + SA ResNet50 23.71M 1.31G 79.92 95.00
    + ECA ResNet50 23.71M 1.31G 79.68 95.05
    + NAM ResNet50 23.71M 1.31G 80.62 95.28
    + CA ResNet50 25.57M 1.36G 80.17 94.94
    + EMA (ours) ResNet50 23.85M 1.32G 80.69 95.59
    Baseline ResNet101 42.70M 2.53G 77.78 94.39
    + CA ResNet101 46.22M 2.54G 80.01 94.78
    + EMA (ours) ResNet101 42.96M 2.53G 80.86 95.75

    Integrating EMA into ResNet-50 improves Top-1 accuracy by 3.43% over the baseline while adding only 0.14M parameters and 0.02G FLOPs. On ResNet-101, EMA achieves 80.86% Top-1 accuracy, outperforming Coordinate Attention (CA) by 0.85% while using 3.26M fewer parameters.

  5. Knowl 5 — ImageNet-1k Classification Performance with MobileNetV2

    data/table

    Validation performance on ImageNet-1k (1000 classes, 224×224224 \times 224 resolution) comparing the Efficient Multi-Scale Attention (EMA) module against Squeeze-and-Excitation (SE), CBAM, and Coordinate Attention (CA) using MobileNetV2 as the backbone. Models were trained for 400 epochs with batch size 256, SGD optimizer (learning rate 0.4, momentum 0.9, weight decay 1×10−51\times 10^{-5}), linear warmup learning rate 1×10−61\times 10^{-6}, and exponential moving average.

    Method Backbone #Param. M-Adds Top-1 (%) Top-5 (%)
    Baseline MobileNetV2 3.50M 300M 72.30 91.02
    + SE MobileNetV2 3.89M 300M 73.50 -
    + CBAM MobileNetV2 3.89M 300M 73.60 -
    + CA MobileNetV2 3.95M 310M 74.30 -
    + EMA (ours) MobileNetV2 3.55M 306M 74.32 91.82

    EMA achieves a Top-1 accuracy of 74.32% on ImageNet-1k with 3.55M parameters, outperforming CA (74.30% with 3.95M parameters) and CBAM (73.60% with 3.89M parameters) while requiring fewer parameters than SE, CBAM, and CA.

  6. Knowl 6 — Object Detection Performance on MS COCO and VisDrone2019

    data/table

    Evaluation of attention modules embedded into YOLOv5s (v6.0) on the MS COCO dataset (input size 640×640640 \times 640, 300 epochs, batch size 50) and into YOLOv5x (v6.0) on the VisDrone2019 dataset (input size 640×640640 \times 640, 300 epochs, batch size 5).

    Model Datasets #Param. FLOPs mAP (0.5) mAP (0.5:0.95)
    YOLOv5s COCO 7.23M 16.5M 56.0% 37.2%
    + CBAM COCO 7.27M 16.6M 57.1% 37.7%
    + SA COCO 7.23M 16.5M 56.8% 37.4%
    + ECA COCO 7.23M 16.5M 57.1% 37.6%
    + CA COCO 7.26M 16.50M 57.5% 38.1%
    + EMA (ours) COCO 7.24M 16.53M 57.8% 38.4%
    YOLOv5x VisDrone 90.96M 314.2M 49.29% 30.0%
    + CBAM VisDrone 91.31M 315.1M 49.40% 30.1%
    + CA VisDrone 91.28M 315.2M 49.30% 30.1%
    + EMA (ours) VisDrone 91.18M 315.0M 49.70% 30.4%

    On MS COCO, adding EMA to YOLOv5s increases mAP(0.5) from 56.0% to 57.8% (+1.8%) and mAP(0.5:0.95) from 37.2% to 38.4% (+1.2%) with a negligible parameter increase of 0.01M. On VisDrone2019, YOLOv5x with EMA achieves 49.70% mAP(0.5) and 30.4% mAP(0.5:0.95), outperforming CBAM and CA while using fewer parameters and FLOPs.

  7. Knowl 7 — Ablation on Group Size and Cross-Spatial Learning in EMA

    data/table

    Ablation study evaluating the contribution of the cross-spatial learning component and the choice of channel group size GG in the EMA module on CIFAR-100 using a ResNet-50 backbone.

    Method #Param. FLOPs Top-1 (%) Top-5 (%)
    + EMA_no 23.84M 1.32G 78.24 94.89
    + EMA_16 24.44M 1.34G 80.35 95.44
    + EMA_32 23.84M 1.32G 80.69 95.59
    • EMA_no: EMA module without cross-spatial learning aggregation.
    • EMA_16: EMA module with group size G=16G = 16.
    • EMA_32: EMA module with group size G=32G = 32.

    Enabling cross-spatial learning (EMA_32) provides a +2.45% gain in Top-1 accuracy over disabling it (EMA_no). Increasing the number of groups from G=16G = 16 to G=32G = 32 reduces model parameters from 24.44M to 23.84M and FLOPs from 1.34G to 1.32G due to grouping more channels into the batch dimension, while simultaneously improving Top-1 accuracy from 80.35% to 80.69%.

Coverage note — None was omitted; all key architectural components, equations, experiments across classification and detection datasets, and ablation results from the paper are fully covered.

References

  1. 1.L. Chen, H. Zhang, J. Xiao, L. Nie, J. Shao, W. Liu, and T. Chua. SCA-CNN: Spatial and channel-wise attention in convolutional networks for image captioning. In CVPR, 2017.
  2. 2.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR, 2018.
  3. 3.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In CVPR, Jun 2018.
  4. 4.Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. CBAM: convolutional block attention module. In ECCV, 2018.
  5. 5.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton, “Imagenet classification with deep convolutional neural networks,” Communications of the ACM, vol. 60, no. 6, pp. 84–90, 2017.
  6. 6.Xiang Li, Xiaolin Hu, and Jian Yang. Spatial groupwise enhance: Improving semantic feature learning in convolutional networks. CoRR, vol. abs/1905.09646, 2019.
  7. 7.Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun, “Shufflenet v2: Practical guidelines for efficient cnn architecture design,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 116–131.
  8. 8.Qibin Hou, Daquan Zhou, and Jiashi Feng. Coordinate Attention for Efficient Mobile Network Design. In CVPR, 2021.
  9. 9.Huajun Liu, Fuqiang Liu, Xinyi Fan, and Dong Huang. Polarized Self-Attention: Towards High-quality Pixel-wise Regression. CoRR, vol. abs/2107.00782, 2021.
  10. 10.Qilong Wang, Banggu Wu, Pengfei Zhu, Peihua Li, Wangmeng Zuo, and Qinghua Hu. ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks. In CVPR, 2020.
  11. 11.Saining Xie, Ross Girshick, Piotr Dollar, Zhuowen Tu, and Kaiming He, “Aggregated residual transformations for deep neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 1492–1500.
  12. 12.Xiang Li, Wenhai Wang, Xiaolin Hu, and Jian Yang,“Selective kernel networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 510–519.
  13. 13.Diganta Misra, Trikay Nalamada, Ajay Uppili Arasanipalai, and Qibin Hou. Rotate to attend: Convolutional triplet attention module. In CVPR, 2021.
  14. 14.Qing-Long Zhang, and Yu-Bin Yang. SA-Net: Shuffle Attention for Deep Convolutional Neural Networks. CoRR, vol. abs/2102.00240, 2021.
  15. 15.Ankit Goyal, Jia Deng, and Vladlen Koltun. Non-deep Networks. CoRR, vol. abs/2110.07641, 2021.
  16. 16.Yichao Liu, Zongru Shao, Yueyang Teng, and Nico Hoffmann. NAM: Normalization-based Attention Module. CoRR, vol. abs/2111.12419, 2021.
  17. 17.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proc. of Computer Vision and Pattern Recognition (CVPR), 2016.
  18. 18.G. Jocher et al., Ultralytics/YOLOv5: V6.0—YOLOv5n ‘nano’ models roboflow integration TensorFlow export OpenCV DNN support, Oct. 2021.
  19. 19.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In NeurIPs, 2012.
  20. 20.Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In CVPR, 2017.
  21. 21.Shang-Hua Gao, Ming-Ming Cheng, Kai Zhao, Xin-Yu Zhang, Ming-Hsuan Yang, and Philip Torr. Res2net: A new multi-scale backbone architecture. arXiv preprint arXiv:1904.01169, 2019.
  22. 22.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the inception architecture for computer vision. In CVPR, 2016.
  23. 23.Xiang Li, Wenhai Wang, Xiaolin Hu, and Jian Yang. Selective kernel networks. In CVPR, 2019.
  24. 24.Hu Zhang, Keke Zu, Jian Lu, Yuru Zou, and Deyu Meng. EPSANet: An Effificient Pyramid Split Attention Block on Convolutional Neural Network. arXiv:2105.14447[cs.CV], 2021.
  25. 25.Jun Fu, Jing Liu, Haijie Tian, Yong Li, Yongjun Bao, Zhiwei Fang, and Hanqing Lu, “Dual attention network for scene segmentation,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 3146–3154.
  26. 26.Luchen Liu, Sheng Guo, Weilin Huang, and Matthew R Scott, “Decoupling category-wise independence and relevance with self-attention for multi-label image classification,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2019, pp. 1682–1686.
  27. 27.Yunpeng Chen, Yannis Kalantidis, Jianshu Li, Shuicheng Yan, and Jiashi Feng, “Aˆ 2-nets: Double attention networks,” Advances in neural information processing systems, vol. 31, 2018.
  28. 28.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He, “Non-local neural networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7794–7803.
  29. 29.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4510–4520, 2018.
  30. 30.Xingkui Zhu, Shuchang Lyu, Xu Wang and Qi Zhao. TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-captured Scenarios. CoRR, vol. abs/2108.11539, 2021.

Citation

MLA
Ouyang, D., et al. “Efficient Multi-Scale Attention Module with Cross-Spatial Learning”. ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5, https://doi.org/10.1109/ICASSP49357.2023.10096516.
APA
Ouyang, D., He, S., Zhang, G., Luo, M., Guo, H., Zhan, J., & Huang, Z. (2023). Efficient Multi-Scale Attention Module with Cross-Spatial Learning. ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5. https://doi.org/10.1109/ICASSP49357.2023.10096516
Chicago
Ouyang, D., S. He, G. Zhang, et al. 2023. “Efficient Multi-Scale Attention Module with Cross-Spatial Learning”. ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5. https://doi.org/10.1109/ICASSP49357.2023.10096516.
Harvard
Ouyang, D. et al. (2023) “Efficient Multi-Scale Attention Module with Cross-Spatial Learning”, ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, pp. 1–5. Available at: https://doi.org/10.1109/ICASSP49357.2023.10096516.
Vancouver
1. Ouyang D, He S, Zhang G, Luo M, Guo H, Zhan J, Huang Z (2023) Efficient Multi-Scale Attention Module with Cross-Spatial Learning. In: ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, pp 1–5

BibTeX

@inproceedings{Ouyang_2023, title={Efficient Multi-Scale Attention Module with Cross-Spatial Learning}, url={http://dx.doi.org/10.1109/ICASSP49357.2023.10096516}, DOI={10.1109/icassp49357.2023.10096516}, booktitle={ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)}, publisher={IEEE}, author={Ouyang, Daliang and He, Su and Zhang, Guozhong and Luo, Mingzhu and Guo, Huaiyong and Zhan, Jian and Huang, Zhijie}, year={2023}, month=June, pages={1–5} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF