CycleMix: A Holistic Strategy for Medical Image Segmentation from Scribble Supervision

Ke ZhangXiahai Zhuang

article2022CVPR110 citations

Proposes a weakly-supervised medical image segmentation framework that combines mixup-based scribble augmentation with global and local cycle-consistency regularization to achieve segmentation accuracy comparable to fully-supervised methods.

Listen

Manual segmentation of medical imaging is essential for clinical decision-making and deep learning model training, yet acquiring complete pixel-level annotations requires extensive expert labor and significant financial expense. Weakly-supervised learning using sparse scribble annotations—where clinical experts draw simple line markings across target anatomical regions—offers a much faster and cheaper alternative. However, training effective models from sparse scribbles remains difficult because standard models struggle to capture correct anatomical boundaries and shape priors.

The article aims to introduce and evaluate CycleMix, a weakly-supervised deep learning framework designed to accurately segment medical images using only sparse scribble annotations. The framework addresses supervision sparsity by pairing a two-step data mixup augmentation with two-level consistency regularization.

To evaluate this approach, the authors conducted experiments using two benchmark cardiac magnetic resonance imaging (MRI) datasets: the cine-MRI ACDC dataset (100 subjects) and the late gadolinium enhancement MSCMRseg dataset (45 subjects). The framework combines two training images and their annotations based on saliency (increments of scribbles) and applies random rectangular occlusions (decrements of scribbles) to improve target localization. To stabilize learning and preserve realistic organ shapes, CycleMix applies a global consistency loss—ensuring image patches segment consistently whether viewed individually or mixed—and a local consistency loss that enforces anatomical connectivity. Performance was evaluated using the Dice similarity coefficient against weakly-supervised baselines, shape-prior methods using additional masks, and fully-supervised models.

The findings show that CycleMix substantially improves scribble-based segmentation performance. First, on scribble annotations alone, CycleMix achieved average Dice scores of 84.8% on ACDC and 80.0% on MSCMRseg, outperforming standard mixup baselines by up to 22.4% and 55.9% in absolute score gains. Second, CycleMix surpassed competing weakly-supervised methods, including advanced generative models that required additional unpaired full-mask images. Third, models trained with CycleMix under scribble supervision matched or marginally exceeded the performance of standard networks trained on 100% fully-annotated masks (84.8% vs. 82.0% on ACDC; 80.0% vs. 75.5% on MSCMRseg). Finally, data sensitivity analyses revealed that incorporating just 20% fully-annotated data alongside scribbles pushed segmentation accuracy above 87%, with overall performance gains leveling off around a 40% full-annotation ratio.

These results demonstrate that clinical institutions can significantly lower data curation costs and project timelines by shifting from exhaustive pixel annotations to quick scribble labeling without compromising segmentation quality. In addition, the framework avoids the computational complexity and data overhead of auxiliary generative networks while preserving essential organ geometry.

Based on these findings, medical AI development teams should adopt scribble-based annotation protocols paired with CycleMix-style mixup and consistency objectives for cardiac MRI segmentation pipelines. When higher precision is required, teams can optimize annotation resources by creating a hybrid dataset containing roughly 20% to 40% fully-annotated cases combined with scribbles, rather than fully annotating the entire cohort.

While confidence in the reported cardiac benchmarks is high, the evaluation is limited to 2D cardiac MRI datasets with relatively small patient cohorts. Further validation across broader imaging modalities, distinct anatomical structures, and larger 3D clinical imaging volumes is recommended before general clinical deployment.

arXiv: 2203.01475
Cover for CycleMix: A Holistic Strategy for Medical Image Segmentation from Scribble Supervision

Abstract

Curating a large set of fully annotated training data can be costly, especially for the tasks of medical image segmentation. Scribble, a weaker form of annotation, is more obtainable in practice, but training segmentation models from limited supervision of scribbles is still challenging. To address the difficulties, we propose a new framework for scribble learning-based medical image segmentation, which is composed of mix augmentation and cycle consistency and thus is referred to as CycleMix. For augmentation of supervision, CycleMix adopts the mixup strategy with a dedicated design of random occlusion, to perform increments and decrements of scribbles. For regularization of supervision, CycleMix intensifies the training objective with consistency losses to penalize inconsistent segmentation, which results in significant improvement of segmentation performance. Results on two open datasets, i.e., ACDC and MSCMRseg, showed that the proposed method achieved exhilarating performance, demonstrating comparable or even better accuracy than the fully-supervised methods. The code and expert-made scribble annotations for MSCMRseg are publicly available at https://github.com/BWGZK/CycleMix.

Table of Contents

  • 1. Introduction
  • 2. Related works
  • 2.1. Learning from scribble supervision
  • 2.2. Mixup augmentations
  • 2.3. Consistency regularization
  • 3. Method
  • 3.1. Mix augmentation of scribble supervision
  • 3.1.1. Increments of scribbles
  • 3.1.2. Decrements of scribbles
  • 3.1.3. Scribble supervision
  • 3.2. Regularization of supervision via cycle consistency
  • 3.2.1. Global consistency
  • 3.2.2. Local consistency
  • 4. Experiments
  • 4.1. Data and evaluation metric
  • 4.2. Experimental setup
  • 4.3. Comparison with different mix-up strategies
  • 4.4. Comparison with weakly-supervised methods
  • 4.5. Ablation study
  • 4.6. Data sensitivity study
  • 4.7. Experiments on fully-annotated data
  • 5. Conclusions
  • References

Knowls

  1. Knowl 1 — CycleMix Framework and Objective Function

    model/method

    CycleMix is a weakly-supervised medical image segmentation framework designed to train neural networks using only sparse scribble annotations. It pairs mixup-based data augmentation of scribble supervision (creating increments and decrements of annotated pixels) with two-level cycle consistency regularization (global and local) to constrain the model without requiring extra fully-annotated masks.

    Given a segmentation network S(⋅)S(\cdot), the overall training objective L\mathcal{L} is formulated as the sum of supervised losses on annotated scribble pixels and unsupervised consistency regularization losses:

    L=(λ1Lunmix+λ2Lmix)+(λ3Lcon-g+λ4Lcon-l)\mathcal{L} = (\lambda_1 L_{\text{unmix}} + \lambda_2 L_{\text{mix}}) + (\lambda_3 L_{\text{con-g}} + \lambda_4 L_{\text{con-l}})

    where:

    • LunmixL_{\text{unmix}} is the cross-entropy loss evaluated on unmixed scribble-annotated training images.
    • LmixL_{\text{mix}} is the cross-entropy loss evaluated on augmented, mixed-and-occluded training images over their combined scribble annotations.
    • Lcon-gL_{\text{con-g}} is the global consistency loss enforcing mix-invariance between segmentations of original images and segmentations of the mixed image.
    • Lcon-lL_{\text{con-l}} is the local consistency loss enforcing anatomical connectivity between predictions and their largest connected components.
    • λ1,λ2,λ3,λ4>0\lambda_1, \lambda_2, \lambda_3, \lambda_4 > 0 are weighting hyperparameters balancing the individual loss components.
  2. Knowl 2 — Two-Step Mix Augmentation of Scribble Supervision

    model/method

    To expand the sparse gradient flow provided by scribbles, CycleMix employs a two-step augmentation pipeline that first combines images to increase annotated regions and then occludes regions to test localization robustness.

    1. Increments of Scribbles via Saliency Mixup: Given two dd-dimensional images with scribble annotations (x1,y1)(x_1, y_1) and (x2,y2)(x_2, y_2), a saliency-guided mixup function M(⋅,⋅)M(\cdot, \cdot) generates mixed samples (x12m,y12m)(x_{12}^m, y_{12}^m):

    x12m=M(x1,x2)=(1−z)⊙Π1Tx1+z⊙Π2Tx2x_{12}^m = M(x_1, x_2) = (1 - z) \odot \Pi_1^T x_1 + z \odot \Pi_2^T x_2

    y12m=M(y1,y2)=(1−z)⊙Π1Ty1+z⊙Π2Ty2y_{12}^m = M(y_1, y_2) = (1 - z) \odot \Pi_1^T y_1 + z \odot \Pi_2^T y_2

    where Π1,Π2∈{0,1}d×d\Pi_1, \Pi_2 \in \{0, 1\}^{d \times d} are transportation matrices, z∈[0,1]dz \in [0, 1]^d is a mixing mask, and ⊙\odot denotes element-wise multiplication. The parameter set {Π1,Π2,z}\{\Pi_1, \Pi_2, z\} is determined by maximizing the mixed saliency:

    {Π1,Π2,z}=arg⁡max⁡Π1,Π2,z[(1−z)⊙Π1Ts(x1)+z⊙Π2Ts(x2)]\{\Pi_1, \Pi_2, z\} = \arg\max_{\Pi_1, \Pi_2, z} \left[ (1 - z) \odot \Pi_1^T s(x_1) + z \odot \Pi_2^T s(x_2) \right]

    where s(x)=∥∇x∥2s(x) = \|\nabla x\|_2 represents the saliency map of image xx.

    1. Decrements of Scribbles via Random Occlusion: To prevent overfitting and promote localized shape learning, a binary rectangular mask 1O\mathbf{1}_O of size n×nn \times n (with random orientation) replaces part of the mixed image and scribble label with background:

    x12o=(1−1O)⊙x12mx_{12}^o = (1 - \mathbf{1}_O) \odot x_{12}^m

    y12o=(1−1O)⊙y12my_{12}^o = (1 - \mathbf{1}_O) \odot y_{12}^m

    1. Supervised Loss Formulation: The supervised loss is computed exclusively on the scribble-annotated pixel subset ΩL\Omega_L using partial cross-entropy Lce(y^,y)=−∑i∈ΩL∑k∈Ky[i,k]log⁡(y^[i,k])L_{\text{ce}}(\hat{y}, y) = -\sum_{i \in \Omega_L} \sum_{k \in K} y[i, k] \log(\hat{y}[i, k]), where KK is the set of class indices and y^=S(x)\hat{y} = S(x) is the predicted probability map:

    Lunmix=12[Lce(S(x1),y1)+Lce(S(x2),y2)]L_{\text{unmix}} = \frac{1}{2} \left[ L_{\text{ce}}(S(x_1), y_1) + L_{\text{ce}}(S(x_2), y_2) \right]

    Lmix=12[Lce(S(x12o),y12o)+Lce(S(x21o),y21o)]L_{\text{mix}} = \frac{1}{2} \left[ L_{\text{ce}}(S(x_{12}^o), y_{12}^o) + L_{\text{ce}}(S(x_{21}^o), y_{21}^o) \right]

  3. Knowl 3 — Global Mix-Invariant Consistency Regularization

    model/method

    Global consistency enforces the mix-invariant property across segmentation predictions, requiring that mixing the segmentations of individual unmixed images yields the same result as segmenting the mixed and occluded image directly:

    (1−1O)⊙M(S(x1),S(x2))=S((1−1O)⊙M(x1,x2))(1 - \mathbf{1}_O) \odot M(S(x_1), S(x_2)) = S\left( (1 - \mathbf{1}_O) \odot M(x_1, x_2) \right)

    The global consistency loss Lcon-gL_{\text{con-g}} penalizes deviations from this equivalence using a symmetrical negative cosine similarity metric:

    Lcon-g=12[Lncs(p12,q12)+Lncs(p21,q21)]L_{\text{con-g}} = \frac{1}{2} \left[ L_{\text{ncs}}(p_{12}, q_{12}) + L_{\text{ncs}}(p_{21}, q_{21}) \right]

    where:

    • p12=(1−1O)⊙M(y^1,y^2)p_{12} = (1 - \mathbf{1}_O) \odot M(\hat{y}_1, \hat{y}_2) is the occluded mix of the independent predictions y^1=S(x1)\hat{y}_1 = S(x_1) and y^2=S(x2)\hat{y}_2 = S(x_2).
    • q12=S(x12o)=S((1−1O)⊙x12m)q_{12} = S(x_{12}^o) = S\left( (1 - \mathbf{1}_O) \odot x_{12}^m \right) is the network prediction on the mixed and occluded input image.
    • p21p_{21} and q21q_{21} are the corresponding quantities for the transposed input order (x2,x1)(x_2, x_1).
    • Lncs(p,q)L_{\text{ncs}}(p, q) is the negative cosine similarity defined over vector flattened representations pp and qq:

    Lncs(p,q)=−p⋅q∥p∥2∥q∥2L_{\text{ncs}}(p, q) = -\frac{p \cdot q}{\|p\|_2 \|q\|_2}

  4. Knowl 4 — Local Connected-Component Consistency Regularization

    model/method

    Mixup transformations can introduce disconnected, fragmented artifacts into segmentation outputs, hindering the network from learning contiguous anatomical shape priors. CycleMix enforces a local connectivity prior based on the anatomical property that target cardiac structures are spatially interconnected.

    The local consistency loss Lcon-lL_{\text{con-l}} minimizes the distance between the unmixed network prediction y^=S(x)\hat{y} = S(x) and its post-processed version retaining only the largest connected component of each target structure:

    Lcon-l=12[Lncs(y^1,C(y^1))+Lncs(y^2,C(y^2))]L_{\text{con-l}} = \frac{1}{2} \left[ L_{\text{ncs}}(\hat{y}_1, C(\hat{y}_1)) + L_{\text{ncs}}(\hat{y}_2, C(\hat{y}_2)) \right]

    where:

    • C(⋅)C(\cdot) is a morphological filtering operator that isolates and outputs the largest connected component for each non-background category in the input prediction probability map.
    • Lncs(p,q)=−p⋅q∥p∥2∥q∥2L_{\text{ncs}}(p, q) = -\frac{p \cdot q}{\|p\|_2 \|q\|_2} is the negative cosine similarity.
  5. Knowl 5 — Experimental Setup and Benchmark Datasets for CycleMix

    experimental setup

    CycleMix is evaluated on two publicly available cardiac MRI segmentation datasets:

    1. ACDC (Automated Cardiac Diagnosis Challenge): 2D cine-MRI images from 100 subjects partitioned into 70 training, 15 validation, and 15 test subjects. For weakly-supervised benchmarking, the 70 training subjects are split into 35 subjects with scribble labels and 35 subjects with full mask annotations. Standard structures evaluated at end-diastolic (ED) and end-systolic (ES) phases are left ventricle (LV), right ventricle (RV), and myocardium (MYO).
    2. MSCMRseg: Late gadolinium enhancement (LGE) MRI images from 45 patients undergoing cardiomyopathy, divided into 25 training, 5 validation, and 20 test subjects. Scribble annotations provide average foreground coverage of 27.7%27.7\% for RV, 31.3%31.3\% for MYO, 24.1%24.1\% for LV, and 3.4%3.4\% for background.

    Implementation Details:

    • Backbone network: 2D UNet+ (implemented in PyTorch).
    • Image preprocessing: in-plane spatial resampling to 1.37×1.37 mm1.37 \times 1.37\text{ mm}, cropped or padded to 212×212212 \times 212 pixels, and intensity-normalized to zero mean and unit variance.
    • Training hyperparameters: learning rate fixed at 10−410^{-4}, trained for 1000 epochs on a single NVIDIA RTX 3090Ti GPU (24 GB).
    • Loss weighting hyperparameters: λ1=1.0,λ2=1.0,λ3=0.05,λ4=1.0\lambda_1 = 1.0, \lambda_2 = 1.0, \lambda_3 = 0.05, \lambda_4 = 1.0.
    • Occlusion mask size: random rotated rectangle of 32×3232 \times 32 pixels.
    • Evaluation metric: Dice similarity coefficient gauging predicted vs. ground-truth masks for LV, MYO, and RV.
  6. Knowl 6 — Comparison of CycleMix with Standard Mixup Strategies under Scribble Supervision

    data/table

    When trained purely on 35 scribble-annotated subjects, naive application of patch-transport mixup methods (Puzzle Mix, Co-mixup) severely degrades segmentation performance due to anatomical shape distortion. CycleMix resolves this degradation through its global and local consistency regularizations, outperforming all mixup baselines as well as fully supervised models trained on 35 full masks.

    Methods ACDC MSCMRseg
    LV MYO RV Avg LV MYO RV Avg
    35 scribbles
    UNetpce+\text{UNet}^+_{\text{pce}} .785±.196.785\pm.196 .725±.151.725\pm.151 .746±.203.746\pm.203 .752.752 .494±.082.494\pm.082 .583±.067.583\pm.067 .057±.022.057\pm.022 .378.378
    MixUp .803±.178.803\pm.178 .753±.116.753\pm.116 .767±.226.767\pm.226 .774.774 .610±.144.610\pm.144 .463±.147.463\pm.147 .378±.153.378\pm.153 .484.484
    Cutout .832±.172.832\pm.172 .754±.138.754\pm.138 .812±.129.812\pm.129 .800.800 .459±.077.459\pm.077 .641±.136.641\pm.136 .697±.149.697\pm.149 .599.599
    CutMix .641±.359.641\pm.359 .734±.144.734\pm.144 .740±.216.740\pm.216 .705.705 .578±.063.578\pm.063 .622±.121.622\pm.121 .761±.105.761\pm.105 .654.654
    Puzzle Mix .663±.333.663\pm.333 .650±.231.650\pm.231 .559±.343.559\pm.343 .624.624 .061±.021.061\pm.021 .634±.084.634\pm.084 .028±.012.028\pm.012 .241.241
    Co-mixup .622±.304.622\pm.304 .621±.214.621\pm.214 .702±.211.702\pm.211 .648.648 .356±.075.356\pm.075 .343±.067.343\pm.067 .053±.022.053\pm.022 .251.251
    CycleMix (ours) .883±.095\mathbf{.883\pm.095} .798±.075\mathbf{.798\pm.075} .863±.073\mathbf{.863\pm.073} .848\mathbf{.848} .870±.061\mathbf{.870\pm.061} .739±.049\mathbf{.739\pm.049} .791±.072\mathbf{.791\pm.072} .800\mathbf{.800}
    35 masks
    UNetF+\text{UNet}^+_F .849±.152.849\pm.152 .792±.140.792\pm.140 .817±.151.817\pm.151 .820.820 .857±.055.857\pm.055 .720±.075.720\pm.075 .689±.120.689\pm.120 .755.755
    Puzzle MixF\text{Puzzle Mix}_F .849±.182.849\pm.182 .807±.088.807\pm.088 .865±.089.865\pm.089 .840.840 .867±.042.867\pm.042 .742±.043.742\pm.043 .759±.039.759\pm.039 .789.789

    CycleMix achieves an average Dice score of 0.8480.848 on ACDC and 0.8000.800 on MSCMRseg, representing improvements of +22.4%+22.4\% and +55.9%+55.9\% over standard Puzzle Mix under scribble supervision, and outperforming the runner-up scribble baseline (CutMix) by +14.6%+14.6\% on MSCMRseg.

  7. Knowl 7 — Comparison of CycleMix with State-of-the-Art Weakly Supervised Methods on ACDC

    data/table

    CycleMix trained solely on 35 scribble-annotated cases outperforms prior weakly-supervised methods on the ACDC dataset, including methods that leverage 35 additional unpaired full segmentation masks to learn shape priors via GANs or autoencoders.

    Methods Data LV MYO RV Avg
    35 scribbles
    UNetpce\text{UNet}_{\text{pce}} scribbles .842.842 .764.764 .693.693 .766.766
    UNetwpce\text{UNet}_{\text{wpce}} scribbles .784.784 .675.675 .563.563 .674.674
    UNetCRF\text{UNet}_{\text{CRF}} scribbles .766.766 .661.661 .590.590 .672.672
    CycleMix (ours) scribbles .883\mathbf{.883} .798.798 .863\mathbf{.863} .848\mathbf{.848}
    35 scribbles + 35 unpaired masks
    UNetD\text{UNet}_D scribbles+masks .404.404 .597.597 .753.753 .585.585
    PostDAE scribbles+masks .806.806 .667.667 .556.556 .676.676
    ACCL scribbles+masks .878.878 .797.797 .735.735 .803.803
    MAAG scribbles+masks .879.879 .817\mathbf{.817} .752.752 .816.816

    CycleMix achieves an overall average Dice score of 84.8%84.8\%, exceeding MAAG (81.6%81.6\%) by +3.2%+3.2\%, and achieves an 11.1%11.1\% gain on the geometrically variable right ventricle structure (RV Dice: 86.3%86.3\% vs. 75.2%75.2\%).

  8. Knowl 8 — Ablation Study of CycleMix Components

    data/table

    An ablation study on the ACDC dataset demonstrates the incremental contribution of each component in CycleMix: unmixed scribble loss (LunmixL_{\text{unmix}}), mixed scribble loss (LmixL_{\text{mix}}), global consistency loss (Lcon-gL_{\text{con-g}}), random rectangular occlusion (1O\mathbf{1}_O), and local connectivity consistency loss (Lcon-lL_{\text{con-l}}).

    Model LunmixL_{\text{unmix}} LmixL_{\text{mix}} Lcon-gL_{\text{con-g}} 1O\mathbf{1}_O Lcon-lL_{\text{con-l}} LV MYO RV Avg
    #1 ✓ ×\times ×\times ×\times ×\times .785±.196.785\pm.196 .725±.151.725\pm.151 .746±.203.746\pm.203 .752.752
    #2 ✓ ✓ ×\times ×\times ×\times .863±.104∗.863\pm.104^* .783±.086∗.783\pm.086^* .782±.173.782\pm.173 .809∗.809^*
    #3 ✓ ✓ ✓ ×\times ×\times .867±.130.867\pm.130 .786±.114.786\pm.114 .837±.097∗.837\pm.097^* .830∗.830^*
    #4 ✓ ✓ ✓ ✓ ×\times .898±.059∗.898\pm.059^* .786±.078.786\pm.078 .847±.132∗.847\pm.132^* .843∗.843^*
    #5 ✓ ✓ ✓ ✓ ✓ .883±.095.883\pm.095 .798±.075∗.798\pm.075^* .863±.073.863\pm.073 .848.848

    Note: Asterisk (∗*) indicates statistically significant improvement over the prior row by Wilcoxon signed-rank test (p≤0.05p \le 0.05).

    Key takeaways:

    • Adding mixed image supervision (LmixL_{\text{mix}}) raises baseline Dice from 75.2%75.2\% to 80.9%80.9\% (+5.7%+5.7\%).
    • Adding global consistency (Lcon-gL_{\text{con-g}}) increases Dice to 83.0%83.0\% (+2.1%+2.1\%).
    • Adding random occlusion (1O\mathbf{1}_O) improves Dice to 84.3%84.3\% (+1.3%+1.3\%).
    • Adding local connected-component consistency (Lcon-lL_{\text{con-l}}) brings overall Dice to 84.8%84.8\%, yielding a statistically significant gain of +1.2%+1.2\% specifically on the myocardium (MYO: .798.798 vs. .786.786).
  9. Knowl 9 — Sensitivity of CycleMix to Annotation Ratios

    data/table

    Varying the ratio of scribble-annotated subjects to fully-annotated subjects across all 70 training subjects of ACDC reveals the label efficiency of CycleMix.

    Method Scribble : Full LV MYO RV Avg
    CycleMix 35 : 00 .883±.095.883\pm.095 .798±.075.798\pm.075 .863±.073.863\pm.073 .848.848
    CycleMix 70 : 00 .880±.115.880\pm.115 .825±.072.825\pm.072 .860±.089.860\pm.089 .855.855
    CycleMix 56 : 14 .898±.075.898\pm.075 .842±.072.842\pm.072 .876±.112.876\pm.112 .872.872
    CycleMix 42 : 28 .911±.063.911\pm.063 .854±.056.854\pm.056 .883±.076.883\pm.076 .883.883
    CycleMix 28 : 42 .902±.080.902\pm.080 .851±.065.851\pm.065 .899±.058.899\pm.058 .884.884
    CycleMix 14 : 56 .906±.065.906\pm.065 .856±.066.856\pm.066 .893±.083.893\pm.083 .885.885
    CycleMix 00 : 70 .919±.065.919\pm.065 .858±.058.858\pm.058 .882±.088.882\pm.088 .886.886
    UNetF+\text{UNet}^+_F 00 : 70 .883±.130.883\pm.130 .831±.093.831\pm.093 .870±.096.870\pm.096 .862.862

    With only 20%20\% full annotations (56:1456:14 ratio), CycleMix attains 87.2%87.2\% average Dice, outperforming the fully supervised baseline UNetF+\text{UNet}^+_F trained on 100%100\% full annotations (7070 cases, 86.2%86.2\% Dice). The performance of CycleMix saturates when full annotations reach 40%40\% (42:2842:28 ratio, 88.3%88.3\% Dice).

Coverage note — Omitted Table 5 comparing fully-supervised variants as it represents a minor derivative experiment beyond the paper's main focus on scribble-supervised learning.

References

  1. 1.Wenjia Bai, Hideaki Suzuki, Chen Qin, Giacomo Tarroni, Ozan Oktay, Paul M Matthews, and Daniel Rueckert. Recurrent neural networks for aortic image sequence segmentation with sparse annotations. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 586–594. Springer, 2018. 2
  2. 2.Christian F Baumgartner, Lisa M Koch, Marc Pollefeys, and Ender Konukoglu. An exploration of 2d and 3d deep learning techniques for cardiac mr image segmentation. In International Workshop on Statistical Atlases and Computational Models of the Heart, pages 111–119. Springer, 2017. 5, 6
  3. 3.Olivier Bernard, Alain Lalande, Clement Zotti, Frederick Cervenansky, Xin Yang, Pheng-Ann Heng, Irem Cetin, Karim Lekadir, Oscar Camara, Miguel Angel Gonzalez Ballester, Gerard Sanroma, Sandy Napel, Steffen Petersen, Georgios Tziritas, Elias Grinias, Mahendra Khened, Varghese Alex Kollerathu, Ganapathy Krishnamurthi, Marc-Michel Rohe, Xavier Pennec, Maxime Sermesant, Fabian Isensee, Paul Jager, Klaus H. Maier-Hein, Peter M. Full, Ivo Wolf, Sandy Engelhardt, Christian F. Baumgartner, Lisa M. Koch, Jelmer M. Wolterink, Ivana Isgum, Yeonggul Jang, Yoonmi Hong, Jay Patravali, Shubham Jain, Olivier Humbert, and Pierre-Marc Jodoin. Deep learning techniques for automatic mri cardiac multi-structures segmentation and diagnosis: Is the problem solved? IEEE Transactions on Medical Imaging, 37(11):2514–2525, 2018. 5
  4. 4.Yigit Baran Can, Krishna Chaitanya, Basil Mustafa, Lisa M. Koch, Ender Konukoglu, and Christian F. Baumgartner. Learning to segment medical images with scribble-supervision alone. In DLMIA/ML-CDS@MICCAI, 2018. 1, 2
  5. 5.Krishna Chaitanya, Neerav Karani, Christian F Baumgartner, Anton Becker, Olivio Donati, and Ender Konukoglu. Semi-supervised and task-driven data augmentation. In International conference on information processing in medical imaging, pages 29–41. Springer, 2019. 2
  6. 6.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4):834–848, 2017. 2
  7. 7.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. arXiv preprint arXiv:2011.10566, 2020. 4
  8. 8.Terrance DeVries and Graham W Taylor. Improved regularization of convolutional neural networks with cutout. arXiv preprint arXiv:1708.04552, 2017. 2, 7
  9. 9.Lee R. Dice. Measures of the amount of ecologic association between species. Ecology, 26(3):297–302, 1945. 5
  10. 10.Jean-Bastien Grill, Florian Strub, Florent Altche, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Remi Munos, and Michal Valko. Bootstrap your own latent: A new approach to self-supervised learning, 2020. 4
  11. 11.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. 2
  12. 12.Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1125–1134, 2017. 2
  13. 13.Zhanghexuan Ji, Yan Shen, Chunwei Ma, and Mingchen Gao. Scribble-based hierarchical weakly supervised learning for brain tumor segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 175–183. Springer, 2019. 1, 2
  14. 14.Anna Khoreva, Rodrigo Benenson, Jan Hosang, Matthias Hein, and Bernt Schiele. Simple does it: Weakly supervised instance and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 876–885, 2017. 1
  15. 15.JangHyun Kim, Wonho Choo, Hosan Jeong, and Hyun Oh Song. Co-mixup: Saliency guided joint mixup with supermodular diversity. In International Conference on Learning Representations, 2021. 2, 3, 5, 7
  16. 16.Jang-Hyun Kim, Wonho Choo, and Hyun Oh Song. Puzzle mix: Exploiting saliency and local statistics for optimal mixup. In International Conference on Machine Learning (ICML), 2020. 2, 3, 5, 7, 8
  17. 17.Samuli Laine and Timo Aila. Temporal ensembling for semi-supervised learning. arXiv preprint arXiv:1610.02242, 2016. 3
  18. 18.Agostina J. Larrazabal, C’esar Mart’inez, Ben Glocker, and Enzo Ferrante. Post-dae: Anatomically plausible segmentation via post-processing with denoising autoencoders. IEEE Transactions on Medical Imaging, 39:3813–3820, 2020. 2, 6, 7
  19. 19.Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun. Scribblesup: Scribble-supervised convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3159–3167, 2016. 1, 2
  20. 20.Sudhanshu Mittal, Maxim Tatarchenko, and Thomas Brox. Semi-supervised semantic segmentation with high-and low-level consistency. IEEE transactions on pattern analysis and machine intelligence, 2019. 1
  21. 21.Yassine Ouali, Celine Hudelot, and Myriam Tami. Semi-supervised semantic segmentation with cross-consistency training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12674–12684, 2020. 3
  22. 22.Deepak Pathak, Philipp Krahenbuhl, and Trevor Darrell. Constrained convolutional neural networks for weakly supervised segmentation. In Proceedings of the IEEE international conference on computer vision, pages 1796–1804, 2015. 1
  23. 23.Nasim Souly, Concetto Spampinato, and Mubarak Shah. Semi supervised semantic segmentation using generative adversarial network. In Proceedings of the IEEE international conference on computer vision, pages 5688–5696, 2017. 1
  24. 24.Nima Tajbakhsh, Laura Jeyaseelan, Qian Li, Jeffrey N Chiang, Zhihao Wu, and Xiaowei Ding. Embracing imperfect datasets: A review of deep learning solutions for medical image segmentation. Medical Image Analysis, 63:101693, 2020. 1, 2
  25. 25.Meng Tang, Abdelaziz Djelouah, Federico Perazzi, Yuri Boykov, and Christopher Schroers. Normalized cut loss for weakly-supervised cnn segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1818–1827, 2018. 5, 7
  26. 26.Meng Tang, Federico Perazzi, Abdelaziz Djelouah, Ismail Ben Ayed, Christopher Schroers, and Yuri Boykov. On regularized losses for weakly-supervised cnn segmentation. In ECCV, 2018. 2, 5
  27. 27.Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. 1, 3
  28. 28.Gabriele Valvano, Andrea Leo, and Sotirios A. Tsaftaris. Learning to segment from scribbles using multi-scale adversarial attention gates. IEEE Transactions on Medical Imaging, pages 1–1, 2021. 2, 5, 6, 7
  29. 29.Vikas Verma, Alex Lamb, Christopher Beckham, Amir Najafi, Ioannis Mitliagkas, David Lopez-Paz, and Yoshua Bengio. Manifold mixup: Better representations by interpolating hidden states. In International Conference on Machine Learning, pages 6438–6447. PMLR, 2019. 2
  30. 30.Dong Wang, Yuan Zhang, Kexin Zhang, and Liwei Wang. Focalmix: Semi-supervised learning for 3d medical image detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3951–3960, 2020. 2
  31. 31.Yunchao Wei, Xiaodan Liang, Yunpeng Chen, Xiaohui Shen, Ming-Ming Cheng, Jiashi Feng, Yao Zhao, and Shuicheng Yan. Stc: A simple to complex framework for weakly-supervised semantic segmentation. IEEE transactions on pattern analysis and machine intelligence, 2016. 1
  32. 32.Qian Yue, Xinzhe Luo, Qing Ye, Lingchao Xu, and Xiahai Zhuang. Cardiac segmentation from lge mri using deep neural network incorporating shape and spatial priors. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 559–567. Springer, 2019. 5
  33. 33.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In International Conference on Computer Vision (ICCV), 2019. 2, 3, 4, 5, 7
  34. 34.Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. International Conference on Learning Representations, 2018. 2, 3, 5, 7
  35. 35.Pengyi Zhang, Yunxin Zhong, and Xiaoqiong Li. Accl: Adversarial constrained-cnn loss for weakly supervised medical image segmentation, 2020. 2, 6, 7
  36. 36.Shuai Zheng, Sadeep Jayasumana, Bernardino Romera-Paredes, Vibhav Vineet, Zhizhong Su, Dalong Du, Chang Huang, and Philip HS Torr. Conditional random fields as recurrent neural networks. In Proceedings of the IEEE international conference on computer vision, pages 1529–1537, 2015. 2, 5, 7
  37. 37.Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE international conference on computer vision, pages 2223–2232, 2017. 3
  38. 38.Xiahai Zhuang. Multivariate mixture model for cardiac segmentation from multi-sequence mri. In MICCAI, 2016. 5
  39. 39.Xiahai Zhuang. Multivariate mixture model for myocardial segmentation combining multi-source images. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(12):2933–2946, 2019. 5

Citation

MLA
Zhang, K., and X. Zhuang. “CycleMix: A Holistic Strategy for Medical Image Segmentation from Scribble Supervision”. arXiv, 2022, http://arxiv.org/abs/2203.01475v2.
APA
Zhang, K., & Zhuang, X. (2022). CycleMix: A Holistic Strategy for Medical Image Segmentation from Scribble Supervision. arXiv. http://arxiv.org/abs/2203.01475v2
Chicago
Zhang, K., and X. Zhuang. 2022. “CycleMix: A Holistic Strategy for Medical Image Segmentation from Scribble Supervision”. arXiv. http://arxiv.org/abs/2203.01475v2.
Harvard
Zhang, K. and Zhuang, X. (2022) “CycleMix: A Holistic Strategy for Medical Image Segmentation from Scribble Supervision”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.01475v2.
Vancouver
1. Zhang K, Zhuang X (2022) CycleMix: A Holistic Strategy for Medical Image Segmentation from Scribble Supervision. arXiv

BibTeX

@article{zhang2022cyclemix,
  title = {CycleMix: A Holistic Strategy for Medical Image Segmentation from Scribble Supervision},
  author = {Zhang, Ke and Zhuang, Xiahai},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.01475v2},
  eprint = {2203.01475}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE