Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation

Yi LuoNima Mesgarani

article2018IEEE/ACM Transactions on Audio Speech and Language Processing2,237 citations

Presents Conv-TasNet, a fully convolutional end-to-end time-domain speech separation network that outperforms ideal time-frequency magnitude masks while achieving low latency and a compact model size suitable for real-time processing.

Listen

Automatic speech separation in single-microphone devices is essential for technologies such as telecommunications, hearing aids, and voice recognition systems. Historically, most methods have converted audio into time-frequency representations (spectrograms) to estimate individual voices. However, this traditional approach decouples signal magnitude from phase, relies on signal transformations that are not optimized for separation, and requires long time windows that introduce substantial processing delays. Consequently, existing models struggle to deliver the accuracy, low latency, and computational efficiency required for real-time applications.

The article evaluates a fully convolutional time-domain audio separation network (Conv-TasNet) to determine whether end-to-end processing directly on raw speech waveforms can improve separation accuracy while reducing latency and model size.

To assess the system, the authors conducted experiments using standard two-speaker and three-speaker benchmarks derived from the Wall Street Journal speech dataset, testing on unseen speakers across 40 hours of training/validation data and 5 hours of test data. The architecture replaces standard frequency transforms with a linear convolutional encoder-decoder and swaps recurrent networks for stacked dilated one-dimensional convolutions using depthwise separable operations. Performance was measured using objective distortion improvements (scale-invariant signal-to-noise ratio and signal-to-distortion ratio), automated perceptual scores, processing time per frame across central and graphics processing units, and subjective listening evaluations with 40 human participants.

The analysis yielded several key findings. First, the non-causal version of the proposed system achieved an accuracy improvement of 15.3 dB in scale-invariant signal-to-noise ratio and 15.6 dB in signal-to-distortion ratio for two-speaker mixtures, significantly outperforming prior models and surpassing theoretical ideal time-frequency magnitude masks. Second, in subjective human listening tests, the system scored an average quality rating of 4.03 out of 5, significantly outperforming the best ideal ratio mask (3.51). Third, the model reduced parameter count to 5.1 million—roughly four to eighteen times smaller than leading alternative networks—while lowering frame processing latency to 2 milliseconds. Fourth, the real-time causal configuration maintained stable separation regardless of shifting mixture start times and processed frames in 0.4 milliseconds on a central processor, roughly five times faster than the 2-millisecond frame duration.

These findings demonstrate that directly modeling audio in the time domain resolves the phase-reconstruction bottleneck that has long constrained speech separation. By significantly lowering computing overhead and latency while improving audio clarity, this architecture makes high-performance speech separation practical on power- and compute-constrained edge hardware, such as smart assistants and wearable hearing devices.

Organizations developing speech processing products should consider adopting time-domain convolutional architectures for single-channel front-ends. When choosing configurations, decision-makers must weigh the trade-off between maximum separation quality in non-causal settings (15.3 dB improvement) and the lower but real-time-capable performance of causal setups (10.6 dB improvement). Further development and pilot testing are recommended to evaluate system robustness in complex, reverberant acoustic environments and to explore multi-microphone expansions.

Confidence in the reported benchmarks is high under controlled clean acoustic conditions with up to three speakers. However, stakeholders should note that the current evaluation assumes isolated clean speech mixtures; performance in real-world scenarios with heavy background noise, room reverberation, and prolonged conversational pauses remains to be verified.

Cover for Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation

Abstract

Single-channel, speaker-independent speech separation methods have recently seen great progress. However, the accuracy, latency, and computational cost of such methods remain insufficient. The majority of the previous methods have formulated the separation problem through the time-frequency representation of the mixed signal, which has several drawbacks, including the decoupling of the phase and magnitude of the signal, the suboptimality of time-frequency representation for speech separation, and the long latency in calculating the spectrograms. To address these shortcomings, we propose a fully-convolutional time-domain audio separation network (Conv-TasNet), a deep learning framework for end-to-end time-domain speech separation. Conv-TasNet uses a linear encoder to generate a representation of the speech waveform optimized for separating individual speakers. Speaker separation is achieved by applying a set of weighting functions (masks) to the encoder output. The modified encoder representations are then inverted back to the waveforms using a linear decoder. The masks are found using a temporal convolutional network (TCN) consisting of stacked 1-D dilated convolutional blocks, which allows the network to model the long-term dependencies of the speech signal while maintaining a small model size. The proposed Conv-TasNet system significantly outperforms previous time-frequency masking methods in separating two- and three-speaker mixtures. Additionally, Conv-TasNet surpasses several ideal time-frequency magnitude masks in two-speaker speech separation as evaluated by both objective distortion measures and subjective quality assessment by human listeners. Finally, Conv-TasNet has a significantly smaller model size and a shorter minimum latency, making it a suitable solution for both offline and real-time speech separation applications.

Table of Contents

  • I Introduction
  • II Convolutional Time-domain Audio Separation Network
  • II-A Time-domain speech separation
  • II-B Convolutional encoder-decoder
  • II-C Estimating the separation masks
  • II-D Convolutional separation module
  • III Experimental procedures
  • III-A Dataset
  • III-B Experiment configurations
  • III-C Training objective
  • III-D Evaluation metrics
  • III-E Comparison with ideal time-frequency masks
  • IV Results
  • IV-A Non-negativity of the encoder output
  • IV-B Optimizing the network parameters
  • IV-C Comparison of Conv-TasNet with previous methods
  • IV-D Subjective and objective quality evaluation of Conv-TasNet
  • IV-E Processing speed comparison
  • IV-F Sensitivity of LSTM-TasNet to the mixture starting point
  • IV-G Properties of the basis functions
  • V Discussion
  • VI Acknowledgments
  • References

Knowls

  1. Knowl 1 — Conv-TasNet Architecture for End-to-End Time-Domain Speech Separation

    model/method

    The Fully-Convolutional Time-Domain Audio Separation Network (Conv-TasNet) performs single-channel speech separation directly on raw waveforms without computing a short-time Fourier transform (STFT). Given a mixture waveform x(t)=∑i=1Csi(t)∈R1×Tx(t) = \sum_{i=1}^C s_i(t) \in \mathbb{R}^{1 \times T} composed of CC clean sources si(t)s_i(t), the system processes audio via three consecutive modules:

    1. Encoder: A 1-D convolutional layer with NN filters of length LL samples and a stride of L/2L/2 samples (50% overlap). It transforms an input segment xk∈R1×L\mathbf{x}_k \in \mathbb{R}^{1 \times L} into an NN-dimensional feature representation w∈R1×N\mathbf{w} \in \mathbb{R}^{1 \times N} via w=H(xkU)\mathbf{w} = \mathcal{H}(\mathbf{x}_k \mathbf{U}), where U∈RN×L\mathbf{U} \in \mathbb{R}^{N \times L} denotes the encoder basis functions and H(⋅)\mathcal{H}(\cdot) is an optional activation (omitted in the standard linear formulation).

    2. Separation Module (TCN): A temporal convolutional network operating on normalized encoder representations to estimate source masks mi∈[0,1]1×N\mathbf{m}_i \in [0, 1]^{1 \times N} for each source i∈{1,…,C}i \in \{1, \dots, C\}. The masked feature representation for source ii is obtained by element-wise multiplication:

    di=w⊙mi\mathbf{d}_i = \mathbf{w} \odot \mathbf{m}_i

    1. Decoder: A 1-D transposed convolutional layer with basis functions V∈RN×L\mathbf{V} \in \mathbb{R}^{N \times L} that reconstructs each estimated source segment as s^i=diV\hat{\mathbf{s}}_i = \mathbf{d}_i \mathbf{V}. Overlapping reconstructed segments are combined via overlap-add to generate the final time-domain estimated waveforms s^i(t)\hat{s}_i(t).
  2. Knowl 2 — Stacked 1-D Dilated Depthwise Separable Convolutional Block Design

    model/method

    The separation network in Conv-TasNet is parameterized as a Temporal Convolutional Network (TCN) composed of RR repeats of MM stacked 1-D convolutional blocks with exponentially increasing dilation factors d∈{1,2,4,…,2M−1}d \in \{1, 2, 4, \dots, 2^{M-1}\}. This exponential dilation structure provides a wide receptive field over time without requiring recurrent layers.

    To minimize model size and computational complexity, each block utilizes depthwise separable convolutions (S-conv) consisting of:

    1. A 1×11 \times 1 pointwise convolution transforming the input from bottleneck dimension BB to hidden dimension HH, followed by a Parametric Rectified Linear Unit (PReLU) activation and a normalization layer (global or cumulative layer normalization).
    2. A depthwise convolution (D-conv) of kernel size PP with dilation factor 2m2^m, convolving each feature channel independently. For an input Y∈RG×M\mathbf{Y} \in \mathbb{R}^{G \times M} and kernel K∈RG×P\mathbf{K} \in \mathbb{R}^{G \times P}:

    D-conv(Y,K)=concat(yj⊛kj),j=1,…,G\text{D-conv}(\mathbf{Y}, \mathbf{K}) = \text{concat}(\mathbf{y}_j \circledast \mathbf{k}_j), \quad j = 1, \dots, G

    where ⊛\circledast denotes standard 1-D convolution. This decreases parameter count relative to a standard convolution by a factor of H×PH+P≈P\frac{H \times P}{H + P} \approx P when H≫PH \gg P. 3. A second PReLU activation and normalization layer. 4. Two parallel linear 1×11 \times 1 convolutions: one serving as the residual output (projecting back to BB channels and added to the block input) and one serving as the skip-connection output (projecting to ScS_c channels).

    The skip-connection outputs from all blocks are summed together and passed to an output 1×11 \times 1 convolution followed by a Sigmoid or Softmax activation to estimate the CC source masks m1,…,mC\mathbf{m}_1, \dots, \mathbf{m}_C.

  3. Knowl 3 — Global and Cumulative Layer Normalization Formulations

    equation

    Conv-TasNet uses two distinct feature normalization techniques within its convolutional blocks depending on the causality requirements of the system:

    Global Layer Normalization (gLN) for non-causal systems normalizes features jointly across both channel dimension NN and temporal dimension TT of a feature tensor F∈RN×T\mathbf{F} \in \mathbb{R}^{N \times T}:

    gLN(F)=F−E[F]Var[F]+ϵ⊙γ+β\text{gLN}(\mathbf{F}) = \frac{\mathbf{F} - \mathbb{E}[\mathbf{F}]}{\sqrt{\text{Var}[\mathbf{F}] + \epsilon}} \odot \boldsymbol{\gamma} + \boldsymbol{\beta}

    where γ,β∈RN×1\boldsymbol{\gamma}, \boldsymbol{\beta} \in \mathbb{R}^{N \times 1} are learnable scaling and bias parameters, ϵ>0\epsilon > 0 is a stability constant, and expectations are calculated over all N×TN \times T elements:

    E[F]=1NT∑n=1N∑t=1TFn,t,Var[F]=1NT∑n=1N∑t=1T(Fn,t−E[F])2\mathbb{E}[\mathbf{F}] = \frac{1}{NT} \sum_{n=1}^N \sum_{t=1}^T \mathbf{F}_{n,t}, \qquad \text{Var}[\mathbf{F}] = \frac{1}{NT} \sum_{n=1}^N \sum_{t=1}^T (\mathbf{F}_{n,t} - \mathbb{E}[\mathbf{F}])^2

    Cumulative Layer Normalization (cLN) for causal real-time systems computes running statistics strictly using frames up to the current frame index kk, avoiding access to future frames. For frame fk∈RN×1\mathbf{f}_k \in \mathbb{R}^{N \times 1} given historical frames ft≤k=[f1,…,fk]∈RN×k\mathbf{f}_{t \le k} = [\mathbf{f}_1, \dots, \mathbf{f}_k] \in \mathbb{R}^{N \times k}:

    cLN(fk)=fk−E[ft≤k]Var[ft≤k]+ϵ⊙γ+β\text{cLN}(\mathbf{f}_k) = \frac{\mathbf{f}_k - \mathbb{E}[\mathbf{f}_{t \le k}]}{\sqrt{\text{Var}[\mathbf{f}_{t \le k}] + \epsilon}} \odot \boldsymbol{\gamma} + \boldsymbol{\beta}

    where:

    E[ft≤k]=1Nk∑n=1N∑t=1kfn,t,Var[ft≤k]=1Nk∑n=1N∑t=1k(fn,t−E[ft≤k])2\mathbb{E}[\mathbf{f}_{t \le k}] = \frac{1}{Nk} \sum_{n=1}^N \sum_{t=1}^k \mathbf{f}_{n,t}, \qquad \text{Var}[\mathbf{f}_{t \le k}] = \frac{1}{Nk} \sum_{n=1}^N \sum_{t=1}^k (\mathbf{f}_{n,t} - \mathbb{E}[\mathbf{f}_{t \le k}])^2

  4. Knowl 4 — Scale-Invariant Signal-to-Noise Ratio (SI-SNR) Optimization Objective

    equation

    Conv-TasNet is trained end-to-end to maximize the scale-invariant source-to-noise ratio (SI-SNR). For an estimated time-domain source waveform s^∈R1×T\hat{\mathbf{s}} \in \mathbb{R}^{1 \times T} and ground-truth clean waveform s∈R1×T\mathbf{s} \in \mathbb{R}^{1 \times T}, both signals are first zero-mean normalized. The target signal component is defined by the orthogonal projection of s^\hat{\mathbf{s}} onto s\mathbf{s}:

    starget:=⟨s^,s⟩∥s∥2s\mathbf{s}_{\text{target}} := \frac{\langle \hat{\mathbf{s}}, \mathbf{s} \rangle}{\|\mathbf{s}\|^2} \mathbf{s}

    enoise:=s^−starget\mathbf{e}_{\text{noise}} := \hat{\mathbf{s}} - \mathbf{s}_{\text{target}}

    SI-SNR:=10log⁡10∥starget∥2∥enoise∥2\text{SI-SNR} := 10 \log_{10} \frac{\|\mathbf{s}_{\text{target}}\|^2}{\|\mathbf{e}_{\text{noise}}\|^2}

    where ⟨s^,s⟩=s^sT\langle \hat{\mathbf{s}}, \mathbf{s} \rangle = \hat{\mathbf{s}} \mathbf{s}^T is the inner product and ∥s∥2=⟨s,s⟩\|\mathbf{s}\|^2 = \langle \mathbf{s}, \mathbf{s} \rangle is signal power. During training, utterance-level Permutation Invariant Training (uPIT) is applied over all source permutations to find the optimal assignment between estimated and target speakers.

  5. Knowl 5 — Two-Speaker Speech Separation Performance on WSJ0-2mix

    data/table

    Conv-TasNet was evaluated on the WSJ0-2mix benchmark (8 kHz audio, mixture SNRs between -5 dB and 5 dB) against time-frequency masking approaches, previous time-domain methods, and theoretical ideal time-frequency magnitude masks (Ideal Binary Mask, IBM; Ideal Ratio Mask, IRM; Wiener Filter-like Mask, WFM). Performance was measured in Scale-Invariant Signal-to-Noise Ratio improvement (SI-SNRi) and Signal-to-Distortion Ratio improvement (SDRi) in dB.

    Method Model size Causal SI-SNRi (dB) SDRi (dB)
    DPCL++ 13.6M ×\times 10.8 –
    uPIT-BLSTM-ST 92.7M ×\times – 10.0
    DANet 9.1M ×\times 10.5 –
    ADANet 9.1M ×\times 10.4 10.8
    cuPIT-Grid-RD 47.2M ×\times – 10.2
    CBLDNN-GAT 39.5M ×\times – 11.0
    Chimera++ 32.9M ×\times 11.5 12.0
    WA-MISI-5 32.9M ×\times 12.6 13.1
    BLSTM-TasNet 23.6M ×\times 13.2 13.6
    Conv-TasNet-gLN 5.1M ×\times 15.3 15.6
    uPIT-LSTM 46.3M ✓ – 7.0
    LSTM-TasNet 32.0M ✓ 10.8 11.2
    Conv-TasNet-cLN 5.1M ✓ 10.6 11.0
    IRM – – 12.2 12.6
    IBM – – 13.0 13.5
    WFM – – 13.4 13.8

    Non-causal Conv-TasNet-gLN achieves 15.3 dB SI-SNRi and 15.6 dB SDRi with 5.1M parameters, surpassing all previous STFT- and time-domain methods as well as all ideal T-F magnitude masks (WFM at 13.8 dB SDRi, IBM at 13.5 dB, IRM at 12.6 dB) while using significantly fewer parameters.

  6. Knowl 6 — Three-Speaker Speech Separation Performance on WSJ0-3mix

    data/table

    Performance of Conv-TasNet on the three-speaker WSJ0-3mix benchmark compared with existing time-frequency models and ideal time-frequency masks:

    Method Model size Causal SI-SNRi (dB) SDRi (dB)
    DPCL++ 13.6M ×\times 7.1 –
    uPIT-BLSTM-ST 92.7M ×\times – 7.7
    DANet 9.1M ×\times 8.6 8.9
    ADANet 9.1M ×\times 9.1 9.4
    Conv-TasNet-gLN 5.1M ×\times 12.7 13.1
    Conv-TasNet-cLN 5.1M ✓ 7.8 8.2
    IRM – – 12.5 13.0
    IBM – – 13.2 13.6
    WFM – – 13.6 14.0

    Non-causal Conv-TasNet-gLN achieves 12.7 dB SI-SNRi and 13.1 dB SDRi, outperforming prior STFT methods by at least 3.7 dB SDRi and matching/exceeding the performance of the Ideal Ratio Mask (13.0 dB SDRi). The causal variant Conv-TasNet-cLN attains 8.2 dB SDRi, exceeding the non-causal uPIT-BLSTM-ST baseline (7.7 dB).

  7. Knowl 7 — Impact of Encoder-Decoder Linearization and Mask Function Choice

    empirical result

    Ablation of encoder-decoder configurations demonstrated that enforcing an explicit autoencoder reconstruction or non-negativity constraint is unnecessary for optimal separation:

    1. Pseudo-inverse Autoencoder: Enforcing V=(UTU)−1UT\mathbf{V} = (\mathbf{U}^T \mathbf{U})^{-1}\mathbf{U}^T and a unit sum constraint ∑i=1Cmi=1\sum_{i=1}^C \mathbf{m}_i = 1 with Softmax resulted in the lowest performance (12.1 dB SI-SNRi / 12.4 dB SDRi).
    2. ReLU Non-Linearity: Applying w=ReLU(xU)\mathbf{w} = \text{ReLU}(\mathbf{x}\mathbf{U}) to enforce non-negative encoder outputs reached 13.0 dB SI-SNRi (Softmax) and 12.9 dB SI-SNRi (Sigmoid).
    3. Unconstrained Linear Encoder-Decoder: A completely linear encoder w=xU\mathbf{w} = \mathbf{x}\mathbf{U} and linear decoder x^=wV\hat{\mathbf{x}} = \mathbf{w}\mathbf{V} paired with a Sigmoid mask estimation function achieved the best performance (13.1 dB SI-SNRi / 13.4 dB SDRi in the baseline comparison setup).

    When the encoder representation is sufficiently overcomplete (N>LN > L), a set of non-negative masks can reconstruct individual clean sources from an unbounded linear representation without requiring an explicit reconstruction constraint on the mixture signal.

  8. Knowl 8 — Discrepancy Between Perceptual Evaluation (PESQ) and Human Subjective Quality (MOS)

    empirical result

    Objective perceptual scoring (PESQ) and subjective listening tests (Mean Opinion Score, MOS, rated 1 to 5 by N=40N=40 normal-hearing listeners on 25 test utterances from WSJ0-2mix) show that PESQ systematically underestimates the quality of time-domain separated speech:

    • PESQ Scores: On WSJ0-2mix, IRM scored 3.74, IBM scored 3.33, WFM scored 3.70, while Conv-TasNet scored 3.24. On the 25 selected samples, IRM scored 3.74 vs Conv-TasNet at 3.22 (and clean reference at 4.50).
    • Subjective Human MOS: Conv-TasNet-gLN achieved an MOS of 4.034.03, which is significantly higher than the IRM MOS of 3.513.51 (p<10−16p < 10^{-16}, two-tailed tt-test), approaching the clean speech rating of 4.234.23.

    Because PESQ relies on STFT magnitude spectrogram comparisons and is sensitive to phase distortions, it underestimates time-domain methods like Conv-TasNet that reconstruct audio without strict adherence to STFT magnitude-phase decoupling.

  9. Knowl 9 — Inference Latency and Real-Time Frame Processing Speed

    data/table

    The execution speed of causal Conv-TasNet-cLN was compared to causal LSTM-TasNet by measuring the Time Per Frame (TPF) required to separate a single segment on a CPU (Intel Core i7-5820K) and GPU (Nvidia Titan Xp):

    Method Frame Length (ms) CPU / GPU TPF (ms)
    LSTM-TasNet 5.0 4.3 / 0.2
    Conv-TasNet-cLN 2.0 0.4 / 0.02

    For real-time implementation, TPF must be less than the frame length. While LSTM-TasNet requires sequential frame processing (CPU TPF of 4.3 ms against a 5.0 ms window), causal Conv-TasNet decouples frame-level sequential dependencies, allowing parallel convolution operations that achieve a CPU TPF of 0.4 ms on 2.0 ms frames (5 times faster than real-time) and a GPU TPF of 0.02 ms.

  10. Knowl 10 — Invariance of Causal Conv-TasNet to Input Starting Point Shifts

    empirical result

    Evaluating speech separation accuracy on the WSJ0-2mix test set with arbitrary sample shifts ss (starting the separation algorithm at sample index ss instead of sample 0) revealed that:

    1. LSTM-TasNet is highly sensitive to the initial sample offset; shifting the starting point by a few samples causes severe oscillations in SDRi due to accumulated recurrent state errors across sequential frames.
    2. Causal Conv-TasNet demonstrates near-constant SDRi across all sample shifts, with substantially lower standard deviation in separation performance across shift variations on the WSJ0-2mix test set.

    This insensitivity occurs because dilated convolutions process local frame contexts independently rather than chaining a continuous recurrent hidden state, preventing single-frame estimation errors from corrupting downstream frames.

  11. Knowl 11 — Tuning Properties of Learned Encoder-Decoder Basis Functions

    empirical result

    Spectral and temporal analysis of the N=512N=512 learned basis vectors in the linear encoder (matrix U\mathbf{U}) and decoder (matrix V\mathbf{V}) of the best-performing Conv-TasNet reveals two distinct physiological and acoustic properties:

    1. Low-Frequency Dominance: More than 60% of the learned filters have center frequencies tuned below 1 kHz, mirroring the non-linear human mel scale and tonotopic organization of mammalian auditory cortex to capture fundamental frequency and pitch cues.
    2. Phase Encoding: Filters tuned to identical center frequencies exhibit varying phase shifts (observed as circularly shifted basis functions in the time domain). This allows Conv-TasNet to explicitly model waveform phase diversity directly, eliminating the phase estimation errors inherent in STFT magnitude masking.
  12. Knowl 12 — Limitations in Context Tracking and Reverberant Acoustic Conditions

    limitation

    The Conv-TasNet framework has two primary documented constraints:

    1. Speaker Tracking Across Pauses: Because Conv-TasNet relies on a finite receptive field determined by its dilation factors (e.g., 1.53 to 3.83 seconds), tracking the identity of individual speakers across long conversational pauses in continuous speech can fail.
    2. Vulnerability to Reverberation: Direct time-domain modeling is sensitive to temporal smear and reflections found in highly reverberant environments, requiring multi-channel extensions (e.g., microphone array inputs) to handle spatial filtering and speaker counts greater than three under noisy, reverberant conditions.

Coverage note — None was omitted. All structural components, mathematical formulations, empirical ablation results, WSJ0-2mix/3mix benchmarks, subjective listening tests, latency evaluations, filter analysis, and stated limitations were fully covered.

References

  1. 1.D. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2018.
  2. 2.X. Lu, Y. Tsao, S. Matsuda, and C. Hori, “Speech enhancement based on deep denoising autoencoder.” in Interspeech, 2013, pp. 436–440.
  3. 3.Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “An experimental study on speech enhancement based on deep neural networks,” IEEE Signal processing letters, vol. 21, no. 1, pp. 65–68, 2014.
  4. 4.——, “A regression approach to speech enhancement based on deep neural networks,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), vol. 23, no. 1, pp. 7–19, 2015.
  5. 5.Y. Isik, J. Le Roux, Z. Chen, S. Watanabe, and J. R. Hershey, “Single-channel multi-speaker separation using deep clustering,” Interspeech 2016, pp. 545–549, 2016.
  6. 6.D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on. IEEE, 2017, pp. 241–245.
  7. 7.M. Kolbæk, D. Yu, Z.-H. Tan, and J. Jensen, “Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 25, no. 10, pp. 1901–1913, 2017.
  8. 8.Z. Chen, Y. Luo, and N. Mesgarani, “Deep attractor network for single-microphone speaker separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on. IEEE, 2017, pp. 246–250.
  9. 9.Y. Luo, Z. Chen, and N. Mesgarani, “Speaker-independent speech separation with deep attractor network,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 26, no. 4, pp. 787–796, 2018. [Online]. Available: http://dx.doi.org/10.1109/TASLP.2018.2795749
  10. 10.Z.-Q. Wang, J. Le Roux, and J. R. Hershey, “Alternative objective functions for deep clustering,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018.
  11. 11.Z.-Q. Wang, J. L. Roux, D. Wang, and J. R. Hershey, “End-to-end speech separation with unfolded iterative phase reconstruction,” arXiv preprint arXiv:1804.10204, 2018.
  12. 12.C. Li, L. Zhu, S. Xu, P. Gao, and B. Xu, “CBLDNN-based speaker-independent speech separation via generative adversarial training,” in Acoustics, Speech and Signal Processing (ICASSP), 2018 IEEE International Conference on. IEEE, 2018.
  13. 13.D. Griffin and J. Lim, “Signal estimation from modified short-time fourier transform,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 32, no. 2, pp. 236–243, 1984.
  14. 14.J. Le Roux, N. Ono, and S. Sagayama, “Explicit consistency constraints for stft spectrograms and their application to phase reconstruction.” in SAPA@ INTERSPEECH, 2008, pp. 23–28.
  15. 15.Y. Luo, Z. Chen, J. R. Hershey, J. Le Roux, and N. Mesgarani, “Deep clustering and conventional networks for music separation: Stronger together,” in Acoustics, Speech and Signal Processing (ICASSP), 2017 IEEE International Conference on. IEEE, 2017, pp. 61–65.
  16. 16.A. Jansson, E. Humphrey, N. Montecchio, R. Bittner, A. Kumar, and T. Weyde, “Singing voice separation with deep u-net convolutional networks,” in 18th International Society for Music Information Retrieval Conference, 2017, pp. 23–27.
  17. 17.S. Choi, A. Cichocki, H.-M. Park, and S.-Y. Lee, “Blind source separation and independent component analysis: A review,” Neural Information Processing-Letters and Reviews, vol. 6, no. 1, pp. 1–57, 2005.
  18. 18.K. Yoshii, R. Tomioka, D. Mochihashi, and M. Goto, “Beyond nmf: Time-domain audio source separation without phase reconstruction.” in ISMIR, 2013, pp. 369–374.
  19. 19.S. Venkataramani, J. Casebeer, and P. Smaragdis, “End-to-end source separation with adaptive front-ends,” arXiv preprint arXiv:1705.02514, 2017.
  20. 20.D. Stoller, S. Ewert, and S. Dixon, “Wave-u-net: A multi-scale neural network for end-to-end audio source separation,” arXiv preprint arXiv:1806.03185, 2018.
  21. 21.Y. Luo and N. Mesgarani, “Tasnet: time-domain audio separation network for real-time, single-channel speech separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2018 IEEE International Conference on. IEEE, 2018.
  22. 22.S.-W. Fu, T.-W. Wang, Y. Tsao, X. Lu, and H. Kawai, “End-to-end waveform utterance enhancement for direct evaluation metrics optimization by fully convolutional neural networks,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), vol. 26, no. 9, pp. 1570–1584, 2018.
  23. 23.S. Pascual, A. Bonafonte, and J. Serra, “Segan: Speech enhancement generative adversarial network,” Proc. Interspeech 2017, pp. 3642–3646, 2017.
  24. 24.O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in International Conference on Medical image computing and computer-assisted intervention. Springer, 2015, pp. 234–241.
  25. 25.J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on. IEEE, 2016, pp. 31–35.
  26. 26.Y. Luo and N. Mesgarani, “Real-time single-channel dereverberation and separation with time-domain audio separation network,” Proc. Interspeech 2018, pp. 342–346, 2018.
  27. 27.F.-Y. Wang, C.-Y. Chi, T.-H. Chan, and Y. Wang, “Nonnegative least-correlated component analysis for separation of dependent sources by volume maximization,” IEEE transactions on pattern analysis and machine intelligence, vol. 32, no. 5, pp. 875–888, 2010.
  28. 28.C. H. Ding, T. Li, and M. I. Jordan, “Convex and semi-nonnegative matrix factorizations,” IEEE transactions on pattern analysis and machine intelligence, vol. 32, no. 1, pp. 45–55, 2010.
  29. 29.C. Lea, R. Vidal, A. Reiter, and G. D. Hager, “Temporal convolutional networks: A unified approach to action segmentation,” in European Conference on Computer Vision. Springer, 2016, pp. 47–54.
  30. 30.C. Lea, M. D. Flynn, R. Vidal, A. Reiter, and G. D. Hager, “Temporal convolutional networks for action segmentation and detection,” in proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 156–165.
  31. 31.S. Bai, J. Z. Kolter, and V. Koltun, “An empirical evaluation of generic convolutional and recurrent networks for sequence modeling,” arXiv preprint arXiv:1803.01271, 2018.
  32. 32.F. Chollet, “Xception: Deep learning with depthwise separable convolutions,” arXiv preprint, 2016.
  33. 33.A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “Mobilenets: Efficient convolutional neural networks for mobile vision applications,” arXiv preprint arXiv:1704.04861, 2017.
  34. 34.D. Wang, “On ideal binary mask as the computational goal of auditory scene analysis,” in Speech separation by humans and machines. Springer, 2005, pp. 181–197.
  35. 35.Y. Li and D. Wang, “On the optimality of ideal binary time–frequency masks,” Speech Communication, vol. 51, no. 3, pp. 230–239, 2009.
  36. 36.Y. Wang, A. Narayanan, and D. Wang, “On training targets for supervised speech separation,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), vol. 22, no. 12, pp. 1849–1858, 2014.
  37. 37.H. Erdogan, J. R. Hershey, S. Watanabe, and J. Le Roux, “Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks,” in Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on. IEEE, 2015, pp. 708–712.
  38. 38.A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “Wavenet: A generative model for raw audio,” CoRR abs/1609.03499, 2016.
  39. 39.L. Kaiser, A. N. Gomez, and F. Chollet, “Depthwise separable convolutions for neural machine translation,” arXiv preprint arXiv:1706.03059, 2017.
  40. 40.K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” in Proceedings of the IEEE international conference on computer vision, 2015, pp. 1026–1034.
  41. 41.J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” arXiv preprint arXiv:1607.06450, 2016.
  42. 42.“Script to generate the multi-speaker dataset using wsj0,” http://www.merl.com/demos/deep-clustering.
  43. 43.D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980, 2014.
  44. 44.E. Vincent, R. Gribonval, and C. Fevotte, “Performance measurement in blind audio source separation,” IEEE transactions on audio, speech, and language processing, vol. 14, no. 4, pp. 1462–1469, 2006.
  45. 45.A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in Acoustics, Speech, and Signal Processing, 2001. Proceedings.(ICASSP’01). 2001 IEEE International Conference on, vol. 2. IEEE, 2001, pp. 749–752.
  46. 46.ITU-T Rec. P.10, “Vocabulary for performance and quality of service,” 2006.
  47. 47.R. R. Sokal, “A statistical method for evaluating systematic relationship,” University of Kansas science bulletin, vol. 28, pp. 1409–1438, 1958.
  48. 48.M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 4510–4520.
  49. 49.“Audio samples for Conv-TasNet,” http://naplab.ee.columbia.edu/tasnet.html.
  50. 50.C. Xu, X. Xiao, and H. Li, “Single channel speech separation with constrained utterance level permutation invariant training using grid lstm,” in Acoustics, Speech and Signal Processing (ICASSP), 2018 IEEE International Conference on. IEEE, 2018.
  51. 51.T.-W. Lee, M. S. Lewicki, M. Girolami, and T. J. Sejnowski, “Blind source separation of more sources than mixtures using overcomplete representations,” IEEE signal processing letters, vol. 6, no. 4, pp. 87–90, 1999.
  52. 52.M. Zibulevsky and B. A. Pearlmutter, “Blind source separation by sparse decomposition in a signal dictionary,” Neural computation, vol. 13, no. 4, pp. 863–882, 2001.
  53. 53.S. Imai, “Cepstral analysis synthesis on the mel frequency scale,” in Acoustics, Speech, and Signal Processing, IEEE International Conference on ICASSP’83., vol. 8. IEEE, 1983, pp. 93–96.
  54. 54.G. L. Romani, S. J. Williamson, and L. Kaufman, “Tonotopic organization of the human auditory cortex,” Science, vol. 216, no. 4552, pp. 1339–1340, 1982.
  55. 55.C. Pantev, M. Hoke, B. Lutkenhoner, and K. Lehnertz, “Tonotopic organization of the auditory cortex: pitch versus frequency representation,” Science, vol. 246, no. 4929, pp. 486–488, 1989.
  56. 56.C. J. Darwin, D. S. Brungart, and B. D. Simpson, “Effects of fundamental frequency and vocal-tract length changes on attention to one of two simultaneous talkers,” The Journal of the Acoustical Society of America, vol. 114, no. 5, pp. 2913–2922, 2003.
  57. 57.J. R. Hershey, S. J. Rennie, P. A. Olsen, and T. T. Kristjansson, “Super-human multi-talker speech recognition: A graphical modeling approach,” Computer Speech & Language, vol. 24, no. 1, pp. 45–66, 2010.
  58. 58.C. Weng, D. Yu, M. L. Seltzer, and J. Droppo, “Deep neural networks for single-channel multi-talker speech recognition,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), vol. 23, no. 10, pp. 1670–1679, 2015.
  59. 59.Y. Qian, X. Chang, and D. Yu, “Single-channel multi-talker speech recognition with permutation invariant training,” arXiv preprint arXiv:1707.06527, 2017.
  60. 60.K. Ochi, N. Ono, S. Miyabe, and S. Makino, “Multi-talker speech recognition based on blind source separation with ad hoc microphone array using smartphones and cloud storage.” in INTERSPEECH, 2016, pp. 3369–3373.
  61. 61.Y. Lei, N. Scheffer, L. Ferrer, and M. McLaren, “A novel scheme for speaker recognition using a phonetically-aware deep neural network,” in Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on. IEEE, 2014, pp. 1695–1699.
  62. 62.M. McLaren, Y. Lei, and L. Ferrer, “Advances in deep neural network approaches to speaker recognition,” in Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on. IEEE, 2015, pp. 4814–4818.
  63. 63.S. Gannot, E. Vincent, S. Markovich-Golan, A. Ozerov, S. Gannot, E. Vincent, S. Markovich-Golan, and A. Ozerov, “A consolidated perspective on multimicrophone speech enhancement and source separation,” IEEE/ACM Transactions on Audio, Speech and Language Processing (TASLP), vol. 25, no. 4, pp. 692–730, 2017.
  64. 64.Z. Chen, J. Li, X. Xiao, T. Yoshioka, H. Wang, Z. Wang, and Y. Gong, “Cracking the cocktail party problem by multi-beam deep attractor network,” in Automatic Speech Recognition and Understanding Workshop (ASRU), 2017 IEEE. IEEE, 2017, pp. 437–444.
  65. 65.Z.-Q. Wang, J. Le Roux, and J. R. Hershey, “Multi-channel deep clustering: Discriminative spectral and spatial embeddings for speaker-independent speech separation,” in Acoustics, Speech and Signal Processing (ICASSP), 2018 IEEE International Conference on. IEEE, 2018.

Citation

MLA
Luo, Y., and N. Mesgarani. “Conv-TasNet: Surpassing Ideal Time–Frequency Magnitude Masking for Speech Separation”. IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 27, no. 8, 2019, pp. 1256–66, https://doi.org/10.1109/TASLP.2019.2915167.
APA
Luo, Y., & Mesgarani, N. (2019). Conv-TasNet: Surpassing Ideal Time–Frequency Magnitude Masking for Speech Separation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 27(8), 1256–1266. https://doi.org/10.1109/TASLP.2019.2915167
Chicago
Luo, Y., and N. Mesgarani. 2019. “Conv-TasNet: Surpassing Ideal Time–Frequency Magnitude Masking for Speech Separation”. IEEE/ACM Transactions on Audio, Speech, and Language Processing 27 (8): 1256–66. https://doi.org/10.1109/TASLP.2019.2915167.
Harvard
Luo, Y. and Mesgarani, N. (2019) “Conv-TasNet: Surpassing Ideal Time–Frequency Magnitude Masking for Speech Separation”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, 27(8), pp. 1256–1266. Available at: https://doi.org/10.1109/TASLP.2019.2915167.
Vancouver
1. Luo Y, Mesgarani N (2019) Conv-TasNet: Surpassing Ideal Time–Frequency Magnitude Masking for Speech Separation. IEEE/ACM Transactions on Audio, Speech, and Language Processing 27:1256–1266

BibTeX

@article{Luo_2019, title={Conv-TasNet: Surpassing Ideal Time–Frequency Magnitude Masking for Speech Separation}, volume={27}, ISSN={2329-9304}, url={http://dx.doi.org/10.1109/TASLP.2019.2915167}, DOI={10.1109/taslp.2019.2915167}, number={8}, journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Luo, Yi and Mesgarani, Nima}, year={2019}, month=Aug, pages={1256–1266} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF