Deep clustering: Discriminative embeddings for segmentation and separation

John R. HersheyZhuo ChenJonathan Le RouxShinji Watanabe

article2015IEEE International Conference on Acoustics, Speech, and Signal Processing1,467 citations

Proposes deep clustering, a framework that trains neural networks to map spectrogram time-frequency bins into discriminative embeddings, enabling speaker-independent audio separation that successfully generalizes to unseen mixtures.

Listen

Separating overlapping voices from a single recording—known as the cocktail party problem—remains a major hurdle for voice interfaces, automated transcription, and telecommunications. Existing deep learning approaches typically rely on assigning fixed output channels to specific signal types or known speakers. When multiple speakers of the same type speak simultaneously, these conventional models struggle because they cannot reliably decide which speaker corresponds to which output channel. Furthermore, traditional spectral clustering techniques that group acoustic features are computationally intensive and perform poorly when scaling to complex mixtures.

The article demonstrates a novel framework called deep clustering for speaker-independent speech separation. The objective is to train a neural network to map the time-frequency elements of an audio mixture into an embedding space where sounds produced by the same speaker naturally cluster together, allowing simple grouping algorithms to separate voices regardless of speaker identity or total speaker count.

To evaluate this approach, the authors constructed multi-speaker datasets using speech from the Wall Street Journal audio corpus, comprising tens of hours of two- and three-speaker mixtures with varied volume balances. The system uses bidirectional recurrent neural networks to map audio spectrogram bins into low-dimensional feature vectors. During training, the objective function pulls features from the same speaker closer while pushing features from different speakers apart, all without requiring predefined speaker labels for fixed output slots. At test time, a standard grouping method, K-means clustering, partitions these feature vectors into distinct source masks to reconstruct individual speech signals.

The experiments show that deep clustering substantially outperforms prior approaches on unknown speakers. For two-speaker mixtures with unfamiliar voices, deep clustering improved the signal-to-distortion ratio by approximately 6.5 dB, whereas conventional deep learning baselines achieved only a 1.2 to 1.3 dB improvement and traditional unsupervised methods reached 3.1 dB. Even when tested against a supervised baseline granted perfect advance knowledge of speaker identities, deep clustering performed roughly 1.4 dB better. The framework also generalized well: a model trained exclusively on two-speaker audio successfully separated three-speaker mixtures, providing up to a 2.8 dB improvement for unknown speakers and reaching a 7.0 dB gain when trained on known three-speaker sets. Model performance remained consistently high across embedding sizes between 20 and 60 dimensions.

These findings indicate that deep clustering successfully solves the assignment ambiguity in single-channel audio separation, making high-quality speaker-independent separation viable for real-world speech systems. By decoupling feature extraction from speaker assignment, organizations can deploy separation models without needing prior training on specific user voices or fixed assumptions about how many individuals are talking. This significantly reduces data collection costs and mitigates performance risks in multi-speaker environments.

Moving forward, stakeholders interested in speech processing applications should consider developing pilot pipelines around the deep clustering framework. Follow-up development should explore alternative neural architectures, such as convolutional networks, expand training data to encompass diverse real-world acoustic environments, and evaluate continuous separation masks to further enhance audio fidelity.

While the results demonstrate clear progress, decision-makers should note that these findings are preliminary and derived from controlled laboratory recordings with downsampled audio. Performance may vary under conditions involving heavy background noise, reverberant spaces, or larger groups of speakers. Nonetheless, the substantial improvement over existing techniques establishes high confidence in the fundamental framework for multi-speaker separation.

arXiv: 1508.04306
  • Paper: On Spectral Clustering: Analysis and an algorithm, Andrew Y. Ng et al. (2001). This paper establishes the foundational spectral clustering algorithm and Laplacian embedding theory that deep clustering directly reformulates and replaces with neural networks.
  • Paper: A tutorial on spectral clustering, Ulrike von Luxburg (2007). This tutorial provides essential theoretical background on graph-partitioning and spectral affinity formulations that motivate the low-rank pairwise affinity objective in deep clustering.
  • Paper: Kernel k-means: spectral clustering and normalized cuts, Inderjit S. Dhillon et al. (2004). This work formalizes the link between normalized cuts and kernel-based clustering objectives, providing foundational math for optimizing affinity-based segmentation.
  • Paper: Self-Tuning Spectral Clustering, Lihi Zelnik-Manor et al. (2004). This study introduces self-tuning pairwise affinities for spectral graph partitioning, directly informing the affinity matrix structure approximated in deep clustering.
  • Paper: A New Learning Algorithm for Blind Signal Separation, Shun-ichi Amari et al. (1995). This paper provides foundational concepts in blind source separation that classic acoustic separation models relied on prior to deep discriminative embeddings.
Cover for Deep clustering: Discriminative embeddings for segmentation and separation

Abstract

We address the problem of acoustic source separation in a deep learning framework we call "deep clustering." Rather than directly estimating signals or masking functions, we train a deep network to produce spectrogram embeddings that are discriminative for partition labels given in training data. Previous deep network approaches provide great advantages in terms of learning power and speed, but previously it has been unclear how to use them to separate signals in a class-independent way. In contrast, spectral clustering approaches are flexible with respect to the classes and number of items to be segmented, but it has been unclear how to leverage the learning power and speed of deep networks. To obtain the best of both worlds, we use an objective function that to train embeddings that yield a low-rank approximation to an ideal pairwise affinity matrix, in a class-independent way. This avoids the high cost of spectral factorization and instead produces compact clusters that are amenable to simple clustering methods. The segmentations are therefore implicitly encoded in the embeddings, and can be "decoded" by clustering. Preliminary experiments show that the proposed method can separate speech: when trained on spectrogram features containing mixtures of two speakers, and tested on mixtures of a held-out set of speakers, it can infer masking functions that improve signal quality by around 6dB. We show that the model can generalize to three-speaker mixtures despite training only on two-speaker mixtures. The framework can be used without class labels, and therefore has the potential to be trained on a diverse set of sound types, and to generalize to novel sources. We hope that future work will lead to segmentation of arbitrary sounds, with extensions to microphone array methods as well as image segmentation and other domains.

Table of Contents

  • 1 Introduction
  • 2 Learning deep embeddings for clustering
  • 3 Speech separation experiments
  • 3.1 Experimental setup
  • 3.2 Training procedure
  • 3.3 Speech separation procedure
  • 4 Results and discussion
  • References

Knowls

  1. Knowl 1 — Deep Clustering Framework for Permutation-Independent Source Separation

    model/method

    Deep clustering is a deep learning framework designed to solve single-channel, speaker-independent audio source separation without suffering from the permutation problem that plagues class-based neural networks.

    Let Xext(orXt,fextforframetextandfrequencyf)X ext{ (or } X_{t,f} ext{ for frame } t ext{ and frequency } f) denote the complex or magnitude spectrogram of an acoustic mixture with N=TimesFN = T imes F time-frequency (T-F) bins. The ground-truth segmentation of the spectrogram into CC sources is represented by a binary indicator matrix Yimes{0,1}N×CY imes \{0, 1\}^{N \times C}, where yi,c=1y_{i,c} = 1 if source cc dominates T-F bin ii, and 00 otherwise. The target pairwise affinity between bins ii and jj is given by the binary matrix A∗=YYTA^* = Y Y^T, which satisfies (YYT)i,j=1(Y Y^T)_{i,j} = 1 if bins ii and jj belong to the same source, and 00 otherwise. Because (YP)(YP)T=YYT(Y P)(Y P)^T = Y Y^T for any C×CC \times C permutation matrix PP, YYTY Y^T provides a permutation-invariant representation of the segmentation.

    A neural network parameterized by θ\theta maps the input feature representation XX to a DD-dimensional embedding matrix V=fθ(X)∈RN×DV = f_\theta(X) \in \mathbb{R}^{N \times D}, where each row vi∈RDv_i \in \mathbb{R}^D is constrained to unit norm (∥vi∥22=1\|v_i\|_2^2 = 1). The estimated pairwise affinity matrix is implicitly represented as A^=VVT\hat{A} = V V^T. The network is trained so that T-F bins dominated by the same source are mapped to nearby embeddings, while bins dominated by different sources are pushed apart.

  2. Knowl 2 — Deep Clustering Objective Function and Efficient Low-Rank Formulation

    equation

    The training loss for deep clustering measures the squared Frobenius norm difference between the estimated affinity matrix VVTV V^T and the target affinity matrix YYTY Y^T:

    CY(V)=∥VVT−YYT∥F2=∑i,j(⟨vi,vj⟩−⟨yi,yj⟩)2C_Y(V) = \|V V^T - Y Y^T\|_F^2 = \sum_{i,j} (\langle v_i, v_j \rangle - \langle y_i, y_j \rangle)^2

    where V∈RN×DV \in \mathbb{R}^{N \times D} is the matrix of unit-norm DD-dimensional embeddings for NN time-frequency bins, Y∈{0,1}N×CY \in \{0,1\}^{N \times C} is the source assignment indicator matrix for CC sources, and ∥⋅∥F\|\cdot\|_F denotes the Frobenius norm. This loss expands to:

    CY(V)=∑i,j:yi=yj(∥vi−vj∥2−1)+∑i,j⟨vi,vj⟩2C_Y(V) = \sum_{i,j: y_i = y_j} (\|v_i - v_j\|^2 - 1) + \sum_{i,j} \langle v_i, v_j \rangle^2

    where the first term pulls embeddings belonging to the same source together, and the second term pushes all embeddings apart to prevent trivial collapse.

    Because N≫DN \gg D and N≫CN \gg C, computing CY(V)C_Y(V) over all N2N^2 pairs directly is computationally prohibitive. The loss is reformulated in low-rank form as:

    CY(V)=∥VTV∥F2−2∥VTY∥F2+∥YTY∥F2C_Y(V) = \|V^T V\|_F^2 - 2\|V^T Y\|_F^2 + \|Y^T Y\|_F^2

    The gradient with respect to VV is efficiently computed as:

    ∂CY(V)∂V=4V(VTV)−4Y(YTV)\frac{\partial C_Y(V)}{\partial V} = 4 V (V^T V) - 4 Y (Y^T V)

    This formulation reduces the computational complexity from O(N2)O(N^2) to O(ND2+NDC)O(N D^2 + N D C).

  3. Knowl 3 — Deep Clustering Inference and Signal Reconstruction

    algorithm

    At inference time, deep clustering predicts the source assignment masks by clustering the learned embeddings using KK-means, and uses the resulting clusters as binary time-frequency masks to reconstruct separated audio streams.

    Input: Complex mixture spectrogram X∈CT×FX \in \mathbb{C}^{T \times F}, trained embedding network fθf_\theta, number of sources CC
    Output: Estimated separated time-domain signals s^1(t),…,s^C(t)\hat{s}_1(t), \dots, \hat{s}_C(t)
    Compute embeddings V=fθ(∣X∣)∈RN×DV = f_\theta(|X|) \in \mathbb{R}^{N \times D} where N=T×FN = T \times F
    Normalize each row viv_i of VV such that ∥vi∥2=1\|v_i\|_2 = 1 for all i∈{1,…,N}i \in \{1, \dots, N\}
    Initialize cluster indicator matrix Y^∈{0,1}N×C\hat{Y} \in \{0, 1\}^{N \times C}
    Estimate Y^\hat{Y} by minimizing the KK-means objective:
        Y^=arg⁡min⁡Y∥V−YM∥F2\hat{Y} = \arg\min_Y \|V - Y M\|_F^2
        where M=(YTY)−1YTV∈RC×DM = (Y^T Y)^{-1} Y^T V \in \mathbb{R}^{C \times D} represents the cluster centroids
    for each source c=1c = 1 to CC do
        Construct binary mask M^c∈{0,1}T×F\hat{M}_c \in \{0, 1\}^{T \times F} such that M^c(t,f)=y^(t,f),c\hat{M}_c(t, f) = \hat{y}_{(t, f), c}
        Apply mask to complex spectrogram: S^c(t,f)=M^c(t,f)⊙X(t,f)\hat{S}_c(t, f) = \hat{M}_c(t, f) \odot X(t, f)
        Reconstruct time-domain signal s^c(t)=iSTFT(S^c)\hat{s}_c(t) = \text{iSTFT}(\hat{S}_c)
    end for
    return s^1(t),…,s^C(t)\hat{s}_1(t), \dots, \hat{s}_C(t)
  4. Knowl 4 — WSJ0-Based Speech Separation Experimental Setup

    experimental setup

    The deep clustering speech separation model is evaluated on single-channel speech mixtures generated from the Wall Street Journal (WSJ0) corpus downsampled to 8 kHz.

    • Mixture Datasets:
      • Two-speaker mixtures: 30 hours of training data and 10 hours of validation data (closed condition, CC) generated by randomly mixing utterances from the WSJ0 training set si_tr_s at signal-to-noise ratios (SNR) between 0 dB and 10 dB. An open condition (OC) evaluation set consists of 5 hours generated from 16 held-out speakers in si_dt_05 and si_et_05.
      • Three-speaker mixtures: 100 mixtures from many speakers under closed condition (MS-CC) and open condition (MS-OC); and a 3-known-speaker set (3S-CC) consisting of 5000 training, 500 validation, and 500 test mixtures drawn from three fixed speakers in si_et_05.
    • Input Features & Segmentation: Log spectral magnitudes from short-time Fourier transforms (STFT) using a 32 ms window, 8 ms hop size, and square-root Hann window. Inputs are segmented into 100-frame half-overlapping chunks.
    • Training Details: Target masks YY are computed via ideal binary masking (IBM). T-F bins with energy lower than −40 dB-40\text{ dB} relative to the maximum mixture magnitude are excluded from the loss. The neural architecture consists of two bidirectional LSTM (BLSTM) layers with 600 cells each, followed by a feedforward projection layer to dimension DD (typically D=40D=40). Training uses stochastic gradient descent with momentum 0.90.9, learning rate 10−510^{-5}, and zero-mean Gaussian noise with variance 0.60.6 added to weights during updates.
  5. Knowl 5 — Two-Speaker Speech Separation Performance Comparison

    data/table

    The performance of deep clustering (DC) was evaluated against supervised/oracle baselines, unsupervised Computational Auditory Scene Analysis (CASA), and class-based BLSTM networks on two-speaker speech separation using Signal-to-Distortion Ratio (SDR) improvement in dB (initial average mixture SDR was 0.2 dB).

    Method CC (dB) OC (dB)
    Oracle NMF 5.1 –
    CASA 2.9 3.1
    DC local KK-means 6.5 6.5
    DC global KK-means 5.9 5.8
    BLSTM stronger 1.3 1.2
    BLSTM permute 1.3 1.3
    BLSTM permute* 1.4 1.2

    DC with local KK-means achieves a 6.5 dB SDR improvement in both closed condition (CC) and open condition (OC), outperforming unsupervised CASA (3.1 dB OC) and oracle NMF (5.1 dB CC). Standard class-based BLSTMs fail on speaker-independent separation (~1.2–1.4 dB SDR improvement) due to the label permutation problem, regardless of whether training targets are assigned based on signal strength (BLSTM stronger) or dynamic output permutation alignment (BLSTM permute and BLSTM permute*).

  6. Knowl 6 — Effect of Embedding Dimension and Activation Function on Separation Performance

    data/table

    SDR improvements (in dB) on two-speaker mixtures across varying embedding dimensions D∈{5,10,20,40,60}D \in \{5, 10, 20, 40, 60\} and output activation functions (tanh vs. logistic):

    CC OC
    Model DC local DC global DC local DC global
    D=5D = 5 -0.8 -1.0 -0.7 -1.1
    D=10D = 10 5.2 4.5 5.3 4.6
    D=20D = 20 6.3 5.6 6.4 5.7
    D=40D = 40 6.5 5.9 6.5 5.8
    D=60D = 60 6.0 5.2 6.1 5.3
    D=40D = 40 logistic 6.6 5.9 6.6 6.0

    At D=5D = 5, the network fails to learn an effective separation space. Performance stabilizes for D≥20D \ge 20, peaking around D=40D = 40. The choice of activation function (tanh vs. logistic) yields virtually identical separation quality (6.5 dB vs. 6.6 dB for local KK-means).

  7. Knowl 7 — Generalization to Three-Speaker Mixtures and Speaker-Dependent Separation

    data/table

    Deep clustering models trained only on two-speaker mixtures generalize zero-shot to three-speaker mixtures by setting the number of clusters K=3K=3 at test time. Performance is also evaluated on three known speakers when trained specifically on that closed set (initial mixture SDR was −3.0 dB-3.0\text{ dB}).

    Zero-Shot (Trained on 2 Speakers) Known Speakers (3S-CC)
    Method MS-CC (dB) MS-OC (dB) Method 3S-CC (dB)
    Oracle NMF 4.4 – Oracle NMF 4.5
    DC local 3.5 2.8 DC local 7.0
    DC global 2.7 2.2 DC global 6.9
    – – – BLSTM stack 6.8

    When trained only on two-speaker mixtures, DC achieves 2.8 dB SDR improvement on open-condition 3-speaker mixtures (MS-OC) without retraining. When trained directly on the three known speakers (3S-CC), DC achieves 7.0 dB SDR improvement, slightly exceeding the class-based stacked BLSTM baseline (BLSTM stack, 6.8 dB), which can only function in the speaker-dependent regime.

Coverage note — None was omitted; all key theoretical formulations, loss functions, network configurations, algorithms, and empirical benchmark results from the paper are fully covered.

References

  1. 1.A. S. Bregman, Auditory scene analysis: The perceptual organization of sound. MIT press, 1990.
  2. 2.F. Weninger, F. Eyben, and B. Schuller, “Single-channel speech separation with memory-enhanced recurrent neural networks,” in Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on. IEEE, 2014.
  3. 3.Y. Wang and D. Wang, “Towards scaling up classification-based speech separation,” IEEE Trans. Audio, Speech, Language Process., vol. 21, no. 7, 2013.
  4. 4.Y. Wang, K. Han, and D. Wang, “Exploring monaural features for classification-based speech segregation,” IEEE Trans. Audio, Speech, Language Process., vol. 21, no. 2, 2013.
  5. 5.J. R. Hershey, S. J. Rennie, P. A. Olsen, and T. T. Kristjansson, “Super-human multi-talker speech recognition: A graphical modeling approach,” Comput. Speech Lang., vol. 24, no. 1, 2010.
  6. 6.P.-S. Huang, M. Kim, M. Hasegawa-Johnson, and P. Smaragdis, “Joint optimization of masks and deep recurrent neural networks for monaural source separation,” arXiv preprint arXiv:1502.04149, 2015.
  7. 7.T. Virtanen, “Speech recognition using factorial hidden markov models for separation in the feature space,” in Proc. Interspeech 2006, Pittsburgh, 2006.
  8. 8.S. J. Rennie, J. R. Hershey, and P. A. Olsen, “Single-channel multitalker speech recognition,” IEEE Signal Process. Mag., vol. 27, no. 6, 2010.
  9. 9.M. Cooke, J. R. Hershey, and S. J. Rennie, “Monaural speech separation and recognition challenge,” Computer Speech & Language, vol. 24, no. 1, 2010.
  10. 10.R. J. Weiss, “Underdetermined source separation using speaker subspace models,” Ph.D. dissertation, Columbia University, 2009.
  11. 11.F. Weninger, H. Erdogan, S. Watanabe, E. Vincent, J. Le Roux, J. R. Hershey, and B. Schuller, “Speech enhancement with lstm recurrent neural networks and its application to noise-robust asr,” in Latent Variable Analysis and Signal Separation. Springer, 2015.
  12. 12.Y. Wang, A. Narayanan, and D. Wang, “On training targets for supervised speech separation,” Audio, Speech, and Language Processing, IEEE/ACM Transactions on, vol. 22, no. 12, 2014.
  13. 13.Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “An experimental study on speech enhancement based on deep neural networks,” Signal Processing Letters, IEEE, vol. 21, no. 1, 2014.
  14. 14.M. P. Cooke, “Modelling auditory processing and organisation,” Ph.D. dissertation, Univ. of Sheffield, 1991.
  15. 15.D. P. W. Ellis, “Prediction-driven computational auditory scene analysis,” Ph.D. dissertation, MIT, 1996.
  16. 16.F. R. Bach and M. I. Jordan, “Learning spectral clustering, with application to speech separation,” JMLR, vol. 7, 2006.
  17. 17.M. Wertheimer, “Laws of organization in perceptual forms,” in A Source book of Gestalt psychology, W. A. Ellis, Ed. Routledge and Kegan Paul, 1938.
  18. 18.K. Hu and D. Wang, “An unsupervised approach to cochannel speech separation,” Audio, Speech, and Language Processing, IEEE Transactions on, vol. 21, no. 1, 2013.
  19. 19.J. Shi and J. Malik, “Normalized cuts and image segmentation,” IEEE Trans. PAMI, vol. 22, no. 8, 2000.
  20. 20.F. Tian, B. Gao, Q. Cui, E. Chen, and T.-Y. Liu, “Learning deep representations for graph clustering,” in Proc. AAAI, 2014.
  21. 21.P. Huang, Y. Huang, W. Wang, and L. Wang, “Deep embedding network for clustering,” in Proc. ICPR, 2014.
  22. 22.C. Fowlkes, S. Belongie, F. Chung, and J. Malik, “Spectral grouping using the nystrom method,” IEEE Trans. PAMI, vol. 26, no. 2, 2004.
  23. 23.M. Meilă, “Local equivalences of distances between clusteringsa geometric perspective,” Machine Learning, vol. 86, no. 3, 2012.
  24. 24.L. Hubert and P. Arabie, “Comparing partitions,” Journal of classification, vol. 2, no. 1, 1985.
  25. 25.M. Meilă, “The stability of a good clustering,” University of Washington Department of Statistics, vol. Technical Report 624, 2014. [Online]. Available: http://www.stat.washington.edu/research/reports/2014/tr624.pdf
  26. 26.J. R. Hershey, Z. Chen, J. L. Roux, and S. Watanabe, “Deep clustering: Discriminative embeddings for segmentation and separation,” Sep. 2015, arXiv:1508.04306. [Online]. Available: http://arxiv.org/abs/1508.04306
  27. 27.P. Smaragdis, “Convolutive speech bases and their application to supervised speech separation,” IEEE Trans. Audio, Speech, Language Process., vol. 15, no. 1, 2007.
  28. 28.J. Le Roux, F. J. Weninger, and J. R. Hershey, “Sparse NMF – half-baked or well done?” MERL, Cambridge, MA, USA, Tech. Rep. TR2015-023, Mar. 2015.
  29. 29.D. Wang, “On ideal binary mask as the computational goal of auditory scene analysis,” in Speech separation by humans and machines. Springer, 2005.
  30. 30.E. Vincent, R. Gribonval, and C. Févotte, “Performance measurement in blind audio source separation,” IEEE Trans. Audio, Speech, Language Process., vol. 14, no. 4, 2006.
  31. 31.K. Hu and D. Wang, “An iterative model-based approach to cochannel speech separation,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2013, no. 1, 2013.
  32. 32.C. Farabet, C. Couprie, L. Najman, and Y. LeCun, “Learning hierarchical features for scene labeling,” IEEE Trans. PAMI, vol. 35, no. 8, 2013.
  33. 33.A. Sharma, O. Tuzel, and M.-Y. Liu, “Recursive context propagation network for semantic scene labeling,” in Proc. NIPS, 2014.

Citation

MLA
Hershey, J. R., et al. “Deep Clustering: Discriminative Embeddings for Segmentation and Separation”. arXiv, 2015, http://arxiv.org/abs/1508.04306v1.
APA
Hershey, J. R., Chen, Z., Roux, J. L., & Watanabe, S. (2015). Deep clustering: Discriminative embeddings for segmentation and separation. arXiv. http://arxiv.org/abs/1508.04306v1
Chicago
Hershey, J. R., Z. Chen, J. L. Roux, and S. Watanabe. 2015. “Deep Clustering: Discriminative Embeddings for Segmentation and Separation”. arXiv. http://arxiv.org/abs/1508.04306v1.
Harvard
Hershey, J.R. et al. (2015) “Deep clustering: Discriminative embeddings for segmentation and separation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1508.04306v1.
Vancouver
1. Hershey JR, Chen Z, Roux JL, Watanabe S (2015) Deep clustering: Discriminative embeddings for segmentation and separation. arXiv

BibTeX

@article{hershey2015deep,
  title = {Deep clustering: Discriminative embeddings for segmentation and separation},
  author = {Hershey, John R. and Chen, Zhuo and Roux, Jonathan Le and Watanabe, Shinji},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1508.04306v1},
  eprint = {1508.04306}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF