Supervised Training of Conditional Monge Maps

Charlotte BunneAndreas KrauseMarco Cuturi

article2022NeurIPS93 citations

Introduces CONDOT, a neural optimal transport framework using partially input convex neural networks to learn context-conditioned Monge maps across probability measures, enabling accurate prediction of cellular responses to unseen drug and genetic perturbation combinations.

Listen

Predicting how complex biological systems respond to interventions is a major hurdle in personalized medicine and drug development. In cancer therapeutics, single-cell sequencing can measure individual cell behaviors before and after interventions, but the measurement process destroys the cells, leaving before and after populations completely unpaired. Optimal transport provides a mathematical foundation for mapping such unpaired distributions, but existing methods typically calculate a standalone map for a single isolated condition. Because real-world biomedical applications involve varying drug dosages, patient-specific covariates, and vast combinations of genetic or therapeutic treatments, there is an urgent need for models that can incorporate contextual variables and generalize to previously unobserved conditions.

The article introduces a framework designed to estimate a global family of context-conditioned optimal transport maps from labeled population pairs. Its core objective is to accurately map source distributions to target distributions across known conditions while demonstrating the ability to predict the outcomes of unseen contexts and composite perturbations.

To achieve this, the researchers parameterized optimal transport maps as the gradients of partially input convex neural networks, which modulate convex potentials via contextual inputs. Contextual factors—ranging from continuous scalars like dosage to discrete genetic perturbations—are processed through specialized embedding modules and permutation-invariant combination networks. To overcome the known training instability of convex neural architectures, the approach introduces an initialization strategy that mimics closed-form affine transport maps between Gaussian approximations of the source and target distributions. The framework was evaluated across single-cell genomics benchmarks, including thousands of gene expression profiles across diverse drug dosages, distinct cell lines, and nearly one hundred distinct CRISPR-guided genetic perturbations.

The experimental findings highlight substantial improvements in predictive performance. When modeling the response of single cells to varying drug dosages, the context-aware model achieved significantly lower distribution discrepancies on out-of-sample dosages compared to baseline models, reducing marker gene perturbation error metrics by more than half relative to standard architectures. In multi-cell-line benchmarks, the method captured cell-type-specific responses where baseline models failed to match target distribution shifts. Furthermore, on large-scale genetic perturbation screens, the architecture accurately predicted cellular responses to completely unseen perturbation combinations, recovering distinct subpopulations that alternative deep-learning baselines missed entirely.

These results establish that embedding contextual variables directly into transport map estimation enables effective multi-task learning and robust out-of-distribution generalization. For drug discovery and clinical development pipelines, this capability offers a viable route to drastically decrease experimental costs, reduce development timelines, and mitigate the risk of missed therapeutic effects by screening combinatorial interventions computationally before conducting laboratory trials.

Organizations evaluating computational perturbation models should consider piloting this architecture for preclinical drug combination screening and dosage interpolation. Incorporating domain-specific action embeddings or mode-of-action representations is recommended to enhance performance when predicting novel compounds. However, practitioners should maintain reasonable caution: the framework relies on the assumption that underlying population shifts follow optimal transport trajectories, and its predictive accuracy decreases as the combinatorial distance from observed training data grows. Future initiatives should focus on validating predictions through targeted wet-lab follow-up experiments and scaling the framework across larger, multi-modal biological datasets.

Cover for Supervised Training of Conditional Monge Maps

Table of Contents

  • 1 Introduction
  • 2 Background on Neural Solvers for the 2-Wasserstein Problem
  • 3 Supervised Training of Conditional Monge Maps
  • 3.1 A Regression Formulation for Conditional OT Estimation
  • 3.2 Integrating Context in Convex Architectures
  • 3.3 Conditional Monge Map Architecture
  • 4 Initialization Strategies for Neural Convex Architectures
  • 5 Evaluation
  • 5.1 Population Dynamics Conditioned on Scalars
  • 5.2 Population Dynamics Conditioned on Covariates
  • 5.3 Population Dynamics Conditioned on Actions
  • 5.3.1 Known Actions
  • 5.3.2 Unknown Actions
  • 5.3.3 Actions in Combination
  • 6 Conclusion
  • Acknowledgments and Disclosure of Funding
  • References

Knowls

  1. Knowl 1 — CONDOT: conditional Monge-map learning from labeled measure pairs

    model/method

    CONDOT learns one context-dependent family of optimal-transport maps from labeled, unpaired population pairs. The training data are NN tuples (ci,(mui,nui))(c_i,(\\mu_i,\\nu_i)), where cic_i belongs to an arbitrary context space mathcalC\\mathcal C and mui,nui\\mu_i,\\nu_i are probability measures on mathbbRd\\mathbb R^d. The model produces a map T_\\theta(c):\\mathbb R^d\\to\\mathbb R^d such that T_\\theta(c_i)_\\#\\mu_i approximates nui\\nu_i, while also producing a map for an unseen context cmathrmnewc_{\\mathrm{new}}. Here T_\\theta(c)_\\#\\mu denotes the push-forward of source measure mu\\mu by the map T_\\theta(c).

    The modeling assumption is that each observed source-to-target transformation is explained by an optimal Monge map and that the maps for different contexts share parameters whose values are modulated by context. CONDOT therefore performs multi-task learning rather than fitting an independent transport map for every pair; it is intended to predict effects at unseen dosages, covariates, actions, and combinations of actions.

  2. Knowl 2 — Partially input-convex neural representation of conditional Monge maps

    model/method

    CONDOT parameterizes a convex potential f_\\theta(x,c) that is convex in the data variable xinmathbbRdx\\in\\mathbb R^d but not necessarily in the context variable cc. The conditional transport map is the gradient

    T_\\theta(c)(x)=\\nabla_x f_\\theta(x,c).

    A KK-layer partially input-convex neural network (PICNN) uses context states uku_k and convex-in-xx states zkz_k. For k=0,ldots,K−1k=0,\\ldots,K-1, with u0=cu_0=c and z0=0z_0=0, its recurrence is

    uk+1=tauk(Vkuk+vk),u_{k+1}=\\tau_k(V_k u_k+v_k),

    zk+1=sigmak!left(Wkzleft(zkodotleft[Wkzuuk+bkzright]right)+Wkxleft(xodotleft[Wkxuuk+bkxright]right)+Wkuuk+bkuright),z_{k+1}=\\sigma_k\\!\\left(W_k^z\\left(z_k\\odot\\left[W_k^{zu}u_k+b_k^z\\right]\\right)+W_k^x\\left(x\\odot\\left[W_k^{xu}u_k+b_k^x\\right]\\right)+W_k^u u_k+b_k^u\\right),

    with output f_\\theta(x,c)=z_K. Here odot\\odot is elementwise multiplication, tauk\\tau_k is an arbitrary context-processing activation, sigmak\\sigma_k is convex and non-decreasing, and all Vk,Wkz,Wkzu,Wkx,Wkxu,WkuV_k,W_k^z,W_k^{zu},W_k^x,W_k^{xu},W_k^u and biases are learned parameters of compatible dimensions. Convexity in xx is guaranteed by the convex non-decreasing activations and nonnegative entries in every latent-to-latent matrix WkzW_k^z; the nonnegativity is enforced by applying softplus or ReLU to precursor matrices, or by penalizing negative entries.

  3. Knowl 3 — Joint optimal-transport training objective for all contexts

    equation

    Because the observed populations are unpaired, CONDOT does not supervise pointwise maps. Instead, it jointly optimizes the conditional convex potential over all labeled measure pairs. For context cic_i, define the Legendre transform of the potential by

    f_\\theta^*(x,c_i)=\\sup_{z\\in\\mathbb R^d}\\{\\langle x,z\\rangle-f_\\theta(z,c_i)\\}.

    Using the paper's dual-objective convention, the learned parameters maximize

    \\theta^*\\in\\arg\\max_\\theta\\sum_{i=1}^N\\left[\\int_{\\mathbb R^d}f_\\theta^*(x,c_i)\\,d\\mu_i(x)+\\int_{\\mathbb R^d}f_\\theta(y,c_i)\\,d\\nu_i(y)\\right].

    Here x,y,zinmathbbRdx,y,z\\in\\mathbb R^d, mui\\mu_i is the source measure, nui\\nu_i is the target measure, and cic_i is the context label. The gradient \\nabla_x f_\\theta(\\cdot,c_i) is consequently trained to be an optimal Monge map for each observed pair. In the practical implementation, the Legendre transform is replaced by the proxy dual construction of Makkuva et al.: two PICNNs, one representing fthetaff_{\\theta_f} and one representing an auxiliary potential gthetagg_{\\theta_g}, share the context embedding and combinator modules, while a regularizer encourages gthetag(cdot,c)g_{\\theta_g}(\\cdot,c) to approximate the Legendre transform of fthetaf(cdot,c)f_{\\theta_f}(\\cdot,c) for every context.

  4. Knowl 4 — Context embeddings and permutation-invariant action combinators

    model/method

    CONDOT separates context processing from the convex transport network. An embedding module E_\\phi maps a raw context into a real-valued vector, and a combinator module C_\\Phi merges one or more embedded contexts before they are supplied to the PICNN. Scalar contexts such as time or dosage can be used directly; small discrete context sets can use one-hot embeddings; and complex actions can use learned continuous embeddings or domain-specific molecular representations. The parameters of the embedding module, combinator, and PICNN are trained end-to-end.

    For a set of actions whose order is irrelevant or unknown, the combinator is permutation-invariant. CONDOT implements this with a Deep Sets-style network, so the representation of an action set does not depend on the order in which its elements are presented. This allows the same conditional Monge-map model to handle individual actions and action combinations, rather than requiring a separately parameterized map for every possible combination.

  5. Knowl 5 — Gaussian closed-form potentials used to initialize transport networks

    equation

    CONDOT initializes convex transport networks using the affine optimal map between Gaussian approximations of a source and target measure. Let mathcalN1=mathcalN(m1,Sigma1)\\mathcal N_1=\\mathcal N(m_1,\\Sigma_1) and mathcalN2=mathcalN(m2,Sigma2)\\mathcal N_2=\\mathcal N(m_2,\\Sigma_2) be distributions on mathbbRd\\mathbb R^d, with Sigma1\\Sigma_1 full rank. Define

    A=\\left[\\Sigma_1^{-1/2}\\left(\\Sigma_1^{1/2}\\Sigma_2\\Sigma_1^{1/2}\\right)^{1/2}\\Sigma_1^{-1/2}\\right]^{1/2},\\qquad b=m_2-A^\\top A m_1.

    A Brenier potential for the Gaussian transport is

    f_{\\mathcal N_1,\\mathcal N_2}^*(x)=\\frac12x^\\top A^\\top A x+b^\\top x+t,

    where the constant is chosen as t=\\frac12b^\\top(A^\\top A)^{-1}b so that the potential is nonnegative and has minimum zero. Equivalently, with \\omega=m_1-(A^\\top A)^{-1}m_2,

    fmathcalN1,mathcalN2∗(x)=frac12left∣A(x−omega)right∣22.f_{\\mathcal N_1,\\mathcal N_2}^*(x)=\\frac12\\left\\|A(x-\\omega)\\right\\|_2^2.

    The associated quadratic layer is qM,m(x)=frac12∣M(x−m)∣22q_{M,m}(x)=\\frac12\\|M(x-m)\\|_2^2; its gradient is the identity when (M,m)=(I,0)(M,m)=(I,0) and transports mathcalN1\\mathcal N_1 to mathcalN2\\mathcal N_2 when (M,m)=(A,omega)(M,m)=(A,\\omega). In experiments, m1,m2,Sigma1,Sigma2m_1,m_2,\\Sigma_1,\\Sigma_2 are replaced by empirical means and covariance matrices of the observed source and target populations.

  6. Knowl 6 — Identity, Gaussian, and context-specific PICNN initialization

    algorithm

    CONDOT uses initialization schemes designed to make the initial gradient meaningful as a transport map rather than starting from an arbitrary convex function.

    For an ICNN, identity initialization inserts qI,0(x)=12∣x∣22q_{I,0}(x)=\frac12\\|x\\|_2^2 into the first hidden state, sets latent-to-latent matrices approximately to uniform averaging matrices mathbf1d2,d1/d1\\mathbf 1_{d_2,d_1}/d_1, sets input matrices approximately to zero, and chooses large positive biases so the derivatives of the activations are approximately one. Gaussian initialization uses the same propagation scheme but inserts qA,omegaq_{A,\\omega} computed from empirical Gaussian source and target moments; its initial gradient therefore approximates the corresponding affine Gaussian transport.

    For a PICNN with training contexts c1,ldots,cJc_1,\\ldots,c_J, the context modulator is initialized as

    u_0(c)=\\operatorname{softmax}(c^\\top M),

    where the columns of MM are the context vectors cjc_j. The initial convex state contains one quadratic potential per training context, z0=[qMj,mj(x)]j=1Jz_0=[q_{M_j,m_j}(x)]_{j=1}^J, with either (Mj,mj)=(I,0)(M_j,m_j)=(I,0) for identity initialization or (Mj,mj)=(Aj,omegaj)(M_j,m_j)=(A_j,\\omega_j) for Gaussian initialization. The first context-to-convex coupling is initialized with W0zu=IW_0^{zu}=I and b0z=0b_0^z=0; later WkzW_k^z matrices are initialized to uniform averaging, while WkxW_k^x and WkxuW_k^{xu} are near zero, bkzb_k^z is near one, and bkub_k^u is near zero. Thus a context close to cjc_j initially selects the quadratic potential associated with the corresponding source-target Gaussian approximation.

  7. Knowl 7 — Givinostat dosage and cell-line evaluation

    data/table

    The dosage and covariate experiments evaluate whether conditional maps predict single-cell populations more accurately than context-free OT or a conditional autoencoder. The Givinostat dataset contains 3,541 cells represented by 1,000 highly variable genes; dosages are 1010, 100100, 1,0001{,}000, and 10,00010{,}000 nM. The model is evaluated both on dosages observed during training and on held-out dosages, and on three cell lines (A549, MCF7, and K562). The reported metrics are MMD and the ell2\\ell_2 distance between perturbation signatures, computed on marker-gene expression features. Lower values are better.

    Could not parse LaTeX table

    The table on page 7 shows that both CONDOT initializations outperform CPA and context-free ICNN OT on the reported dosage and cell-line metrics. The out-of-sample dosage results demonstrate interpolation to dosages absent from training, with Gaussian initialization obtaining the lowest reported out-of-sample MMD and ell2(PS)\\ell_2(PS) among the compared methods.

  8. Knowl 8 — Context-aware prediction of known genetic perturbations

    empirical result

    For known-action prediction, CONDOT was trained on pooled CRISPR single-cell RNA-sequencing data containing 98,419 profiles, 92 genetic perturbations, and 1,500 highly variable genes per cell. Each perturbation was represented by a one-hot context embedding, and the predicted target population was compared with context-free ICNN OT and the conditional perturbation autoencoder CPA.

    The scatter plots on page 8 compare Wasserstein losses for CONDOT and ICNN OT under mode-of-action and one-hot embeddings, while a further plot compares CONDOT with CPA. CONDOT captures perturbation-specific responses more accurately than both baselines, including effects that are subtle in the high-dimensional gene-expression space. The improvement over context-free ICNN OT indicates that conditioning the Monge map on the perturbation is useful beyond predicting an average perturbation effect.

  9. Knowl 9 — Mode-of-action embeddings enable prediction for unseen perturbations

    model/method

    To represent a perturbation that was not observed as a context during training, CONDOT constructs a mode-of-action embedding from marginal access to the target population of each individual perturbation. It computes pairwise Wasserstein distances between the individual target populations and applies multidimensional scaling (MDS) to obtain vectors in which perturbations with similar effects are close together. These vectors replace one-hot identifiers for the action-conditioning network.

    Using this embedding, CONDOT predicted responses for the unseen perturbations BAK1, FOXF1, MAP2K6, and MAP4K3. The reported Wasserstein losses for these unknown actions were similar to the losses for perturbations observed during training, whereas one-hot representations cannot produce a meaningful representation for an unseen action. This experiment demonstrates generalization through learned effect-based geometry rather than through an unseen discrete identifier.

  10. Knowl 10 — Generalization to unseen combinations of genetic perturbations

    empirical result

    CONDOT predicts combination perturbations from observations of individual perturbations. Two combination mechanisms were evaluated: addition of individual one-hot embeddings and a learned permutation-invariant Deep Sets combinator applied to mode-of-action embeddings. The train/test splits were made progressively harder by reducing the number of perturbation combinations observed during training while retaining individual perturbations.

    Across the known-perturbation and unknown-combination evaluations shown on pages 8–10, CONDOT had lower regularized Wasserstein and MMD losses than ICNN OT and CPA. Performance decreased as fewer combinations were available during training, but CONDOT remained superior to both baselines. In the UMAP visualization on page 10 for the KLF1+MAP2K6 combination, CONDOT's predictions align with the observed perturbed population and cover its subpopulations, whereas ICNN OT and CPA fail to reproduce some of those subpopulations.

Coverage note — No substantial contributed material was omitted; appendix-level proxy-dual implementation details, metric definitions, and supplementary plots were treated as supporting details of the included training and experimental results rather than separate knowls.

References

  1. 1.D. Alvarez-Melis, Y. Schiff, and Y. Mroueh. Optimizing Functionals on the Space of Probabilities with Input Convex Neural Networks. arXiv Preprint arXiv:2106.00774, 2021.
  2. 2.B. Amos, L. Xu, and J. Z. Kolter. Input Convex Neural Networks. In International Conference on Machine Learning (ICML), volume 34, 2017.
  3. 3.M. Arjovsky, S. Chintala, and L. Bottou. Wasserstein Generative Adversarial Networks. In International Conference on Machine Learning (ICML). PMLR, 2017.
  4. 4.Y. Brenier. Décomposition polaire et réarrangement monotone des champs de vecteurs. CR Acad. Sci. Paris Sér. I Math., 305, 1987.
  5. 5.C. Bunne, S. G. Stark, G. Gut, J. S. del Castillo, K.-V. Lehmann, L. Pelkmans, A. Krause, and G. Ratsch. Learning Single-Cell Perturbation Responses using Neural Optimal Transport. bioRxiv, 2021.
  6. 6.C. Bunne, Y.-P. Hsieh, M. Cuturi, and A. Krause. Recovering Stochastic Dynamics via Gaussian Schrödinger Bridges. arXiv Preprint arXiv:2202.05722, 2022a.
  7. 7.C. Bunne, L. Meng-Papaxanthos, A. Krause, and M. Cuturi. Proximal Optimal Transport Modeling of Population Dynamics. In International Conference on Artificial Intelligence and Statistics (AISTATS), volume 25, 2022b.
  8. 8.R. Caruana. Multitask Learning. Machine Learning, 28(1), 1997.
  9. 9.Y. Chandak, G. Theocharous, J. Kostas, S. Jordan, and P. Thomas. Learning Action Representations for Reinforcement Learning. In International Conference on Machine Learning (ICML), 2019.
  10. 10.Y. Chen, Y. Shi, and B. Zhang. Optimal Control Via Neural Networks: A Convex Approach. In International Conference on Learning Representations (ICLR), 2019.
  11. 11.S. Chopra, R. Hadsell, and Y. LeCun. Learning a similarity metric discriminatively, with application to face verification. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), volume 1. IEEE, 2005.
  12. 12.M. Cuturi. Sinkhorn Distances: Lightspeed Computation of Optimal Transport. In Advances in Neural Information Processing Systems (NeurIPS), volume 26, 2013.
  13. 13.M. Cuturi and G. Peyré. Semidual Regularized Optimal Transport. SIAM Review, 60(4):941–965, 2018.
  14. 14.M. Cuturi, L. Meng-Papaxanthos, Y. Tian, C. Bunne, G. Davis, and O. Teboul. Optimal Transport Tools (OTT): A JAX Toolbox for all things Wasserstein. arXiv Preprint arXiv:2201.12324, 2022.
  15. 15.J. De Leeuw and P. Mair. Multidimensional Scaling Using Majorization: SMACOF in R. Journal of Statistical Software, 31, 2009.
  16. 16.A. Dixit, O. Parnas, B. Li, J. Chen, C. P. Fulco, L. Jerby-Arnon, N. D. Marjanovic, D. Dionne, T. Burks, R. Raychowdhury, et al. Perturb-Seq: Dissecting Molecular Circuits with Scalable Single-Cell RNA Profiling of Pooled Genetic Screens. Cell, 167(7), 2016.
  17. 17.J. Fan, A. Taghvaei, and Y. Chen. Scalable Computations of Wasserstein Barycenter via Input Convex Neural Networks. In International Conference on Machine Learning (ICML), 2021.
  18. 18.M. Gelbrich. On a Formula for the L2 Wasserstein Metric between Measures on Euclidean and Hilbert Spaces. Mathematische Nachrichten, 147(1), 1990.
  19. 19.A. Genevay, L. Chizat, F. Bach, M. Cuturi, and G. Peyré. Sample Complexity of Sinkhorn Divergences. In International Conference on Artificial Intelligence and Statistics (AISTATS), volume 22, 2019.
  20. 20.A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. Smola. A Kernel Two-Sample Test. Journal of Machine Learning Research, 13, 2012.
  21. 21.D. Ha, A. Dai, and Q. V. Le. Hypernetworks. arXiv preprint arXiv:1609.09106, 2016.
  22. 22.T. Hashimoto, D. Gifford, and T. Jaakkola. Learning Population-Level Diffusions with Generative Recurrent Networks. In International Conference on Machine Learning (ICML), volume 33, 2016.
  23. 23.C.-W. Huang, R. T. Q. Chen, C. Tsirigotis, and A. Courville. Convex Potential Flows: Universal Probability Distributions with Optimal Transport and Convex Optimization. In International Conference on Learning Representations (ICLR), 2021.
  24. 24.J.-C. Hütter and P. Rigollet. Minimax estimation of smooth optimal transport maps. The Annals of Statistics, 49(2), 2021.
  25. 25.L. Jacob, J. She, A. Almahairi, S. Rajeswar, and A. Courville. W2GAN: Recovering an Optimal Transport Map with a GAN. In arXiv Preprint, 2018.
  26. 26.H. Janati, T. Bazeille, B. Thirion, M. Cuturi, and A. Gramfort. Multi-subject MEG/EEG source imaging with sparse multi-task regression. NeuroImage, 220, 2020.
  27. 27.L. Kantorovich. On the transfer of masses (in Russian). In Doklady Akademii Nauk, volume 37, 1942.
  28. 28.D. P. Kingma and J. Ba. Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations (ICLR), 2014.
  29. 29.A. Korotin, V. Egiazarian, A. Asadulaev, A. Safin, and E. Burnaev. Wasserstein-2 Generative Networks. In International Conference on Learning Representations (ICLR), 2020.
  30. 30.A. Korotin, L. Li, A. Genevay, J. M. Solomon, A. Filippov, and E. Burnaev. Do Neural Optimal Transport Solvers Work? A Continuous Wasserstein-2 Benchmark. Advances in Neural Information Processing Systems (NeurIPS), 34, 2021.
  31. 31.S. Kummar, H. X. Chen, J. Wright, S. Holbeck, M. D. Millin, J. Tomaszewski, J. Zweibel, J. Collins, and J. H. Doroshow. Utilizing targeted cancer therapeutic agents in combination: novel approaches and urgent requirements. Nature Reviews Drug discovery, 9(11), 2010.
  32. 32.M. Lotfollahi, F. A. Wolf, and F. J. Theis. scGen predicts single-cell perturbation responses. Nature Methods, 16(8), 2019.
  33. 33.M. Lotfollahi, M. Naghipourfar, F. J. Theis, and F. A. Wolf. Conditional out-of-distribution generation for unpaired data using transfer VAE. Bioinformatics, 36, 2020.
  34. 34.M. Lotfollahi, A. K. Susmelj, C. De Donno, Y. Ji, I. L. Ibarra, F. A. Wolf, N. Yakubova, F. J. Theis, and D. Lopez-Paz. Compositional perturbation autoencoder for single-cell response modeling. bioRxiv, 2021.
  35. 35.R. K. Mahabadi, S. Ruder, M. Dehghani, and J. Henderson. Parameter-efficient Multi-task Fine-tuning for Transformers via Shared Hypernetworks. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2021.
  36. 36.A. Makkuva, A. Taghvaei, S. Oh, and J. Lee. Optimal transport mapping via input convex neural networks. In International Conference on Machine Learning (ICML), volume 37, 2020.
  37. 37.R. J. McCann. A Convexity Principle for Interacting Gases. Advances in Mathematics, 128(1), 1997.
  38. 38.L. McInnes, J. Healy, and J. Melville. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv Preprint, 2018.
  39. 39.A. Mead. Review of the Development of Multidimensional Scaling Methods. Journal of the Royal Statistical Society: Series D (The Statistician), 41(1), 1992.
  40. 40.T. Mikolov, K. Chen, G. Corrado, and J. Dean. Efficient Estimation of Word Representations in Vector Space. International Conference on Learning Representations (ICLR), Workshop Track, 2013a.
  41. 41.T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. Distributed Representations of Words and Phrases and their Compositionality. Advances in Neural Information Processing Systems (NeurIPS), 26, 2013b.
  42. 42.T. Mikolov, W.-t. Yih, and G. Zweig. Linguistic Regularities in Continuous Space Word Representations. In Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL-HLT), 2013c.
  43. 43.R. B. Mokhtari, T. S. Homayouni, N. Baluch, E. Morgatskaya, S. Kumar, B. Das, and H. Yeger. Combination therapy in combating cancer. Oncotarget, 8(23):38022, 2017.
  44. 44.P. Mokrov, A. Korotin, L. Li, A. Genevay, J. Solomon, and E. Burnaev. Large-Scale Wasserstein Gradient Flows. Advances in Neural Information Processing Systems (NeurIPS), 2021.
  45. 45.G. Monge. Mémoire sur la théorie des déblais et des remblais. Histoire de l’Académie Royale des Sciences, pages 666–704, 1781.
  46. 46.T. M. Norman, M. A. Horlbeck, J. M. Replogle, A. Y. Ge, A. Xu, M. Jost, L. A. Gilbert, and J. S. Weissman. Exploring genetic interaction manifolds constructed from rich single-cell phenotypes. Science, 365(6455), 2019.
  47. 47.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research, 12, 2011.
  48. 48.G. Peyré and M. Cuturi. Computational Optimal Transport. Foundations and Trends in Machine Learning, 11(5-6), 2019. ISSN 1935-8245.
  49. 49.A.-A. Pooladian and J. Niles-Weed. Entropic estimation of optimal transport maps. arXiv preprint arXiv:2109.12004, 2021.
  50. 50.N. Prasad, K. Yang, and C. Uhler. Optimal Transport using GANs for Lineage Tracing. arXiv preprint arXiv:2007.12098, 2020.
  51. 51.A. Rambaldi, C. M. Dellacasa, G. Finazzi, A. Carobbio, M. L. Ferrari, P. Guglielmelli, E. Gattoni, S. Salmoiraghi, M. C. Finazzi, S. Di Tollo, et al. A pilot study of the Histone-Deacetylase inhibitor Givinostat in patients with JAK2V617F positive chronic myeloproliferative neoplasms. British journal of haematology, 150(4):446–455, 2010.
  52. 52.D. Rezende and S. Mohamed. Variational Inference with Normalizing Flows. In International Conference on Machine Learning (ICML), 2015.
  53. 53.J. Richter-Powell, J. Lorraine, and B. Amos. Input Convex Gradient Networks. arXiv preprint arXiv:2111.12187, 2021.
  54. 54.P. Rigollet and A. J. Stromme. On the sample complexity of entropic optimal transport. arXiv preprint arXiv:2206.13472, 2022.
  55. 55.D. Rogers and M. Hahn. Extended-Connectivity Fingerprints. Journal of Chemical Information and Modeling, 50(5), 2010.
  56. 56.Y. Rong, Y. Bian, T. Xu, W. Xie, Y. Wei, W. Huang, and J. Huang. Self-Supervised Graph Transformer on Large-Scale Molecular Data. Advances in Neural Information Processing Systems (NeurIPS), 2020.
  57. 57.F. Santambrogio. Optimal Transport for Applied Mathematicians. Birkhäuser, NY, 55(58-63):94, 2015.
  58. 58.G. Schiebinger, J. Shu, M. Tabaka, B. Cleary, V. Subramanian, A. Solomon, J. Gould, S. Liu, S. Lin, P. Berube, et al. Optimal-Transport Analysis of Single-Cell Gene Expression Identifies Developmental Trajectories in Reprogramming. Cell, 176(4), 2019.
  59. 59.M. A. Schmitz, M. Heitz, N. Bonneel, F. Ngole, D. Coeurjolly, M. Cuturi, G. Peyré, and J.-L. Starck. Wasserstein dictionary learning: Optimal transport-based unsupervised nonlinear dictionary learning. SIAM Journal on Imaging Sciences, 11(1):643–678, 2018.
  60. 60.P. Schwaller, A. C. Vaucher, R. Laplaza, C. Bunne, A. Krause, C. Corminboeuf, and T. Laino. Machine intelligence for chemical reaction space. Wiley Interdisciplinary Reviews: Computational Molecular Science, page e1604, 2022.
  61. 61.S. R. Srivatsan, J. L. McFaline-Figueroa, V. Ramani, L. Saunders, J. Cao, J. Packer, H. A. Pliner, D. L. Jackson, R. M. Daza, L. Christiansen, et al. Massively multiplex chemical transcriptomics at single-cell resolution. Science, 367(6473), 2020.
  62. 62.V. Stathias, A. M. Jermakowicz, M. E. Maloof, M. Forlin, W. Walters, R. K. Suter, M. A. Durante, S. L. Williams, J. W. Harbour, C.-H. Volmar, et al. Drug and disease signature integration identifies synergistic combinations in glioblastoma. Nature Communications, 9(1), 2018.
  63. 63.G. Tennenholtz and S. Mannor. The Natural Language of Actions. In International Conference on Machine Learning (ICML), 2019.
  64. 64.C. Villani. Topics in Optimal Transportation, volume 58. American Mathematical Soc., 2003.
  65. 65.F. A. Wolf, P. Angerer, and F. J. Theis. SCANPY: large-scale single-cell gene expression data analysis. Genome biology, 19(1), 2018.
  66. 66.K. D. Yang and C. Uhler. Scalable Unbalanced Optimal Transport using Generative Adversarial Networks. International Conference on Learning Representations (ICLR), 2019.
  67. 67.K. D. Yang, K. Damodaran, S. Venkatachalapathy, A. C. Soylemezoglu, G. Shivashankar, and C. Uhler. Predicting cell lineages using autoencoders and optimal transport. PLoS Computational Biology, 16(4), 2020.
  68. 68.M. Zaheer, S. Kottur, S. Ravanbakhsh, B. Poczos, R. R. Salakhutdinov, and A. J. Smola. Deep Sets. In Advances in Neural Information Processing Systems (NeurIPS), volume 30, 2017.

Citation

MLA
Bunne, C., et al. “Supervised Training of Conditional Monge Maps”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 6859–72, https://proceedings.neurips.cc/paper_files/paper/2022/file/2d880acd7b31e25d45097455c8e8257f-Paper-Conference.pdf.
APA
Bunne, C., Krause, A., & Cuturi, M. (2022). Supervised Training of Conditional Monge Maps. Advances in Neural Information Processing Systems, 35, 6859–6872. https://proceedings.neurips.cc/paper_files/paper/2022/file/2d880acd7b31e25d45097455c8e8257f-Paper-Conference.pdf
Chicago
Bunne, C., A. Krause, and M. Cuturi. 2022. “Supervised Training of Conditional Monge Maps”. Advances in Neural Information Processing Systems 35: 6859–72. https://proceedings.neurips.cc/paper_files/paper/2022/file/2d880acd7b31e25d45097455c8e8257f-Paper-Conference.pdf.
Harvard
Bunne, C., Krause, A. and Cuturi, M. (2022) “Supervised Training of Conditional Monge Maps”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 6859–6872. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/2d880acd7b31e25d45097455c8e8257f-Paper-Conference.pdf.
Vancouver
1. Bunne C, Krause A, Cuturi M (2022) Supervised Training of Conditional Monge Maps. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 6859–6872

BibTeX

@inproceedings{bunne2022supervised,
  title = {Supervised Training of Conditional Monge Maps},
  author = {Bunne, Charlotte and Krause, Andreas and Cuturi, Marco},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {6859-6872},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/2d880acd7b31e25d45097455c8e8257f-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors