A Linear Non-Gaussian Acyclic Model for Causal Discovery

Shohei ShimizuPatrik O. HoyerAapo HyvärinenAntti Kerminen

article2006JMLR2,076 citations

Introduces the LiNGAM framework, which uses independent component analysis and non-Gaussian error distributions to identify the complete, directed causal structure of continuous observational data without requiring prior variable ordering.

Listen

Understanding cause-and-effect relationships is essential for predicting the consequences of organizational interventions and strategic policy decisions. However, conducting direct, controlled experiments is often prohibitively expensive, unethical, or logistically impossible, forcing leaders to rely on observational data. Traditional statistical discovery methods for continuous data assume standard bell-curve (Gaussian) distributions, which generally fail to identify the true directional flow of cause and effect and leave decision-makers with multiple indistinguishable, competing models.

The article demonstrates that full causal discovery—identifying both the exact directional order and the numerical strength of relationships without prior knowledge of time sequencing—is mathematically possible for continuous observational data. It establishes this capability under the framework of a Linear Non-Gaussian Acyclic Model, which relies on the condition that external disturbance or noise variables do not follow a bell-curve distribution.

To accomplish this, the article introduces an analytical algorithm based on independent component analysis, a statistical technique that separates mixed signals into independent, non-bell-curve components. The method resolves core mathematical ambiguities by pairing variables to disturbances and ordering them into a directional chain where causes strictly precede effects. The authors also formulated hypothesis tests to prune statistically insignificant connections and assess overall model fit. The credibility of the framework was tested across 1,000 synthetic trials across various system sizes (3 to 100 variables) and sample scopes (200 to 10,000 data points), as well as on 22 real-world economic and environmental time-series datasets.

Key findings show that non-Gaussian distributions provide sufficient information to uniquely determine the full causal structure and parameter values without any pre-specified variable ordering. In extensive simulations, the algorithm estimated underlying causal connection strengths near-flawlessly as sample sizes grew. When pruning unnecessary connections at sample sizes of 5,000 and above, the method achieved statistical power between 90% and 97% for correctly identified connections, keeping total classification errors below 10%. In real-world time-series evaluations, the method successfully recovered the true forward time order in well-behaved data, identified reverse causality in financial random-walk series, and flagged zero-effect relationships where no true dependencies existed.

These findings mean organizations can extract directional causal insights directly from historical and observational data, significantly reducing the cost, time, and operational risk associated with running large-scale physical experiments. Unlike traditional covariance-based techniques that yield ambiguous results, this approach provides a single, actionable model of causal mechanisms. This allows decision-makers to simulate interventions and evaluate compliance or performance impacts with much higher precision.

Practitioners should consider using this analytical discovery method as an exploratory first step to formulate data-driven causal hypotheses prior to committing capital to physical trials. When applying the model, analysts must implement the recommended statistical pruning and model-fit tests to filter out weak or spurious links. Further analysis and validation against domain knowledge are necessary before executing critical operational decisions.

Confidence in the results is high when the underlying assumptions hold: the relationships must be linear, free of unobserved common causes (confounders), and driven by non-bell-curve noise. Readers should exercise caution, as computational optimization can occasionally settle into local errors, and violations of key assumptions—such as non-linear dynamics or missing confounding factors—will degrade model reliability.

  • Paper: The max-min hill-climbing Bayesian network structure learning algorithm, Ioannis Tsamardinos et al. (2006). Introduces the Max-Min Hill-Climbing algorithm, offering a premier hybrid baseline for continuous and discrete DAG structure discovery to contrast against ICA-based identification.
  • Paper: Learning Bayesian Networks with the bnlearn R Package, Marco Scutari (2009). Implements constraint-based and score-based network structure learning algorithms in R, providing the standard computational suite used alongside LiNGAM for causal graph benchmarking.
  • Paper: Invariant Risk Minimization, Martin Arjovsky et al. (2019). Extends linear structural equation modeling to multi-environment settings by learning invariant causal representations under intervention and distribution shift.
  • Paper: Identifying Weight-Variant Latent Causal Models, Yuhang Liu et al. (2026). Generalizes linear causal model identifiability to latent variable structures where causal parameters vary across contexts, building upon linear structural discovery foundations.
  • Paper: Counterfactual Fairness, Matt J. Kusner et al. (2017). Applies structural equation modeling and latent causal factor identification to enforce fairness constraints in downstream predictive models.
Cover for A Linear Non-Gaussian Acyclic Model for Causal Discovery

Abstract

In recent years, several methods have been proposed for the discovery of causal structure from non-experimental data. Such methods make various assumptions on the data generating process to facilitate its identification from purely observational data. Continuing this line of research, we show how to discover the complete causal structure of continuous-valued data, under the assumptions that (a) the data generating process is linear, (b) there are no unobserved confounders, and (c) disturbance variables have non-Gaussian distributions of non-zero variances. The solution relies on the use of the statistical method known as independent component analysis, and does not require any pre-specified time-ordering of the variables. We provide a complete Matlab package for performing this LiNGAM analysis (short for Linear Non-Gaussian Acyclic Model), and demonstrate the effectiveness of the method using artificially generated data and real-world data.

Table of Contents

  • 1. Introduction
  • 2. Linear Causal Networks
  • 3. Model Identification Using Independent Component Analysis
  • 4. LiNGAM Discovery Algorithm
  • 5. Permutation Algorithms for Large Dimensions
  • 5.1 Permuting the Rows of W
  • 5.2 Permuting B to Get a Causal Order
  • 6. Statistical Tests for Pruning Edges
  • 6.1 Wald Test for Examining Significance of Edges
  • 6.2 A Chi-Square Test for Evaluating the Overall Fit of the Estimated Model
  • 6.2.1 MOMENT STRUCTURES OF MODELS
  • 6.2.2 SOME TEST STATISTICS TO EVALUATE A MODEL FIT
  • 6.2.3 A DIFFERENCE CHI-SQUARE TEST FOR MODEL COMPARISON OF NESTED MODELS
  • 6.3 A Method for Pruning Edges
  • 7. Simulations
  • 7.1 Estimation of B
  • 7.2 Pruning Edges
  • 8. Examples With Real-World Data
  • 9. Conclusions
  • Acknowledgments
  • Appendix A. Proof of Uniqueness of Row Permutation
  • Appendix B. An Example of the Permutation Problem
  • Appendix C. ML Derivation of Objective Function for Finding the Correct Row Permutation
  • Appendix D. Asymptotic Variance of ICA
  • Appendix E. Exact Form of J = ∂σ 2 ( τ ) / ∂τ T
  • References

Knowls

  1. Knowl 1 — Linear Non-Gaussian Acyclic Model

    model/method

    The Linear Non-Gaussian Acyclic Model (LiNGAM) describes a data generating process over mm continuous observed random variables x=[x1,…,xm]Tx = [x_1, \dots, x_m]^T defined by structural linear equations:

    xi=∑k(j)<k(i)bijxj+ei+ci,x_i = \sum_{k(j) < k(i)} b_{ij} x_j + e_i + c_i,

    where:

    • k(i)k(i) is a causal ordering function of the variables such that no later variable causes an earlier variable, ensuring the directed graph of relations is acyclic (DAG).
    • bijb_{ij} represents the linear connection strength from xjx_j to xix_i.
    • cic_i is an optional constant intercept term.
    • eie_i are continuous disturbance (noise) terms with non-zero variances and non-Gaussian distributions.
    • The disturbance variables e1,…,eme_1, \dots, e_m are mutually independent, satisfying p(e1,…,em)=∏i=1mpi(ei)p(e_1, \dots, e_m) = \prod_{i=1}^m p_i(e_i), which implies causal sufficiency (absence of unobserved confounders).

    In matrix notation, subtracting the means yields x=Bx+ex = Bx + e, where BB is an m×mm \times m coefficient matrix with zeros on the diagonal that can be permuted to strict lower triangular form by simultaneous row and column permutations corresponding to k(i)k(i).

  2. Knowl 2 — Identifiability of LiNGAM Causal Structure via ICA

    theoretical result

    Under the LiNGAM assumptions, the full causal network structure and all connection strengths BB are uniquely identifiable from observational data without prior knowledge of the causal ordering.

    From the mean-centered model x=Bx+ex = Bx + e, solving for xx gives:

    x=(I−B)−1e=Ae,x = (I - B)^{-1} e = A e,

    where A=(I−B)−1A = (I - B)^{-1} is the mixing matrix and W=A−1=I−BW = A^{-1} = I - B is the demixing matrix. Because the disturbances eie_i are non-Gaussian and independent, standard linear Independent Component Analysis (ICA) identifies AA (and WW) up to column/row permutations and component scalings: WICA=PICAWW_{\text{ICA}} = P_{\text{ICA}} W, where PICAP_{\text{ICA}} is a random permutation matrix.

    Because BB is strictly lower triangular under the true causal ordering PdP_d, the demixing matrix can be written as W=PdMPdTW = P_d M P_d^T, where MM is lower triangular with non-zero diagonal entries. For any lower triangular matrix MM with non-zero diagonal entries, a permutation P1MP2TP_1 M P_2^T has only non-zero entries on the main diagonal if and only if P1=P2P_1 = P_2. Consequently, there exists exactly one row permutation of WICAW_{\text{ICA}} that produces non-zero entries along the main diagonal, uniquely resolving ICA's permutation indeterminacy and identifying the causal ordering and parameter matrix BB.

  3. Knowl 3 — LiNGAM Discovery Algorithm

    algorithm

    The LiNGAM discovery algorithm estimates the causal ordering and the connection strength matrix BB from an observational data matrix XX of size m×nm \times n (with m≪nm \ll n).

    Input: Continuous data matrix XX of size m×nm \times n
    Output: Estimated connection matrix B^\hat{B}, causal permutation matrix PP
    Subtract sample mean from each row of XX
    Apply ICA decomposition to obtain X=ASX = A S, where SS contains independent components
    Compute demixing matrix W=A−1W = A^{-1}
    Find row permutation of WW yielding W~\tilde{W} that minimizes ∑i=1m1/∣W~ii∣\sum_{i=1}^m 1/|\tilde{W}_{ii}|
    Divide each row ii of W~\tilde{W} by its diagonal element W~ii\tilde{W}_{ii} to obtain W~′\tilde{W}'
    Compute initial connection matrix estimate B^=I−W~′\hat{B} = I - \tilde{W}'
    Find permutation matrix PP to minimize upper triangular sum of squares ∑i≤j(PB^PT)ij2\sum_{i \le j} (P \hat{B} P^T)_{ij}^2
    Set upper triangular and diagonal elements of PB^PTP \hat{B} P^T to 0
    return B^\hat{B}, PP
  4. Knowl 4 — Row Permutation and Causal Ordering for High Dimensions

    algorithm

    For high-dimensional networks where exhaustive search is infeasible, row permutation of WW and causal ordering of B^\hat{B} are solved via linear assignment and iterative pruning.

    1. Row Permutation of WW: The maximization of diagonal magnitudes is formulated as a linear assignment problem minimizing ∑i=1mCϕ(i),i\sum_{i=1}^m C_{\phi(i), i}, where Cij=1/∣Wij∣C_{ij} = 1 / |W_{ij}| and ϕ\phi is a permutation. This is solved with worst-case computational complexity O(m3)O(m^3).

    2. Causal Ordering of B^\hat{B} via Iterative Pruning:

    Input: Estimated connection matrix B^\hat{B} of size m×mm \times m
    Output: Permuted strictly lower triangular matrix B~\tilde{B}
    Set the m(m+1)/2m(m+1)/2 smallest absolute value elements of B^\hat{B} to 0
    loop
        Test DAGness: search for a row with all zeros, append index to ordering list pp, remove that row and column from matrix, repeat until empty or no zero row found
        if DAGness test succeeds then
            Apply permutation pp to rows and columns of B^\hat{B} to produce B~\tilde{B}
            return B~\tilde{B}
        else
            Set the next smallest absolute value element of B^\hat{B} to 0
        end if
    end loop
  5. Knowl 5 — Wald Test for Edge Significance

    model/method

    To prune weak or spurious edges in the estimated network, a Wald test evaluates the null hypothesis H0:bij=0H_0: b_{ij} = 0 (equivalent to H0:wij=0H_0: w_{ij} = 0) against H1:bij≠0H_1: b_{ij} \ne 0.

    The Wald test statistic is:

    Waldij=w^ij2avar(w^ij),\text{Wald}_{ij} = \frac{\hat{w}_{ij}^2}{\text{avar}(\hat{w}_{ij})},

    where w^ij\hat{w}_{ij} is the estimated element of WW and avar(w^ij)\text{avar}(\hat{w}_{ij}) is its asymptotic variance derived from the estimating equations of the ICA estimator (e.g., FastICA).

    Under H0H_0, the statistic Waldij\text{Wald}_{ij} asymptotically follows a chi-square distribution with 1 degree of freedom (χ2(1)\chi^2(1)). An edge is judged non-significant if the probability of observing a value greater than or equal to Waldij\text{Wald}_{ij} exceeds a significance threshold α\alpha (e.g., 0.05).

  6. Knowl 6 — Chi-Square Goodness-of-Fit and Difference Tests

    equation

    Overall model fit to observational second-order moments is evaluated by comparing empirical moments m2=1n∑j=1nvec+(xjxjT)m_2 = \frac{1}{n} \sum_{j=1}^n \text{vec}^+(x_j x_j^T) against model-implied moments σ2(τ)=vec+{E(xxT)}\sigma_2(\tau) = \text{vec}^+\{E(xx^T)\}, where τ\tau contains the free parameters of BB and disturbance variances E(ei2)E(e_i^2), and vec+(⋅)\text{vec}^+(\cdot) vectorizes the non-duplicate entries of a symmetric matrix.

    The weighted least-squares distance is:

    F(τ^)={m2−σ2(τ^)}TM^{m2−σ2(τ^)},F(\hat{\tau}) = \{m_2 - \sigma_2(\hat{\tau})\}^T \hat{M} \{m_2 - \sigma_2(\hat{\tau})\},

    where M^=V^−1−V^−1J^(J^TV^−1J^)−1J^TV^−1\hat{M} = \hat{V}^{-1} - \hat{V}^{-1} \hat{J} (\hat{J}^T \hat{V}^{-1} \hat{J})^{-1} \hat{J}^T \hat{V}^{-1}, V^\hat{V} is the sample covariance matrix of m2m_2, and J^=∂σ2(τ)/∂τT∣τ=τ^\hat{J} = \partial \sigma_2(\tau)/\partial \tau^T |_{\tau=\hat{\tau}}.

    The Yuan-Bentler corrected test statistic for overall fit is:

    T2=nF(τ^)1+F(τ^),T_2 = \frac{n F(\hat{\tau})}{1 + F(\hat{\tau})},

    which asymptotically follows χ2(u−v)\chi^2(u - v) under H0:E(m2)=σ2(τ)H_0: E(m_2) = \sigma_2(\tau), where uu is the number of distinct second moments and vv is the number of free parameters in τ\tau.

    For two nested models with qq and q−1q-1 edges, the difference statistic T2(q)−T2(q−1)T_2(q) - T_2(q-1) asymptotically follows χ2(1)\chi^2(1) to test whether pruning an edge significantly degrades overall model fit.

  7. Knowl 7 — Combined Statistical Pruning Procedure

    algorithm

    To prune the fully connected initial DAG estimate while maintaining overall model validity, individual Wald tests are combined with nested model difference tests and chi-square goodness-of-fit tests.

    Input: Estimated lower triangular matrix B~\tilde{B}, significance level α\alpha (e.g., 0.05)
    Output: Pruned DAG connection matrix B~pruned\tilde{B}_{\text{pruned}}
    Identify all non-significant edges in strictly lower triangular part of B~\tilde{B} using individual Wald tests at level α\alpha
    Order candidate non-significant edges by ascending significance (smallest Wald statistic first)
    for each candidate non-significant edge do
        Create a candidate model with this edge set to 0
        Compute nested difference statistic T2(q)−T2(q−1)T_2(q) - T_2(q-1) against preceding model
        Compute overall fit statistic T2(q−1)T_2(q-1) of candidate model
        if difference test accepts null (fit change not significant) and overall fit test accepts null (model fits data) then
            Prune the edge (keep entry set to 0)
        else
            Retain the edge in the model
        end if
    end for
    return B~pruned\tilde{B}_{\text{pruned}}
  8. Knowl 8 — Empirical Edge Pruning Performance on Synthetic Data

    data/table

    Across 1000 simulation trials on synthetic linear acyclic models with non-Gaussian disturbances generated via power non-linearities, the combined statistical pruning method (evaluated at α=0.05\alpha = 0.05) demonstrates increasing edge identification accuracy and power as sample size nn increases.

    Condition True Pos. False Neg. True Neg. False Pos. Sum of False Pos./Neg.
    Dim. = 5, n=1000n = 1000 8101 (90.5%) 849 (9.5%) 921 (87.7%) 129 (12.3%) 978 (9.8%)
    Dim. = 5, n=5000n = 5000 8556 (95.6%) 394 (4.4%) 943 (89.8%) 107 (10.2%) 501 (5.0%)
    Dim. = 5, n=10000n = 10000 8691 (97.1%) 259 (2.9%) 972 (92.6%) 78 (7.4%) 337 (3.4%)
    Dim. = 10, n=1000n = 1000 27825 (80.6%) 6698 (19.4%) 8171 (78.0%) 2306 (22.0%) 9004 (20.0%)
    Dim. = 10, n=5000n = 5000 31623 (91.6%) 2900 (8.4%) 9350 (89.2%) 1127 (10.8%) 4027 (8.9%)
    Dim. = 10, n=1000n = 1000 32477 (94.1%) 2046 (5.9%) 9466 (90.4%) 1011 (9.6%) 3057 (6.8%)

    True positives and negatives were counted strictly within the lower triangular support. For n≥5000n \ge 5000, statistical power exceeds 90%90\% and total error rates (false positives plus false negatives) remain under 10%10\% across both 5- and 10-variable networks.

  9. Knowl 9 — Causal Discovery in Autoregressive Time Series

    empirical result

    A stationary autoregressive process of order pp, Xt=∑k=1pϕkXt−k+atX_t = \sum_{k=1}^p \phi_k X_{t-k} + a_t with independent non-Gaussian white noise ata_t, conforms to LiNGAM when represented over sliding time windows [Xt,Xt+1,…,Xt+m−1][X_t, X_{t+1}, \dots, X_{t+m-1}] as multivariate samples (though order p≥2p \ge 2 introduces latent confounding for the first pp variables in a window).

    Empirical evaluation over 22 real-world economic and environmental time series revealed four distinct outcome behaviors:

    1. Correct Causal Ordering (5 cases): Observed in stationary series with non-Gaussian innovations (e.g., monthly precipitation PEAS), correctly recovering AR parameters.
    2. Reverse Causal Ordering (9 cases): Observed in non-stationary random walks Xt=Xt−1+atX_t = X_{t-1} + a_t (e.g., daily IBM stock prices), where time direction is mathematically symmetric and unidentifiable, yielding estimated first AR coefficients near 1 in either direction.
    3. Zero Connection Matrix (1 case): Observed in independent white noise series (e.g., daily precipitation PRECIP), where all coefficients are statistically non-significant.
    4. Inconsistent Ordering (7 cases): Caused by strong model violations, such as unmodelled non-linearities, strong seasonal/trend non-stationarity, or external confounding (e.g., atmospheric CO2\text{CO}_2 MLCO2).
  10. Knowl 10 — Limitations and Failure Modes of LiNGAM

    limitation

    The LiNGAM framework is subject to several core limitations and failure modes:

    • Gaussianity: If disturbances eie_i are Gaussian, ICA mixing matrices are identifiable only up to orthogonal rotations, leaving the causal direction indistinguishable beyond the Markov equivalence class.
    • Unobserved Confounding: Latent common causes ff alter the disturbance structure to e=Gf+e′e = G f + e', introducing statistical dependencies among disturbance terms that violate mutual independence p(e1,…,em)=∏pi(ei)p(e_1, \dots, e_m) = \prod p_i(e_i).
    • Non-stationarity and Random Walks: Processes with unit roots (Xt=Xt−1+atX_t = X_{t-1} + a_t) create symmetrical forward-backward statistical representations, rendering temporal direction unidentifiable.
    • Non-convex ICA Optimization: Algorithms such as FastICA can become trapped in local optima, producing inconsistent estimates across different random initializations when sample size is low or model assumptions fail.

Coverage note — Omitted the explicit tensor and matrix derivative derivations for the asymptotic variance formulas (Appendices D and E) and the maximum likelihood derivation of the linear assignment objective (Appendix C), as these serve as mathematical derivations supporting the stated test statistics and algorithms rather than standalone conceptual contributions.

References

  1. 1.Y. Benjamini and Y. Hochberg. Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B, 57:289–300, 1995.
  2. 2.K. A. Bollen. Structural Equations with Latent Variables. John Wiley & Sons, 1989.
  3. 3.G. E. P. Box and G. M. Jenkins. Time Series Analysis: forecasting and control. Holden-Day, Oakland, California, USA, revised edition, 1976.
  4. 4.P. J. Brockwell and R. A. Davis. Time Series: Theory and Methods. Springer-Verlag, New York, USA, 1987.
  5. 5.M. W. Browne. Asymptotically distribution-free methods for the analysis of covariance structures. British Journal of Mathematical and Statistical Psychology, 9:665–672, 1984.
  6. 6.R. E. Burkard and E. Cela. Linear assignment problems and extensions. In P. M. Pardalos and D. Z. Du, editors, Handbook of Combinatorial Optimization - Supplement Volume A, pages 75–149. Kluwer, 1999.
  7. 7.J.-F. Cardoso and B. H. Laheld. Equivariant adaptive source separation. IEEE Trans. on Signal Processing, 44:3017–3030, 1996.
  8. 8.J.-F. Cardoso and A. Souloumiac. Blind beamforming for non Gaussian signals. IEE Proceedings-F, 140(6):362–370, 1993.
  9. 9.P. Comon. Independent component analysis – a new concept? Signal Processing, 36:287–314, 1994.
  10. 10.Y. Dodge and V. Rousson. On asymptotic properties of the correlation coefficient in the regression setting. The American Statistician, 55(1):51–54, 2001.
  11. 11.B. Efron and R. Tibshirani. An Introduction to the Bootstrap. Chapman & Hall, New York, 1993.
  12. 12.D. Geiger and D. Heckerman. Learning gaussian networks. In Proceedings of the 10th Annual Conference on Uncertainty in Artificial Intelligence (UAI-94), pages 235–243, 1994.
  13. 13.V. P. Godambe. Estimating functions. Oxford University Press, New York, 1991.
  14. 14.J. Himberg, A. Hyvärinen, and F. Esposito. Validating the independent components of neuroimaging time-series via clustering and visualization. Neuroimage, 22:1214–1222, 2004.
  15. 15.Y. Hochberg. A sharper Bonferroni procedure for multiple tests of significance. Biometrika, 4: 800–802, 1988.
  16. 16.Y. Hochberg and A. C. Tamhane. Multiple comparison procedures. John Wiley & Sons, New York, 1987.
  17. 17.S. Holm. A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6:65–70, 1979.
  18. 18.P. O. Hoyer, S. Shimizu, A. Hyvärinen, Y. Kano, and A. J. Kerminen. New permutation algorithms for causal discovery using ICA. In Proceedings of International Conference on Independent Component Analysis and Blind Signal Separation, Charleston, SC, USA, pages 115–122, 2006a.
  19. 19.P. O. Hoyer, S. Shimizu, and A. J. Kerminen. Estimation of linear, non-gaussian causal models in the presence of confounding latent variables. In Proc. the third European Workshop on Probabilistic Graphical Models (PGM2006), 2006b. In press.
  20. 20.L. Hu, P. M. Bentler, and Y. Kano. Can test statistics in covariance structure analysis be trusted? Psychological Bulletin, 112:351–362, 1992.
  21. 21.R. J. Hyndman. Time series data library, 2005. URL http://www-personal.buseco.monash.edu.au/~hyndman/TSDL/. [June 2005].
  22. 22.A. Hyvärinen. One-unit contrast functions for independent component analysis: A statistical analysis. In Neural Networks for Signal Processing VII (Proceedings of IEEE Workshop on Neural Networks for Signal Processing), pages 388–397, 1997.
  23. 23.A. Hyvärinen. Fast and robust fixed-point algorithms for independent component analysis. IEEE Trans. on Neural Networks, 10(3):626–634, 1999.
  24. 24.A. Hyvärinen, J. Karhunen, and E. Oja. Independent Component Analysis. Wiley Interscience, 2001.
  25. 25.M. Kawanabe and K. R. Müller. Estimating functions for blind separation when sources have variance dependencies. Journal of Machine Learning Research, 6:453–482, 2005.
  26. 26.National Statistics, 2005. URL http://www.statistics.gov.uk/. [June 2005].
  27. 27.J. Pearl. Causality: Models, Reasoning, and Inference. Cambridge University Press, 2000.
  28. 28.D. T. Pham and P. Garrat. Blind separation of mixture of independent sources through a quasi-maximum likelihood approach. Signal Processing, 45:1457–1482, 1997.
  29. 29.S. Shimizu, A. Hyvärinen, P. O. Hoyer, and Y. Kano. Finding a causal ordering via independent component analysis. Computational Statistics & Data Analysis, 50(11):3278–3293, 2006a.
  30. 30.S. Shimizu, A. Hyvärinen, Y. Kano, and P. O. Hoyer. Discovery of non-gaussian linear causal models using ICA. In Proc. the 21st Conference on Uncertainty in Artificial Intelligence (UAI-2005), pages 526–533, 2005.
  31. 31.S. Shimizu, A. Hyvärinen, Y. Kano, P. O. Hoyer, and A. J. Kerminen. Testing significance of mixing and demixing coefficients in ICA. In Proceedings of International Conference on Independent Component Analysis and Blind Signal Separation, Charleston, SC, USA, pages 901–908, 2006b.
  32. 32.S. Shimizu and Y. Kano. Use of non-normality in structural equation modeling: Application to direction of causation. Journal of Statistical Planning and Inference, 2006. In press.
  33. 33.R. J. Simes. An improved Bonferroni procedure for multiple tests of significance. Biometrika, 73: 751–754, 1986.
  34. 34.P. Spirtes, C. Glymour, and R. Scheines. Causation, Prediction, and Search, 2nd ed. MIT Press, 2000.
  35. 35.Statistical Software Information, 2005. URL http://www-unix.oit.umass.edu/~statdata/. [June 2005].
  36. 36.P. Tichavský, Z. Koldovský, and E. Oja. Performance analysis of the FastICA algorithm and Cramèr-Rao bounds for linear independent component analysis. IEEE Trans. on Signal Processing, 54(4):1189–1203, 2006.
  37. 37.K-H. Yuan and P. M. Bentler. Mean and covariance structure analysis: Theoretical and practical improvements. Journal of the American Statistical Association, 92(438):767–774, 1997.

Citation

MLA
Shimizu, S., et al. “A Linear Non-Gaussian Acyclic Model for Causal Discovery”. 2006, http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.163.2163.
APA
Shimizu, S., Hoyer, P. O., Hyvärinen, A., & Kerminen, A. (2006). A Linear Non-Gaussian Acyclic Model for Causal Discovery. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.163.2163
Chicago
Shimizu, S., P. O. Hoyer, A. Hyvärinen, and A. Kerminen. 2006. “A Linear Non-Gaussian Acyclic Model for Causal Discovery”. Preprint. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.163.2163.
Harvard
Shimizu, S. et al. (2006) “A Linear Non-Gaussian Acyclic Model for Causal Discovery”. Available at: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.163.2163.
Vancouver
1. Shimizu S, Hoyer PO, Hyvärinen A, Kerminen A (2006) A Linear Non-Gaussian Acyclic Model for Causal Discovery.

BibTeX

@article{shimizu2006linear,
  title = {A Linear Non-Gaussian Acyclic Model for Causal Discovery},
  author = {Shimizu, Shohei and Hoyer, Patrik O. and Hyvärinen, Aapo and Kerminen, Antti},
  year = {2006},
  url = {http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.163.2163}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/