Learning Bayesian Networks with the bnlearn R Package

Marco Scutari

article2009Journal of Statistical Software (2010), 35(3), 1-222,030 citations

Presents the bnlearn R package, providing a unified and computationally scalable framework for learning the structure of discrete and continuous Bayesian networks through constraint-based and score-based algorithms.

Listen

Modern data analytics across healthcare, genetics, and operational monitoring increasingly depend on understanding complex probabilistic relationships within high-dimensional datasets. Bayesian networks represent these interactions as directed acyclic graphs, mapping direct dependencies while breaking down complex global probability distributions into manageable local calculations. However, identifying the true underlying network structure from massive datasets is computationally intensive, and existing software implementations typically restrict analysts by tying specific learning algorithms to rigid statistical tests or data formats.

The article demonstrates the implementation and capabilities of bnlearn, a software package developed for the R environment that provides a flexible, highly optimized framework for learning Bayesian network structures from both discrete and continuous data. The author set out to evaluate whether decoupling learning algorithms from underlying statistical metrics, combined with algorithmic enhancements and parallel execution, delivers a more versatile and scalable solution for network structure learning.

To evaluate the system, the article conducts comparative analyses on both synthetic and empirical benchmarks, including a 37-variable medical monitoring benchmark containing 20,000 observations (the ALARM network) and an educational dataset of 88 students evaluated across five continuous examination subjects. The evaluation covers both constraint-based approaches—which map networks using conditional independence tests—and score-based greedy search algorithms. The package incorporates computational optimizations, such as backtracking to reduce redundant statistical testing by roughly half, caching to speed score comparisons, and parallel computing capabilities.

The findings show that constraint-based algorithms (including Grow-Shrink and variants of the Incremental Association algorithm) and greedy hill-climbing consistently recover the true structural relationships within a few arcs. On the ALARM dataset, the default optimized constraint-based algorithms completed execution in approximately 13 to 19 seconds, while hill-climbing required roughly 72 seconds. Crucially, the choice of statistical metric heavily impacts accuracy: employing Monte Carlo permutation tests on sparse discrete data reduced structural learning errors from twelve missed connections under standard parametric tests down to only five. Furthermore, the analysis shows that integrating prior domain expertise through whitelisting or blacklisting specific connections effectively resolves structural ambiguities and optimizes network orientation.

These results demonstrate that analysts can achieve substantial performance gains and lower computational costs without sacrificing model accuracy. The capacity to mix arbitrary statistical criteria with different learning algorithms allows teams to tailor models directly to the mathematical properties of their data, mitigating risks of misidentification in sparse or high-dimensional environments. For enterprise and research workflows, this flexibility reduces development cycle times while supporting rigorous causal inference and predictive modeling.

Organizations analyzing complex multivariate data should adopt flexible workflows that decouple network search algorithms from independence tests, prioritizing permutation-based tests when dealing with small sample sizes or sparse contingency tables. In addition, practitioners should systematically incorporate known operational constraints and expert rules into the model configuration to guide edge direction and eliminate spurious dependencies. Users should note that while the package provides advanced structural learning and diagnostics, the current implementation handles discrete and continuous datasets separately under multinomial or multivariate normal assumptions, rather than mixed hybrid networks within a single model.

arXiv: 0908.3817cran/bnlearn

No sufficiently relevant recommendations were found.

Cover for Learning Bayesian Networks with the bnlearn R Package

Abstract

bnlearn is an R package which includes several algorithms for learning the structure of Bayesian networks with either discrete or continuous variables. Both constraint-based and score-based algorithms are implemented, and can use the functionality provided by the snow package to improve their performance via parallel computing. Several network scores and conditional independence algorithms are available for both the learning algorithms and independent use. Advanced plotting options are provided by the Rgraphviz package.

Table of Contents

  • 1 Introduction
  • 2 Bayesian networks
  • 3 Structure learning algorithms
  • 4 Package implementation
  • 4.1 Structure learning algorithms
  • 4.2 Conditional independence tests
  • 4.3 Network scores
  • 4.4 Arc whitelisting and blacklisting
  • 5 A simple example
  • 5.1 Loading the package
  • 5.2 Learning a Bayesian network from data
  • 5.3 Network analysis and manipulation
  • 5.4 Debugging utilities and diagnostics
  • 6 Practical examples
  • 6.1 The ALARM network
  • 6.2 The examination marks data set
  • 7 Other packages for learning Bayesian networks
  • 8 Conclusions
  • References

Knowls

  1. Knowl 1 — Structure Learning Algorithm Suite and Execution Modes in bnlearn

    model/method

    The bnlearn package provides several structure learning algorithms for Bayesian networks with discrete or continuous variables, along with multiple algorithmic execution modes:

    1. Constraint-Based Algorithms:

      • Grow-Shrink (gs): Learns the network structure based on the Grow-Shrink Markov blanket detection algorithm by performing a forward inclusion phase followed by a backward false-positive removal phase.
      • Incremental Association (iamb): Learns the Markov blanket via a two-phase scheme (forward selection followed by false-positive elimination) to reconstruct the graph skeleton.
      • Fast Incremental Association (fast.iamb): A variant of IAMB that employs speculative stepwise forward selection to reduce the total number of conditional independence tests.
      • Interleaved Incremental Association (inter.iamb): A variant of IAMB that interleaves forward selection and backward elimination steps to prevent false positives during Markov blanket estimation.
      • Max-Min Parents and Children (mmpc): A forward selection algorithm for neighborhood detection based on maximizing the minimum association measure observed with any subset of previously selected variables; it learns the undirected skeleton of the Bayesian network without orienting edges.
    2. Score-Based Algorithms:

      • Hill-Climbing (hc): A greedy local search across the directed acyclic graph (DAG) space. It supports random restarts (restart), network perturbation operations (perturb), and preseeded initial structures (start) to escape local score maxima.
    3. Implementations and Optimizations:

      • Optimized (Default): Uses backtracking across node pairs in constraint-based learning to roughly halve the required number of independence tests. In score-based learning, it utilizes score caching, score decomposability, and score equivalence to avoid redundant score evaluations.
      • Unoptimized: Faithfully runs the textbook algorithm without caching or backtracking, useful for benchmarking.
      • Parallel: Uses the snow R package to parallelize conditional independence tests across multi-core machines or clusters.
  2. Knowl 2 — Discrete Conditional Independence Tests

    equation

    For discrete random variables XX and YY given a conditioning set of discrete variables ZZ, conditional independence tests evaluate the 3-dimensional contingency table of observed cell counts {nijk}\{n_{ijk}\} for i=1,…,Ri = 1, \dots, R (levels of XX), j=1,…,Cj = 1, \dots, C (levels of YY), and k=1,…,Lk = 1, \dots, L (configurations of ZZ), where n=∑i,j,knijkn = \sum_{i,j,k} n_{ijk} is the sample size, ni+k=∑j=1Cnijkn_{i+k} = \sum_{j=1}^C n_{ijk}, n+jk=∑i=1Rnijkn_{+jk} = \sum_{i=1}^R n_{ijk}, and n++k=∑i=1R∑j=1Cnijkn_{++k} = \sum_{i=1}^R \sum_{j=1}^C n_{ijk}.

    • Mutual Information Test (mi, mc-mi): MI(X,Y∣Z)=∑i=1R∑j=1C∑k=1Lnijknlog⁡(nijkn++kni+kn+jk)MI(X, Y \mid Z) = \sum_{i=1}^{R} \sum_{j=1}^{C} \sum_{k=1}^{L} \frac{n_{ijk}}{n} \log \left( \frac{n_{ijk} n_{++k}}{n_{i+k} n_{+jk}} \right) The statistic 2n⋅MI(X,Y∣Z)2n \cdot MI(X, Y \mid Z) is proportional to the log-likelihood ratio test statistic G2G^2 and deviance. Under the null hypothesis of conditional independence, it asymptotically follows a χ2\chi^2 distribution with (R−1)(C−1)L(R-1)(C-1)L degrees of freedom (mi), or can be tested nonparametrically via Monte Carlo permutations (mc-mi).

    • Pearson's X2X^2 Test (x2, mc-x2): X2(X,Y∣Z)=∑i=1R∑j=1C∑k=1L(nijk−mijk)2mijk,where mijk=ni+kn+jkn++kX^2(X, Y \mid Z) = \sum_{i=1}^{R} \sum_{j=1}^{C} \sum_{k=1}^{L} \frac{(n_{ijk} - m_{ijk})^2}{m_{ijk}}, \quad \text{where } m_{ijk} = \frac{n_{i+k} n_{+jk}}{n_{++k}} Evaluated against an asymptotic χ2\chi^2 distribution with (R−1)(C−1)L(R-1)(C-1)L degrees of freedom (x2) or via Monte Carlo permutation testing (mc-x2).

    • Fast Mutual Information Heuristic (fmi): Computes MI(X,Y∣Z)MI(X, Y \mid Z) as above, but truncates the statistic to zero if n(R−1)(C−1)L<5\frac{n}{(R-1)(C-1)L} < 5, avoiding unreliable tests in sparse contingency tables.

    • AIC-Based Independence Test (aict): Rejects the null hypothesis of conditional independence when: MI(X,Y∣Z)>(R−1)(C−1)LnMI(X, Y \mid Z) > \frac{(R - 1)(C - 1)L}{n} which corresponds to an increase in the global network AIC score upon adding an edge between XX and YY.

  3. Knowl 3 — Continuous Conditional Independence Tests Based on Partial Correlation

    equation

    For continuous random variables XX and YY conditioned on a subset of continuous variables ZZ under the assumption of multivariate normality, conditional independence tests evaluate functions of the sample partial correlation coefficient ρXY∣Z\rho_{XY \mid Z}:

    • Linear Correlation Student's tt Test (cor, mc-cor): Tests ρXY∣Z=0\rho_{XY \mid Z} = 0 using Student's tt distribution with n−∣Z∣−2n - |Z| - 2 degrees of freedom (cor) or a Monte Carlo permutation test (mc-cor), where nn is sample size and ∣Z∣|Z| is the cardinality of ZZ.

    • Fisher's ZZ Test (zf, mc-zf): Applies Fisher's transformation to the partial correlation: Z(X,Y∣Z)=12n−∣Z∣−3log⁡(1+ρXY∣Z1−ρXY∣Z)Z(X, Y \mid Z) = \frac{1}{2} \sqrt{n - |Z| - 3} \log \left( \frac{1 + \rho_{XY \mid Z}}{1 - \rho_{XY \mid Z}} \right) Under the null hypothesis of conditional independence, Z(X,Y∣Z)Z(X, Y \mid Z) asymptotically follows a standard normal distribution N(0,1)\mathcal{N}(0, 1) (zf), or can be tested via Monte Carlo permutation (mc-zf).

    • Continuous Mutual Information Test (mi-g, mc-mi-g): MIg(X,Y∣Z)=−12log⁡(1−ρXY∣Z2)MI_g(X, Y \mid Z) = -\frac{1}{2} \log \left( 1 - \rho_{XY \mid Z}^2 \right) The scaled statistic 2n⋅MIg(X,Y∣Z)2n \cdot MI_g(X, Y \mid Z) is asymptotically distributed as χ2\chi^2 with 1 degree of freedom (mi-g) or evaluated via Monte Carlo permutation (mc-mi-g).

  4. Knowl 4 — Network Scoring Functions for Discrete and Continuous Bayesian Networks

    equation

    For a Bayesian network structure GG on variables V={X1,…,Xv}\mathbf{V} = \{X_1, \dots, X_v\}, score-based learning algorithms evaluate candidate graphs using discrete or continuous network scores:

    1. Discrete Scores:

      • Log-Likelihood (loglik) and Likelihood (lik): log⁡L(X1,…,Xv)=∑i=1v∑j=1Li∑k=1Rinijklog⁡(nijk∑k′=1Rinijk′)\log L(X_1, \dots, X_v) = \sum_{i=1}^v \sum_{j=1}^{L_i} \sum_{k=1}^{R_i} n_{ijk} \log \left( \frac{n_{ijk}}{\sum_{k'=1}^{R_i} n_{ijk'}} \right) where RiR_i is the number of states of node XiX_i, LiL_i is the number of parent configurations of XiX_i, and nijkn_{ijk} is the count of observations where Xi=kX_i = k and its parents ΠXi\Pi_{X_i} are in state jj.
      • Akaike Information Criterion (aic) and Bayesian Information Criterion (bic): AIC=log⁡L(X1,…,Xv)−d,BIC=log⁡L(X1,…,Xv)−d2log⁡nAIC = \log L(X_1, \dots, X_v) - d, \qquad BIC = \log L(X_1, \dots, X_v) - \frac{d}{2} \log n where d=∑i=1v(Ri−1)Lid = \sum_{i=1}^v (R_i - 1) L_i is the number of free parameters in the network and nn is the total sample size. BIC is equivalent to the Minimum Description Length (MDL) score. Both AIC and BIC are score-equivalent.
      • Bayesian Dirichlet Equivalent Score (bde): The logarithm of the Dirichlet posterior density under uniform prior hyperparameters, satisfying score equivalence across Markov equivalent DAGs.
      • K2 Score (k2): The logarithm of the Dirichlet posterior density with uniform parameter priors: K2=∏i=1vK2(Xi),K2(Xi)=∏j=1Li(Ri−1)!(∑k=1Rinijk+Ri−1)!∏k=1Rinijk!K2 = \prod_{i=1}^{v} K2(X_i), \quad K2(X_i) = \prod_{j=1}^{L_i} \frac{(R_i - 1)!}{\left( \sum_{k=1}^{R_i} n_{ijk} + R_i - 1 \right)!} \prod_{k=1}^{R_i} n_{ijk}! Unlike bde, aic, and bic, the K2 score is not score-equivalent.
    2. Continuous Score:

      • Bayesian Gaussian Equivalent Score (bge): A score-equivalent Gaussian posterior density assuming a Wishart prior distribution for continuous multivariate normal networks.
  5. Knowl 5 — Arc Whitelisting and Blacklisting Semantics

    model/method

    In bnlearn, prior structural constraints can be enforced across all learning algorithms via the whitelist and blacklist arguments, which take sets of directed arc pairs (A,B)(A, B) representing A→BA \to B:

    • Bidirectional Whitelist: If both A→BA \to B and B→AB \to A are whitelisted, an edge connecting AA and BB must exist in the learned network, but its orientation (A→BA \to B, B→AB \to A, or undirected A−BA - B) is left to the learning algorithm.
    • Unidirectional Whitelist: If A→BA \to B is whitelisted but B→AB \to A is not, the directed arc A→BA \to B is strictly enforced in the learned graph. This automatically blacklists both the reverse arc B→AB \to A and the undirected edge A−BA - B.
    • Bidirectional Blacklist: If both A→BA \to B and B→AB \to A are blacklisted, no edge between AA and BB can appear in the network, precluding A→BA \to B, B→AB \to A, and A−BA - B.
    • Unidirectional Blacklist: If A→BA \to B is blacklisted but B→AB \to A is not, the directed arc A→BA \to B and the undirected edge A−BA - B are prohibited, but the reverse arc B→AB \to A remains eligible for inclusion.
    • Precedence Rule: If an arc is specified in both whitelist and blacklist, the whitelist takes precedence and the arc is removed from the blacklist.
  6. Knowl 6 — Structure Learning Performance on the ALARM Network

    data/table

    The ALARM network benchmark evaluates constraint-based and score-based learning algorithms on a sample of n=20,000n = 20{,}000 observations generated from the true ALARM network (3737 discrete variables with 2 to 4 levels each, and 4646 true arcs). Benchmarks were executed on an Intel Core 2 Duo machine with 1GB of RAM.

    Metric gs iamb fast.iamb inter.iamb hc
    Independence tests / Network comparisons 1727 2874 2398 3106 2841
    Learned arcs (directed / undirected) 42 (29/13) 43 (29/14) 45 (32/13) 43 (30/13) 53 (53/0)
    Execution time (seconds) 13.54360 17.63735 14.80890 18.59825 72.38705

    All constraint-based algorithms (gs, iamb, fast.iamb, inter.iamb) successfully reconstruct the ALARM network within a few arcs and directions, requiring roughly 1,7001{,}700 to 3,1003{,}100 conditional independence tests and executing in 13.5–18.6 seconds. gs requires the fewest tests (1727) and has the lowest runtime (13.54 s). Hill-climbing (hc with BIC) performs 2841 score comparisons, completely directs all 53 learned arcs (including several extraneous arcs beyond the true 46), and requires substantially longer computation time (72.39 s).

  7. Knowl 7 — Accuracy Gains of Monte Carlo Permutation Tests on Sparse Discrete Data

    empirical result

    When learning discrete Bayesian network structures from data with high-cardinality conditioning sets, the resulting conditional contingency tables frequently suffer from severe sparsity, causing asymptotic χ2\chi^2 distribution assumptions to break down.

    In structure learning on the 37-variable ALARM network data (n=20,000n = 20{,}000) using the Grow-Shrink (gs) algorithm:

    • Using the asymptotic parametric Pearson's X2X^2 test (test = "x2") results in 1212 true network arcs being missed.
    • Using the Monte Carlo permutation Pearson's X2X^2 test (test = "mc-x2" with B=10,000B = 10{,}000 permutations) misses only 55 true arcs.

    Nonparametric permutation tests substantially improve edge recovery accuracy over parametric asymptotic tests in discrete Bayesian networks when data sparsity violates asymptotic assumptions.

  8. Knowl 8 — Continuous Structure Learning and Conditional Independence on Examination Marks Data

    empirical result

    Structure learning on the continuous marks dataset (n=88n = 88 student scores across Mechanics MECH, Vectors VECT, Algebra ALG, Analysis ANL, and Statistics STAT) demonstrates consistency across different continuous conditional independence tests and structure learning algorithms under multivariate normality:

    1. Conditional Independence Tests: Testing the conditional independence of (MECH,VECT)(\text{MECH}, \text{VECT}) and (ANL,STAT)(\text{ANL}, \text{STAT}) given ALG\text{ALG} yields non-significant pp-values across all parametric and nonparametric test formulations:

      • MECH⊥ ⁣ ⁣ ⁣⊥ANL∣ALG\text{MECH} \perp \!\!\! \perp \text{ANL} \mid \text{ALG}: Student's tt linear correlation test yields ρ=0.0352\rho = 0.0352 (df=85df = 85, p=0.7459p = 0.7459).
      • STAT⊥ ⁣ ⁣ ⁣⊥VECT∣ALG\text{STAT} \perp \!\!\! \perp \text{VECT} \mid \text{ALG}:
        • Student's tt linear correlation (cor): ρ=0.0527\rho = 0.0527 (df=85df = 85, p=0.628p = 0.628).
        • Fisher's ZZ asymptotic test (zf): p=0.6289p = 0.6289.
        • Monte Carlo permutation linear correlation (mc-cor): p=0.6332p = 0.6332.
        • Continuous mutual information asymptotic test (mi-g): p=0.6209p = 0.6209.
        • Continuous mutual information Monte Carlo test (mc-mi-g): p=0.6226p = 0.6226.
    2. Learned Network Structures: All constraint-based algorithms (gs, iamb, fast.iamb, inter.iamb), neighborhood discovery (mmpc), and score-based search (hc) learn consistent graphical models showing that ANL and STAT are conditionally independent of MECH and VECT given ALG, identifying algebra as the central bridging skill across the examination subjects.

Coverage note — Standard theoretical definitions of Bayesian network factorizations (Pearl 1988, Korb and Nicholson 2004), the generic Inductive Causation (IC) steps, and basic syntax for R helper functions (such as `empty.graph` and `modelstring`) were omitted as standard background and software syntactic wrappers.

References

  1. 1.Acid S, de Campos LM, Fernandez-Luna J, Rodriguez S, Rodriguez J, Salcedo J (2004). "A Comparison of Learning Algorithms for Bayesian Networks: A Case Study Based on Data from An Emergency Medical Service." Artificial Intelligence in Medicine, 30, 215–232.
  2. 2.Bach FR, Jordan MI (2003). "Learning Graphical Models with Mercer Kernels." In "Advances in Neural Information Processing Systems (NIPS) 15," pp. 1009–1016. MIT Press.
  3. 3.Beinlich I, Suermondt HJ, Chavez RM, Cooper GF (1989). "The ALARM Monitoring System: A Case Study with Two Probabilistic Inference Techniques for Belief Networks." In "Proceedings of the 2nd European Conference on Artificial Intelligence in Medicine," pp. 247–256. Springer-Verlag. URL http://www.cs.huji.ac.il/labs/compbio/Repository/Datasets/alarm/alarm.htm.
  4. 4.Boettcher SG, Dethlefsen C (2003). "deal: A Package for Learning Bayesian Networks." Journal of Statistical Software, 8(20), 1–40. ISSN 1548-7660. URL http://www.jstatsoft.org/v08/i20.
  5. 5.Chickering DM (1995). "A Transformational Characterization of Equivalent Bayesian Network Structures." In "UAI '95: Proceedings of the Eleventh Annual Conference on Uncertainty in Artificial Intelligence," pp. 87–98. Morgan Kaufmann.
  6. 6.Chickering DM (2002). "Optimal Structure Identification with Greedy Search." Journal of Machine Learning Research, 3, 507–554.
  7. 7.Cooper GF, Herskovits E (1992). "A Bayesian Method for the Induction of Probabilistic Networks from Data." Machine Learning, 9(4), 309–347.
  8. 8.Daly R, Shen Q (2007). "Methods to Accelerate the Learning of Bayesian Network Structures." In "Proceedings of the 2007 UK Workshop on Computational Intelligence," Imperial College, London.
  9. 9.Edwards DI (2000). Introduction to Graphical Modelling. Springer.
  10. 10.Friedman N, Linial M, Nachman I (2000). "Using Bayesian Networks to Analyze Expression Data." Journal of Computational Biology, 7, 601–620.
  11. 11.Friedman N, Pe'er D, Nachman I (1999). "Learning Bayesian Network Structure from Massive Datasets: The "Sparse Candidate" Algorithm." In "Proceedings of Fifteenth Conference on Uncertainty in Artificial Intelligence (UAI)," pp. 206–221. Morgan Kaufmann.
  12. 12.Geiger D, Heckerman D (1994). "Learning Gaussian Networks." Technical report, Microsoft Research, Redmond, Washington. Available as Technical Report MSR-TR-94-10.
  13. 13.Gentleman R, Whalen E, Huber W, Falcon S (2010). graph: A package to handle graph data structures. R package version 1.26.0.
  14. 14.Gentry J, Long L, Gentleman R, Falcon S, Hahne F, Sarkar D (2010). Rgraphviz: Provides Plotting Capabilities for R Graph Objects. R package version 1.26.0.
  15. 15.Good P (2005). Permutation, Parametric and Bootstrap Tests of Hypotheses. Springer, 3rd edition.
  16. 16.Heckerman D, Geiger D, Chickering DM (1995). "Learning Bayesian Networks: The Combination of Knowledge and Statistical Data." Machine Learning, 20(3), 197–243. Available as Technical Report MSR-TR-94-09.
  17. 17.Højsgaard S (2010). gRain: Graphical Independence Networks. R package version 0.8.5.
  18. 18.Højsgaard S, Dethlefsen C, Bowsher C (2010). gRbase: A package for graphical modelling in R. R package version 1.3.4.
  19. 19.Højsgaard S, Lauritzen SL (2008). gRc: Inference in Graphical Gaussian Models with Edge and Vertex Symmetries. R package version 0.2.2.
  20. 20.Holmes DE, Jain LC (eds.) (2008). Innovations in Bayesian Networks: Theory and Applications, volume 156 of Studies in Computational Intelligence. Springer.
  21. 21.Kalisch M, Bühlmann P (2007). "Estimating High-Dimensional Directed Acyclic Graphs with the PC-Algorithm." Journal of Machine Learning Research, 8, 613–66.
  22. 22.Korb K, Nicholson A (2004). Bayesian Artificial Intelligence. Chapman and Hall.
  23. 23.Kullback S (1959). Information Theory and Statistics. Wiley.
  24. 24.Lam W, Bacchus F (1994). "Learning Bayesian Belief Networks: An Approach Based on the MDL Principle." Computational Intelligence, 10, 269–293.
  25. 25.Legendre P (2000). "Comparison of Permutation Methods for the Partial Correlation and Partial Mantel Tests." Journal of Statistical Computation and Simulation, 67, 37–73.
  26. 26.Mardia KV, Kent JT, Bibby JM (1979). Multivariate Analysis. Academic Press.
  27. 27.Margaritis D (2003). Learning Bayesian Network Model Structure from Data. Ph.D. thesis, School of Computer Science, Carnegie-Mellon University, Pittsburgh, PA. Available as Technical Report CMU-CS-03-153.
  28. 28.Moore A, Wong W (2003). "Optimal Reinsertion: A New Search Operator for Accelerated and More Accurate Bayesian Network Structure Learning." In "Proceedings of the 20th International Conference on Machine Learning (ICML '03)," pp. 552–559. AAAI Press.
  29. 29.Neapolitan RE (2003). Learning Bayesian Networks. Prentice Hall.
  30. 30.Pearl J (1988). Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference. Morgan Kaufmann.
  31. 31.R Development Core Team (2009). R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria. ISBN 3-900051-07-0, URL http://www.R-project.org.
  32. 32.Rissanen J (1978). "Modeling by Shortest Data Description." Automatica, 14, 465–471.
  33. 33.Spirtes P, Glymour C, Scheines R (2001). Causation, Prediction and Search. MIT Press.
  34. 34.Tierney L, Rossini AJ, Li N, Sevcikova H (2008). snow: Simple Network of Workstations. R package version 0.3-3.
  35. 35.Tsamardinos I, Aliferis CF, Statnikov A (2003). "Algorithms for Large Scale Markov Blanket Discovery." In "Proceedings of the Sixteenth International Florida Artificial Intelligence Research Society Conference," pp. 376–381. AAAI Press.
  36. 36.Tsamardinos I, Brown LE, Aliferis CF (2006). "The Max-Min Hill-Climbing Bayesian Network Structure Learning Algorithm." Machine Learning, 65(1), 31–78.
  37. 37.Verma TS, Pearl J (1991). "Equivalence and Synthesis of Causal Models." Uncertainty in Artificial Intelligence, 6, 255–268.
  38. 38.Whittaker J (1990). Graphical Models in Applied Multivariate Statistics. Wiley.
  39. 39.Witten IH, Frank E (2005). Data Mining: Practical Machine Learning Tools and Techniques. Morgan Kaufmann, 2nd edition.
  40. 40.Yaramakala S, Margaritis D (2005). "Speculative Markov Blanket Discovery for Optimal Feature Selection." In "ICDM '05: Proceedings of the Fifth IEEE International Conference on Data Mining," pp. 809–812. IEEE Computer Society, Washington, DC, USA.

Citation

MLA
Scutari, M. “Learning Bayesian Networks with the Bnlearn R Package”. Journal of Statistical Software (2010), 35(3), 1-22, 2009, http://arxiv.org/abs/0908.3817v2.
APA
Scutari, M. (2009). Learning Bayesian Networks with the bnlearn R Package. Journal of Statistical Software (2010), 35(3), 1-22. http://arxiv.org/abs/0908.3817v2
Chicago
Scutari, M. 2009. “Learning Bayesian Networks with the Bnlearn R Package”. Journal of Statistical Software (2010), 35(3), 1-22. http://arxiv.org/abs/0908.3817v2.
Harvard
Scutari, M. (2009) “Learning Bayesian Networks with the bnlearn R Package”, Journal of Statistical Software (2010), 35(3), 1-22 [Preprint]. Available at: http://arxiv.org/abs/0908.3817v2.
Vancouver
1. Scutari M (2009) Learning Bayesian Networks with the bnlearn R Package. Journal of Statistical Software (2010), 35(3), 1-22

BibTeX

@article{scutari2009learning,
  title = {Learning Bayesian Networks with the bnlearn R Package},
  author = {Scutari, Marco},
  year = {2009},
  journal = {Journal of Statistical Software (2010), 35(3), 1-22},
  url = {http://arxiv.org/abs/0908.3817v2},
  eprint = {0908.3817}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF