Orange: data mining toolbox in python

Janez DemšarTomaž CurkAleš ErjavecČrt GorupTomaž HočevarMitar MilutinovičMartin MožinaMatija PolajnarMarko ToplakAnže Starič

article2013JMLR2,196 citations

Presents Orange, an open-source Python data mining library that enables rapid prototyping and interactive data analysis by combining high-level scriptable components with fast C++ implementations for core machine learning tasks.

Listen

Modern data analysis requires software environments that support rapid prototyping and interactive exploration without sacrificing computational efficiency. Scripting languages such as Python provide accessible syntax for exploratory analysis, but interpreted execution can lead to performance bottlenecks. The article presents Orange, a long-standing, open-source machine learning and data mining toolbox designed to provide a flexible, component-based scripting framework that balances execution speed and ease of use.

The toolbox uses a dual-layer architecture where performance-critical procedures are implemented in low-level code, while high-level modeling and assembly occur in Python. The core comprises nearly 200 C++ classes incorporating specialized external libraries for linear algebra, convex hulls, and support vector machines. The Python scripting interface sits on top of this layer, allowing developers and analysts to easily compose, wrap, and extend algorithms. Code reliability is validated using over 1,500 automated regression tests derived from practical examples, complemented by formal unit testing.

Orange provides a comprehensive suite of data mining capabilities, spanning data preprocessing, classification, regression, association rules, clustering, projections, and evaluation routines. Unlike standard Python scientific libraries that rely almost exclusively on numerical arrays, Orange uses rich data structures that natively support symbolic, string, and metadata attributes. This design retains variable names, accommodates symbolic learning methods, and automatically applies learned transformations directly to new data during prediction. Demonstration workflows in the article illustrate that users can rapidly build complex pipelines, such as stacked ensemble models and feature selection wrappers, in only a few lines of code.

These findings indicate that component-based toolboxes can reduce the development overhead and code complexity typically associated with advanced analytics. By allowing seamless transitions between low-level operations and high-level scripting, organizations can streamline exploratory data analysis and accelerate prototype deployment. Furthermore, the inclusion of metadata handling reduces data preparation errors and improves workflow maintainability compared to purely matrix-based tools.

Moving forward, the development roadmap includes migrating the library to Python 3 and transitioning from the standalone C++ core to modern numerical libraries, while maintaining backward compatibility for the scripting interface. The primary operational constraint noted is the system's reliance on Python 2.6 and 2.7. Stakeholders can confidently deploy the current version across major operating systems, but should account for planned core architecture updates when scheduling long-term integration and development work.

  • Paper: Scikit-learn: Machine Learning in Python, Fabian Pedregosa et al. (2011). Provides the foundational Python machine learning library design and standard API patterns that contextualize Orange's rich metadata and component-based data mining framework.
  • Paper: Torch7: A Matlab-like Environment for Machine Learning, R. Collobert et al. (2011). Examines the dual-layer design pattern of bridging low-level compiled routines with high-level scripting languages for machine learning execution efficiency.
  • Paper: MOA: Massive Online Analysis, A. Bifet et al. (2010). Introduces modular, component-based workbench architectures and experimental evaluation routines for machine learning data streams.
  • Paper: The CN2 Induction Algorithm, Peter Clark et al. (1989). Establishes classical rule induction algorithms that Orange directly supports within its symbolic and interpretable data mining toolset.
  • Paper: Induction of Decision Trees, J. R. Quinlan (1986). Presents top-down decision tree induction, serving as foundational background for the symbolic and discrete attribute modeling featured in Orange.
  • Paper: A Brief Introduction to Boosting, R. Schapire (1999). Details ensemble boosting principles that inform the construction and evaluation of complex model pipelines in data mining toolboxes.
  • Paper: Isolation-Based Anomaly Detection, Fei Tony Liu et al. (2012). Introduces isolation forest algorithms for anomaly detection, representing essential core modeling components utilized in comprehensive data mining suites.
Cover for Orange: data mining toolbox in python

Abstract

Orange is a machine learning and data mining suite for data analysis through Python scripting and visual programming. Here we report on the scripting part, which features interactive data analysis and component-based assembly of data mining procedures. In the selection and design of components, we focus on the flexibility of their reuse: our principal intention is to let the user write simple and clear scripts in Python, which build upon C++ implementations of computationally-intensive tasks. Orange is intended both for experienced users and programmers, as well as for students of data mining.

Table of Contents

  • 1. Introduction
  • 2. Toolbox Overview
  • 3. Scripting Examples
  • 4. Code Design
  • 5. Availability, Requirements and Plans for the Future
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Orange Component Hierarchy and Functional Scope

    model/method

    The Orange machine learning and data mining library is structured as a hierarchical component toolbox. Low-level routines (such as data filtering, probability assessment, and feature scoring) serve as building blocks for higher-level modeling and analytical algorithms. The toolbox is partitioned into several primary branches:

    • Data management and preprocessing: Data input/output, filtering, sampling, missing-value imputation, feature manipulation (discretization, continuization, normalization, scaling, scoring), and feature selection.
    • Classification: Supervised algorithms including decision trees, random forests, instance-based learning, Bayesian methods, rule induction, and support vector machines via integration with LIBSVM.
    • Regression: Linear and lasso regression, partial least squares regression, regression trees, regression forests, and multivariate adaptive regression splines (via the Earth library).
    • Association: Mining of association rules and frequent itemsets.
    • Ensembles: Wrapper modules providing bagging, boosting, forest trees, and stacked generalization.
    • Clustering: Unsupervised partitioning including kk-means and hierarchical clustering methods.
    • Evaluation: Cross-validation and sampling procedures, scoring metrics (such as Area Under the ROC Curve, AUC), and prediction reliability estimators.
    • Projections: Dimensionality reduction and manifold mapping techniques including principal component analysis (PCA), multi-dimensional scaling (MDS), and self-organizing maps (SOM).
  2. Knowl 2 — Symbolic and Metadata-Aware Data Representation in Orange

    model/method

    Orange employs a dedicated data model (Orange.data.Table and Orange.data.Domain) designed for interactive and symbolic data analysis rather than purely numerical arrays. Its characteristics include:

    • Heterogeneous feature support: Seamless combination of symbolic, categorical, string, and numerical variables alongside instance-level metadata.
    • Name-based indexing: Direct access to variables and attribute values via descriptive names rather than integer indices.
    • Variable transformation mapping: Variables encapsulate mapping functions that store transformations (e.g., continuization or normalization) applied during training data preparation. When a trained model makes predictions on new, unseen data, these stored mapping functions automatically reapply the required transformations without requiring manual data pipeline reconfiguration.
  3. Knowl 3 — Extensible Learner and Classifier Architecture in Orange

    model/method

    Orange provides an object-oriented scripting interface allowing users to implement custom learners and wrappers by subclassing Orange.classification.PyLearner and returning Orange.classification.PyClassifier objects. Any custom learning algorithm implements a __call__(self, data, weights=None) method that takes an input data table and optional sample weights, applies processing or base models, and returns an encapsulated predictor.

    An example of this design is a feature subset selection learner (FSSLearner) that wraps an underlying base learner, ranks features using information gain (Orange.feature.scoring.InfoGain), creates a reduced data domain containing only the top mm features, and delegates training to the base learner:

    class FSSLearner(Orange.classification.PyLearner):
        def __init__(self, base_learner, m=5):
            self.m = m
            self.base_learner = base_learner
    
        def __call__(self, data, weights=None):
            gain = Orange.feature.scoring.InfoGain()
            best = sorted(data.domain.features, key=lambda x: -gain(x, data))[:self.m]
            domain = Orange.data.Domain(best + [data.domain.class_var])
            new_data = Orange.data.Table(domain, data)
            model = self.base_learner(new_data, weights)
            return Orange.classification.PyClassifier(classifier=model)
    
  4. Knowl 4 — Two-Layer Implementation Architecture and Quality Assurance in Orange

    model/method

    Orange is designed with a hybrid two-tier architecture balancing execution speed and scripting flexibility:

    1. C++ Core Engine: Consists of nearly 200 self-contained C++ classes that handle fundamental data structures, preprocessing algorithms, and computationally intensive tasks without invoking Python callbacks. It incorporates external libraries including LIBSVM, LIBLINEAR, Earth (for spline regression), QHull (for convex hulls), and a subset of BLAS.
    2. Python Layer: Provides high-level abstractions, workflow assembly, and integration with the scientific Python ecosystem (using NumPy for linear algebra, NetworkX for graph structures, and Matplotlib for plotting), as well as the backend for graphical visual programming.
    3. Quality Assurance: System correctness and stability are maintained via automated execution of over 1,500 regression tests extracted directly from documentation code examples, combined with strict unit tests.
  5. Knowl 5 — Performance Results of Model Stacking and Feature Selection in Orange

    empirical result

    Cross-validation evaluations conducted via Orange's evaluation and scoring modules (Orange.evaluation.testing.cross_validation and Orange.evaluation.scoring.AUC) yielded the following Area Under the ROC Curve (AUC) metrics:

    • Titanic Survival Dataset (2,2012,201 instances):

      • Naive Bayes learner (NaiveLearner): AUC≈0.7149\text{AUC} \approx 0.7149
      • Support Vector Machine learner (SVMLearner): AUC≈0.7319\text{AUC} \approx 0.7319
      • Stacked learner combining Naive Bayes and SVM (StackedClassificationLearner): AUC≈0.7636\text{AUC} \approx 0.7636
      • Stacked learner evaluated exclusively on the female passenger subset (470470 instances): AUC≈0.8124\text{AUC} \approx 0.8124
    • Promoters Dataset (106106 instances, 5757 features):

      • Standard Naive Bayes classifier: AUC≈0.9330\text{AUC} \approx 0.9330
      • Wrapped Naive Bayes classifier with information gain feature subset selection (FSSLearner selecting top m=5m=5 features): AUC=0.9450\text{AUC} = 0.9450

Coverage note — No substantial contributed material was omitted. General background discussions on the history of Python in machine learning and routine distribution/licensing details were omitted in accordance with the extraction criteria.

References

  1. 1.D. Albanese, R. Visintainer, S. Merler, S. Riccadonna, G. Jurman, and C. Furlanello. mlpy: Machine learning Python. CoRR, abs/1202.6548, 2012.
  2. 2.C. B. Barber, D. P. Dobkin, and H. T. Huhdanpaa. The Quickhull algorithm for convex hulls. ACM Trans. on Mathematical Software, 22(4), 1996.
  3. 3.L. S. Blackford, A. Petitet, R. Pozo, K. Remington, R. C. Whaley, J. Demmel, J. Dongarra, I. Duff, S. Hammarling, and G. Henry. An updated set of basic linear algebra subprograms (BLAS). ACM Transactions on Mathematical Software, 28(2):135–151, 2002.
  4. 4.C.-C. Chang and C.-J. Lin. LIBSVM: A library for support vector machines. ACM Transactions on Intelligent Systems and Technology, 2:27:1–27:27, 2011.
  5. 5.Aric A. Hagberg, Daniel A. Schult, and Pieter J. Swart. Exploring network structure, dynamics, and function using NetworkX. In Proceedings of the 7th Python in Science Conference (SciPy2008), pages 11–15, Pasadena, CA USA, 2008.
  6. 6.J. D. Hunter. Matplotlib: A 2D graphics environment. Computing In Science & Engineering, 9(3): 90–95, 2007.
  7. 7.E. Jones, T. Oliphant, P. Peterson, et al. SciPy: Open source scientific tools for Python, 2001–. URL http://www.scipy.org/.
  8. 8.F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, and V. Dubourg. scikit-learn: Machine learning in Python. The Journal of Machine Learning Research, 12:2825–2830, 2011.
  9. 9.F. Rong-En, C.Kai-Wei, H. Cho-Jui, W. Xiang-Rui, and L. Chih-Jen. LIBLINEAR: A library for large linear classification. Journal of Machine Learning Research, 9:1871–1874, 2008.
  10. 10.T. Schaul, J. Bayer, D. Wierstra, Y. Sun, M. Felder, F. Sehnke, T. Rückstieß, and J. Schmidhuber. PyBrain. Journal of Machine Learning Research, 11:743–746, 2010.
  11. 11.D. H. Wolpert. Stacked generalization. Neural Networks, 5(2):241–259, 1992.

Citation

MLA
Demšar, J., et al. “Orange: Data Mining Toolbox in Python”. Journal of Machine Learning Research, vol. 14, no. 71, 2013, pp. 2349–53, https://www.jmlr.org/papers/v14/demsar13a.html.
APA
Demšar, J., Curk, T., Erjavec, A., Gorup, Č., Hočevar, T., Milutinovič, M., Možina, M., Polajnar, M., Toplak, M., Starič, A., Štajdohar, M., Umek, L., Žagar, L., Žbontar, J., Žitnik, M., & Zupan, B. (2013). Orange: Data Mining Toolbox in Python. Journal of Machine Learning Research, 14(71), 2349–2353. https://www.jmlr.org/papers/v14/demsar13a.html
Chicago
Demšar, J., T. Curk, A. Erjavec, et al. 2013. “Orange: Data Mining Toolbox in Python”. Journal of Machine Learning Research 14 (71): 2349–53. https://www.jmlr.org/papers/v14/demsar13a.html.
Harvard
Demšar, J. et al. (2013) “Orange: Data Mining Toolbox in Python”, Journal of Machine Learning Research, 14(71), pp. 2349–2353. Available at: https://www.jmlr.org/papers/v14/demsar13a.html.
Vancouver
1. Demšar J, Curk T, Erjavec A, et al (2013) Orange: Data Mining Toolbox in Python. Journal of Machine Learning Research 14:2349–2353

BibTeX

@article{JMLR:v14:demsar13a,
  author  = {Janez Dem{\v{s}}ar and Toma{\v{z}} Curk and Ale{\v{s}} Erjavec and {\v{C}}rt Gorup and Toma{\v{z}} Ho{\v{c}}evar and Mitar Milutinovi{\v{c}} and Martin Mo{\v{z}}ina and Matija Polajnar and Marko Toplak and An{\v{z}}e Stari{\v{c}} and Miha {\v{S}}tajdohar and Lan Umek and Lan {\v{Z}}agar and Jure {\v{Z}}bontar and Marinka {\v{Z}}itnik and Bla{\v{z}} Zupan},
  title   = {Orange: Data Mining Toolbox in Python},
  journal = {Journal of Machine Learning Research},
  year    = {2013},
  volume  = {14},
  number  = {71},
  pages   = {2349--2353},
  url     = {http://jmlr.org/papers/v14/demsar13a.html}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/