CodaLab Competitions: An Open Source Platform to Organize Scientific Challenges

Adrien PavaoIsabelle GuyonAnne-Catherine LetournelDinh-Tuan TranXavier BaróHugo Jair EscalanteSergio EscaleraTyler ThomasZhen Xu

article2023JMLR163 citations

Presents an open-source framework that enables researchers to design, host, and scale reproducible machine learning competitions using custom Docker environments, multi-phase evaluations, and external compute resources.

Listen

Data science and machine learning competitions serve as powerful drivers of technological breakthroughs, yet organizing these challenges often requires substantial computing resources, strict data privacy safeguards, and specialized infrastructure. Many existing challenge platforms are proprietary, lack customization, or cannot support complex computational workflows. The article evaluates and demonstrates CodaLab Competitions, an open-source web platform designed to facilitate crowdsourced problem-solving through customizable, reproducible, and secure scientific challenges.

The article describes the system architecture and operational model of CodaLab Competitions, drawing on its deployment across hundreds of scientific challenges in disciplines such as computer vision, physics, and natural language processing. The platform is built using standard open-source tools, including the Django web framework, PostgreSQL databases, MinIO storage, RabbitMQ message queues, and Docker containers. The evaluation highlights how this infrastructure coordinates participants, automated scoring scripts, and backend compute workers to evaluate competitive entries.

The key findings highlight four main operational strengths. First, the platform supports both prediction-based result submissions and full server-side code execution in isolated Docker containers, which standardizes testing conditions and ensures reproducibility. Second, it utilizes a modular compute worker architecture that allows challenge organizers to attach their own external CPU or GPU hardware to dedicated queues. This mechanism shifts significant computing costs away from the primary host. Third, keeping datasets directly on external compute workers enables private, secure evaluations on confidential information, making the system suitable for sensitive medical and industrial domains. Fourth, the platform provides advanced configuration options, including multi-phase competitions with automatic code forwarding and multi-score leaderboards.

These capabilities mean that research institutions and enterprises can organize large-scale machine learning competitions without paying high commercial platform fees or compromising proprietary data. The decentralized execution model lowers operational risks and prevents server overload by distributing expensive computing tasks across organizer-supplied infrastructure. Furthermore, containerized code execution ensures fairness, as all participant models run against identical environments.

Organizations planning data challenges should consider adopting CodaLab Competitions or self-hosting instances using its open-source codebase under the Apache 2.0 license. To handle the platform's growing user base, future development efforts should focus on decentralized file storage, cross-instance platform federation, dedicated benchmark hosting, and formal security accreditations for handling protected health data. The findings are based on extensive real-world usage across academic and industrial challenges, although real-world performance and security remain dependent on each organizer's specific infrastructure deployment, network firewalls, and storage configurations.

Pavao et al (2023).pdf

No sufficiently relevant recommendations were found.

Cover for CodaLab Competitions: An Open Source Platform to Organize Scientific Challenges

Abstract

CodaLab Competitions is an open source web platform designed to help data scientists and research teams to crowd-source the resolution of machine learning problems through the organization of competitions, also called challenges or contests. CodaLab Competitions provides useful features such as multiple phases, results and code submissions, multi-score leaderboards, and jobs running inside Docker containers. The platform is very flexible and can handle large scale experiments, by allowing organizers to upload large datasets and provide their own CPU or GPU compute workers.

Table of Contents

  • 1. Introduction
  • 2. Key Concepts and Features
  • 3. Technical Aspects
  • 4. Conclusion and Ongoing Developments
  • Acknowledgments
  • Authors' Contributions
  • References

Knowls

  1. Knowl 1 — CodaLab Competitions System Architecture and Backend Components

    model/method

    CodaLab Competitions is structured around three primary functional blocks:

    1. Front-End and API Core: Built using Python's Django web framework. The front-end renders user-facing views via Django's templating engine, while the back-end API is implemented with the Django REST Framework to manage core domain entities (including users, competitions, and submissions).

    2. Data Storage Layer: Relational state metadata is maintained in a PostgreSQL database, while all large binary assets (such as datasets, submission archives, and competition bundles) are stored in an S3-compatible MinIO object storage system.

    3. Asynchronous Evaluation Pipeline: Job orchestration operates on a producer-consumer model. Tasks are dispatched through a RabbitMQ message broker and managed via Celery worker clients, executing submission jobs and system background tasks asynchronously across distributed infrastructure.

  2. Knowl 2 — Sandboxed Ingestion and Scoring Execution Pipeline

    model/method

    CodaLab Competitions supports both static predictions (results submissions) and full algorithmic code execution (code submissions) within Docker containers to isolate execution and prevent host corruption:

    • Environment Specification: Organizers define execution environments via standard Docker Hub image identifiers and tags, ensuring reproducible dependency sets across submissions.
    • Ingestion Program: An optional organizer-defined executable that orchestrates the participant's submitted code. It implements an organizer-specified API (e.g., executing fit and predict routines, or stepping through interactive reinforcement learning environments and data-generating simulators) to produce predicted outputs in a controlled manner.
    • Scoring Program: A mandatory program that processes submission outputs or direct predictions, compares them against reference ground-truth data, computes evaluation metrics, and outputs scores, logs, runtimes, and diagnostic plots.

    Outputs produced by this pipeline populate the public or private competition leaderboard.

  3. Knowl 3 — Decentralized Worker Queues and Private Data Isolation

    model/method

    To handle large computational workloads and sensitive datasets, CodaLab Competitions implements a modular queueing model for compute workers:

    • External Worker Attachments: While a default worker pool is provided on the public instance, organizers can instantiate custom queues and link their own dedicated CPU or GPU compute nodes (hosted locally or on cloud providers).
    • Confidential Data Security: Organizers can store confidential or medical training/testing data directly on their attached private compute workers rather than uploading them to public cloud storage. Submitted algorithms are sent to the worker, executed against local data, and evaluated without exposing sensitive underlying datasets to participants or central platform storage.
  4. Knowl 4 — Competition Bundle Structure and Lightweight Cloning

    definition

    A competition bundle is a standardized ZIP archive that encapsulates all assets necessary to initialize and configure a machine learning competition on CodaLab Competitions:

    • Contents: Configuration files (phase dates, rules, resource limits), HTML documentation pages, submission evaluation/ingestion scripts, and associated datasets.
    • Online Editing and Archival: Bundles can be uploaded to spawn competitions, edited directly in the platform's web interface, and re-exported as updated ZIP archives for sharing and archival.
    • Lightweight Bundles: To enable rapid competition cloning without duplicate network transfers of massive files, the platform provides a light bundle format containing only configuration metadata that references existing server-side database records for previously uploaded programs and datasets.
  5. Knowl 5 — Multi-Phase Progression and Multi-Score Leaderboard Evaluation

    model/method

    CodaLab Competitions provides configurable multi-stage workflows and multi-metric ranking capabilities:

    • Phased Workflows: A competition can be partitioned into distinct sequential phases (e.g., an open development/feedback phase followed by a final blind test phase). Each phase can define independent start/end dates, submission rate limits, evaluation datasets, and scoring programs. Submissions from earlier phases can be configured to forward automatically to subsequent phases.
    • Multi-Score Aggregation: Leaderboards can record multiple custom sub-scores from scoring programs. The platform supports computing the aggregate participant rank as the average of the participant's ranks across individual sub-scores.
  6. Knowl 6 — Comparative Feature Analysis of Data Science Competition Platforms

    data/table

    A comparative evaluation of CodaLab Competitions against other scientific competition and benchmarking platforms across ten operational dimensions:

    Criteria AICrowd CodaLab CrowdAnalytiX EvalAI Kaggle RAMP Tianchi
    Code-sharing ×\times ✓ ×\times ✓ ✓ ✓ ✓
    Code submission ✓ ✓ ×\times ✓ ✓ ✓ ✓
    Active community ⋆⋆\star\star ⋆⋆⋆\star\star\star ⋆\star ⋆⋆\star\star ⋆⋆⋆\star\star\star ⋆\star ⋆⋆⋆\star\star\star
    Custom metrics ✓ ✓ ✓ ✓ ✓ ✓ ?
    Staged challenge ✓ ✓ ×\times ✓ ×\times ✓ ×\times
    Private evaluation ×\times ✓ ×\times ✓ ×\times ×\times ×\times
    Open-source ✓ ✓ ×\times ✓ ×\times ✓ ×\times
    Human evaluation ×\times ×\times ×\times ✓ ×\times ×\times ×\times
    RL-friendly ✓ ✓ ×\times ✓ ×\times ×\times ×\times
    Run for free ×\times ✓ ×\times ✓ ✓ ×\times ?

    The comparison highlights that CodaLab Competitions and EvalAI uniquely support the full combination of open-source licensing, private/isolated worker evaluation, staged multi-phase competitions, reinforcement learning execution, and free hosting.

Coverage note — None was omitted; all primary architectural components, execution pipelines, data structures, and comparative analyses presented in the paper are covered.

References

  1. 1.Anthony Goldbloom and Ben Hamner. Kaggle. 2010. URL https://www.kaggle.com/docs/competitions.
  2. 2.Alibaba Group. Tianchi. 2014. URL https://tianchi.aliyun.com/.
  3. 3.Richard Isdahl and Odd Erik Gundersen. Out-of-the-box reproducibility: A survey of machine learning platforms. In 15th International Conference on eScience, eScience 2019, San Diego, CA, USA, September 24-27, 2019, pages 86–95. IEEE, 2019. doi: 10.1109/eScience.2019.00017. URL https://doi.org/10.1109/eScience.2019.00017.
  4. 4.Balázs Kégl, Alexandre Boucaud, Mehdi Cherti, Akin Osman Kazakci, Alexandre Gramfort, Guillaume M Lemaitre, Joris Van den Bossche, Djalel Benbouzid, and Camille Marini. The RAMP framework: from reproducibility to transparency in the design and optimization of scientific workflows. International Conference On Machine Learning, July 2018. URL https://hal.archives-ouvertes.fr/hal-02072341. Poster.
  5. 5.Divyabh Mishra. Crowdanalytix. 2012. URL https://www.crowdanalytix.com/.
  6. 6.Sharada Mohanty, Shivam Khandelwal, and Marcel Salathé. Aicrowd. 2016. URL https://github.com/AIcrowd/AIcrowd.
  7. 7.Benedict Neo. 12 data science and ai competitions to advance your skills in 2021. Towards Data Science, 2021. URL https://towardsdatascience.com/12-data-science-ai-competitions-to-advance-your-skills-in-2021-32e3fcb95d8c.
  8. 8.Humphrey. Quill and David. Penney. John Harrison, Copley medallist and the L.20,000 longitude prize / by H. Quill ; [line drawings by David Penney]. Antiquarian Horological Society [Ticehurst] ([New House, High St., Ticehurst, Wadhurst, Sussex TN5 7AL]), 1976. ISBN 0901180130.
  9. 9.David Rousseau and Andrey Ustyuzhanin. Machine learning scientific competitions and datasets. 2020. URL https://arxiv.org/abs/2012.08520.
  10. 10.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015. doi: 10.1007/s11263-015-0816-y.
  11. 11.Zhen Xu, Sergio Escalera, Adrien Pavao, Magali Richard, Wei-Wei Tu, Quanming Yao, Huan Zhao, and Isabelle Guyon. Codabench: Flexible, easy-to-use, and reproducible meta-benchmark platform. Patterns, 3(7):100543, 2022. ISSN 2666-3899. doi: https://doi.org/10.1016/j.patter.2022.100543. URL https://www.sciencedirect.com/science/article/pii/S2666389922001465.
  12. 12.Deshraj Yadav, Rishabh Jain, Harsh Agrawal, Prithvijit Chattopadhyay, Taranjeet Singh, Akash Jain, Shivkaran Singh, Stefan Lee, and Dhruv Batra. Evalai: Towards better evaluation systems for AI agents. CoRR, abs/1902.03570, 2019. URL http://arxiv.org/abs/1902.03570.

Citation

MLA
Pavao, A., et al. “CodaLab Competitions: An Open Source Platform to Organize Scientific Challenges”. Journal of Machine Learning Research, vol. 24, no. 198, 2023, pp. 1–6, https://www.jmlr.org/papers/v24/21-1436.html.
APA
Pavao, A., Guyon, I., Letournel, A.-C., Tran, D.-T., Baro, X., Escalante, H. J., Escalera, S., Thomas, T., & Xu, Z. (2023). CodaLab Competitions: An Open Source Platform to Organize Scientific Challenges. Journal of Machine Learning Research, 24(198), 1–6. https://www.jmlr.org/papers/v24/21-1436.html
Chicago
Pavao, A., I. Guyon, A.-C. Letournel, et al. 2023. “CodaLab Competitions: An Open Source Platform to Organize Scientific Challenges”. Journal of Machine Learning Research 24 (198): 1–6. https://www.jmlr.org/papers/v24/21-1436.html.
Harvard
Pavao, A. et al. (2023) “CodaLab Competitions: An Open Source Platform to Organize Scientific Challenges”, Journal of Machine Learning Research, 24(198), pp. 1–6. Available at: https://www.jmlr.org/papers/v24/21-1436.html.
Vancouver
1. Pavao A, Guyon I, Letournel A-C, Tran D-T, Baro X, Escalante HJ, Escalera S, Thomas T, Xu Z (2023) CodaLab Competitions: An Open Source Platform to Organize Scientific Challenges. Journal of Machine Learning Research 24:1–6

BibTeX

@article{JMLR:v24:21-1436,
  author  = {Adrien Pavao and Isabelle Guyon and Anne-Catherine Letournel and Dinh-Tuan Tran and Xavier Baro and Hugo Jair Escalante and Sergio Escalera and Tyler Thomas and Zhen Xu},
  title   = {CodaLab Competitions: An Open Source Platform to Organize Scientific Challenges},
  journal = {Journal of Machine Learning Research},
  year    = {2023},
  volume  = {24},
  number  = {198},
  pages   = {1--6},
  url     = {http://jmlr.org/papers/v24/21-1436.html}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/