Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Delta Lake tables

Delta Lake tables are structured datasets stored using the open-source Delta Lake storage format over distributed file systems or cloud object stores. Built upon columnar Apache Parquet data files paired with an ordered transaction log, they provide full ACID transaction guarantees, schema enforcement, and data versioning capabilities such as time travel. This design enables scalable metadata operations, data layout optimization, and reliable reads and writes for both batch and streaming workloads across diverse big data processing engines.

1 item

Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores

Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores

Michael Armbrust, Tathagata Das, Sameer Paranjpye, Reynold Xin, Shixiong Zhu, Ali Ghodsi, Burak Yavuz, Mukul Murthy, Joseph Torres, Liwen Sun, Peter Boncz, Mostafa Mokhtar, Herman van Hovell, Adrian Ionescu, Alicja Luszczak, Michal Switakowski, Michal Szafranski, Xiao Li, Takuya Ueshin, Pieter Senster, Matei Zaharia

OrganizationsCentrum Wiskunde InformaticaDatabricksStanford UniversityUniversity of California Berkeley

Why you should read this

Introduces Delta Lake, a high-performance ACID table storage layer that overcomes the consistency and performance limitations of cloud object stores to enable the lakehouse paradigm for exabyte-scale data processing and analytics.

Cloud object stores such as Amazon S3 are some of the largest and most cost-effective storage systems on the planet, making them an attractive target to store large data warehouses and data lakes. Unfortunately, their implementation as key-value stores makes it difficult to achieve ACID transactions and high performance: metadata operations such as listing objects are expensive, and consistency guarantees are limited. In this paper, we present Delta Lake, an open source ACID table storage layer over cloud object stores initially developed at Databricks. Delta Lake uses a transaction log that is compacted into Apache Parquet format to provide ACID properties, time travel, and significantly faster metadata operations for large tabular datasets (e.g., the ability to quickly search billions of table partitions for those relevant to a query). It also leverages this design to provide high-level features such as automatic data layout optimization, upserts, caching, and audit logs. Delta Lake tables can be accessed from Apache Spark, Hive, Presto, Redshift and other systems. Delta Lake is deployed at thousands of Databricks customers that process exabytes of data per day, with the largest instances managing exabyte-scale datasets and billions of objects.

Added

2026-04-18