DuckLake
github.com/duckdb/ducklake- Category
- Developer Tools
- Rank
- No. 2090Tools index
- Pricing
- Open Source
- Type
- TOOL
- GitHub
- 3.2k stars
- Date
About
DuckLake is an open lakehouse format built on SQL and Parquet. It stores metadata in a catalog database and data in Parquet files, and the DuckLake extension lets DuckDB read and write that data directly. Keeping metadata in a database gives the format a single integrated catalog and table layer.
What it does
A DuckDB extension that turns an attached database into a versioned table store. Table rows land as columnar files in a directory or bucket you choose, while every table definition, file list and change record sits in a separate relational store: a local DuckDB file, PostgreSQL or SQLite. Once attached, ordinary SQL handles creates, inserts, updates and schema changes, and every commit becomes a numbered snapshot you can query back to or diff against.
Why it's ranked here
Maintained under the DuckDB Foundation and MIT licensed, it installs with one command from the official extension repository. The feature surface matches what teams usually reach for heavier stacks to get: time travel by snapshot number, a change feed per table, schema evolution, compaction, snapshot expiry and orphan file cleanup, all exposed as SQL calls. The test suite runs the full DuckDB core tests with this format as the storage backend, plus separate passes against PostgreSQL and SQLite catalogs.
What's good
Small writes do not spray tiny files: inserts under a row limit, ten rows by default, are kept inside the catalog and flushed to files later on demand. Commits retry on conflict with exponential backoff, ten attempts by default, and the PostgreSQL backend recognises primary key clashes on its own metadata tables as retryable regardless of server language. Snapshots record author and commit message, and a catalog can be set to refuse commits without one. Bucket partitioning uses an Iceberg compatible hash.
Tradeoffs
It is a DuckDB extension, so reads and writes go through DuckDB; there is no standalone client here. The PostgreSQL catalog cannot natively hold nested types, dates, timestamps, text or binary statistics values, so those are stored as text or raw bytes and cast back. Writing Iceberg style deletion vectors instead of positional delete files is marked experimental and off by default. Housekeeping is manual: expiring snapshots and deleting old or orphaned files are explicit calls, not background jobs.
How to use it well
Fits analysts and small data teams who already live in DuckDB and want versioned, shareable tables on object storage without running a separate metastore. Start with a local DuckDB file as the catalog, move to PostgreSQL when several writers need to share it, and schedule the compaction, expiry and cleanup calls yourself. Require commit messages if the history should be auditable. It does not replace a query engine for other runtimes; anything outside DuckDB needs its own reader for the format.
Technical notes+
src/ducklake_extension.cpp registers the storage extension under the ducklake prefix and sets global options: ducklake_max_retry_count (10), ducklake_retry_wait_ms (100), ducklake_retry_backoff (1.5), ducklake_default_data_inlining_row_limit (10, 0 disables), ducklake_target_file_size and the experimental ducklake_write_deletion_vectors (puffin instead of Parquet positional deletes). It also registers table functions for snapshots, table_changes, merge_adjacent_files, rewrite_data_files, cleanup_old_files, cleanup_orphaned_files, expire_snapshots, flush_inlined_data and add_data_files, plus a murmur3_32 scalar. src/storage/ducklake_inline_data.cpp buffers rows per thread and falls back to pass-through once the global count exceeds the limit. src/metadata_manager/postgres_metadata_manager.cpp maps unsupported types to VARCHAR or BYTEA and treats ducklake_*_pkey violations as retryable. CMakeLists.txt requires C++17 and roaring via vcpkg.json.
Observed
- License
- MIT (Stichting DuckDB Foundation) per LICENSE
- Language
- C++17, built with CMake per CMakeLists.txt
- Install surface
- DuckDB extension installed with INSTALL ducklake; nightly builds from core_nightly
- Interface
- SQL only, via ATTACH 'ducklake:...' inside DuckDB
- Catalog backends
- DuckDB file, PostgreSQL and SQLite, each with a test config per docs/README.md
- Data inlining default
- 10 rows per insert kept in the catalog, set in src/ducklake_extension.cpp
- Deletion vectors
- Iceberg V3 puffin deletion vectors marked EXPERIMENTAL, default off
- Native dependency
- roaring bitmaps via vcpkg.json
- Test policy for new code
- AGENTS.md requires minimal sqltests for new code
Read from docs/README.md, vcpkg.json, CMakeLists.txt, src/ducklake_extension.cpp, src/storage/ducklake_catalog.cpp, src/functions/ducklake_snapshots.cpp, src/metadata_manager/postgres_metadata_manager.cpp, src/storage/ducklake_inline_data.cpp, AGENTS.md, LICENSE.
Tags
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.