Overview of Talk Python to Me #562: DuckLake: The Lakehouse That's Just SQL and Parquet
In this episode, Michael Kennedy talks with Pedro Hollanda and Guillermo Sanchez-Doinas about DuckLake, DuckDB’s open table format for lakehouse-style data management. The core idea is simple: keep data in Parquet files, but move the metadata into a real SQL database. That design removes a lot of the complexity and file-chasing common in other lakehouse formats like Iceberg and Delta, while improving performance, transactional behavior, and operational simplicity.
Main Takeaways
- DuckLake is intentionally minimal: metadata in a database, data in Parquet.
- Instead of navigating multiple metadata files, DuckLake uses one SQL query to the catalog to determine which data files to read.
- DuckLake is designed to be open and engine-agnostic: any system that can read Parquet and speak to a supported catalog can participate.
- The project emphasizes simplicity, transactional correctness, and practical performance over elaborate metadata stacks.
- With Quack as a client-server protocol/catalog setup, DuckLake can handle high contention and much better transaction throughput than traditional open table formats.
What DuckLake Is and Why It Exists
The Problem It Solves
Traditional open table formats often rely on:
- JSON / Avro / other metadata files
- multiple round trips to discover the right Parquet files
- lots of small files and metadata growth over time
- complicated ACID semantics layered on top of file-based metadata
DuckLake was created to simplify this. The metadata lives in a proper database catalog, so the system can use standard SQL queries and ACID transactions directly instead of reconstructing them from file formats.
Core Design
DuckLake’s architecture is essentially:
- Storage: Parquet files in S3, Blob Storage, or similar object storage
- Catalog: a SQL database such as PostgreSQL, SQLite, DuckDB, or Quack
- Compute: DuckDB, Spark, DataFusion, or any engine that can read the format
The key advantage is that the catalog can answer “which files should I read?” with a single SQL query.
DuckDB Background and Embedded Databases
Why DuckDB Matters
The guests explain DuckDB as part of a broader shift back toward embedded databases:
- no separate server to provision
- no network protocol overhead for every query
- shared memory between app and database
- easier integration with Python data tools like NumPy and Pandas
- very strong analytical performance due to columnar storage and vectorized execution
Columnar vs Row-Oriented Storage
They give a clear explanation of why DuckDB excels at analytics:
- Row-oriented storage is better for transactional workloads and point updates.
- Column-oriented storage is better for scans, aggregates, and analytical queries.
- Columnar storage improves:
- selective reads
- compression
- cache efficiency
- vectorized query execution
DuckLake Architecture
1. Storage
Data is stored as Parquet files in object storage.
2. Catalog
The catalog stores all metadata about:
- table schemas
- snapshots
- file lists
- statistics
- transaction state
This can be backed by:
- DuckDB
- SQLite
- PostgreSQL
- Quack
3. Compute
Any compute engine that understands the format can query the data by consulting the catalog and then reading the relevant Parquet files.
Why DuckLake Is Different from Iceberg and Delta
Simpler Metadata Model
Instead of juggling multiple metadata layers, DuckLake uses the database catalog directly.
Fewer Round Trips
To read a table snapshot, DuckLake typically needs:
- one SQL query to the catalog
- reads of the relevant Parquet files
That’s much simpler than formats that require repeated metadata traversal.
ACID Is Built In
Because the catalog is already a transactional SQL database, DuckLake doesn’t need to simulate ACID from files. The transaction semantics come from the database itself.
Quack and High-Contention Workloads
What Quack Is
Quack is DuckDB’s newer client-server protocol and catalog layer for more distributed setups.
Why It Matters
With Quack, DuckLake can:
- support concurrent clients more naturally
- retry transactions server-side
- reduce round-trip costs during conflicts
- achieve much higher throughput under contention
The guests mentioned that in a high-contention environment:
- PostgreSQL catalog setup can fall to around 5 transactions/sec
- Quack-based setup can reach about 200 transactions/sec
That makes DuckLake unusually strong for an open table format in transactional scenarios.
Data Inlining and Small File Problems
The Small File Problem
One major pain point in lakehouse systems is generating many tiny files from frequent inserts or streaming workloads.
DuckLake’s Solution: Data Inlining
DuckLake can temporarily store small amounts of data inside the catalog database itself before flushing them out to Parquet.
This helps:
- avoid lots of tiny files
- reduce object storage latency costs
- preserve transactional semantics
- improve write performance dramatically
The guests mentioned this can yield very large performance improvements compared to raw Iceberg workflows.
Frozen DuckLake
A Practical Pattern
A particularly interesting pattern discussed was the “frozen DuckLake”:
- use DuckDB as an embedded catalog
- do batch updates locally
- upload the DuckDB catalog file to object storage
- serve many read-only consumers from that frozen snapshot
This works especially well when updates are infrequent, such as daily batch jobs, but reads are frequent.
Production Readiness
Status
DuckLake has moved into a production-ready phase, with work focused on:
- bug fixing
- stability
- checkpointing
- compaction
- orphan file cleanup
- removing old files
- compatibility with Iceberg import/export workflows
What Was Needed for 1.0
A major focus was making sure long-lived systems could:
- compact data
- checkpoint state
- manage file growth
- avoid degradation as metadata accumulates
The guests emphasized that real companies are already using DuckLake in production.
Getting Started
Best Way to Start
Their recommendation is to start with:
- DuckDB in-process
That gives the simplest “one line and you’re running” experience.
Other Options
You can also use:
- PostgreSQL for production catalogs
- SQLite for lightweight local setups
- Quack for distributed/client-server workflows
Practical Advice
- Start small and local with DuckDB.
- Check the DuckLake website for tutorials and examples.
- If you hit issues, file a reproducible GitHub issue with steps and scripts.
Notable Themes and Quotes
Simplicity Wins
A recurring theme was that many database and lakehouse systems get overly complicated, and DuckLake’s advantage is that it collapses metadata into a normal SQL database.
Embedded Systems Are Back
The episode strongly reinforces the idea that embedded databases are a big deal again because they dramatically reduce:
- operational overhead
- latency
- deployment complexity
Playful Branding, Serious Engineering
The DuckDB ecosystem keeps a whimsical tone—DuckLake, Quack, “frozen DuckLake”—but the engineering goals are serious:
- lower cost
- better performance
- simpler operations
- more open interoperability
Bottom Line
DuckLake is best understood as a lakehouse format stripped down to the essentials:
- Parquet for data
- SQL database for metadata
- simple, transactional, and open by design
If Iceberg or Delta feel too metadata-heavy, DuckLake offers a compelling alternative that aims to be easier to implement, faster in practice, and more operationally friendly—especially when paired with DuckDB or Quack.
