#562: DuckLake: The Lakehouse That's Just SQL and Parquet

Summary of #562: DuckLake: The Lakehouse That's Just SQL and Parquet

by Michael Kennedy

1h 11m•September 10, 2026

Overview of Talk Python to Me #562: DuckLake: The Lakehouse That's Just SQL and Parquet

In this episode, Michael Kennedy talks with Pedro Hollanda and Guillermo Sanchez-Doinas about DuckLake, DuckDB’s open table format for lakehouse-style data management. The core idea is simple: keep data in Parquet files, but move the metadata into a real SQL database. That design removes a lot of the complexity and file-chasing common in other lakehouse formats like Iceberg and Delta, while improving performance, transactional behavior, and operational simplicity.

Main Takeaways

  • DuckLake is intentionally minimal: metadata in a database, data in Parquet.
  • Instead of navigating multiple metadata files, DuckLake uses one SQL query to the catalog to determine which data files to read.
  • DuckLake is designed to be open and engine-agnostic: any system that can read Parquet and speak to a supported catalog can participate.
  • The project emphasizes simplicity, transactional correctness, and practical performance over elaborate metadata stacks.
  • With Quack as a client-server protocol/catalog setup, DuckLake can handle high contention and much better transaction throughput than traditional open table formats.

What DuckLake Is and Why It Exists

The Problem It Solves

Traditional open table formats often rely on:

  • JSON / Avro / other metadata files
  • multiple round trips to discover the right Parquet files
  • lots of small files and metadata growth over time
  • complicated ACID semantics layered on top of file-based metadata

DuckLake was created to simplify this. The metadata lives in a proper database catalog, so the system can use standard SQL queries and ACID transactions directly instead of reconstructing them from file formats.

Core Design

DuckLake’s architecture is essentially:

  • Storage: Parquet files in S3, Blob Storage, or similar object storage
  • Catalog: a SQL database such as PostgreSQL, SQLite, DuckDB, or Quack
  • Compute: DuckDB, Spark, DataFusion, or any engine that can read the format

The key advantage is that the catalog can answer “which files should I read?” with a single SQL query.

DuckDB Background and Embedded Databases

Why DuckDB Matters

The guests explain DuckDB as part of a broader shift back toward embedded databases:

  • no separate server to provision
  • no network protocol overhead for every query
  • shared memory between app and database
  • easier integration with Python data tools like NumPy and Pandas
  • very strong analytical performance due to columnar storage and vectorized execution

Columnar vs Row-Oriented Storage

They give a clear explanation of why DuckDB excels at analytics:

  • Row-oriented storage is better for transactional workloads and point updates.
  • Column-oriented storage is better for scans, aggregates, and analytical queries.
  • Columnar storage improves:
    • selective reads
    • compression
    • cache efficiency
    • vectorized query execution

DuckLake Architecture

1. Storage

Data is stored as Parquet files in object storage.

2. Catalog

The catalog stores all metadata about:

  • table schemas
  • snapshots
  • file lists
  • statistics
  • transaction state

This can be backed by:

  • DuckDB
  • SQLite
  • PostgreSQL
  • Quack

3. Compute

Any compute engine that understands the format can query the data by consulting the catalog and then reading the relevant Parquet files.

Why DuckLake Is Different from Iceberg and Delta

Simpler Metadata Model

Instead of juggling multiple metadata layers, DuckLake uses the database catalog directly.

Fewer Round Trips

To read a table snapshot, DuckLake typically needs:

  1. one SQL query to the catalog
  2. reads of the relevant Parquet files

That’s much simpler than formats that require repeated metadata traversal.

ACID Is Built In

Because the catalog is already a transactional SQL database, DuckLake doesn’t need to simulate ACID from files. The transaction semantics come from the database itself.

Quack and High-Contention Workloads

What Quack Is

Quack is DuckDB’s newer client-server protocol and catalog layer for more distributed setups.

Why It Matters

With Quack, DuckLake can:

  • support concurrent clients more naturally
  • retry transactions server-side
  • reduce round-trip costs during conflicts
  • achieve much higher throughput under contention

The guests mentioned that in a high-contention environment:

  • PostgreSQL catalog setup can fall to around 5 transactions/sec
  • Quack-based setup can reach about 200 transactions/sec

That makes DuckLake unusually strong for an open table format in transactional scenarios.

Data Inlining and Small File Problems

The Small File Problem

One major pain point in lakehouse systems is generating many tiny files from frequent inserts or streaming workloads.

DuckLake’s Solution: Data Inlining

DuckLake can temporarily store small amounts of data inside the catalog database itself before flushing them out to Parquet.

This helps:

  • avoid lots of tiny files
  • reduce object storage latency costs
  • preserve transactional semantics
  • improve write performance dramatically

The guests mentioned this can yield very large performance improvements compared to raw Iceberg workflows.

Frozen DuckLake

A Practical Pattern

A particularly interesting pattern discussed was the “frozen DuckLake”:

  • use DuckDB as an embedded catalog
  • do batch updates locally
  • upload the DuckDB catalog file to object storage
  • serve many read-only consumers from that frozen snapshot

This works especially well when updates are infrequent, such as daily batch jobs, but reads are frequent.

Production Readiness

Status

DuckLake has moved into a production-ready phase, with work focused on:

  • bug fixing
  • stability
  • checkpointing
  • compaction
  • orphan file cleanup
  • removing old files
  • compatibility with Iceberg import/export workflows

What Was Needed for 1.0

A major focus was making sure long-lived systems could:

  • compact data
  • checkpoint state
  • manage file growth
  • avoid degradation as metadata accumulates

The guests emphasized that real companies are already using DuckLake in production.

Getting Started

Best Way to Start

Their recommendation is to start with:

  • DuckDB in-process

That gives the simplest “one line and you’re running” experience.

Other Options

You can also use:

  • PostgreSQL for production catalogs
  • SQLite for lightweight local setups
  • Quack for distributed/client-server workflows

Practical Advice

  • Start small and local with DuckDB.
  • Check the DuckLake website for tutorials and examples.
  • If you hit issues, file a reproducible GitHub issue with steps and scripts.

Notable Themes and Quotes

Simplicity Wins

A recurring theme was that many database and lakehouse systems get overly complicated, and DuckLake’s advantage is that it collapses metadata into a normal SQL database.

Embedded Systems Are Back

The episode strongly reinforces the idea that embedded databases are a big deal again because they dramatically reduce:

  • operational overhead
  • latency
  • deployment complexity

Playful Branding, Serious Engineering

The DuckDB ecosystem keeps a whimsical tone—DuckLake, Quack, “frozen DuckLake”—but the engineering goals are serious:

  • lower cost
  • better performance
  • simpler operations
  • more open interoperability

Bottom Line

DuckLake is best understood as a lakehouse format stripped down to the essentials:

  • Parquet for data
  • SQL database for metadata
  • simple, transactional, and open by design

If Iceberg or Delta feel too metadata-heavy, DuckLake offers a compelling alternative that aims to be easier to implement, faster in practice, and more operationally friendly—especially when paired with DuckDB or Quack.