On this page
Planned Iceberg tables in Loam are designed and not built. This post explains the projects and the plan.
AI applications generate analytics data: every LLM call with its tokens, cost and latency, traces, eval results, product events. Loam's plan is to keep that data as Apache Iceberg tables in your bucket, catalogued by Lakekeeper and written with iceberg-rust, so DuckDB, Trino, Spark, Sail, ClickHouse, Snowflake and PyIceberg read the same tables without Loam.
| Repositories | apache/iceberg, apache/iceberg-rust, lakekeeper/lakekeeper |
| Licenses | All Apache-2.0 |
| Versions checked | Iceberg 1.11.0 (Java reference), iceberg-rust 0.10.1, Lakekeeper 0.13.6 |
| In Loam | Planned |
How Iceberg works
An Iceberg table is a tree of immutable files in object storage, with one mutable pointer at the top.
A writer creates data files, manifests, a manifest list and a new metadata file, then asks the catalog to swap the table's pointer from the metadata it started from to the new one. If another writer got there first, the swap fails and the writer retries on top of the new state. Readers load one metadata file and see one consistent snapshot. Statistics in manifests let a reader prune files without opening them.
Iceberg's other ideas matter too: hidden partitioning (a partition is a transform of a column, like day(ts), so queries filter on the column and never on a partition field), schema evolution by column id, time travel through snapshots, and branches and tags. Row-level changes started as position deletes and equality deletes in v2; v3 adds deletion vectors (compact bitmaps of deleted row positions stored in Puffin files), a variant type for semi-structured data and row lineage. v3 is shipping in Snowflake, Databricks and AWS.
iceberg-rust: what it can write
iceberg-rust is the Apache project's Rust implementation. It reads tables well and has REST, Glue, Hive Metastore, S3 Tables and SQL catalogs. For writes, the Iceberg project's own status table lists Rust as supporting appending data files only: no row deltas, no delete files, no file rewrites, no snapshot expiry. That table can lag the main branch, but it matches what we found. Reads have a gap too: applying Iceberg v3 deletion vectors landed on main in September 2026 but is not in a release yet, so a reader of keyed tables with deletion vectors needs a newer build.
Loam needs row-level writes for keyed tables, so the plan has two parts:
- use RisingWave's fork of iceberg-rust, which adds
RowDeltaandRewriteFilesand equality and position deletes, with the aim of converging on upstream; - build a deletion-vector writer and
RowDeltafor upstream iceberg-rust and contribute them, since Loam's keyed tables need v3 deletion vectors anyway.
Compaction would come from nimtable/iceberg-compaction, from the RisingWave ecosystem.
Lakekeeper: the catalog
Lakekeeper is an Iceberg REST catalog written in Rust, with access control, credential vending (short-lived storage credentials handed to query engines per table) and audit. Its authorization uses OpenFGA. Its storage backend is Postgres. It is Apache-2.0, maintained by Vakamo, which also sells a commercial edition.
Why Lakekeeper: the REST catalog protocol is what every engine speaks now, it is Rust like Loam, and its design helped shape Loam's own metastore: one catalog trait, several database backends. Loam's MetaStore trait follows that model, and Loam's planned OpenFGA authorizer adapts Lakekeeper's OpenFGA authorization model and its migration and reconcile patterns, keeping its notice. Loam and Lakekeeper would share one OpenFGA store, so a grant on a namespace covers both collections and tables.
What Loam adds on top
Table semantics declared in SQL, all mapping onto standard Iceberg:
CREATE TABLE llm_calls (
ts TIMESTAMP(3) WITH TIME ZONE, tenant STRING, model STRING,
prompt_tokens INT, completion_tokens INT, cost_usd DOUBLE, latency_ms INT
)
PARTITIONED BY (day(ts))
SORTED BY (tenant, model, ts)
WITH (retention = '180 days' ON ts);| Kind | Semantics | Implementation |
|---|---|---|
| Append-only | Rows are appended | Data files, with the declared partition spec and sort order |
Keyed (PRIMARY KEY) | One live row per key, latest write wins | A primary-key index plus v3 deletion vectors on upsert |
Versioned (VERSION BY col) | The greatest version wins; late, older rows are dropped | As keyed, comparing versions at apply |
| Aggregating | Rows are partial aggregate states merged by key | State columns merged at read and in compaction |
| Retention | Old rows removed | Partition drops, or deletion vectors |
Freshness. Writes enter through Loam's log like any other write and are committed to Iceberg by a link at a cadence of 10 to 60 seconds. External engines see those snapshots. Loam's own queries also merge the log tail, so they see rows seconds after they are written.
A hot tier for Iceberg, following StarRocks' shared-data design: a file index for pruning without I/O, cached Parquet pages on NVMe, and, for pinned partitions, sorted hot projections with a sparse primary-key index and skip indexes, the MergeTree idea applied as a cache over unchanged Iceberg files.
Engines around it. Sail is a named reader in the Iceberg exit gate; Sail's main branch already reads Puffin deletion vectors, so what we once budgeted as a contribution becomes a verification. RisingWave writes to Loam's tables through its Iceberg sink via Lakekeeper.
Limits we plan around
- iceberg-rust writes are append-only upstream today, which is why we depend on a fork for row deltas until our writer lands upstream.
- Commit cadence is the freshness external engines see. Only Loam's own queries get the tail.
- One writer class per table. A table is written either by Loam (links and ingest) or by an external engine, not both.
- Lakekeeper needs Postgres for its own catalog storage, one more service in a self-hosted install; it can share a server with other Postgres users.
- Aggregate state columns are opaque bytes to external engines.
The next post covers the API surface Loam's planned function runtime borrows: Dapr, and its secrets building block.