Loam is pre-alpha: the engine core runs today; Live and Durable are in progress. See the roadmap

Blog/What Loam is built on

Lance: the columnar format that holds Loam’s documents and vectors

How Lance lays out fragments, versions and indexes for fast random access on object storage, why Loam stores collections in it instead of Parquet, the detached versions we commit, and the churn we pin against.

Lance: the columnar format that holds Loam’s documents and vectors
On this page
  1. What Lance is
  2. Why Loam uses it
  3. How Loam uses it
  4. What we do not use
  5. Limits

Every Loam collection has a Lance dataset at its core. It holds each document's source JSON, its dense and sparse vectors, and the system columns Loam needs, and it carries the IVF vector indexes that serve cold vector search. We described how Lance and Tantivy share one commit point in One manifest, two formats; this post is about Lance itself.

Repositorylance-format/lance (moved from lancedb/lance)
Docslance.org
LicenseApache-2.0
Version12.0.0 (September 2026); Loam pins =12.0.0 and file format 2.1
In LoamAvailable

What Lance is

Lance is a columnar data format and table format designed for AI data: vectors, embeddings, images and long text, read by training loops and search indexes. It has two layers.

The file format. Like Parquet, a Lance file stores columns separately. Unlike Parquet, it has no row groups, and it is designed for random access: fetching a few hundred scattered rows by position is cheap, because each column's pages can be located and read individually. That is exactly the access pattern of search (look up the top-k documents by row id) and of shuffled training reads. File format versions are tracked explicitly (data_storage_version); stable versions stay readable, and newer ones are opt-in.

The table format. A dataset is a directory:

lance/
  data/           # data files
  _deletions/     # deletion files: which rows of a fragment are gone
  _transactions/  # one record per commit
  _indices/       # vector and scalar indexes
  _versions/      # one manifest per version

Rows are grouped into fragments, and each fragment can hold several data files, each contributing a subset of columns. So adding a column, or backfilling embeddings for an existing one, writes new files for that column only and is mostly a metadata change. Deletes write deletion files instead of rewriting data. Every commit writes a new manifest listing the fragments, files and indexes of that version, which gives versioning, time travel and ACID commits. With stable row ids turned on, a row keeps its id when compaction rewrites its fragment, so indexes and external references survive compaction.

Indexes. Lance builds vector indexes (IVF with product quantization, IVF with HNSW, and others), scalar indexes (B-tree, bitmap, label lists) and a full-text index.

Why Loam uses it

A collection needs a durable home for documents and vectors on object storage with four properties: columnar (so vector search reads only vector columns), random access by row id (so fetching the top-k is a few ranged GETs), versioned (so a reader sees a consistent snapshot while writers commit), and open (so the data is readable without Loam).

  • Parquet is columnar and open, but its row groups make random access expensive and it has no table-level versioning of its own.
  • A vector database's internal format (Qdrant's, Milvus's) is built for local disk and a single engine.
  • Lance has all four, plus vector indexes, and it is readable by pylance, Polars, DuckDB's Lance extension and Ray. It is also built on DataFusion and Arrow, the same stack as Loam's query engine.

How Loam uses it

The schema. Each collection's dataset has _pk (the canonical primary-key bytes), _source (the document's JSON), _ingest_partition and _ingest_offset (the log record that last wrote the document), one fixed-size-list column per dense vector, one struct column per sparse vector, and Lance's stable row id. Typed fields (keywords, numbers, dates) are deliberately not Lance columns: they live in Tantivy and can always be re-derived from _source. So evolving a collection's field schema never touches Lance, and a new vector column is added lazily, all nulls, by a metadata-only commit.

Batched writes. Loam does not use Lance's per-write commit path for small writes; every commit writes a manifest, and many tiny commits are slow and contend. Writes land in Loam's own log first, and a worker writes large fragments from a batch and commits once.

Detached versions. This is the least obvious choice. Lance datasets normally advance along one mainline of versions. Loam instead commits every change as a detached version, built from exactly the Lance version the parent collection manifest names. The mainline holds only an empty version 1. The collection manifest (Loam's own, compare-and-swapped in the metastore) records which detached version belongs to it. The payoff: a crashed or fenced writer's Lance version is simply never referenced and never blocks the next commit, and there is exactly one linearization point for Lance, Tantivy and delete bitmaps together.

Reads. Cold vector search uses Lance's IVF index; hot collections use an HNSW graph built from the same vectors. Documents are fetched with coalesced take_rows calls by stable row id. Row ids flow through Loam's operators as one roaring-bitmap space.

Garbage collection is Loam's. Lance's own cleanup never runs, because it does not know which detached versions are still referenced by retained collection manifests. Loam computes reachability from the manifests it keeps (24 hours of history by default, for time travel and pinned reads) and deletes the rest.

External readers get a pinned view only through scan pinning: because the mainline is empty, opening the dataset root shows no rows. A scan plan from Loam names the detached version and the fragments to read, which a Ray Data, Polars, PySpark or PyTorch reader then reads directly from the bucket. Credential vending and the Python adapters

What we do not use

  • Lance's full-text index. Text lives in Tantivy splits, which give Loam BM25 statistics, analyzers and aggregations compatible with Elasticsearch (next post).
  • File format 2.2. Lance 12 defaults to 2.2; Loam sets 2.1 explicitly until 2.2 is evaluated.
  • Lance's own distributed serving. LanceDB Inc.'s serving, caching and indexing layer for its enterprise product is closed; Loam's serving layer is its own.

Limits

  • Fast API churn. Lance is vendor-led by LanceDB Inc. and in active development; its SDK and API versioning is separate from the file format's. Loam pins an exact version and keeps Lance behind a trait boundary.
  • It pins DataFusion and Arrow. Lance 12 depends on DataFusion 54 and arrow 58, which sets those versions for all of Loam (DataFusion post).
  • Small commits are expensive, which is why Loam batches through its own log.
  • Detached versions are an unusual mode. Whether pylance opens a detached version by id is an open question we check before the Python adapters ship.

The next post covers the other half of every collection: full text, in Tantivy.

More from the blog