On this page
Loam has one firm rule about other engines: it integrates with them over protocols and open formats, and never links them into its binary. Engines built on DataFusion or Arrow move in version lockstep, and Loam's versions are pinned by Lance; some engines embed a Python interpreter or need a JVM; one of the four below is GPL-licensed. Running them as separate processes keeps Loam's binary, license and upgrade cadence its own.
This post covers four such engines. None of them runs in Loam today.
| Engine | Role for Loam | License | Status |
|---|---|---|---|
| Sail | PySpark jobs, and a reader of Loam's Iceberg tables | Apache-2.0 | Planned |
| RisingWave | Streaming SQL beside Loam; the target for Flink SQL jobs | Apache-2.0 (some Premium features behind a license key) | Planned |
| Neon | Postgres on the bucket, with branches per agent workspace | Apache-2.0 | Under evaluation |
| WeSQL | MySQL on the bucket | GPL-2.0-only | Not adopted |
Sail: Spark Connect in Rust
What it is. Sail (v0.7.1) is a drop-in replacement for Spark's compute, written in Rust on Arrow and DataFusion. It implements the Spark Connect protocol, the gRPC protocol with which PySpark 3.5 and later can talk to a remote Spark server, so PySpark clients connect to Sail with SparkSession.builder.remote("sc://…") and do not know the difference. It runs locally or in a Kubernetes cluster mode (a server, a driver per session that launches worker pods, and shuffle through object storage since 0.7), and runs Python UDFs in an embedded CPython.
What it covers. Spark SQL and the DataFrame API; Python UDFs, pandas UDFs, UDTFs and the mapInPandas family; Iceberg v1 to v3 with copy-on-write writes (merge-on-read writes and deletion-vector reads are still in progress, so tables that use deletion vectors need checking before Sail reads them), through REST, Glue, Unity and Hive catalogs; S3, GCS, ADLS and HDFS.
What it does not. RDDs and SparkContext (Spark Connect does not carry them at all), Java and Scala UDFs, MLlib and pandas-on-Spark; Structured Streaming is not ready. Sail's own CI shows about 91% of Spark 3.5's Connect tests passing and about 71% of Spark 4.2's. Its performance claims are its own, derived from TPC-H runs by the vendor, and its compatibility checker does not verify behavioural parity.
For Loam. Two roles. Sail is a named reader in the Iceberg exit gate, so PySpark curation jobs on stateless compute read Loam's keyed tables correctly. And in the jobs proposal (platform part 6), Loam runs one Sail server per namespace behind an authenticating sc:// proxy, with Apache Spark on Kubernetes as the fallback for jobs Sail cannot run. Sail's crates are on DataFusion 55 and arrow 59 against Loam's 54 and 58, and its maintainers declined Rust-level Lance integration for the same reason, so Sail is always a separate process.
RisingWave: streaming SQL on object storage
What it is. RisingWave (v3.1.0) is a streaming database: you define sources, materialized views and sinks in Postgres-dialect SQL, and it maintains the views incrementally as data arrives. It is written in Rust but is a service, not a library: a Postgres-wire frontend, a meta service, compute nodes and compactors. Its state lives in Hummock, an LSM-tree store on object storage, with epoch-based checkpoints. It reads from Kafka, Pulsar, Kinesis and NATS, and reads and writes Iceberg, including through Lakekeeper.
For Loam. RisingWave is Loam's companion stream processor: it keeps Loam out of stateful stream processing (windowed joins with checkpointed state) while giving users joins, windows and materialized views. Before Loam's Kafka gateway exists, RisingWave can write into Loam through its Elasticsearch sink (Loam's _bulk subset), its HTTP sink and, later, its Iceberg sink; once the Kafka gateway exists, it reads Loam streams directly. In the jobs proposal it is also where Flink SQL jobs move, after a dialect port, since RisingWave does not speak Flink SQL.
Limits. Some features are "Premium" behind a license key, with a free tier capped at a small size; Loam would use none of them (the Iceberg REST catalog is not Premium; the Glue catalog is). It does not replace Flink's DataStream API or its long tail of connectors.
Neon: Postgres whose storage is the bucket
What it is. Neon splits Postgres in three:
- Compute: stateless Postgres with a modified storage manager that asks for pages over the network instead of reading local files.
- Safekeepers: a Paxos-replicated service that accepts the compute's WAL, so a commit is durable once a quorum has it, and offloads WAL to object storage.
- Pageserver: ingests the WAL, stores page versions as immutable layer files in object storage, and serves any page as of any WAL position. Its local disk is a cache.
A storage broker connects them, and a storage controller handles placement and failover. Because every page is versioned, a branch is a new timeline whose ancestor is an existing one at a chosen WAL position: copy-on-write, with no data copied.
For Loam. Our spike ran it on RustFS: branches in 36 ms, isolated both ways, logical replication whose slot survived a compute replacement, OpenFGA's and GlitchTip's migrations unchanged, and full recovery from the bucket after the pageserver's disk was wiped. The proposal has Loam act as Neon's control plane, route Postgres connections by database name, give every agent workspace its own branch, and feed changes into collections. Platform part 7 has the details.
Why it is only under evaluation. The public repository has been nearly inactive since August 2025, after the Databricks acquisition; its Postgres is 16.9 and 17.5 from May 2025 with Neon's patches; the control plane and the proxy's authentication backend were never open source. Loam would own every fix. So the design keeps plain Postgres as a working exit, and a Loam-maintained fork is being designed as a separate track before anything is committed.
WeSQL: MySQL on the bucket, not adopted
What it is. WeSQL is a MySQL 8.0 distribution whose storage engine, SmartEngine (an LSM tree), keeps data, WAL and binlog in object storage. It is GPL-2.0-only, inherited from MySQL, has no releases, and recommends itself for development and testing. In August 2026 the project removed its multi-replica Raft mode after more than a year without commits.
What we found. Basic SQL worked on RustFS once RustFS was configured for virtual-hosted-style URLs. But SmartEngine does not support foreign keys, so Forgejo's migrations failed; ENGINE=InnoDB is silently rewritten; and after a SIGKILL without the local volume, committed rows written after the last object-store snapshot were lost, even though their binlog had been archived to the bucket. That last finding may be a configuration issue, and we would re-test it, but with the image's defaults recent commits live only on local disk.
Decision. Not adopted. Had it been, it could only ever have run as an unmodified separate process: linking or modifying GPL-2.0-only code would put Loam's code under the GPL.
What the four have in common
Each one does a job Loam deliberately does not: Spark's API, stateful stream processing, Postgres OLTP, MySQL OLTP. Each keeps its durable state in object storage, which is why they fit next to a bucket-native platform. And each is a separate process with its own license, reached over Spark Connect, the Postgres wire protocol, Iceberg or Loam's own APIs.
The last post in this series covers the smaller libraries inside Loam that carry more weight than their size suggests.