§12 Roadmap, Testing & Risks
Milestones M0–M6 and exit gates (v1.0 = M2, v1.1 = M2.x cloud and BYOC), testing strategy, risk register
Status: Approved · 2026-09-22 · revised 2026-09-25 (architecture review, D42–D50) · revised 2026-09-26 (metastore backends, router, tenancy and erasure, D58–D70, §18; FoundationDB dropped, the stream API core and OTLP logs in M2, Kafka in M5, one router for every resource kind, consistency tokens on every backend, D71–D76) · revised 2026-09-26 (turbopuffer gap analysis: write backpressure, filter writes and a limits page before launch; conditional writes, branches, ranking expressions, recall, analyzers, vector types, CMEK and more in M2; sharding in M2.x; D86–D103) · amended 2026-09-27 (M1.3 done) · amended 2026-10-02 (parallel tracks CLI, RT, FL and CN, GT, AP and MT from §30–§38, D281–D459; D347 reverses D45 for the House pending Q333; the owner rulings D400–D406)
Build order follows the pain point and the shortest path to production: hybrid retrieval first (Qdrant + a targeted Elasticsearch subset + Arrow Flight SQL), then production hardening to v1.0 (followed by cloud and BYOC as v1.1), then native graph for GraphRAG, then analytics on Iceberg, then Kafka compatibility and changelog streams, then scale. The internal log exists from M0 because everything depends on it, and its native stream API core ships in v1.0 (D72). The protocol footprint is deliberately narrow (D42): Operon speaks Qdrant, a framework-driven subset of Elasticsearch, Arrow Flight SQL and its own native REST/gRPC API, plus OTLP for logs (M2, D73) and the Kafka wire protocol for streams (M5, D74). The engine does not emulate Neo4j or ClickHouse; through the separate loam-fabric service, Loam House may serve the ClickHouse HTTP interface over a declared surface (D347, proposed pending Q333).
1. Milestones
| M | Name | Scope | Exit gates |
|---|---|---|---|
| M0 | Foundation | operon dev/standalone; meta on openraft; object_store abstraction with fault-injection wrapper; internal stream engine (standard class, leaderless sequencing, segmenter); foyer cache (H0/H1); worker leases + task framework; link framework (exactly-once apply); PkIndex on SlateDB; GC; DST harness | kill -9 at every step of write/commit/segment paths → no acknowledged-data loss, no torn state; object-store PUT/GET/412/409 fault matrix passes; linearizability check on sequencer & manifest-pointer CAS |
| M1 | Collections (Qdrant + Elasticsearch subset + Flight SQL) | Lance + Tantivy splits under one manifest; upserts/deletes; the MetaStore semantic trait (D47); tail indexes; native hybrid API + Python/TS SDK; Arrow Flight SQL with DoPut bulk ingest into collections and streams (D49); Qdrant API Phase A with sparse vectors (exact, IDF; D31); Elasticsearch API Phase A trimmed to what the framework suites send (D48); scan pinning for external readers (D53); to_arrow()/to_polars() in the Python SDK (D54); multi-target collection aliases (D57); hot tier: pinned splits + HNSW artifacts built by qdrant-edge (R20); affinity routing (M1.3, done 2026-09-27); W0 MCP server. Added 2026-09-26 (turbopuffer gap analysis): a per-collection unapplied-data budget with write backpressure (HTTP 429 / RESOURCE_EXHAUSTED with Retry-After; M1.3, D86); native delete-by-filter and patch-by-filter shared by the native API, ES _delete_by_query/_update_by_query and the Qdrant filter writes (M1.5, D87); a per-response performance block (M1.6, D92); an OpenAPI spec for the native REST API (M1.6, D101); a recall endpoint (M1.7, D92); a published limits page (M1.7, D88) | LangChain + LlamaIndex vector-store tests (ES and Qdrant backends, sparse and hybrid tests included) pass unmodified; the ADBC Flight SQL drivers (Python and Go) and Spice's Flight SQL connector pass against Operon (D49, D56); BEIR nDCG@10 within 1 pt of ES BM25; Recall@10 within 1% of Qdrant at equal hot-tier latency; results identical with hot tier on/off; every limit on the published limits page is exercised by a test at the limit and one past it, and the page is rendered from the enforced values (D88); a write that arrives while a collection is over its unapplied budget gets 429 with Retry-After, and a strong read at the budget is served from the live tail (D86); a filter write never touches a document written after its pin, and its cursor continues at the same snapshot (D87); the recall endpoint agrees with the M1.7 recall harness on the same collection within 0.01 (D92) |
| M2 | Production hardening → v1.0 | Metastore (§18 §2–§5): first, the MetaStore contract amendment: per-group commit_wal, bounded-skew stamps with GC claims, documented read orders (D59), the namespace on bare-id calls (D70), paginated list methods and a scoped change feed (D63), each with conformance and linearizability cases; then operon-meta-postgres and operon-meta-dynamodb (D58); a metastore fault matrix per backend and the floci, Alternator and nightly AWS deployment jobs (D60); catalog-scale fixes: an incremental catalog cache, dirty-set maintenance, snapshot builds that do not block apply, snapshots past 5 GiB, bounded-load placement (D63); one router for every resource kind, with placement keys for stream partitions and consumer coordination (D75). Tenancy: orgs and the ControlStore; AuthN on every surface: API keys (loam_<key_id>_<secret>, D65), TLS and mTLS; AuthZ through the Authorizer trait with built-in RBAC (roles over namespaces and collections, D66); quotas (request rate, ingest bytes, concurrent queries, storage, metadata operations; D65). GDPR erasure path (D68). Native stream API core (D72, §02 §7): HTTP and gRPC produce with idempotent producer ids, long-poll fetch, streaming subscribe, named consumers with committed offsets, stream create, list, describe and drop with retention settings; OTLP logs ingest over OTLP/HTTP and OTLP/gRPC (D73, §02 §7.1). Observability: Prometheus metrics, OpenTelemetry traces, structured logs, a live diagnostic-dump endpoint; multi-node clusters (3-node metastore over the network, node registry, affinity routing from M1.3); Kubernetes operator and Helm chart; zero-downtime rolling upgrades with format-version checks; restore from bucket; RustFS as the default self-hosted store and provider conformance CI (S3, GCS, Azure, RustFS; D61). AI data ecosystem workstream (§17): retained dataset tags (D52); credential vending for direct fragment reads; Ray Data datasource/datasink, Polars scan_loam() (experimental), loamdb.torch and the PySpark/Sail data source (D54). Turbopuffer gap adoption (D86–D103): the unapplied budget as a quota, with Prometheus metrics and a role-gated override (D86); conditional writes decided in log order, serving ES if_seq_no, Qdrant update_filter, a native condition and the Read Committed re-check of filter writes (D89, D87); collection branches in constant time with lineage GC (D90); one ranking-expression IR for the native API, ES function_score and Qdrant formula (D91); sampled continuous recall (D92); ES language analyzers, folding, pre-tokenized text and per-field BM25 k1/b (D93); f16, i8 and u8 vectors (D94); customer-managed keys by envelope encryption, which also gives crypto-shredding (D96); online index changes with background backfill (D97); cost-weighted per-collection concurrency (D98); prefix and cursor on list calls (D99); audit events (D100); a Go SDK (D101); usage counters in logical bytes (D103) | Unauthenticated and cross-tenant requests rejected on every surface (native, Flight SQL, Qdrant, ES, MCP, OTLP); requests over quota get HTTP 429 or gRPC RESOURCE_EXHAUSTED, and a revoked API key is refused within the key-cache TTL; the metastore conformance suite, the linearizability checks and the M0/M1 crash and fault gates pass identically on openraft, Postgres and DynamoDB (the full suite on floci, the single-item subset on Alternator), and each backend's metastore fault matrix matches its blessed table; the Store provider suite passes on RustFS, and the S3 fault matrix over RustFS matches the in-memory table; the nightly AWS deployment job is green (D60); after an erasure completes, the erased key is unreadable through every surface and absent from every object and object version in the bucket (byte scan), every retained or tagged manifest and the caches, and the erasure log records it (D68); no acknowledged stream write is lost and idempotent producers write no duplicate across node kills, a named consumer resumes from its committed offset across a node restart, and fetch returns records in partition order (D72); a write acknowledged with a consistency token is visible when read with that token on every other node and on openraft, Postgres and DynamoDB, including from a node routing with a stale membership view (D76); Fluent Bit's opentelemetry output, the OpenTelemetry Collector and Vector ship logs into a stream over OTLP with no plugin, and a link makes them searchable (D73); a catalog of 100k namespaces adds no per-change catalog re-read (the catalog cache updates incrementally); a rolling upgrade of a 3-node cluster under load loses no acknowledged write and fails no read beyond client retries; operator end-to-end tests on kind (deploy, scale, upgrade, node loss); restore-from-bucket drill; 24 h chaos soak (node kills, store faults) green; a tagged collection read through Ray Data, Polars, PySpark (Spark 4 and Sail) and loamdb.torch returns identical rows, and a torch run resumed mid-epoch on a different GPU count runs equal step counts on every rank and sees each unpadded sample exactly once per epoch (§17 §5.5); conditional writes pass a per-key linearizability check over offset order on every backend, and link apply and the tail agree on every op's outcome across crashes (D89); a branch never sees the source's later writes nor the reverse, and GC after dropping the source keeps every object a branch references (D90); function_score and Qdrant formula fixtures match the Elasticsearch and Qdrant oracles (D91); each added language analyzer produces ES _analyze's tokens on a fixed corpus (D93); f16, i8 and u8 collections pass the recall gate's method at their own width (D94); with a namespace key set, a byte scan of the bucket finds no plaintext of the corpus, destroying the key makes that namespace's bytes unreadable in WAL objects, segments, Lance files, splits and noncurrent versions while other namespaces in the same WAL objects still read, and a KMS outage fails only that namespace's requests, retryably (D96); a query on a field being re-indexed answers 409 until the backfill completes and never a partial result (D97); a query that finds its collection's 16 slots taken waits at most 800 ms, then gets 429 (D98) |
| M2.x | Cloud and BYOC → v1.1 | operon-meta-remote (a MetaStore client over gRPC with idempotency keys and a streamed change feed) and the hosted operon-control control plane, with the hosted ControlStore and the namespace directory pushed to gateways; both BYOC modes, names in clear (BYOC-managed-meta and BYOC-local-meta, D64); nodes cache only the namespaces they own (D63); OpenFGA authorization through the Authorizer trait, sharing one store with Lakekeeper (D67, default); (per-chunk envelope encryption moved to M2 with CMEK, D96); single-collection sharding, 1–256 shards fixed at creation (D95); cross-bucket, cross-region and cross-org collection copy with a re-key, and the asynchronous operations API (D90); console SSO, OIDC/JWT, per-org IP allowlists, PrivateLink and Private Service Connect (D102); audit exports and the console view (D100); billing on logical bytes (D103) | operon-meta-remote passes the metastore conformance suite and its fault matrix; a BYOC-managed-meta cluster keeps serving cached reads while the control plane is unreachable, fails writes with retryable errors, and resumes without losing an acknowledged write; a BYOC-local-meta cluster runs with the control plane down; captured control-plane traffic contains no document, vector, text or bucket credential from the ingested corpus; no OpenFGA grant makes a subject from another org effective (tenant fence); a token read through a gateway with a stale directory sees the token's writes (D76); a sharded collection returns the same results as the same data in one shard (BM25 scores bit-identical, D78), a write spanning shards is all-or-nothing across crashes, and a token read sees it on every shard (D95); a copy into another region reads back identical rows under the target's key (D90) |
| M3 | Graph (native GraphRAG) + durable execution | Mapped graphs over collections (over Iceberg tables once M4 lands); ID maps; CSR/CSC sidecars; edge overlay; ExpandExec (1–2 hop neighbor traversal, filtered) and shortest path as DataFusion operators; graph_expand / graph_neighbors SQL table functions; an expand stage in the native hybrid search API (seed by vector/BM25 → expand → rerank); graphs placed by (ns, graph) (D75); leiden/pagerank/wcc table functions; Operon-native graph-store adapters for LightRAG and the LlamaIndex property graph (D44). Resonate surface Phase A (§14): forked Resonate gateway + blob server on operon-store, namespace auth, origin-affinity routing | Our LightRAG and LlamaIndex property-graph adapters pass those frameworks' storage-backend test suites; expand results identical to a naive-adjacency reference (networkx) on fixture graphs, with the hot tier on and off; a GraphRAG retrieval benchmark (seed → 2 hops → rerank) tracked; Resonate TS + Python SDK suites pass unmodified; Resonate differential/linearizability harness passes against 3 gateways over one bucket with fault injection |
| M4 | Analytics (Iceberg) | Lakekeeper integration; Iceberg writes incl. DV writer; keyed tables; stream→table links; MVs with mergeable states; SQL over Flight SQL and the native API; Iceberg hot tier T0–T3, with tables placed by (ns, table) (D75); arrow segment encoding for table implicit streams; Resonate Phase B (§14: change stream, system.durable_* tables, execution graph, cluster-wide timer shards) | ClickBench (hot) over Flight SQL median within 2–3× ClickHouse OSS (ClickHouse is a performance reference, not a protocol); TPC-H SF100 completes; Spark, Sail, Trino, DuckDB and ClickHouse read Operon tables (keyed tables with deletion vectors included; Sail's DV reads contributed upstream, D55) through Lakekeeper's Iceberg REST catalog; Ray Data, Polars and the torch loader read tagged table snapshots; query results match DuckDB over the same Iceberg tables |
| M5 | Streams: Kafka gateway and changelog streams | Kafka wire-protocol gateway (D74, §02 §7.2), staged: Produce, Fetch, ListOffsets, Metadata and ApiVersions; idempotent producers on the native producer ids; consumer groups with committed offsets (classic protocol, then KIP-848); no Kafka transactions; the RisingWave companion integration over it (D22); Flight DoGet replay (Flight DoPut ingest into streams ships in M1, D49; the native stream API core in M2, D72); express WAL class; zone-aware routing; changelog streams from keyed collections/tables (upsert and full modes, fenced appends; D21), read through the native API | The Kafka client test suites (librdkafka, franz-go or the Java client) pass against Loam for each stage; RisingWave, Flink and Kafka Connect read and write Loam streams end to end; Jepsen-style tests (no lost acknowledged writes, no reordering, no duplicates with idempotent producers, native and Kafka) across node and AZ kills; changelog streams checked against a full-rebuild diff across crashes; native-API throughput benchmark tracked against AutoMQ OSS/Kafka on the same hardware |
| M6 | Scale & reliability | quorum WAL (Raft journals); datafusion-distributed; meta sharding: ShardedMetaStore over openraft groups (multi-Raft), Postgres databases, TiKV keyspaces (D124) or DynamoDB, metadata-only namespace moves with gateway buffering, size-class placement keys, openraft's catalog out of the monolithic snapshot (D63); operon-meta-tidb over the MySQL protocol (D58)operon-meta-tikv in R1, D124); multi-tenancy at scale (fair share, 1M namespaces); region DR. (OpenFGA moved to M2.x, D67; single-collection shards moved to M2.x, D95) | Failover RTO < 5 s (quorum) with RPO 0; chaos suite green for 72 h; 1M-namespace test; tenant-isolation tests; region restore drill; (the TiKV backend's conformance suite and fault matrix move to R1, D124); a namespace move under load loses no acknowledged write and fails no request beyond client retries, and a token read during the move sees the token's writes (D76) |
| W | Agent workspaces (§15), parallel track | W0 (with M1): MCP server. W1 (after M3): Git on the bucket with O(1) forks, code-index links, credential vending, agent telemetry. W2: operon-sandbox with copy-on-write environment images, package caches, registry proxy; the 100-agent fleet demo (§16), staged as α after M3 + W1 and β after M4 + W2. W3 (Phase C): jj backend, REAPI, VM snapshots. W1's repositories and W2's sccache wiring, crates proxy and lazy workspace mount are delivered by track GT (§36); W1's remaining scope (code-index links, credential vending, agent telemetry, the MCP gateway, session workflows) is unchanged. (§29's WeSQL milestones W1–W4 are a different track, in the plans README) | Per phase in §15 §13 |
| GT | Loam Git, build cache and crates mirror (§36), parallel track | GT1 (bucket WAL git core with group commit, git-remote-loam); GT2 (Smart HTTP v2 for stock clients, compaction, partial clone); GT3 (sccache direct and gateway paths with trust classes, the crates.io mirror); GT4 (loam-vfs scopes) and GT5 (replica caches, gossip) not yet planned. Carries W1's repository scope and parts of W2 | GT1: no lost acknowledged push under faults and 8 concurrent pushers, linearizable RefLog histories, ≥ 30 pushes/s per hot repository at 100 ms PUT latency; GT2: the git/gitoxide/libgit2/JGit matrix and the upload-pack differential; GT3: warm hit rate ≥ 90% of cacheable compilations, untrusted builds cannot write the trusted cache, cargo fetch with egress limited to Loam |
| R | Loam Live and the TiKV metastore (§20; TiDB SQL dropped by D260), parallel track | R1 (beside M1): operon-tikv; operon-meta-tikv (D124); one Live app in one keyspace with documents, indexes, QuickJS queries and mutations, the commit journal, live queries over the connect-rust sync API, a generated TypeScript client. R2: the multi-tenant router and keyspaces, mandatory BR log backup (PITR) to object storage (D131), the ControlStore on Live (D125), actions and scheduling, Python and Go clients. R3: the collections bridge (D129), auth through the unified auth plan (D111), Swift and Kotlin clients. R4: TiDB Operator v2 for PD and TiKV only (D179) and the Helm chart, BYOC (D127) | R1: the reactive correctness and transaction checkers pass under the fault matrix; the metastore conformance suite passes on TiKV; a TypeScript client sees a live update after a mutation; MySQL and Live clients share one playground without seeing each other's data. Later gates per §20 §18 |
| CLI | The loams CLI, installer and agent bootstrap (§30), parallel track | CLI1 (beside M1, working name operon): operon-cli with the output and exit-code contract, profiles, local stacks over operon dev/standalone with the engine registry, the H1 cache flags and sudo operon storage prepare for NVMe, .env.loam, embedded tested docs and snippets, the stdio bootstrap MCP server and mcp install for four agents. CLI2 (after D33's rename): the server split and the cli variant, prebuilt standard/full variants, cargo-dist releases with attestations and a minisign-signed manifest, install.sh at loams.dev, self-update, pkg add. CLI3 (after the unified auth plan): keys, login, agent tokens, companions, cloud stacks | CLI1: every client command's JSON output matches its schema snapshot, and every failure exits with its documented code; a stack created by the CLI serves Qdrant, Elasticsearch and Flight SQL clients on the ports in .env.loam; storage prepare refuses all nine unsafe device cases and the loop-device job puts H1 cache files on the mount; an rmcp client lists the seven bootstrap tools over stdio, and the canary secret appears in no tool output, JSON output or log; mcp install round-trips every agent's config with other entries preserved; every snippet runs green against operon dev. CLI2: install.sh passes its bats suite on dash, bash and zsh across Ubuntu, Debian, Fedora and macOS; self-update refuses a tampered archive and an unknown signing key; the cli variant's dependency tree contains no engine crate; the variant guard holds |
| RT | Loam Router and verification (§31), parallel track | RT0: specs (ShardMap, ReshardCutover), the compatibility inventory, operon-sqlrouter kernels, Lean partition proofs. RT1: the deterministic simulator (operon-detsim), the shard map on TiKV, PgDog rendering and adapters, trace validation, single-shard routing through PgDog. RT2: Lean merge/limit/aggregate and the oracle, the cross-shard differential, Postgres 2PC under D306, the in-doubt monitor, the change stream and a checked split. RT3: Vitess v24 with WeSQL (D302, D320). RT4: failover and commit specs, the multi-instance cutover, the full fault catalog. RT5: live resharding, nemesis runs, the compatibility matrix | Per phase in §31 §17; RT1 needs §28 P3, RT2 §28 P2b, RT3 §29 W2, RT4 §29 W3 and §28 P4c |
| FL | Loam Flow, the Event Fabric and Loam House (§32, §33), parallel track | FL1: Iggy + Fluss + Lakekeeper dev stack, CloudEvents on both, loam-fabric ingest with dedup, fluss_sink/loam_sink, tiering to Iceberg. FL2: chDB behind ClickHouse HTTP 8123 over Fluss and Iceberg, Tier 1 engines, chsurface-1.0, the differential harness. CN1: the connector registry and 21 ★ connectors. Later: FL3 Flow routes and MVs, FL4 graph alignment with M3, FL5 distribution and the native protocol, FL6 hard engines, CN2/CN3 | FL1: no acknowledged event lost across kills of Iggy, Fluss, tiering, the connectors runtime and ingest, no duplicates inside the dedup window. FL2: 100 % of the declared corpus minus approved allowlist entries, Rust and Python drivers, strict allowlist. CN1: every ★ capability proven by contract tests; Postgres CDC end to end converges |
| AP | The cordis console, Loams Desktop and the mobile apps (§37), parallel track | AP0: the app protos (instance, devices, approvals, operations, notifications, errors) and operon-apps-mock. AP1a: the console as a cordis v4 application with manifests, the loams.yml catalog, gated Connect services, typed slots, sandboxed third-party plugins and live reload; today's pages as plugins. AP1: Loams Desktop on Tauri 2 over the loams CLI, with a Rust network bridge, PKCE and the keychain, signed updates. AP2/AP3: native Android (Compose, connect-kotlin) and iOS (SwiftUI, connect-swift) for approvers, with pinned pairing, biometric-bound decisions and sealed push. AP4 (after the auth plan): the server side and the push gateway | AP0: buf lint and buf breaking pass; every service answers over Connect, gRPC and gRPC-Web from the mock; decision proofs are verified and replays refused; watch streams resume without loss. AP1a: the Playwright parity suite matches main; disabling a plugin removes its pages and streams without reload; a third-party sample cannot fetch and its calls carry an attenuated token. AP1: a terminal-created stack appears and can be stopped in the app; no secret appears in any command output or log; the capability allowlist check passes; a signed update installs. AP2/AP3: every shared scenario passes; pairing by QR pins a self-signed server; a pushed approval needs biometrics and is decided once; no payload is readable by the gateway |
| MT | Knative, Authentik and GitOps for self-hosted Loam (§38), parallel track | MT1 Authentik identity (RFC 8693 exchange, groups to OpenFGA, the Enterprise guard, the showcase off Keycloak); MT2 Knative for http-port with per-namespace tenancy and no meter, Eventing as an adapter; MT3 waves, health checks, Flux layout, k3s and CKE profiles | k3d e2e: sign in through Authentik and call the API with an exchanged token; an http-port function scales from zero under gVisor; the waves reach Healthy from an empty cluster on Argo CD and on Flux |
| Phase C | Research and deferred surfaces | Kafka transactions, on demand (Q4); SPFresh incremental IVF; DiskANN hot tier; WCOJ/factorized graph joins; region failover; Vortex hot encoding | — |
2. Testing strategy
- Deterministic simulation testing (DST) for meta, sequencer, journals, link apply and manifest commits — simulated network, clock, disk and object store (evaluate
madsimvsturmoil; Iggy and FoundationDB/TigerBeetle as practice references). Every merged PR runs a DST seed sweep. As built in M0 (D28): a seeded in-process simulation (operon-sim), not a bit-exact deterministic one: seeded workloads, network partitions (Router), worker crashes and object-store faults (FaultyStore::random) on a single-threaded runtime, with real time and real I/O underneath, and a Wing–Gong–Lowe linearizability check of the recorded histories. A failing seed reports its full schedule but may not replay exactly; a full DST port stays an option. Amended 2026-10-01 (D313, §31 §13): Loam's router control plane is written as sans-I/O machines and is simulated bit-exactly byoperon-detsim(one seeded RNG, simulated time, network and fault models, trace hashes, shrinking, swarm runs): 2 000 fresh seeds per scenario plus the regression corpus on every PR that touches it, 1 000 000 nightly. D28's seeded simulation stays for the engine. - Object-store fault injection: an
object_storewrapper injecting latency, 5xx, throttling (503 SlowDown), 412/409 on conditional writes, partial reads, and crashes between PUT and commit. From M2 the S3 fault matrix also runs over RustFS on every PR and must match the in-memory table; aStoreprovider conformance suite runs on RustFS per PR and on S3, GCS and Azure nightly (D61, §18 §4.4). Track GT'sWalStoreandRefLogrun under the sameFaultyStore, with the linearizability checker over ref-log histories (GT1). - Crash-consistency tests: kill -9 at instrumented failpoints (
failcrate) across write, segment, commit, compaction and GC paths. - Property-based tests (
proptest): WAL/segment encode/decode, offset index, manifest evolution, deletion bitmap algebra, CSR build vs. naive adjacency, tail merge vs. full rebuild. - Metastore conformance (D47, D60): one test suite written against
trait MetaStore— sequencer, catalog, manifest CAS, leases and fencing, with the linearizability checker and the relaxed contract of D59 — runs against every backend: openraft single-node and 3-node; Postgres and DynamoDB from M2 (DynamoDB's full suite on floci, the single-item subset on ScyllaDB Alternator 6.2.3 viaBackend::capabilities(), real DynamoDB nightly); TiKV (operon-meta-tikv) from R1 ontiup playground(D124; the M6 TiDB backend is superseded, and there is no TiDB to test, D260). Each backend also has a metastore fault matrix (faultsBeforeSend,AfterApply,Undetermined,Conflict,Throttle,Race,Delay, injected in process) with a blessed expected table, and nightly toxiproxy soaks. The seeded simulation and the crash/fault gates run on each backend, and so does the consistency-token gate: a token read on every other node and backend, during a move and with a stale directory, sees the token's writes (D76, §18 §3.5). Container services run one at a time (§18 §4). The suite is theoperon-meta-conformancecrate; each backend expandsmetastore_conformance!in its tests (M1.2a). - Differential testing:
- Hot tier on/off must return identical results (random disabling in CI).
- Operon vs. reference engines: ES (BM25 rankings on fixed corpora), Qdrant (recall), a naive-adjacency reference (graph expansion on fixture graphs), DuckDB over the same Iceberg tables (analytical query results), DataFusion-on-Parquet baseline.
- Sharded SQL (§31 §14.2): the same SQL stream on an unsharded engine and through PgDog (2 and 4 shards) or vtgate, with the Lean oracle as the third opinion for merges and aggregates; documented router deviations live in an allowlist with their sources.
- Loam House (§32 §10, D349): the declared ClickHouse surface against a CI-only reference
clickhouse-serverpinned tochdb-core's ClickHouse version, with a strict allowlist (D348).
- Jepsen-style tests for native stream semantics (M2 gates the API core on lost writes, duplicates, committed-offset resume and partition order; M5 adds node and AZ kills) and cross-object consistency tokens (write via the native API, Flight
DoPut, Qdrant or ES → strong read in another surface must reflect it); changelog streams checked against a full-rebuild diff (no lost or duplicated change across crashes).- Durable execution reuses Resonate's own harness: the differential test against the executable oracle, the linearizability search, and the Lean trace checker run against Operon's server plugin. Its trace-checking approach is also a reference for our own DST checks.
- Sharded SQL (RT5, D315):
operon-nemesisruns the simulator's workloads and checkers (bank, list-append, split, liveness) against PgDog, Vitess, Loam Postgres, WeSQL, TiKV and RustFS with kills, pauses, partitions and clock skew.
- Conformance suites per surface (client libraries, ADBC drivers and framework integrations listed in each milestone gate; the Kafka client test suites in M5) — these define compatibility scope.
- Security tests (from M2): authentication and authorization on every surface, tenant isolation, TLS/mTLS configuration matrix.
- Cloud deployment test (nightly, from M2; D60, D62): floci runs Loam from its shipped infrastructure-as-code module with S3, DynamoDB, IAM enforcement, and workers as Lambda functions triggered by EventBridge Scheduler and SQS; lease takeover after a killed worker, least-privilege IAM and an end-to-end erasure are checked. Lambda workers are a CI harness, not a production deployment.
- Benchmarks in CI (nightly): BEIR, VectorDBBench, a GraphRAG retrieval benchmark, ClickBench, TPC-H, native-stream throughput; tracked for regressions, including S3 request counts per operation (cost regressions are bugs).
- Limits tests (from M1.7, D88): one table in code holds every enforced limit; a test renders the published limits page from it, and each limit has a test at the value (accepted) and one past it (refused with the documented error). From M2 the quota and concurrency limits (D86, D98) join the table.
- Formal specifications and kernels (from RT0, D308–D312): TLA+ specs of the protocols Loam builds or orchestrates (
ShardMap,ReshardCutover,CrossShardCommit,PrimaryFailover;RouterSessionoptional) inspec/tla/router, checked by TLC and Apalache on every PR that touches them, with expected-violation variants for unsafe configurations; trace validation of simulation and real-run traces against the specs, and a test that every spec action has a code emitter; Lean 4 proofs of the pure kernels (key-range partition, merge,LIMIT/OFFSET, aggregate decomposition) compiled into an oracle that the Rust mirror and the routers' cross-shard results are diffed against; compatibility inventories of what each bought router asks of its engine (conformance/router/, D309).
3. Risk register
| # | Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|---|
| 1 | Scope: a multi-model engine drifts into several products | Medium | High | Protocol footprint narrowed to Qdrant, an ES subset, Flight SQL and the native API (D42); strict milestone order; compat scope defined by conformance suites, not feature requests; v1.0 is M1 + hardening, useful on its own |
| 2 | Lance vendor control / API churn | Medium | High | Pin format 2.1 and crate versions; CollectionStore trait boundary; contribute upstream; Vortex/own format as long-term fallback |
| 3 | Lance small-commit cost | High (if misused) | Medium | Never commit per write; batch via Operon log |
| 4 | iceberg-rust write gaps (DVs, RowDelta) | Certain | Medium | RisingWave fork + own DV writer, upstreamed; append-only tables unblocked meanwhile |
| 5 | Lakekeeper requires Postgres | Medium | Medium | Verify pluggable backend; the M2 Postgres metastore backend means many deployments already run Postgres, which Lakekeeper can share |
| 6 | Meta Raft becomes the bottleneck at high partition/namespace counts | Medium | High | Batch proposals per node flush; coalesce offset commits; MetaStore trait (D47) lets deployments move to Postgres or DynamoDB (M2) or TiKV (track R1, D124); the O(N) catalog readers and snapshot stalls are fixed in M2 and the metastore is sharded by namespace in M6 (D63, §18 §5) |
| 7 | Cold-query latency disappoints (0.5–1 s) | Medium | Medium | Affinity routing, published hot artifacts, prewarm API, clear "pin" UX; document the cold/warm contract |
| 8 | Cross-AZ costs on quorum class | Certain | Low–Medium | Default classes avoid it; placement hints; document pricing |
| 9 | Qdrant fork maintenance (large codebase) | Medium | Medium | Take minimal subset (HNSW, quantization, filter planner); evaluate qdrant-edge boundary first |
| 10 | Quickwit fork divergence | Medium | Low–Medium | Fork only storage/directories/DSL crates; pin; periodic rebase |
| 11 | Competitive: HelixDB / LanceDB / Milvus / Databend / Apache Fluss / Spice move into the same slot (Spice sells SQL + vector + BM25 search for agents on DataFusion, as a cache over other systems) (Databend now markets analytics + search + vector "for agents" on S3; Fluss is an ASF streaming lakehouse with Lance tiering "for AI") | Medium | High | Speed to v1.0; differentiate on hybrid retrieval (vector + BM25 + graph expansion in one planned query) + Qdrant/ES compatibility + Flight SQL + Iceberg + fully open serving layer (Databend has ELv2 parts) + S3-only state (Fluss keeps data on tablet-server disks) + durable execution; consider collaboration with HelixDB |
| 12 | Compatibility long tail (ES DSL) | Medium | Medium | ES scope is what the framework suites send (D48); usage-driven additions only; clear, documented unsupported-feature errors |
| 13 | S3 provider behavior differences (conditional writes, Express/Rapid semantics) | Medium | Medium | Store provider conformance suite on RustFS per PR and on real S3/GCS/Azure nightly; the S3 fault matrix over RustFS must match the in-memory table (M2 gate, D61) |
| 14 | Correctness bugs in tail merge / consistency tokens | Medium | High | DST + differential tests + Jepsen-style cross-surface checks; the token gate runs on every backend and node, during moves and with stale directories (D76) |
| 15 | Resonate upstream churn or abandonment (seed-stage vendor, git-only crates, protocol still evolving) | Medium | Medium | Pin a git revision; Operon code only behind the ResonateServer trait; the protocol is formally specified, so a fork stays checkable; Phase A has no Operon-specific storage format |
| 16 | Scope: durable execution adds another surface | Medium | Medium | Phase A is integration of a forked server, not new storage; Phase B is gated on M4; no workflow DSL or SDK of our own (§14 §7) |
| 17 | Scope: agent workspaces (§15) add a Git server, a filesystem layer and a registry proxy | High | High | Run as a parallel track after M1 ships; W0 (MCP) alone first; buy nydus, gitoxide and runtimes; no runtime, POSIX FS or GitHub UI of our own |
| 18 | Adopters expect a Kafka-compatible endpoint (Connect, Flink, existing producers) before M5 | Medium | Medium | v1.0 has native HTTP/gRPC produce, Flight DoPut and OTLP logs (D49, D72, D73), and Fluent Bit, Vector and RisingWave's sinks write to Loam with no plugin (§02 §7.3); Iceberg tables are writable by external stream processors via Lakekeeper; the Kafka gateway ships in M5 (D74) |
| 19 | Operon-native graph adapters (LightRAG, LlamaIndex) need upkeep as the frameworks change | Medium | Low–Medium | Contribute the adapters upstream; gate on the frameworks' own storage-backend suites so drift shows up in CI |
| 20 | Pluggable metastore backends diverge in semantics | Medium | High | One MetaStore conformance suite with the linearizability checker for every backend, plus a blessed fault matrix per backend (§2 item 5, D60); backends expose semantic operations, not raw KV (D47); the contract every caller may rely on is the relaxed one (D59) |
| 21 | DataFusion/Arrow version lockstep with ecosystem engines (Lance, Sail, Spice, Polars' own Arrow) | High | Medium | Integrate over Flight, Iceberg REST, Lance files and Arrow PyCapsule only (D51); no engine crates linked; pin versions and test each integration in CI |
| 22 | Callers come to depend on the openraft backend's stronger guarantees (cross-partition atomicity of one WAL object, one monotonic clock, snapshot reads), which DynamoDB and a sharded metastore do not give | Medium | High | D59 is written into the trait docs first (M2); conformance and linearizability cases check the relaxed contract only; M2's contract task audits every caller (append_many is the known one, §18 §3.1); the DynamoDB backend runs every gate |
| 23 | Emulator and CI-target fidelity: floci is seven months old and never produces conflicts or throttling; Alternator has no transactions; RustFS 1.0 is new, with open regressions and unverified list-after-write, single-disk fsync and SSE ETag behaviour | Medium | Medium | Pin versions; the conformance and provider suites are the arbiter; conflicts, throttling, 409 and 503 are injected in process; real DynamoDB and S3 nightly, and each emulator release is checked against them before the pin moves (D60, D61) |
| 24 | Router complexity: a directory, a sharded metastore and namespace moves add distributed state and new failure modes | Medium | High | Staged (§18 §5.7): M2 fixes only the O(N) readers and snapshots; ownership stays a soft hint, so a stale directory routes suboptimally, never wrongly; moves copy metadata only and are fenced by the namespace lease; the trait changes land in M2 so M6 adds no signature change |
| 25 | Scope: v1.0 grows with a second new metastore backend (DynamoDB), tenancy, erasure, the catalog-scale fixes, the native stream API core and OTLP logs ingest (D72, D73) | Medium | High | The contract task comes first and bounds the backend work; the stream API builds on the M0.3 routes and the existing sequencer, and replay, express and changelog streams stay in M5; BYOC, the hosted control plane and OpenFGA are M2.x (v1.1), and crypto-shredding comes with customer-managed keys (D96); sharding is M6, and the TiKV backend is in track R (D124; it supersedes the M6 TiDB backend) |
| 26 | An erasure misses a copy of the data (a retained or tagged manifest, a WAL object shared with other namespaces, a cache, a noncurrent version) | Medium | High | The inventory in §18 §9 drives the M2 erasure gate (a byte scan of the bucket and the caches); noncurrent object versions are deleted by version id and verified, with lifecycle rules as a backstop; crypto-shredding with customer-managed keys in M2 (D96, amending D69) covers what a scan cannot |
| 27 | Kafka protocol fidelity: clients depend on details a partial broker gets wrong (API version negotiation, error codes, idempotent-producer sequence rules, group rebalance timing, compacted internal topics for Kafka Connect) | High | Medium | Stage the APIs; ApiVersions advertises only what is implemented; each stage gates on the clients' own test suites (librdkafka, franz-go or the Java client) and on RisingWave, Flink and Kafka Connect end to end; no transactions in M5; Nisshi and the Kafka protocol spec as references (D74) |
| 28 | A slow or stalled link lets the unapplied backlog grow: strong reads fall back to range tails that re-read the log per query, then fail with Unavailable | Medium | High | A per-collection unapplied budget refuses writes with 429 and Retry-After before the live tail overflows (M1.3 Task 15, D86); writes with the bulk-load override (Operon-Backpressure: off, up to 4× the budget) can still overflow it, and strong reads then use range tails or answer Unavailable; the backlog is reported on every write response and in CollectionInfo; M2 makes it a quota with metrics and alerts; sharding (M2.x, D95) removes the one-link-task-per-collection ceiling |
| 29 | Filter writes on a leaderless WAL: in M1 a key that changes between the pin and its batch is still deleted or patched, which users may not expect | Medium | Medium | Documented on the limits page and in the API reference, as ES _delete_by_query without version checks behaves; the pin and cursor keep the matched set fixed (D87); M2's conditional ops re-check the filter at apply, giving Read Committed (D89) |
| 30 | Customer-managed keys: a revoked, deleted or unreachable KMS key makes a namespace unreadable, and the key-caching layer adds a new failure mode | Medium | High | Revocation and loss are documented as data loss (as turbopuffer documents them); key-encryption keys are cached in memory for an hour, so a short KMS outage does not stall a warm namespace; a KMS failure affects only that namespace, retryably; the M2 gate checks shredding and isolation by byte scan (D96) |
| 31 | Scope: v1.0 grows again with the turbopuffer gap adoption (D86–D103) | High | High | Only backpressure, filter writes, the limits page, the performance block, the recall endpoint and the OpenAPI spec enter M1, each as one self-contained task; sharding, copy, SSO, private networking and billing are M2.x; each M2 item has its own gate, so any one of them can slip to M2.x without blocking the rest |
| 32 | Sharded collections break a single-collection assumption somewhere (global BM25 statistics, one manifest per collection, one link task, pins and scan plans) | Medium | High | Shards stay inside one collection's partition group, so writes stay atomic (D59); statistics are summed across shards before scoring (D78), and the M2.x gate compares a sharded collection with the same data in one shard, scores bit-identical (D95); pins and scan plans list every shard's manifest |
| 33 | The binary name collides with another tool on users' machines | Low | Medium | The binary is loams (D401), not loam; ~/.loam/bin first on PATH; loams version identifies itself |
| 34 | Release signing key compromise or loss (Q282) | Low | High | GitHub release environment with required reviewers; two embedded public keys for rotation; attestations as an independent check |
| 35 | storage prepare formats the wrong disk | Low | High | Never outside an explicit sudo command; nine refusal rules; typed confirmation; pure planner tested on fixtures; loop-device CI |
| 36 | PgDog's v0.1.x maturity for sharded production databases (§31 row 1) | Medium | High | Differential and nemesis tests; beta until Q305; unsharded fallbacks |
| 37 | Vitess on WeSQL fails the compatibility gate, or Vitess drops MySQL 8.0 support after v24 (§31 rows 2–3) | Medium | High | RT3 inventory first; fixes in the WeSQL fork; pin v24; Q302 rebase to 8.4; D317 fallback |
| 38 | Specs or simulation models drift from code and engines (§31 rows 6–7) | Medium | Medium | Trace validation, the action-coverage test, contract suites against model and reality |
| 39 | Two durable logs (the Loam WAL and the Event Fabric) confuse users and double operations (D331, D405) | Medium | Medium | One home log per record; Flow routes as the documented path; bridges with stated guarantees |
| 40 | Iggy pre-1.0 protocol churn and local-disk storage (D332) | Medium | Medium | Pinned sets; bounded retention; the tiered-storage contribution |
| 41 | ClickHouse "compatible" read as 100 % (D347) | Medium | Medium | The declared surface page with published pass rates; the strict allowlist (D348) |
| 42 | Scope: track GT adds a Git server, a WAL, a build cache and a registry mirror | Medium | High | Buy sccache and stock git (as processes); one WAL protocol for every writer; GT1 ships a serverless helper before any server; GT4–GT5 stay unplanned until GT2's measurements |
| 43 | cordis has one maintainer and v4 is a release candidate (§37 §14 risk 1, Q427) | Medium | High | Exact pins behind @loams/cordis, a patch log, a vendoring trigger (30 days with a patch, 90 days of upstream silence); a small used surface |
| 44 | A malicious console plugin (§37 §5.6) | Medium | High | Third-party plugins only in opaque-origin frames with connect-src 'none' and attenuated, audited tokens; host-decided tiers; owner-only installs; approval gates in protected environments |
| 45 | A stolen phone approves a destructive operation (§37 §7.3) | Low | High | Decision proofs need biometrics or the passcode per signature; enrolment changes void the key; revocation from any session |
| 46 | connect-kotlin is beta (§37 §14 risk 4) | Medium | Medium | Thin use (unary and server streams); shared conformance scenarios; deliberate upgrades |
| 47 | Authentik moves a feature Loam uses into its Enterprise tree, or changes the licence | Low–Medium | Medium | The guard (D458) on every bump; IdP-agnostic gateway, Keycloak as fallback; pin minors |
4. Next 90 days (from 2026-09-25)
- M1.2a: introduce
trait MetaStore(with the shared metastore types and errors inoperon-common), make the openraftMetaClientits first implementation, move every downstream crate onto the trait, and re-run the M0 and M1.1 gates. - M1.2–M1.3: the query engine (DataFusion operators, tail merge, sparse/dense fusion, BM25), Flight SQL with
DoPutingest, the hot tier and affinity routing. M1.3 adds the unapplied-data budget and write backpressure as its Task 15 (D86). Status (2026-09-27): M1.2a, M1.2 and M1.3 are done: split merges, Lance compaction, hot HNSW artifacts, pinned splits, fragment prefetch, the metastore over the network,operon cluster, rendezvous routing with read forwarding, the hot on/off differential harness and write backpressure (D104–D110). - M1.4–M1.6: the Qdrant gateway, the trimmed Elasticsearch gateway (with native delete-by-filter and patch-by-filter, D87), the SDKs (with the performance block and the OpenAPI spec, D92, D101) and the MCP server, in parallel.
- M1.7: the published limits page and the recall endpoint (D88, D92), the M1 exit gates, then the rename to Loam (D33).
- M2 kickoff: the
MetaStorecontract amendment first (D59, D70; §18 §3.4), then the Postgres and DynamoDB metastore backends with their CI targets, AuthN/AuthZ with API keys and quotas, telemetry and the Kubernetes operator. - Parallel tracks (added 2026-10-02): track CLI: CLI1 beside M1 (it depends only on
main), CLI2 after the rename PR. Track FL: FL1 (after M1.2 and #170/#171), then FL2 once Q333 is answered; CN1 beside FL2. Track AP: AP0 now (it depends only onmain); AP1a and AP2/AP3 against the mock; AP1 after AP1a's host and CLI1. Tracks RT, GT and MT start as their plans' dependencies land: RT0 needs nothing new; GT1 needs no other plan, but whether track GT starts now or in W1's slot after M3 is the owner's Q395, which gates GT1's start; MT1 needs §19's M2 identity work. All interleave on the one-build machine (D127).