§13 Decision Log

Decisions made so far and open questions

Living document. Newest decisions at the bottom of each table.

Decisions

IDDateDecisionRationaleStatus
D12026-09-22Object storage is the only durable source of truth; compute is statelessturbopuffer/WarpStream/ClickHouse Cloud model; cost; operabilityApproved; amended by D130 (holds for the retrieval engine, not for Loam Live on TiKV)
D22026-09-22OLTP (Postgres) is out of scopeMillisecond interactive transactions conflict with S3-native design; Neon needed a Paxos tier; users keep a small Postgres and stream CDC inApproved (user); amended by D130 (Loam Live, an OLTP product line on TiKV, is outside the retrieval engine)
D32026-09-22Scope = combined Kafka + Elasticsearch + Qdrant + Neo4j + ClickHouseUser's target pain point: AI apps deploying all of theseSuperseded by D42 (2026-09-25)
D42026-09-22The log is the spine; every write lands in a stream; other objects are link-maintained materializationsZero-ETL, consistency tokens, one durability mechanismApproved (§01); amended by D89 (conditional ops are decided in log order)
D52026-09-22Five object kinds: stream, table, collection, graph, linkEach replaces one system (links replace connector glue)Approved (§01)
D62026-09-22Iceberg for tables (via Lakekeeper), Lance for collectionsUser requirement (Iceberg for analytics); Lance is the only Rust substrate with columns + vectors + versioningApproved (§01)
D72026-09-22Tantivy for full-text (Quickwit storage/DSL/aggregation crates forked), not Lance FTSMaturity, aggregations, ES DSL translation already exists in QuickwitApproved (§01 amendment)
D82026-09-22Vectors: Lance IVF durable tier + Qdrant-derived HNSW hot tierLance wins storage/cost/versioning; Qdrant wins serving latency/filtered recall/freshnessApproved (user); amended by D94 (f16, i8 and u8 vector columns)
D92026-09-22Hot-tier model applied uniformly, including Iceberg + Lakekeeper (T0 file index, T1 Parquet cache, T2 hot projections, T3 tail)User request; ClickHouse-like latency and freshness on open IcebergApproved (user, §04)
D102026-09-22Metastore = embedded Raft (openraft) by default; pluggable FoundationDB/PostgresKafka-rate metadata cannot run on S3 CASApproved (§01); amended by D47 (semantic MetaStore trait), D58 (backends: Postgres and DynamoDB in M2, TiDB in M6), D71 (FoundationDB dropped) and D124 (a TiKV backend in R1 replaces TiDB); D260: no TiDB backend in any milestone, TiKV is the scale-out backend
D112026-09-22Apache-2.0 license; no AGPL/BSL/SSPL/ELv2 dependenciesBig-company adoptionApproved
D122026-09-22WAL classes standard / express (2-of-3 zonal buckets) / quorum (Raft journals). Renames the zonal class shown in the approved §01 diagram to express and makes it multi-AZ durableAutoMQ-grade reliability in OSS without stateful broker disks; AutoMQ OSS only has S3 WALApproved (§02)
D132026-09-22Compatibility scope defined by external conformance suites (client libs, framework integrations)Prevents unbounded compat long tailApproved; limits are defined the same way, by tests (D88)
D142026-09-22Build order: M0 foundation → M1 collections → M2 graph → M3 Kafka → M4 analytics → M5 scaleFollows the AI-app pain point (ES+Qdrant+Neo4j first)Superseded by D46 (2026-09-25)
D152026-09-22DataFusion as the single query engine; datafusion-distributed (not Ballista)Extensibility, ecosystem, interactive distributed executionApproved (§05)
D162026-09-23Metastore Raft: openraft pinned to =0.10.0-alpha.34; the local Raft log, vote and snapshot pointer in redb; log entries and snapshot bodies encoded with postcardopenraft 0.10 alphas change APIs between releases; redb is pure Rust and ACID (RocksDB would add a C++ build)Approved (M0.2 plan)
D172026-09-23Meta snapshots live only in object storage, one set per node: meta/snapshots/<node_id>/<term>-<index>.snap (amends §01 §6)Nodes snapshot independently; per-node paths let each node delete its previous snapshot without breaking another node's restartApproved (M0.2 plan)
D182026-09-23Metastore ids are dense u64 counters allocated by the state machine; stream ids are unique across the clusterapply must be deterministic (no random ULIDs); a cluster-unique stream id makes (stream, partition) a complete keyApproved (M0.2 plan)
D192026-09-24Durable execution via the Resonate protocol (§14): fork Resonate's Rust gateway + blob server (pinned git revision) and run it in the gateway role over operon-store, state under ns/<ns>/durable/, no metastore traffic. Phase A in M2, Phase B (change stream, search tables, execution graph, timer shards) in M4Agents need durable runs next to their memory; Resonate is Apache-2.0, formally specified, already S3-native on the same object_store crate, and plugin-based, so this is buy, not buildApproved (user); Phase A moved to M3 (D46); amended by D138–D145 (§21: embedded in the binary, SQLite/TiDB/TiKV backends, track D)
D202026-09-24arrow segment encoding for schema'd streams (idea from Apache Fluss); the encoding field is reserved in the first WAL/segment format (M0.3), arrow ships in M4Column pruning and no decode for links and tails; reserving the field now avoids a format breakApproved (user)
D212026-09-24Changelog streams from keyed tables and collections (upsert / full with before images; idea from Apache Fluss), written by link apply with fenced appends; M3CDC out of Operon and incremental consumers that need deletes; the PK index already locates before imagesApproved (user); moved to M5 and read through the native API (D43)
D222026-09-24RisingWave is the supported companion stream processor, run alongside Operon (not embedded): it reads topics and changelog streams (upsert / debezium-json wire formats) and writes Iceberg tables via Lakekeeper or Kafka topics; externally written keyed tables have one writer classKeeps Operon out of stateful stream processing (§00 §7) while giving users joins/windows/MVs; RisingWave is Apache-2.0 Rust and already speaks Kafka + Iceberg RESTDeferred past v1.0 with the Kafka gateway (D43); returns in M5 with the Kafka gateway (D74)
D232026-09-24Operon as the state plane for agent sandboxes (§15): Git on the bucket with O(1) forks, copy-on-write environment images (nydus), package caches, registry proxy, MCP server; runtimes are integrated, never builtSandboxes are stateless compute; their state (code, deps, checkpoints, memory, traces) fits Operon's bucket model, and agents already speak GitApproved (user)
D242026-09-24Agent sessions: every session is a Resonate durable execution and is stored in Operon (workspace branch, harness session files, sessions stream → tables and a searchable collection); operon-sandbox runtimes behind a Runtime trait (microsandbox default, Firecracker fleet, gVisor for Kubernetes without KVM, process for dev); MCP surfaces target the 2026-07-28 stateless spec, with a gateway that retrieves tool definitions via hybrid search + tool graph and exposes find_tools / call_toolCrash-proof sessions without repeated model calls; Rust microVMs with fork; the stateless spec makes MCP horizontally scalable and forbids per-connection tool listsApproved (user) (§15, §16)
D252026-09-24WAL objects live at the cluster level: wal/<class>/<node_id>/<ulid>.wal, not under ns/<ns>/One WAL object holds chunks from many namespaces (§02 §3 step 1); per-namespace objects would multiply PUTs by active namespaces. Per-namespace client-side encryption (Phase B) will need per-namespace WAL objects or per-chunk envelope keysApproved (M0.3 plan, ruling 4); amended by D96 (per-chunk envelope keys in WAL format version 2, M2)
D262026-09-24A segment is one offset-index entry; its per-batch index lives in the segment footer (WAL-backed entries stay one per chunk)Keeps metastore state proportional to segments, not batches (§02 §5); readers pay one cached footer read per segmentApproved (M0.3 plan, ruling 5)
D272026-09-24WAL commit window (CommitWal carries the object's creation time; commits older than 15 min are rejected as stale, dedupe records pruned after 30 min) and deterministic retired-object tracking in the metastore (live chunk counts per WAL object; a retired set for unreferenced WAL objects and trimmed segments, collected by GC)Bounded dedupe memory without ever committing a retried object twice; reachability GC (§03 §7) without listing the bucketApproved (M0.3 plan, rulings 6 and 7)
D282026-09-24M0's deterministic-simulation harness is a seeded in-process simulation, not a full deterministic simulator (madsim/turmoil): operon-sim runs a 3-node metastore over the in-process Router, log writers, a reader and a worker on one single-threaded runtime, with seeded workloads, node isolation, worker crashes and FaultyStore::random store faults, and checks histories with its own Wing–Gong–Lowe linearizability checkeropenraft, redb and object_store do real I/O and spawn blocking work, so bit-exact replay would need a madsim port of every dependency; a seed sweep on every PR finds the same classes of bugsApproved (M0.4 plan, rulings 1 and 6); scope amendment proposed 2026-10-01 by D313 (bit-exact DST for sans-I/O control-plane machines, §31)
D292026-09-24The kill -9 crash gate aborts at named failpoints (fail crate, compiled in only with the failpoints feature) in a child operon dev process, plus a random-time SIGKILL loop under loadAborting at a named point covers every step of the write, commit, segment, link, GC, snapshot and retention paths deterministically; random kills cover points nobody named; release builds carry no failpointsApproved (M0.4 plan, ruling 2)
D302026-09-24One link-apply task per link in M0, covering all source partitions (task lease task/link/<link_id>)Enough for exactly-once; splitting a link into (link, partition range) tasks (§09 §3) is a throughput optimisation for M1, and changes only the task key formatApproved (M0.4 plan, ruling 3)
D312026-09-25Qdrant sparse vectors are in M1 (pulled from Phase B): named sparse vectors with Qdrant's idf modifier, exact dot-product scoring over the live documents of the read snapshot, fusable with dense and text retrievers; stored as a Lance column plus postings and weights inside the Tantivy splits, with no hot artifact. ES sparse_vector, sparse recommend/discover/context/MMR and the optimized posting-file index (§06 §6) stay Phase BThe LangChain and LlamaIndex Qdrant gates include 172 sparse/hybrid server tests; the split already gives postings, deletes, merges, the tail and pinned files, so the only new code is the scorer; qdrant-edge's sparse index is private and local-onlyApproved (owner; M1 overview A14, A26–A30, R22)
D322026-09-25The ES gateway serves the endpoints elasticsearch-py's wipe_cluster test fixture calls (empty snapshots, data streams, templates, cluster settings, pending tasks) and wildcard index deletes by default (destructive_requires_name = false, configurable), so the elasticsearch-py client suite is gated in M1Without them every elasticsearch-py server test errors in setupSuperseded by D48 (2026-09-25)
D332026-09-25The product is renamed Loam, package name loamdb (PyPI, npm, crates.io), in one mechanical rename after M1 completes; M1 keeps the working names (operon-*, operon-client, @operon/client) and publishes nothingoperon on PyPI and the @operon npm scope are takenApproved (owner; resolves Q1; M1 overview A32); package name superseded 2026-10-01 by D400 (loams on crates.io, PyPI and npm @loams), binary loams (D401), repository ostrium-labs/loams (D406)
D34 (number assigned at merge)2026-09-24Lance detached versions are the R7 mechanism: a collection's Lance dataset has one mainline version (the empty version 1); every later Lance commit is detached (_versions/d<id>.manifest), built from exactly the parent collection manifest's lance_version, and the Operon manifest CAS is the only linearization pointLance's mainline commit rebases onto newer versions even with no retries, and version paths are deterministic, so a zombie's or crashed writer's version would be absorbed or would block every later commit; detached versions have unique paths and never rebaseApproved (M1.1 plan, ruling 1)
D35 (number assigned at merge)2026-09-24Operon GC owns Lance deletion; Lance cleanup is never run (nor auto-cleanup). CollectionGcRoots computes Lance reachability from the Lance manifests of the retained chain and collects only Lance's known file classesLance cleanup lists only mainline manifests and would delete every file only detached versions reference once it is 7 days oldApproved (M1.1 plan, ruling 2; overview A15)
D36 (number assigned at merge)2026-09-24Stable row ids join Lance and Tantivy: splits store the Lance stable row id in a _rowid fast field, and typed fields exist only in Tantivy (Lance holds _source, system columns and vectors)Row ids survive Lance compaction, so splits and the PK index never change when Lance compacts; Tantivy filters, sorts and aggregates, and ES fields may be single- or multi-valued, which has no natural Arrow typeApproved (M1.1 plan, rulings 4 and 8)
D37 (number assigned at merge)2026-09-24The PK index is derived state with a watermark (the manifest and applied offsets it reflects), written only after the manifest CAS and repaired from per-commit PK delta objects (or rebuilt from Lance)The CAS is the commit point; the PK index may lag it by one commit, and the stream may be trimmed, so the deltas cannot be re-derived from itApproved (M1.1 plan, ruling 7)
D38 (number assigned at merge)2026-09-24Collection time travel: every manifest is kept for time_travel_retention (24 h) after it is superseded, plus the last keep_manifests, with every object they reference; implicit streams are trimmed only below the oldest retained manifestPinned reads (ES PIT, Qdrant snapshots) need a manifest's objects and the stream tail after its offsets for as long as it is retained, with no metastore holdApproved (M1.1 plan, ruling 12); amended by D90 (GC reachability counts every branch of a lineage, M2)
D39 (number assigned at merge)2026-09-24Lucene-compatible analyzers (standard, english, simple, whitespace, keyword), with the original Porter stemmer ported from Porter's C reference, not Porter2The BEIR gate (within 1 nDCG@10 of ES) depends on tokenization and stemming matching LuceneApproved (M1.1 plan, ruling 25); amended by D93 (ES language analyzers, folding, pre-tokenized text and per-field k1 and b in M2)
D40 (number assigned at merge)2026-09-24JSON field layout: a JSON field n is five Tantivy fields (n, _text.n, _date.n, _null.n, _count.n)Exact match, full text, date ranges, Exists, IsNull, IsEmpty and ValuesCount on any path without a per-key schema or backfillApproved (M1.1 plan, ruling 26)
D41 (number assigned at merge)2026-09-24Sparse vectors are a Lance column plus split postings and weights: Lance _sparse_<j> (Struct<indices, values>, the source of truth), split fields _sparse.<name> (u64 postings) and _sparse_w.<name> (bytes fast weights)Reuses split writes, delete bitmaps, merges, the tail index and pinned splits unchanged; the Phase B posting-file index replaces the two split fieldsApproved (M1.1 plan, ruling 27; overview R22)
D422026-09-25Protocol footprint narrowed (architecture review §3). Operon speaks: its native REST/gRPC API, Arrow Flight SQL, the Qdrant REST + gRPC API, and a targeted Elasticsearch subset. The Kafka wire protocol is deferred past v1.0 (D43); Neo4j Bolt/Cypher is dropped (D44); the ClickHouse HTTP interface and dialect are dropped (D45). Replaces the scope in D3Emulating five products' protocols is a five-product trap; Qdrant and ES carry the AI-framework adoption, Flight SQL is the clean native path for every language, and Iceberg already gives external engines SQL accessApproved (owner; architecture review); the Kafka gateway is planned for M5 (D74)
D432026-09-25Kafka gateway deferred past v1.0; streams are reached through the native streaming API (HTTP and gRPC produce with idempotent producer ids, streaming subscribe, long-poll fetch, named consumers with committed offsets) and Flight DoPut/DoGet, in M5. Changelog streams (D21) are read through the native API. The RisingWave companion integration (D22) moves to Phase C with the Kafka gateway. The internal log, WAL classes and express are unchangedRebuilding a Kafka broker (consumer-group rebalancing, KIP-848, transactions) is a project of its own and not on the path to v1.0Approved (owner; architecture review); the native streaming API core moves to M2 (D72); the Kafka gateway and the RisingWave integration move from Phase C to M5 (D74)
D442026-09-25Graph is native, without Cypher or Bolt: mapped graphs over collections and tables, CSR/CSC sidecars cached in the hot tier, ExpandExec (1–2 hops) and shortest path as DataFusion operators, reached through SQL table functions (graph_expand, graph_neighbors) and an expand stage in the native hybrid search API; M3 gates on Operon-native graph-store adapters for LightRAG and the LlamaIndex property graph. Native (Cypher-first) graphs, Bolt, the Cypher subset, Neo4j procedure/APOC shims and the LDBC/openCypher TCK gates are droppedGraphRAG needs vector/BM25 seeds → 1–2 hops → rerank, in one planned query; Cypher parsing and recursive execution have a fraction of the demand of vector/searchApproved (owner; architecture review)
D452026-09-25No ClickHouse surface: analytics are Iceberg tables (Lakekeeper), queried through Flight SQL and the native API, and by DuckDB, Trino, Spark or ClickHouse itself through the Iceberg REST catalog. The ClickHouse HTTP interface, dialect, function-compat layer and MergeTree-engine DDL are dropped; ClickBench stays as a performance referenceData is already open Iceberg; a ClickHouse server inside Operon would be redundantApproved (owner; architecture review); reversal for the House only proposed 2026-10-01 by D347 (§32, Q333)
D462026-09-25Build order and v1.0: M0 foundation → M1 collections (Qdrant + ES subset + Flight SQL) → M2 production hardening = v1.0 (AuthN/Z, tenant quotas, Prometheus/OTel, multi-node, Postgres metastore, Kubernetes operator, rolling upgrades) → M3 native graph + Resonate Phase A → M4 Iceberg analytics + Resonate Phase B → M5 native streams → M6 scale. Agent-workspace W1 follows M3. Replaces D14Shortest path to production; hardening before new surfacesApproved (owner; architecture review §6); M2 gains the native stream API core and OTLP logs ingest (D72, D73); M5 gains the Kafka gateway (D74); Resonate Phase A leaves M3 for track D (D145)
D472026-09-25Pluggable metastore behind a semantic trait (Lakekeeper model): trait MetaStore in operon-common exposes domain operations (WAL commit and segment swap, catalog and schema evolution, manifest-pointer CAS, leases and fencing, and the rest of the as-built client surface), not raw KV. The embedded openraft MetaClient is the default implementation; downstream crates depend on Arc<dyn MetaStore> from M1.2a. operon-meta-postgres ships in M2; operon-meta-fdb in M6. Amends D10A raw KV trait would make every backend reinvent sequencers, fencing and CAS loops; every enterprise already runs managed PostgresApproved (owner; architecture review §5); the FoundationDB clause is superseded by D58 (TiDB in M6) and D71 (FoundationDB dropped); the contract is relaxed by D59
D482026-09-25Elasticsearch Phase A is what the framework suites send: document APIs, _bulk, _search with the core Query DSL, knn, hybrid + RRF, _delete_by_query, the fixed script_score vector scripts and minimal index admin, as used by the LangChain and LlamaIndex ES suites and BEIR. The elasticsearch-py client-suite gate and its wipe_cluster endpoints are dropped, as are aggregations, point in time and highlighting, unless a gated test needs them (_msearch stays because the BEIR harness sends every query through it). Replaces D32No full ES DSL parity; scope is defined by the conformance suites that matter for AI apps (D13)Approved (owner; architecture review §3); amended by D87 (native filter writes; _update_by_query with recognised scripts joins Phase A) and D91 (function_score in M2)
D492026-09-25Arrow Flight SQL is a core surface: Flight DoPut bulk ingest into collections and streams ships with Flight SQL in M1.2, and the ADBC Flight SQL drivers (Python and Go) are an M1 exit gateZero-copy Arrow ingest and query for every language; the high-throughput half of the native ingest that replaces KafkaApproved (owner; architecture review §3)
D502026-09-25Positioning: "The Unified Hybrid Retrieval Engine on Object Storage" — StarRocks-style stateless compute over open formats with a RAM + NVMe hot tier, for vector + text + GraphRAG retrievalFocus the product identity on the AI retrieval workload rather than on replacing five systemsApproved (owner; architecture review §6)
D512026-09-25Integrate with the AI data ecosystem over protocols and open formats, never by embedding other engines (§17). Loam is a source and sink for Ray Data, Polars, PySpark/Sail, PyTorch/JAX loaders and Spice; the Spice runtime, Sail's crates, the Polars crates, TileDB and Petastorm are not dependenciesEngines built on DataFusion or Arrow move in version lockstep (Lance 12 pins DataFusion 54 / arrow 58; Sail's main is on 55; Spice ships forks of both); Polars carries a second Arrow implementation; TileDB's core pins older versions and its Rust bindings have no license; Petastorm is deprecated upstreamApproved (owner; §17); amendment proposed 2026-10-01 by D343 (a Loam-owned service binary may link an engine library, §32)
D522026-09-25Retained dataset tags (M2): a named tag pins one collection manifest (later also an Iceberg snapshot) until the tag is deleted; GC roots include every tagged manifest and its objects. A tag can only be taken on a manifest whose applied offsets cover the requested consistency token, so reading a tag never needs the stream tail and tags never hold back implicit-stream trimming (D38)Reproducible training runs and eval sets need snapshots that outlive the 24 h time_travel_retention; idea from TileDB's timestamped fragments and Lance/Iceberg tagsApproved (owner; §17); amended by D90 (a tag can seed a branch)
D532026-09-25Scan pinning (M1.2): the native API resolves a collection (at the current manifest, a manifest version, a consistency token or, from M2, a tag) into a pinned scan plan: the manifest version, the Lance dataset URI with its detached version id, the fragment list and whether a tail exists. External readers (Ray, Polars, PySpark, the torch loader) read those fragments directly or fall back to FlightLoam commits Lance as detached versions (D34), so ray.data.read_lance(uri) and similar readers see only the empty mainline; they need Loam to hand out the pinned versionApproved (owner; §17)
D542026-09-25Python ecosystem integrations in the SDK, as optional extras: to_arrow() and to_polars() on results (Arrow PyCapsule, zero-copy) in M1.6; in M2 (the "AI data ecosystem" workstream, shipping in v1.0 with auth and credential vending): a Ray Data datasource and datasink (built on lance-ray), an experimental Polars IO plugin scan_loam(), loamdb.torch (deterministic shuffling that is elastic across GPU counts, mid-epoch resume, pinned tags; built on Lance's PyTorch loader), and a PySpark Python data source format("loam") that runs on Apache Spark 4 and SailAI labs run their data pipelines on Ray, PySpark and PyTorch; meeting them there beats building distributed compute, dataframes or loadersApproved (owner; §17)
D552026-09-25Sail is a named Iceberg reader in the M4 gate, and M4 budgets an upstream contribution of Iceberg deletion-vector and delete-file reads to SailPySpark-shaped curation jobs on stateless Sail compute can then read Loam's keyed tables correctly; Sail is Rust on DataFusion and the DV reader mirrors the writer M4 builds anywayApproved (owner; §17)
D562026-09-25Spice interoperability: Spice's Flight SQL connector is a client in the M1.7 Flight SQL gate; Loam's SQL search functions use Spice's names where they overlap (vector_search, text_search, and rrf for reciprocal-rank fusion; rerank reserved for M3's rerank stage)Spice federates and accelerates data for agent apps; being a first-class Flight SQL source reaches its users, and shared function names lower switching cost. Spice is also a competitor (risk #11)Approved (owner; §17)
D572026-09-25Collection aliases may name several collections (M1.5): reads through a multi-target alias fan out and merge as multi-index searches do; writes need exactly one write target (is_write_index), as in Elasticsearch. The catalog change keeps the M1.1 on-disk encoding readable (new command variants are appended; existing records do not change shape)LangChain's test_cache.py puts one alias on two indices before every test, and M1.7 gates every server-backed testApproved (owner)
D582026-09-26Metastore backends (§18 §2): operon-meta-postgres and operon-meta-dynamodb ship in M2 and gate v1.0. Postgres follows Lakekeeper's patterns (READ COMMITTED with explicit row locks in a fixed order, CAS as a conditional UPDATE with the row count checked, a commit-token row in the same transaction for unknown outcomes, separate read and write pools, sqlx with offline query checks). DynamoDB uses one table with the item layout and method mappings of §18 §2.3 (lazy partition heads, janitor-driven drops, ClientRequestToken on application retries). operon-meta-tidb (MySQL protocol through sqlx, not tikv-client) ships in M6 and takes FoundationDB's M6 role. Supersedes D47's FoundationDB clauseMost enterprises run Postgres, and Lakekeeper proves its patterns; DynamoDB is the AWS-native, serverless option and the store for the hosted control plane (M2.x); TiDB is scale-out, strongly consistent SQL with a managed offeringApproved (owner); FoundationDB dropped from the roadmap by D71; the operon-meta-tidb clause is superseded by D124 (operon-meta-tikv in R1); D260 confirms: no TiDB backend in any milestone
D592026-09-26Relaxed MetaStore contract (§18 §3): (a) commit_wal is atomic per partition group (a DynamoDB transaction group or a metastore shard, each committed idempotently through its own (object, group) record), not across all of a WAL object's partitions, and one stream's chunks of one object stay in one group while they fit; (b) the single monotonic metastore clock becomes bounded-skew stamps (the existing ClockSkew bound; WAL commit records pruned after 2·window + max_clock_skew) plus explicit GC-claim records, checked by every command that makes an object reachable; (c) composite reads over DynamoDB follow documented safe read orders. Written into the trait docs and checked by the conformance suite and the linearizability checker in M2's first task, before any new backend. The openraft backend may keep its stronger behaviour; callers may rely only on the relaxed contract. Amends D27 (pruning) and D47DynamoDB transactions stop at 100 items and a clock item would be one hot key (~500 transactions/s); a sharded metastore (D63) cannot give cross-shard atomicity or one clock at all; Kafka promises per-partition atomicity too, and ES _bulk less (each operation succeeds or fails on its own)Approved (owner)
D602026-09-26CI for metastore backends (§18 §4): floci (MIT, floci/floci:2.1.0) runs the full 49-case conformance suite for DynamoDB, plus a nightly "AWS deployment" job (workers as Lambda functions triggered by EventBridge Scheduler and SQS, metastore on DynamoDB, storage on S3, IAM enforced); ScyllaDB Alternator 6.2.3 (AGPL-3.0, an unmodified CI-only service, never linked or distributed, so consistent with D11) runs the single-item subset, selected by a Backend::capabilities() flag; DynamoDB Local is optional; real DynamoDB runs a nightly smoke test. Each backend has a metastore fault matrix with a blessed expected table, modelled on the S3 fault matrix, with in-process injection (a run_tx hook for sqlx, one SDK interceptor for DynamoDB). Postgres and TiDB run as containers; container services run one at a timefloci verifiably honours ClientRequestToken and supports transactions, and adds S3, Lambda, SQS and IAM; Alternator has no transactions but real LWT timeouts; LocalStack's free tier ended; a blessed matrix turns every fault outcome into a reviewed diffApproved (owner)
D612026-09-26RustFS replaces MinIO as the default local and self-hosted object store in the docs and dev tooling (rustfs/rustfs:1.0.x, Apache-2.0), and is a per-PR CI target for a new Store provider conformance suite and for the S3 fault matrix, which must produce the same expected table as the in-memory store. M2's provider list becomes S3, GCS, Azure, RustFS. The suite must check what is unverified on RustFS: list-after-write (also across pages), fsync durability in single-disk mode (restart and kill -9 of the container), and ETag stability with SSE on (§18 §4.4)MinIO's community edition is archived (AGPL-3.0); RustFS 1.0 is Apache-2.0 and its conditional writes return S3's 412s under a write lockApproved (owner)
D622026-09-26Lambda workers are a CI harness first: the AWS deployment job (D60) runs workers as Lambda functions; production workers stay on containers and Kubernetes until the harness proves the model out, by a later decisionKeeps a new deployment surface out of v1.0 while the lease and fencing model is exercised under Lambda's timeoutsApproved (owner)
D632026-09-26A router for millions of namespaces (§18 §5): stateless gateways; a versioned directory (org → namespaces; namespace name → {id, shard, state, size class}) pushed to gateways; ShardedMetaStore routing every call by namespace; metadata-only namespace moves with gateway buffering; size-class placement keys on M1.3's rendezvous hashing with bounded load; ownership as a soft hint; nodes caching only the namespaces they own through a scoped change feed; paginated list methods; dirty-set-driven maintenance. Staged: M2 removes what breaks first (the whole-catalog re-read on every change, (None) rescans, missing pagination, snapshot builds that block apply, the 5 GiB single-PUT snapshot) and adds bounded load; M2.x adds owned-namespace caches and the pushed directory; M6 adds the sharded metastore, moves and size-class keysEvery node holds the whole catalog today and re-reads it on every metastore change; bulk bytes already sit under ns/<id>/, so resharding moves only metadata; ideas from Neki (proprietary, docs only), PgDog (AGPL-3.0, reference only), turbopuffer and WarpStreamApproved (owner); amended by D75 (one router for every resource kind), D95 (collection shards in M2.x) and D99 (a prefix on paginated lists)
D642026-09-26BYOC in both modes, names in clear (§18 §8), in M2.x / v1.1 "cloud and BYOC": BYOC-managed-meta (the WarpStream model: operon-meta-remote, a MetaStore client over gRPC with idempotency keys, talks to the hosted operon-control control plane) and BYOC-local-meta (the turbopuffer model: the metastore runs in the customer's VPC and the control plane only drives a pull-based ops agent). The control plane may see namespace and collection names, schemas, object paths, offsets, pointers and lease keys; never documents, vectors, text or bucket credentialsCustomers differ on whether a hosted service may sit on their write path; names in clear keep the control plane debuggableApproved (owner)
D652026-09-26Multi-tenant auth and quotas (§18 §6): org → namespaces → collections; a ControlStore trait separate from MetaStore (orgs, API keys, role bindings, quotas, usage rollups; later the directory); API keys loam_<key_id>_<secret> with only a hash of the secret stored; quotas for request rate, ingest bytes, concurrent queries, storage bytes and metadata operations; over-quota requests get HTTP 429 or gRPC RESOURCE_EXHAUSTED. API keys, RBAC and quotas ship in M2 (v1.0); the hosted ControlStore with BYOC in M2.xA billing and isolation unit above namespaces; a separate trait lets BYOC host the control store remotely; metadata-operation quotas protect a shared metastoreApproved (owner); amended by D86 (unapplied-data quota), D98 (cost-weighted per-collection concurrency) and D103 (usage in logical bytes)
D662026-09-26An Authorizer trait (§18 §7): check, batch_check, filter_visible, lifecycle hooks and bootstrap; implementations AllowAll, built-in RBAC (M2) and OpenFGA. Depend on openfga-client 0.6 (Apache-2.0); copy and adapt Lakekeeper's v4.12 .fga model and its migration and reconcile patterns, keeping Lakekeeper's NOTICE; add a tenant fence to the model; tuple writes go through a transactional outboxLakekeeper's model is proven and Apache-2.0; its write-before-commit tuple handling orphans tuples, which an outbox avoidsApproved (owner)
D672026-09-26OpenFGA timing (default): OpenFGA moves from M6 to M2.x, with the control plane, and Loam shares one OpenFGA store with Lakekeeper (M4), Lakekeeper's types unchanged and Loam's added beside them. Amends §10 §4's "OpenFGA is M6"Multi-org tenancy and BYOC want delegated grants, and Lakekeeper needs an OpenFGA deployment anywayApproved (owner, 2026-09-27; D114)
D682026-09-26A GDPR erasure path in M2 (§18 §9): an erase API (by key or filter) that deletes at once and records an erasure request in the metastore; a forced purge (a compaction that materializes the deletions, a re-indexing merge of the affected splits, rebuilt hot artifacts and PK state, and time travel dropped before the erasure point); a stream trim past the erasure offset; GC with explicit cache eviction; an erasure log of key hashes; documented lifecycle rules for versioned buckets. Overrides D38's retention for erased dataDeleted bytes otherwise survive in WAL objects, segments, Lance fragments, splits, retained and tagged manifests, hot artifacts, caches and noncurrent versionsApproved (owner)
D692026-09-26Erasure policy (default, owner-overridable): an erasure rewrites a tagged manifest onto a purged copy, and the tag records that it was rewritten (erasure wins over bit-exact reproducibility; amends D52); the completion deadline is 30 days (the expected time is the bound in §18 §9 item 7); per-chunk envelope encryption (crypto-shredding) moves forward from Phase B to M2.xKeeps tagged datasets usable while honouring erasure; a per-chunk data key is needed because WAL objects span namespaces (D25)Approved (owner, 2026-09-27; D114); the encryption clause is decided by D96 (approved): per-chunk envelope encryption ships in M2 with CMEK, and crypto-shredding is the destruction of the namespace key
D702026-09-26Ids under sharding (default) (§18 §10): ids stay unsharded dense counters (no format change); every MetaStore call that takes a bare CollectionId, StreamId or LinkId gains the NamespaceId in M2's first task, before the sharded metastore, and lease calls take a namespace or cluster scope. No M1.7 format change is needed: object paths already carry ns/<namespace_id>/Shard-carrying ids would change every stored reference and make namespace moves rewrite idsApproved (owner, 2026-09-27; D114)
D712026-09-26FoundationDB is dropped from the roadmap (§18 §2.1): no FoundationDB metastore backend is planned in any milestone or on the Phase C list. The backends are openraft (default), Postgres and DynamoDB (M2), the remote client (M2.x) and TiDB (M6). Amends D10, D47 and D58TiDB, DynamoDB and Postgres plus the sharded metastore (D63) cover its roles; it needs native libfdb_c at the cluster's API version on every host and has no managed offeringApproved (owner); its "TiDB (M6)" backend is replaced by TiKV (D124, D260)
D722026-09-26The native stream API core ships in v1.0 (M2) (§02 §7): HTTP and gRPC produce with idempotent producer ids, building on the M0.3 HTTP produce route (carried into M1.2's native REST) and M1.2's Flight DoPut; long-poll fetch; streaming subscribe; named consumers with committed offsets; stream create, list, describe and drop with retention settings. Changelog streams (D21), the express WAL class, Flight DoGet replay, zone-aware routing and the Jepsen-style gate stay in M5. The M2 gate: no lost acknowledged writes, no duplicates with idempotent producers, a consumer resumes from its committed offset across a node restart, fetch returns records in partition order. Supersedes D43's placement of the API core in M5; the rest of D43's M5 scope standsLog and event ingest is needed in v1.0 (OTLP logs, D73), and the core builds on what exists (the sequencer, the M0.3 routes, DoPut); what stays in M5 needs new storage or protocol workApproved (owner)
D732026-09-26OTLP logs ingest in v1.0 (M2) (§02 §7.1): OTLP/HTTP (protobuf and JSON) and OTLP/gRPC, logs only, each LogRecord appended to a stream as one record (and optionally linked into a collection). Fluent Bit's opentelemetry output, the OpenTelemetry Collector and Vector then ship logs with no custom plugin. The no-code integrations that exist or are planned are listed in §02 §7.3: Fluent Bit http → the native produce endpoint (with a plain-JSON body, M2); Fluent Bit es → the ES _bulk subset (M1.5); RisingWave's Elasticsearch, HTTP and Iceberg (M4, through Lakekeeper) sinks → Loam as a sink target. Traces and metrics stay with W1 (§15)Every log shipper speaks OTLP; mapping it onto streams reuses the M2 stream API core and needs no plugin of oursApproved (owner)
D742026-09-26Kafka wire-protocol gateway in M5 (§02 §7.2), as the long-term source-compatibility protocol: Loam streams become readable and writable by RisingWave, Flink, Spark, Kafka Connect, Debezium, Fluent Bit's kafka output and Vector. Staged within M5: (1) Produce, Fetch, ListOffsets, Metadata, ApiVersions; (2) idempotent producers mapped onto the native producer ids (D72); (3) consumer groups with committed offsets, the classic group protocol first, then KIP-848 server-side assignment. No Kafka transactions in M5; a later milestone adds them on demand (Q4). The RisingWave companion integration (D22) returns in M5 with it. M5 gates on the Kafka client test suites (librdkafka, franz-go or the Java client). Rejected: a Kinesis-compatible subset (reaches fewer tools, and would become redundant once Kafka exists); a RisingWave-only native connector (Kafka covers RisingWave). Supersedes D43's deferral of the Kafka gateway and of D22 to Phase CWarpStream, AutoMQ and Bufstream show that a Kafka-compatible log on object storage is what stream tools connect to; Loam's WAL already stores Kafka RecordBatch v2 bytes and its sequencer already dedupes by producer id, so the gateway maps onto existing machinery; Nisshi (formerly Tansu; Apache-2.0, Rust) is a reference broker on S3 and PostgresApproved (owner); deferral in favour of Iggy's Kafka gateway proposed 2026-10-01 by D332 (Q331)
D752026-09-26One router for every resource kind (§18 §5.2–§5.3): metastore sharding is by namespace whatever the kind, so a namespace's streams, collections, Iceberg tables (M4) and graphs (M3) live in one shard and operations across kinds stay atomic; the directory knows only namespaces. The placement key is (ns, kind, id[, shard]) on the same rendezvous hashing with bounded load: stream partitions by (ns, stream, partition) for fetch and subscribe, collections by (ns, collection), Iceberg tables by (ns, table), graphs by (ns, graph); size classes still apply (small namespaces place by ns alone). Writes need no routing: the leaderless WAL lets any log node append, and the metastore orders the writes. Consumer-group coordination (native named consumers; Kafka groups in M5) runs on the rendezvous owner of (ns, group), with committed offsets in the metastore; conditional commits keep it correct when the owner is wrong. Workers own maintenance, and quotas are keyed, the same way. Amends D63One placement and sharding scheme instead of one per kind; affinity where caches pay (tails, segments, hot tiers, adjacency), and none where the leaderless WAL makes it unnecessaryApproved (owner); amended by D95 (the collection key's shard lands in M2.x)
D762026-09-26Consistency tokens are a hard guarantee on every backend, not a required client input (§01 §5, §18 §3.5): a token {(stream, partition, offset)} means the same on openraft, Postgres, DynamoDB and TiDB, under the D59 relaxations, during namespace moves and under stale routing, because offsets come from the metastore, per-partition order is kept, and any serving node merges the tail up to the token. A conformance-suite and router gate checks it: a write acknowledged with a token, then read with that token on every other node and backend, during a move and with a stale directory, is visible (staged M2, M2.x, M6). Clients may omit tokens: default single-object reads are strong, and eventual skips the tail; tokens are needed only for reads through derived objects (tables, graphs, or another collection fed by a link). Amends D59 and D63Cross-surface read-your-writes is what the log-as-spine design promises (D4); it must not weaken as backends and sharding are addedApproved (owner)
D772026-09-24The tail overlay rule and range tails: a tail entry for key k in partition p at offset o overrides the durable state iff o >= manifest.applied[p]; entries below applied are covered and dropped, so one live tail serves every manifest at or after its base. Reads the live tail cannot serve (a pinned upper bound below its head, an overflowed tail) use a range tail built from the log over exactly (applied, upper]Latest-by-key state cannot answer "state at offset T" below its head; comparing with applied keeps the overlay correct for any adopted manifest without a rebuild per commitApproved (M1.2 plan, ruling 1); amended by D86 (Eventual reads only the live tail and never builds a range tail)
D782026-09-24BM25 statistics are global and live-only: over every split of the manifest plus the tail, document counts and document frequencies exclude deleted and shadowed docs, and token totals sum the midpoint of each live doc's fieldnorm bucket; candidates are rescored with a canonical two-clause form of the query so scores are bit-identical across split layoutsLucene's counting of deleted docs until a merge would make a score depend on where a document's history lives, and the tail-merge property test could never passApproved (M1.2 plan, ruling 2; row 15.1)
D792026-09-24Operon's exact vector kernel scores every vector result (f64 accumulation in index order, rounded to f32); exact search (exact, a metric override, Manhattan, or a candidate set under the brute-force threshold) never uses an index, Lance's kernels or a hot artifactR12 needs bit-identical exact results with the hot tier on and off and exact scores on approximate paths; SIMD kernels differ in summation order between librariesApproved (M1.2 plan, ruling 3); extended by D92 (ground truth of the recall endpoint and of sampled recall)
D802026-09-24DBSF is Qdrant's distr_norm, unclamped: f32 Welford mean and sample variance, (s − (μ − 3σ)) / 6σ; a one-hit list or σ = 0 scores 0.5M1.4's contract compares against Qdrant, which does not clampApproved (M1.2 plan, ruling 4; overview A20)
D812026-09-24Two ranking modes and their aggregation domains: score mode fuses the retrievers' top-k lists; field mode (the first sort key is a field or the primary key) takes every match of one text retriever and the filter, as ES sorts every match. Aggregations run over the text matches with one text retriever, over the fused candidates with any other retriever, and over the filter's matches with noneES semantics for query + sort and for aggregations with knnApproved (M1.2 plan, ruling 6); amended by D91 (expression ranking over these domains, M2)
D822026-09-24The hot switch is a task-local scope set from the Operon-Hot header or operon-hot gRPC metadata; the service reads it once per call and uses no hot tier when it is off, and the structures used are reported in hot_used and Operon-Hot-UsedThe §6.7 method signatures stay the same for every gateway, and the switch reaches forwarded readsApproved (M1.2 plan, ruling 11)
D832026-09-24Pins hold nothing: pin() returns the manifest version and the high watermarks from one linearizable read; a pinned read serves them while the manifest is retained (D38) and is NotFound { kind: "pin" } afterwards; manifest version 0 is served by a range tail while the log still starts at 0ES _delete_by_query and Qdrant snapshots need a stable view, and amendment A6 replaced the metastore holdApproved (M1.2 plan, ruling 14; overview A6)
D842026-09-24Sparse search is exhaustive and exact with live-only IDF statistics: every candidate of the query indices' postings (minus deleted, shadowed and filtered docs) is scored with Qdrant's f32 dot product in ascending index order; N and df count live documents of the manifest's splits and the tail, or of an idf_corpusOne deterministic formula over one live-doc set keeps results identical hot or cold and across split layouts; Qdrant's server counts deleted points until optimization, its local mode does notApproved (M1.2 plan, ruling 21; overview R22)
D852026-09-24Scan plans rely on Lance's deletion files and report the tail with a pin; token scans wait for durability (D53 as built): the Lance version a manifest names holds exactly its live rows, so the plan lists deletion files and no split bitmaps; writes after applied are reported (tail, tail_records, per-partition offsets) with a pin that reads them through Flight SQL or the native API; at: {"token"} waits up to consistency_wait for a manifest covering the token, else TimeoutDeletions are already materialized where external readers look; the tail is the one state a Lance reader cannot see, so the plan says so; waiting for the tail would not help a reader that cannot read itApproved (M1.2 plan, rulings 22 and 24)
D862026-09-26Unapplied-data budget and write backpressure (gap analysis G1). Each collection has a budget on its records past applied (Σ_p (hwm_p − applied_p)) and on their log bytes (the offset-index entries above applied): max_unapplied_records 1 000 000 and max_unapplied_bytes 128 MiB by default, with max_unapplied_bytes ≤ tail.max_bytes / 2, so a backlog at the budget fits the live tail. A collection write (native REST, Flight DoPut, Qdrant, ES) that arrives while the backlog is at or over either budget is refused with ServiceError::ResourceExhausted: HTTP 429 with Retry-After (whole seconds, estimated from the link's recent apply rate, 1–30 s), gRPC RESOURCE_EXHAUSTED. Every collection write response carries Operon-Unapplied-Records and Operon-Unapplied-Bytes, and CollectionInfo reports both and a backpressure state. A per-request override (Operon-Backpressure: off) serves bulk loads up to 4× the budget; above the budget, strong reads may fall back to range tails or answer Unavailable, and Eventual keeps serving. M1.3 Task 15 builds the budget, the refusal, the headers and the counters. M2 makes the budget a quota (per namespace and per collection in the ControlStore), exports the Prometheus metrics and restricts the override to a role. Amends D77 (Eventual never builds a range tail: it reads the durable state and the live tail the node holds, at most tail.max_bytes, and reports the records it did not see) and D65 (a sixth quota, unapplied data per collection, enforced at write admission)M1.2 bounds tail memory, but a slow or stalled link lets the backlog grow without telling the writer: strong reads then fall back to range tails that re-read the log on every query, and past 1 GiB they fail with Unavailable. turbopuffer returns 429 at 2 GB or 1M unindexed rows and offers disable_backpressure for bulk loads (limits, write)Approved (owner; turbopuffer gap analysis)
D872026-09-26Native delete-by-filter and patch-by-filter (G2). CollectionService::delete_by_filter and patch_by_filter, served natively at POST …/collections/{c}/documents/delete_by_filter and …/documents/patch_by_filter. ES _delete_by_query, a new ES _update_by_query and the Qdrant filter writes (delete, set/overwrite/delete/clear_payload and delete_vectors with a filter; M1.4 Ruling 13) all call them. _update_by_query accepts only recognised scripts (assignments from params and remove calls), compiled to one patch; a request without a script, or with any other script, is refused. Semantics: the filter is evaluated at one pin (Strong, or AtLeast(token) when the caller sends a token), so documents written after the pin are never touched. Matched keys are paged by primary key and written as plain Delete or Patch { upsert: None } ops in batches of 1 000; each batch is one atomic write and passes D86's admission (a refused batch waits for Retry-After within the request deadline). The result carries matched, affected, rows_remaining, a cursor (the pin and the last key) that continues at the same snapshot, and one consistency token covering every batch. Limits per call: 5 000 000 rows for deletes and 50 000 for patches; over the limit the call fails before writing anything, unless allow_partial is set. In M1 a key that changes between the pin and its batch is still deleted or patched, as ES _delete_by_query without version checks does (M1.5 Ruling 6); M2 re-checks the filter at apply through D89's conditional ops, which gives Read Committed. M1.5 Task 9a; the SDKs expose it in M1.6 Task 11. Amends D48 (_update_by_query with recognised scripts joins Phase A; _delete_by_query becomes a caller) and the M1 overview §6.7–§6.8 (A46)Filter writes are how the frameworks delete (LlamaIndex delete_nodes(filters), LangChain clear()), and three gateway-local loops would differ in semantics and limits. turbopuffer caps them at 5M and 50k rows with rows_remaining and *_allow_partial, and re-checks the filter in a second phase (write, guarantees)Approved (owner; turbopuffer gap analysis)
D882026-09-26A published limits page, enforced by tests (G22). docs/guides/limits.md lists every enforced limit and the error past it: record size (16 MiB), native request body (16 MiB; ES 100 MiB), ops per write (10 000), string primary keys (512 bytes), namespace and collection names ([A-Za-z0-9._-], 1–255 and 1–222 bytes, no leading _), fields per collection (max_fields, default 1 000, at most 100 000), dense dimensions (65 536; ES 4 096), partitions per collection, k and offset + limit (100 000), retrievers per request (16), fusion depth (4), query clauses (1 024), query nesting depth (32, new), keys per get and rows per scroll (10 000), SQL result rows, filter-write rows (D87) and the unapplied budget (D86). It also gives guidance that M1 does not enforce: documents per collection (the largest collection the M1.7 benchmarks ran), collections per namespace, and request rates (M2 quotas, D65). One table in code (operon_query::limits::LIMITS) is the source: a test renders the page from it and fails on drift, and each enforced limit has a test at the limit (accepted) and one past it (refused with the documented error). M1.7 Task 11. The counterpart of D13: compatibility is defined by suites, and limits by testsLimits are scattered across plans and code; max top-k, request size and id length are undocumented. turbopuffer publishes about 40 limits with footnotes (limits), and users design against themApproved (owner; turbopuffer gap analysis)
D892026-09-26Conditional writes are decided in log order (G3). A write op may carry a condition: a filter over the key's current document, with $new.<path> references to the op's own values. A missing document makes an upsert apply and a patch or delete skip, as in turbopuffer. The condition is evaluated at apply, in offset order, by one deterministic function that link apply and the tail both run, so every node and the durable state agree on each op's outcome. The write's response reports each op's outcome (applied or condition_failed) once the serving node's tail has reached the write's own offsets (a strong read of those keys; no wait for indexing). This serves ES if_seq_no/if_primary_term (the key's last-op offset, M1.5 Ruling 4) and op_type=create, Qdrant update_filter, a native condition, and D87's filter re-check. M2: the design and conformance cases first (a per-key linearizability model over offset order), then the surfaces. Rejected: checking the condition on a per-key owner before the WAL append, which needs a leader per key, which the leaderless WAL deliberately lacks (§02 §3). Amends D4 (an op's effect may depend on the state before it in its partition's order)The partition offset is already Loam's one serialization point, and the tail and link apply already fold ops in that order. turbopuffer evaluates conditions atomically with the write (serializable) and uses them for version checks and insert-if-absent (write, guarantees)Approved (owner; turbopuffer gap analysis)
D902026-09-26Collection branches and copies (G4, G7). M2: POST …/collections/{c}/branch {name, at} creates a collection in the same namespace from a retained manifest (current, a manifest version, a token or a tag, D52) in constant time. Its first manifest references the source's Lance version, splits and delete bitmaps; it gets its own implicit stream; after creation, writes to either collection are never visible in the other. GC reachability is computed per lineage (a source and every branch that shares its objects), so a shared object lives while any retained manifest in the lineage references it, and dropping the source keeps what its branches use. An erasure (D68) covers every lineage member that holds the key. Where a branch's Lance commits live (the source's dataset directory, isolated by detached versions, D34, or a directory of its own) is decided in the M2 plan. M2.x: copy into another namespace, bucket, region or org, as an asynchronous operation (Prefer: respond-async → 202 with Location: /v1/operations/{id}; results kept 1 h), which copies the live manifest's objects and re-keys them under the target's key (D96). Amends D52 (a tag can seed a branch, not only pin a manifest) and D38 (retention and GC count every branch of a lineage)Immutable manifests and reachability GC make a copy-on-write branch mostly a catalog operation; branches serve eval sets, index experiments and tenant forks. turbopuffer branches in constant time (p50 440 ms, unlimited branches) and copies across regions and orgs with a re-key (branching, write)Approved (owner; turbopuffer gap analysis)
D912026-09-26One ranking-expression IR (G9). The search IR gains rank: RankExpr, an expression over score (the retriever's or the fused score), numeric and date fields (field, with missing), constants, saturate(x, midpoint, exponent), decay(x, origin, scale, offset, decay, gauss|exp|linear), distance(x, origin) (recency on dates), if_match(query, then, else) (rank by filter), and sum, max, min, product and weighted. It is evaluated in f64 in a fixed order and rounded to f32, so results are identical hot and cold. Its domain follows D81: every match with one text retriever and no vector retriever, else the fused candidates. Fusion (RRF, DBSF, weighted; D80) is its input, not a second mechanism. ES function_score (weight, field_value_factor, gauss/exp/linear, per-function filter, score_mode, boost_mode, max_boost, min_score), Qdrant formula and the native API compile to it. M2. Amends D48 (function_score moves from Phase B to M2) and D81 (expression ranking over D81's domains)Recency and popularity boosts are what RAG ranking asks for after hybrid fusion, and one IR keeps three surfaces consistent. turbopuffer's rank_by has Sum, Max, Product, Saturate, Decay and Dist, and ranks by filter (query)Approved (owner; turbopuffer gap analysis)
D922026-09-26A recall endpoint and per-response performance stats (G10, G11). M1.6 Task 10: every native SearchResponse carries performance: server_total_ms, queue_ms, planning_ms, execution_ms, manifest_version, tail_records (unindexed records the query scanned), stale_records (Eventual only, D86), rows_scanned per retriever (text documents scored, vector candidates, brute-force rows), cache (H1 hit and miss bytes and hit_ratio) and object_store_requests (lower bounds: Lance's reads on its own tasks are not attributed, M1.7 Ruling 9), beside the existing hot_used; the SQL response carries the timings, and both carry Server-Timing. M1.7 Task 10: POST /v1/namespaces/{ns}/collections/{c}/recall {vector, num, k, filter, params} (num default 25, at most 100; k default 10) samples num stored vectors, runs each through the normal ANN path and through the exact kernel, and returns the mean and minimum recall@k and the mean hit counts of both; the M1.7 recall harness cross-checks it. M2: sampled continuous recall (a configurable fraction of vector queries, default 1 %, re-run exactly in the background under a CPU budget), exported as a metric per collection. Extends D79: the exact kernel is also the ground truth of the recall endpoint and of the sampled metricUsers need to see why a query was slow or cold, and whether ANN recall has drifted, without an offline harness. turbopuffer returns cache and timing fields on every query and offers POST …/_debug/recall, with 1 % of live traffic sampled continuously (query, recall)Approved (owner; turbopuffer gap analysis)
D932026-09-26Language analyzers and BM25 parameters (G13, G14). M2 adds the ES language analyzers that Lucene builds from Snowball stemmers and stop-word lists (the list and each chain are fixed in the M2 plan and checked token for token against ES _analyze); the asciifolding and lowercase filters and max_token_length for custom analyzers; pre-tokenized text (an array of tokens indexed verbatim); and per-field BM25 k1 and b in the schema (defaults 1.2 and 0.75), so ES similarity BM25 settings map instead of being refused. Loam computes BM25 itself when it rescores (D78); the candidate stage prunes with Tantivy's fixed-parameter bounds, so the M2 plan must show the slack still covers other parameters, or turn pruning off for them (verify). CJK tokenizers stay later (§06 §9). Amends D39Non-English corpora are common in RAG. turbopuffer ships 18 stemming languages, ASCII folding and tunable k1 and b (fts, fts tuning)Approved (owner; turbopuffer gap analysis)
D942026-09-26f16, i8 and u8 vectors (G18). VectorElement gains F16, I8 and U8 beside F32. Lance stores the column at that width (FixedSizeList<Float16 | Int8 | UInt8>); IVF indexes and hot artifacts build over it (qdrant-edge's datatypes are checked in the M2 plan); the exact kernel (D79) widens each element to f64, so exact scores stay deterministic. Queries may send f32 values: they are rounded to nearest for f16, and refused, not clamped, when out of range for i8 or u8. Qdrant datatype: float16 | uint8 and ES element_type: byte map onto them and stop being refused. M2. Amends D8 (both tiers store the declared width) and the M1 overview §6.2f16 halves storage and scan bytes, and int8 embeddings come straight from several embedding APIs. turbopuffer offers [N]f16 and [N]i8 (write)Approved (owner; turbopuffer gap analysis)
D952026-09-26Single-collection sharding moves to M2.x (G23). A collection may be created with shards (1–256, fixed at creation; partitions a multiple of it). Shard s owns a fixed set of the implicit stream's partitions and has its own manifest chain, Lance dataset, splits, link task and hot artifacts under collections/<cid>/shards/<s>/, placed by (ns, collection, shard) (D75). A query fans out to every shard's owner and merges: BM25 statistics are summed across shards before scoring, so D78 holds per collection, and vector scores stay exact (D79). A write spanning shards is one multi-partition append, atomic because a collection's partitions stay in one partition group (D59) while its chunks fit one group (32 on DynamoDB); a larger write follows the M2 contract task's ruling for append_many (§18 §3.1: per-group atomicity or a transaction of its own), and that ruling bounds shards and partitions on each backend. A strong or token read is consistent at one token across all shards; Eventual may see shards at different points. Changing the shard count is a copy (D90). Moves from M6 to M2.x. Amends D63 (collection shards leave M6's staging) and D75 (the collection key's shard lands in M2.x)One link task per collection bounds ingest (M1 overview R9), and one Lance dataset bounds collection size. turbopuffer shards a namespace 1–256 ways with a shared WAL, atomic writes and snapshot-consistent strong reads (sharding)Approved (owner; turbopuffer gap analysis)
D962026-09-26Customer-managed keys by envelope encryption (G28). A namespace may name a KMS key at creation (AWS KMS, GCP Cloud KMS or Azure Key Vault behind a KeyProvider trait; a file provider for tests). Objects under ns/<id>/ are written with the provider's per-object KMS setting pointing at that key (S3 SSE-KMS key id, GCS kmsKeyName, Azure encryption scope), so external readers with a grant on the key still read Lance fragments (D53). WAL objects mix namespaces (D25), so their chunks are envelope-encrypted: each chunk (one partition's batches, so one namespace) is sealed with AES-256-GCM under its own data key, wrapped by a per-namespace key-encryption key. That key comes from the KMS once per namespace, node and hour (GenerateDataKey), is held in memory only, and is stored wrapped at ns/<id>/keys/<ulid>.key; the chunk header carries the wrapped data key and the key's id (WAL format version 2; version-1 objects still read). Caches key entries by namespace, encrypt a CMEK namespace's NVMe entries under a per-process key, and drop a namespace's entries when its key is revoked. Destroying or revoking the namespace key makes all of its bytes unreadable (WAL chunks, segments, Lance files, splits, noncurrent versions), while other namespaces in the same WAL objects read normally: that is crypto-shredding (D69), with no second mechanism. Namespaces without a key keep bucket-default encryption and pay nothing. M2. Rejected: WAL objects grouped by KMS key (multiplies PUTs by active CMEK namespaces, the cost D25 avoids, and leaves caches and segments uncovered) and SSE-KMS alone (it cannot cover a multi-namespace object). Amends D25 (per-chunk envelope keys, not per-namespace WAL objects), D69 (per-chunk envelope encryption moves from M2.x to M2, with CMEK) and §10 §3CMEK is an enterprise requirement, and the note on D25 left the conflict open; one scheme gives per-namespace keys, crypto-shredding and a free default. turbopuffer sets a KMS key per namespace on the first write, with no latency cost, and revoking the key loses the data (encryption)Approved (owner; turbopuffer gap analysis)
D972026-09-26Online index changes on existing fields (G5). PATCH …/collections/{c}/schema may change a field's indexed, fast and analyzer, type an existing JSON path as a field, and change BM25 parameters (D93). The tail and new splits use the new schema at once; a background task re-derives every older split from _source with it (M1.3's re-indexing merge, Ruling 4), tracked by each split's schema_version. Until every split is rebuilt, a query that depends on the changed field answers 409 index_building with progress, never a partial result, and CollectionInfo reports the backfill. M2. Amends the M1 overview A4 (schema changes were additive, with no backfill)Users change mappings after loading data, and the re-indexing merge already rebuilds splits from _source. turbopuffer changes filterable, full-text settings and more in place, and answers 409 until the index is built (write)Approved (owner; turbopuffer gap analysis)
D982026-09-26Cost-weighted per-collection concurrency (G25). D65's concurrent-queries quota becomes a semaphore per collection on its owner, 16 slots by default: a text, filter or ANN query takes 1 slot, an exact or brute-force vector query 2, and an aggregation, group_by or SQL scan 4. A query that cannot get its slots waits up to 800 ms, then gets 429 with Retry-After: 1; the time waited is performance.queue_ms (D92). Limits are set per namespace and collection in the ControlStore, and a pinned collection (M1.3) may be given more. M2. Amends D65One heavy query class otherwise starves a shared node. turbopuffer's per-namespace semaphore (16 slots; aggregations take 4 and kNN 2; an 800 ms wait) is a tested design (limits)Approved (owner; turbopuffer gap analysis)
D992026-09-26Prefix and cursor on list calls (G26). GET /v1/namespaces, GET …/collections, and the alias and stream lists take prefix, cursor and page_size (default 100, at most 1 000), and D63's paginated MetaStore list methods take (prefix, after, limit). M2. Amends D63Deployments with a namespace per tenant list by tenant prefix. turbopuffer lists with prefix, cursor and page_size ≤ 1 000 (namespaces)Approved (owner; turbopuffer gap analysis)
D1002026-09-26Audit events (G29). M2: the engine writes one audit record per admin action or security event to the stream _audit of a system namespace: API keys created and revoked, role bindings changed, failed authentications (rate-limited), namespaces and collections created and dropped, erasure requests and completions, key (CMEK) changes and backpressure overrides. A record holds the actor, org, action, resource, outcome, time, request id and source address, never document contents. Retention is 30 days by default; the org's admin role reads it through the stream API. M2.x: a console view and exports to object storage, HTTPS, Datadog, Splunk and Sentinel. Amends §10 §4 (Audit)Enterprise buyers ask for an audit trail before v1.0. turbopuffer keeps 30 days and streams to SIEMs (audit-logs)Approved (owner; turbopuffer gap analysis)
D1012026-09-26An OpenAPI spec, then a Go SDK (G31). M1.6 Task 12 publishes docs/api/openapi.json (OpenAPI 3.1) for the native REST API; a test fails when a route the router serves is missing from it or the reverse, and every wire-fixture request and response validates against it. M2: a Go SDK over the native API with the Python and TypeScript SDKs' contract (the fixture corpus, retries, tokens), and the native gRPC protos, published with the gRPC surfaceOther languages reach Loam through the Qdrant, ES and ADBC clients, but the native API (tokens, pins, scan plans, filter writes) needs a machine-readable contract. turbopuffer publishes an OpenAPI spec and six SDKs (api-overview)Approved (owner; turbopuffer gap analysis)
D1022026-09-26SSO, IP allowlists and private networking (G30), in M2.x (cloud and BYOC): console SSO (SAML 2.0 and OIDC), OIDC/JWT for API callers, per-org IP allowlists enforced at the gateways, and AWS PrivateLink and GCP Private Service Connect endpoints for hosted clusters. Amends §10 §4 ("OIDC/JWT validation is a later addition" becomes M2.x)Enterprise procurement checks these for a hosted service. turbopuffer offers SSO and IP allowlists on its Scale plan and PrivateLink/PSC on Enterprise (security, private-networking)Approved (owner; turbopuffer gap analysis)
D1032026-09-26Metering on logical bytes (G32). From M2 the engine counts, per namespace, the logical bytes written, the logical bytes stored (live documents, and bytes held only by tags and branches), the bytes queried (scanned) and returned, the queries, and hot-tier GB-hours, and rolls them up into the ControlStore's usage records (D65). M2.x: the hosted service bills on these units; prices live in the cloud product docs, not in the engine docs. With metering on, the performance block (D92) reports a query's billable bytesLogical bytes are predictable for users and independent of Loam's physical layout (compaction, merges, hot copies). turbopuffer bills by logical bytes written, stored and queried (pricing-log)Approved (owner; turbopuffer gap analysis)
D1042026-09-24HNSW artifacts are keyed by stable row id and stay valid across maintenance (M1.3 Rulings 1–2). An artifact built by qdrant-edge from manifest s is current at v ≥ s iff no LinkApply commit in (s, v] inserted or deleted rows; compaction, merges, index builds and hot commits never make it stale. A query at a newer manifest uses a view: the artifact minus the rows deleted since (must_not has_id), plus an appendable delta index of the rows inserted since; a view with more than 100 000 exclusions is not served until the rebuild lands. Artifacts live under hot/hnsw/<column>/<source_version:020>-<ulid>/Row ids survive compaction and merges (R8), and inserted rows form a row-id range, so no Lance scan is needed per version; "source version = manifest version" would otherwise never hold after a hot commitApproved (M1.3 plan)
D1052026-09-24Split merges re-index from _source (M1.3 Ruling 4): the merged split holds its inputs' live docs, rebuilt with the inputs' own schema in ascending row-id order; only splits of one schema version merge, and a split with ≥ 30 % deleted docs is rewritten alone. BM25 over live docs (M1.2 Ruling 2) makes a merge change no scoreM1.1's delete bitmaps are external to Tantivy, and a Tantivy segment merge would break the row-id order that row_id_ranges and the row locator requireApproved (M1.3 plan)
D1062026-09-24Lance compaction is a detached Rewrite with pre-assigned fragment ids (M1.3 Ruling 5): Lance's planner and rewriter write the files; Operon assigns fragment ids from max_fragment_id + 1 and commits Operation::Rewrite through LanceCommitter::commit, never commit_compaction; on rebase a group is kept only if its fragments are unchangedcommit_compaction reserves fragment ids with a mainline commit, which R7 forbids, and the detached path does no conflict resolution, so a group whose fragment gained deletions would resurrect rowsApproved (M1.3 plan)
D1072026-09-24The metastore over the network (M1.3 Ruling 10): the same openraft MetaClient, networked, with a local replica on every node (voters on meta nodes, learners elsewhere). Local reads stay local; writes and read indexes go to the leader over HTTP (postcard), and the client waits until its replica has applied the index. It passes the MetaStore conformance suite over HTTP. The internal routes are private to the openraft backend, not a remote MetaStore protocol (D64)Many MetaStore reads are Local and its change notification feeds long-polls; a replica keeps them local and leaves the trait implementation unchanged (D47)Approved (M1.3 plan)
D1082026-09-24Node registry as leases, rendezvous ownership (M1.3 Rulings 12–13): each node holds node/<id> (TTL 10 s) with its descriptor; owners of a placement key (ns, kind, id[, shard]) are the top r live query nodes by xxh3_64 score, one per zone first; a node that failed a forward is suspect for 5 s; a read is forwarded at most once and falls back to local executionLiveness and fencing come from leases with no new command; rendezvous moves only the keys whose top node changes; ownership is a hint, so correctness never depends on it (§04 §5). Extensible to D75's placement keysApproved (M1.3 plan)
D1092026-09-24Hot configuration is catalog state; auto-promotion is lease-based and off by default (M1.3 Rulings 7–8): SetCollectionHot stores {vectors, text, fragments} per collection in a serde-skipped map, written to a new snapshot format version only while non-empty (Ruling 20). With auto-promotion on, the owner holds hot-promote/<ns>/<cid> while the collection is hot, and workers treat it as a pin. --hot-pin-all is an in-process overrideWorkers must see pins and promotions to build artifacts; a lease expires by itself when heat drops; the encoding only grows (A44)Approved (M1.3 plan)
D1102026-09-24Fragment prefetch is H1, not H2 (M1.3 Ruling 9): LanceEnv reads through a caching object_store wrapper over the range cache, and prefetching a pinned collection's Lance files fills it. It is not reported in Operon-Hot-Used, and Operon-Hot: off does not bypass itLance's dataset and session are shared across requests, so no per-request switch below Lance is possible; H1 returns the same checksummed bytes (A13)Approved (M1.3 plan)
D1112026-09-27Auth and TLS are deferred to one unified auth plan, written after M1. It covers the native API, the Qdrant and Elasticsearch gateways, MCP, the console, multi-tenant API keys and quotas (D65), OpenFGA (D66, D67) and TLS; §10 §4 is its starting point (Q30). M1 ships no auth or TLS on any listener. The new M1 gateways (Qdrant, Elasticsearch, MCP) bind 127.0.0.1 by default; the existing native and cluster listeners keep their current bind defaults until the unified plan. MCP has its own listener, 127.0.0.1:8083 by default (--mcp-listen, [gateways] mcp.listen), serving the endpoint at /mcp; /mcp is never served on the native REST listener, so binding native REST to a non-loopback address never exposes MCP. Amends §01 §3.3 and the M1.6 plan (Ruling 19), which mounted MCP on native REST. Every listener bound to a non-loopback address logs a startup warning that it is unauthenticatedOne plan keeps credentials, scopes, quotas and TLS consistent across every surface, instead of a partial scheme per gateway that the plan would then replace. Loopback defaults keep an unauthenticated gateway off the network unless the operator asks for it, and the warning makes that choice visible. A path on a shared socket cannot be loopback-only while the socket is not, so MCP needs its own; 8083 follows native REST, gRPC and Flight SQL (8080–8082) and clashes with no documented portApproved (owner); narrowing proposed 2026-10-02 by D451 (MT1 plus per-listener adoption, §38)
D1122026-09-27The Qdrant gateway binds 127.0.0.1:6333 (REST) and 127.0.0.1:6334 (gRPC) by default on dev, standalone and cluster; a deployment sets --qdrant-listen and --qdrant-grpc-listen. Closes the owner question on M1.4's default addresses (M1.4 Task 2, T2-7)Follows D111: the gateway is unauthenticated in M1, and the Qdrant clients' local default (localhost:6333) still reaches itApproved (owner)
D1132026-09-27A skew limit on stamps and a GC safety margin (§18 §3.2). A node measures its own clock skew and refuses to issue stamps (so proposes no stamped command) while the measured skew exceeds max_stamp_skew (default 30 s); it resumes once the skew is back under the limit. This is a node-side check, tighter than the meta leader's ClockSkew refusal at max_clock_skew (default 5 min, §10 §2), which stays. GC adds gc_safety_margin (default 10 min) on top of every retention window (time_travel_retention, the GC grace period, WAL commit-record pruning and stream retention) before it deletes, in addition to the max_clock_skew terms of §18 §3.2–§3.3. Both values are configurable. Lands with the M2 contract task (D59)A node with a bad clock is stopped before it issues stamps far off, rather than only when the generous leader bound catches it. The margin covers skew within the bounds and small clock errors for 10 minutes of extra storage, and keeps an early deletion from ever deciding correctnessApproved (owner)
D1142026-09-27The defaults of D67, D69 and D70 are accepted as they stand: OpenFGA in M2.x with one store shared with Lakekeeper (D67); tagged manifests rewritten onto a purged copy and a 30-day completion deadline (D69; its encryption clause stays decided by D96); unsharded ids with the namespace on every bare-id call (D70). Closes Q25The owner delegated these; each default was chosen with its rationale in §18, and nothing since argues against itApproved (owner)
D1152026-09-27Erasure-log retention and restore after erasure (§18 §9, §10 §6). An erasure-log record identifies the erased keys and ids (as D68's keyed hashes) and the erasure time, never content. It is kept for as long as any snapshot, backup or time-travel version that predates the erasure still exists, plus 30 days. Storage: each record is held in two places that no restore rolls back together: in the ControlStore, next to the org's erasure-log HMAC key, and as an immutable object under _erasure/<org_id>/ in the cluster bucket, written with a conditional put (If-None-Match: *) so it is never overwritten (§01 §6, §18 §9). The log is part of every ControlStore and metadata backup, and the object copy lives outside every snapshot it must override, so it survives a metastore or ControlStore restore. A restore does not roll the erasure log back. Restore order: (1) load the union of the records in the restored ControlStore and the _erasure/ objects; (2) replay every erasure newer than the restore point's data; (3) only then serve traffic; (4) if either source cannot be read, refuse to serve (the node stays not ready). Answers Q27 (replay, not rebuild) and sets the retention §18 §9 item 6 left to the M2 planReplay is the only answer that also works for backups and snapshots Loam cannot rewrite, and it needs only the log. Keeping each record until the last copy that predates it is gone means a restore can never bring back erased data, and the 30 days cover restores and audits after that. A log held only in the store being restored would be rolled back with it; the write-once object copy in the bucket is the record no restore can undoApproved (owner)
D1162026-09-27Loam is an AI-native cloud (§20 §1): the retrieval engine (§00–§18), a reactive application database (Loam Live, D117), SQL through TiDB (D123), streams (D72, D74) and AI-gateway integration, sold as one cloud. The retrieval engine's positioning (D50) is unchanged inside itThe owner's direction of 2026-09-27 ("loam will a ai native cloud"). AI applications need an application database, SQL and retrieval together; offering all of them on shared infrastructure, with app data flowing into retrieval without ETL, is the productApproved (owner); the "SQL through TiDB (D123)" part is superseded by D260 (no TiDB; MySQL wire access is open, Q260)
D1172026-09-27Loam Live, a Convex-style reactive database on TiKV (§20; working name "Loam Reactive", "Loam Live" proposed). TiKV provides ACID transactions (Percolator), Raft replication, MVCC and timestamps (PD TSO). Loam builds: (a) a reactive query engine that tracks each query's read set as index ranges and point keys, invalidates it from committed writes, reruns the query and pushes the new result; (b) server functions: queries, mutations as single TiKV transactions retried on conflict, and actions (R2; Resonate a later option for durable actions); (c) a protobuf sync API served by connect-rust (Connect, gRPC, gRPC-Web) with generated clients for web, iOS, Android, Python and Go; (d) a namespace router mapping tenants and apps to keyspaces or key ranges; (e) a bridge from Live tables into Loam collections so app data becomes searchable (vector, full-text, hybrid). Concepts are taken from Convex's public docs; Convex's FSL-licensed code is never copiedConvex's developer model (live queries, transactional functions) is what app builders want, and TiKV supplies the hard distributed-storage parts under Apache-2.0. The collections bridge is the differentiator: no reactive database also offers Loam's hybrid retrieval over the same dataApproved (owner); the name is proposed
D1182026-09-27Mutation transactions and isolation (§20 §5). A mutation is one TiKV optimistic transaction at a TSO start timestamp; a write conflict reruns the whole (deterministic) function, up to 8 attempts or its deadline; an optional idempotency key recorded in the same transaction makes client retries exactly once. Queries read one snapshot at the subscription manager's tick. Isolation is snapshot isolation, with documents read by id promoted into the lock set (lock_keys), so write skew is impossible on point reads and possible on index-range reads; serializable range reads are Q31. Commits default to async commit with 1PC, which cut commit p50 by roughly 30–50% against 2PC in the 2026-09-27 spike (indicative; loaded machine). A commit_mode switch falls back to 2PC for any component whose R1 gate fails with them, because tikv-client 0.4.0's async commit does not set max_commit_ts. Limits per mutation (R1): 8 MiB written, 16 000 documents written, 32 000 scanned, 4 096 index ranges, 1 s JavaScript CPU, 10 s wall clockTiKV's Percolator gives snapshot isolation; Convex promises serializability. Promotion closes the common read-modify-write case at the cost of one lock per read document; range serializability needs either guard keys (which serialize inserts into a bucket) or journal validation (which needs a proof), so it is decided with measurements in R2. Async commit can give a commit timestamp above a later start timestamp unless min_commit_ts comes from a fresh TSO (verify for client-rust)Approved (owner) for the transaction model; the isolation detail is proposed
D1192026-09-27Invalidation from a sharded, sequenced commit journal (§20 §5.3, §8). Each mutation writes, in its own transaction, one journal entry (the documents it wrote with their removed and added index keys) at seq = head + 1 of one of the app's journal shards (16 by default, up to 1 024) and advances that shard's head. A tailer per (node, app) reads the shard heads and new entries at each TSO tick and matches the written keys against the read-set index. A janitor deletes entries once every consumer has passed them and they are older than 10 min. TiKV CDC is a secondary path, not the invalidation sourceThe entries visible at tick T and not at t are exactly head(t) < seq ≤ head(T), so no update is missed and no timing assumption is needed. The 2026-09-27 spike showed that TiKV CDC does serve a txn-API keyspace from Rust with kv_api=TiDB (TxnKV is refused with a misleading compatibility error). But it needs copied private tikv-client stubs and a hand-written, region-by-region subscriber, and commit order waits on resolved ts, which advances about once a second. The journal invalidates within milliseconds and supports serializability validation (Q31). CDC stays a validated fallback, and a candidate feed for the collections bridge, where ~1 s is acceptable. The journal costs two keys per mutation and also gives the collections bridge dense producer sequences (D129)Proposed (§20)
D1202026-09-27Server functions run in QuickJS through rquickjs 0.14 (MIT) in R1 (§20 §6). One runtime per (node, app, deployment), 64 MiB memory limit, the interrupt handler enforcing the CPU limit; deterministic Date.now() and Math.random() from the start timestamp, and crypto.getRandomValues()/randomUUID() never seeded (they throw in queries and mutations and use the OS CSPRNG in actions); a fresh context per invocation, so no module state survives a call; no timers, network or filesystem in queries and mutations. Users write TypeScript; the CLI bundles it with esbuild (MIT). Deployed bundles live in object storage with a pointer in the app's catalog. V8 (v8/deno_core, MIT) and wasmtime (Apache-2.0 WITH LLVM-exception) stay options (Q35)Convex-style functions are small and I/O-bound, so an interpreter is fast enough; QuickJS compiles from C in seconds on a machine that builds one crate graph at a time, and has CPU and memory limits built in. V8 costs a ~100 MB prebuilt library or a V8 build; wasmtime needs a JS engine compiled to WASM for TypeScript anyway. The function API is engine-neutralProposed (§20)
D1212026-09-27The sync API is protobuf over connect-rust 0.9 (§20 §7): package loam.live.v1, service LiveService with Watch (a server-streamed session of versioned Transitions, heartbeats every 15 s), ModifyQuerySet, Query, Mutate (returns the commit timestamp) and Deploy. A session's version is (query-set version, identity version, ts); a client applies a Transition only from its current version; every query in a session is evaluated at one timestamp; a reconnecting client resumes and receives full results. The listener binds 127.0.0.1:7710 and, because Mutate and Deploy are unauthenticated in R1 (D111), refuses any non-loopback address until the unified auth plan covers the Live API. Clients are generated (protobuf-es and connect-es, connect-go, connect-swift, connect-kotlin, connect-python; Apache-2.0) with a thin hand-written reactive layer per platformA server stream plus unary calls works on HTTP/1.1 and HTTP/2 and over Connect, gRPC and gRPC-Web, which bidi does not (connect-rust sends no response on HTTP/1.1 until the request body ends, connect-rust/docs/guide.md:862). connect-rust passes the Connect conformance suite and serves through tower/axum; versioned transitions follow Convex's published sync modelProposed (§20); the transport is the owner's choice
D1222026-09-27Multi-tenancy by TiKV API v2 keyspaces with size classes (§20 §9). The TiKV cluster runs storage.api-version = 2 with enable-ttl from creation (a store holding TxnKV data cannot switch from v1). Small apps share a pool of keyspaces under an app-id key prefix; large apps and every SQL-enabled app get a dedicated keyspace. A directory (org, app) → {keyspace, prefix, class, state} lives in the ControlStore and is pushed to Live nodes as a live query; apps move between classes by a fenced copy and a directory flip. Loam runs the MVCC GC loop for its txn-API keyspaces through PD's keyspace GC-state RPCs (Q32). Amended by Q32's answer: Loam's GC loop is the cluster's MVCC GC worker: it resolves locks in every keyspace and advances the one cluster-wide safe point through PD's cluster GC RPCs. Router in R2; R1 has one app in one keyspaceKeyspaces give per-tenant key ranges, region boundaries, GC state and 2^24 ids on one cluster, but each costs at least one region, so millions of apps need shared keyspaces, as §18 §5.3's size classes do for namespaces. client-rust's gc() only moves the cluster-level safe point, so nothing advances a txn-API keyspace's safe point unless Loam does TiKV v8.5.8 has one cluster-wide safe point and a keyspace-mode TiDB neither computes it nor resolves locks, so Loam must do both for every keyspaceApproved (owner) for the router; the size classes are proposed; the GC clause is amended by Q32's answer (R1 Task 0: on v8.5.8, MVCC GC is cluster-wide and Loam's GC loop is the cluster's GC worker)
D1232026-09-27MySQL protocol through unmodified TiDB (§20 §10). TiDB's SQL layer (Apache-2.0) runs against the same TiKV cluster with keyspace-name set, in its own keyspace per SQL-enabled tenant. A TiDB process serves exactly one keyspace, so each SQL tenant gets its own TiDB pool; a shared TiDB with resource groups gives quotas but no data isolation and is not used for tenant SQL. Keyspace mode requires TiKV API v2 and the keyspace pre-created in PD. Loam builds no MySQL server. Live tables and SQL tables do not see each other in R1–R3. R1 runs one keyspace-mode TiDB in the dev playground; per-tenant pools come in R2 and scale to zero behind a proxy (TiProxy, verify) in R4TiDB is a mature MySQL-compatible SQL layer that already runs on TiKV keyspaces (tidb/pkg/store/driver/tikv_driver.go:190-202,252). Keyspace mode was verified on classic tiup playground v8.5.8 (2026-09-27 spike): TiDB and a Rust client ran in separate keyspaces with isolation confirmed. Q33 is narrowed to tidb-operator v2's support for it on classic clustersApproved (owner); superseded by D260 (2026-09-29): no TiDB; MySQL wire access is open (Q260)
D1242026-09-27operon-meta-tikv, a MetaStore backend over tikv-client 0.4, moves from M6 to R1 and becomes the metadata backend for Loam cloud (§20 §11). One transaction per call: optimistic for single-record writes, pessimistic with get_for_update on partition heads in key order for commit_wal, swap_segment and trim. commit_wal runs one transaction per partition group (at most 1 024 chunks and 4 MiB of metastore writes, under TiKV's Raft entry limit). A call that fits one group is atomic across all its partitions (stronger than D59). A larger call is split without ever splitting one stream's chunks, and each group commits idempotently through its own commit record. Success is returned only after every group commits, and a retry recommits only the missing groups: D59's per-group atomicity. clock_ms is the TSO physical time, one monotonic clock with no hot row. Fences and GC claims are checked in the same transaction. Unknown outcomes are resolved by a commit-token key read at a fresh timestamp. watch_changes polls per-scope change counters. It passes the 49-case conformance suite with linearizability histories and its own fault matrix. Supersedes D58's operon-meta-tidb (TiDB over sqlx, M6) and the tikv-client entry on the avoid list of §11. Postgres and DynamoDB stay v1.0 backends (D58); openraft + redb stays the default for operon dev and single node (D10). Gaps: MVCC GC (Q32), tikv-client's pre-1.0 status and its tonic 0.10 / prost 0.12 in the 0.4.0 release (0.12 / 0.13 on master; the workspace is on 0.14), per-call latency of a TSO fetch plus commit, and client gotchas found in the spike (the TSO stream dies after a PD stall and needs a supervisor; no pessimistic lock retry; async commit's missing max_commit_ts; no keyspace means no API v2). The first upstream PR candidates are reconnecting the TSO stream and exposing the generated proto modulesD58 rejected tikv-client because it needed a PD and TiKV cluster and rebuilt what TiDB's SQL layer provides; Loam Live now runs that cluster and builds the KV layer (operon-tikv) once for both. TiKV meets the strict contract (a snapshot per call, atomic multi-partition commits, a monotonic clock), so M2's relaxation costs this backend littleApproved (owner)
D1252026-09-27The Loam cloud ControlStore runs on Loam Live (§20 §9.5): §19's orgs, members, teams, projects, environments, agents, service accounts, API keys, role bindings, quotas, usage rollups, the namespace directory and billing accounts are tables of a system app _control in a fixed keyspace (loam_control). The console and the gateways read them through live queries; billing reports use the usage rollups exported to Iceberg (M4). The ControlStore trait (D65) stays, with the metastore-backed implementation for OSS and single node. R2The console and the gateways' directory are the heaviest control-plane readers and both want pushed changes, which a live query gives and TiDB SQL does not; one Rust data-access stack instead of Live plus sqlx and the MySQL dialect (Q24); dogfooding hardens Live fastest. Rejected: TiDB SQL, which is more mature and better for ad-hoc reporting, but needs polling or TiCDC for pushesProposed (§20)
D1262026-09-27PD, TiKV, TiDB and TiCDC run unmodified from official releases (§20 §15): tidb-operator (v2 CRDs) on Kubernetes, tiup playground (pinned to a v8.5.x release) in dev and CI. Loam forks none of them. Licenses: TiKV, PD, TiDB, TiCDC, tidb-operator, tiup, kvproto, tikv-client all Apache-2.0. Fixes Loam needs in tikv-client (keyspace GC, exposed protos) are offered upstream; until they merge, Loam generates its own stubs from the vendored kvproto protos rather than patching the crateUnmodified upstream keeps upgrades and security fixes cheap, and matches the buy-over-build preference; all are Apache-2.0, compatible with D11Approved (owner); the TiDB, TiCDC and TiProxy parts are superseded by D260 (PD and TiKV only, still unmodified)
D1272026-09-27Track R, "as soon as possible", beside M1 (§20 §18). The build machine builds one crate graph at a time, so R tasks interleave with M tasks and the playground runs only between builds. R1: operon-tikv; operon-meta-tikv with conformance and fault matrix; the keyspace GC loop; one Live app in one keyspace with documents, indexes, built-in and QuickJS queries and mutations, the commit journal, live queries over the server-streamed sync API, a generated TypeScript client with a reactive layer; TiDB SQL in its own keyspace in the dev playground. R2: the multi-tenant router and keyspaces, the ControlStore on Live, actions and scheduling, index backfill, Q31, multi-node sessions, Python and Go clients, per-tenant TiDB pools. R3: the collections bridge, auth through the unified auth plan (D111), Swift and Kotlin clients, React hooks. R4: tidb-operator and the Helm chart, TiDB pools that scale to zero, backups, BYOC, durable actions. First plan: docs/plans/2026-09-27-r1-reactive-core.mdThe owner asked for this "as soon as possible"; a separate track keeps M1's gates and plans intact. R1 is the smallest slice that proves the three claims: TiKV as the metastore, a live query end to end, and SQL on the same clusterApproved (owner); R1's "TiDB SQL in dev" item is dropped by D260
D1282026-09-27One protobuf toolchain for Loam Live and the native stream API (§20 §18): buffa messages and connect-rust services on the server (generated in build.rs with connectrpc-build and the system protoc), buf with protobuf-es and the Connect generators for clients. Proposed M1.6 amendment: where an SDK covers a protobuf service (the native gRPC protos of D101, the stream API of D72), it wraps the generated client instead of hand-writing transport; M1.6's REST SDKs, and the zero-runtime-dependency rule of its Ruling 4, are unchanged. M1.6 is not rewrittenOne toolchain means one set of generated clients per platform for every protobuf surface. Generated Connect clients depend on @bufbuild/protobuf and @connectrpc/connect, which conflicts with M1.6 Ruling 4, so the amendment is limited to protobuf surfaces. The workspace then carries buffa beside prost (Qdrant) and tikv-client's prost 0.13Approved (owner) for the shared toolchain; the M1.6 amendment is proposed
D1292026-09-27The collections bridge (§20 §12). A table declared searchable gets a Loam collection; a bridge task per (app, table) tails the commit journal (D119), reads the changed documents at each tick, and appends DocOps to the collection's implicit stream with idempotent producer id = (app, shard) and sequence = journal seq (D72), so a crash between append and checkpoint cannot duplicate. Embeddings come from the collection's ingest path, never from the mutation. The bridge records per collection the highest tick fully appended and its consistency token, so a search with after_ts waits for the bridge and reads with that token (D76). It runs under worker leases (§09), honours the unapplied-data budget (D86) and holds a GC barrier while it reads. R3Makes application data searchable (vector, full-text, hybrid) without ETL, reusing the log, links, tokens and idempotent producers the retrieval engine already hasProposed (§20)
D1302026-09-27D1 and D2 keep holding for the retrieval engine and do not apply to Loam Live (§20 §17). The retrieval engine keeps object storage as its only source of truth and stays out of OLTP. Loam Live is an OLTP product line whose source of truth is TiKV (Raft replication on local disks); the Live role itself stays stateless. Amends D1 and D2; mandatory BR log backup of Live keyspaces to object storage (D131) softens the D1 tensionThe owner's direction adds an application database; stating the boundary keeps the retrieval engine's design principles intact and makes clear which guarantees each product line givesProposed (§20)
D1312026-09-27TiDB's object-storage, vector and full-text features (§20 §10.4). (1) The next-gen TiDB kernel (object storage as the single source of truth) is not used: TiDB's side is open (the nextgen build tag, tidb/pkg/config/kerneltype/doc.go), but its shared-storage TiKV engine is not public (tikv/tikv's cloud-engine branch is stale since 2022-09-26; master has none), so self-hosted TiKV keeps its row store on local disks with Raft. (2) TiFlash (Apache-2.0: columnar replicas, a disaggregated mode on S3, the VECTOR type with HNSW indexes built as columnar indexes on TiFlash) is an optional add-on for SQL tenants in R4, run unmodified; Loam does not depend on it. (3) TiDB full-text is not used: in the open-source build a boolean MATCH … AGAINST is rewritten to ILIKE '%term%' (no scores, no stop words, no word boundaries, no index; tidb/pkg/planner/core/fulltext_to_like.go), and scoring positions need a TiFlash full-text path not found in public TiFlash. (4) BR log backup (PITR) to object storage is mandatory for every Live and metastore keyspace, into the tenant's or operator's bucket, from R2: a continuous log backup task plus periodic full backups, and a tested restore before a cluster serves external traffic (Q37). (5) tiup 1.17.1's --mode tidb-x / tidb-cse (next-gen, S3-backed) are unverified in binary source and license and are evaluated under Q38; the binary-only tici* components (alpha, no public repository) are not usable. Loam's own engine stays the vector and full-text engine for Live data, through the collections bridge (D129). Amends D130's rationaleLoam's retrieval engine is S3-native and gives scored hybrid retrieval with the hot tier and the Qdrant and Elasticsearch surfaces, which is the differentiator; TiDB's open full-text is a LIKE fallback, and TiFlash sees SQL tables only. Mandatory log backup puts a restorable copy of every Live keyspace in object storage, which softens the D1 tension D130 accepts, using open TiKV components (tikv/components/backup-stream, external_storage)Proposed (owner question, 2026-09-27; §20 §10.4); the TiFlash add-on (it serves TiDB tables only) is dropped by D260; BR log backup of TiKV keyspaces stays
D132 (number assigned at merge)2026-09-27The Qdrant gateway surface is Qdrant 1.19.1, REST and gRPC, including the legacy search, recommend and discover routes, their batches and groups, and the matching deprecated gRPC methods, which the 1.19.1 server still serves; GET / reports version 1.19.1. Verified in CI against the official Python qdrant-client 1.15.1 (REST, legacy methods) and 1.19.1 (query API, REST and gRPC)LangChain's lockfile pins client 1.15.1, whose Qdrant class calls search; LlamaIndex runs on 1.19; both suites must run unmodified against OperonApproved (M1.4 plan, Ruling 1; Task 10)
D133 (number assigned at merge)2026-09-27Vendored Qdrant protos with tonic-prost-build, and a hand-written REST model: nine protos from qdrant/qdrant v1.19.1 (Apache-2.0, qdrant.proto minus six internal imports) compiled with the protoc Lance already needs; REST types are a hand-written serde model of the Phase A subset, which gRPC converts into, so each operation has one executorQdrant's Rust client lags the server's protos and pulls client builders and TLS stacks into a server; the 1.19 OpenAPI omits the legacy routes, and generators get its anyOf precedence wrongApproved (M1.4 plan, Rulings 2 and 3)
D134 (number assigned at merge)2026-09-27Qdrant payload filters always read the catch-all payload Json field, text and datetime conditions included; payload indexes add lenient typed payload_index.<key> fields that filters never read (kept for payload_schema and hot-tier payload links)Qdrant filters any key without an index, and fields added later are not backfilled (overview A4), so reading typed fields would make results depend on when an index was createdApproved (M1.4 plan, Rulings 5 and 6; overview A2–A4)
D135 (number assigned at merge)2026-09-27Gateway-side scoring for non-nearest Qdrant queries and grouping: recommend (best_score, sum_scores), discover, context and MMR apply Qdrant's formulas to the union of IR candidate searches (candidate_k per example); groups follow Qdrant's collect-then-fill driver over the compiled queryThe IR has no custom-scorer retriever; Qdrant's own custom scoring is approximate too (inside an HNSW walk), and every returned score is exactApproved (M1.4 plan, Rulings 10 and 11)
D136 (number assigned at merge)2026-09-27Sparse vectors are served in M1 (owner decision 2026-09-25): config with modifier, values, nearest in query and prefetches, fusion, params.idf, exact scoring, so LangChain's sparse and hybrid cases are gated; sparse recommend/discover/context/MMR and multivectors answer 501The IR scores sparse vectors exactly as Qdrant does; Qdrant's sparse custom scoring is a full scan, which candidate-bounded rescoring would not matchApproved (M1.4 plan, Rulings 15 and 21; overview A14, A26)
D137 (number assigned at merge)2026-09-27Qdrant divergences settled by the owner: a write request over 10,000 listed operations is 400, asking to split it (O1); planning reads inside a batch see the pre-batch state (O2); sum_scores with negatives only and an empty context are accepted as Qdrant accepts them, and negatives-only best_score stays (T8-7); the group fill phase ORs the integer and string keys, where Qdrant ANDs them (T9-2). Each is listed in §06 §8's divergencesAtomicity over oversized batches; the fill OR differs only for groups of mixed key types and fills them where Qdrant's AND fills noneApproved (owner, M1.4 plan "Owner rulings")
D1382026-09-27The Resonate server is embedded in the operon/loam binary (§21 §3.1–§3.2), in a new crate operon-durable behind the cargo feature durable (on by default). It is built with Resonate's public composition API (resonate_base::build, Running::start/stop), never resonate_base::run, which installs a global tracing subscriber and waits on signals itself. Configuration comes from Loam's flags through resonate_plugin::Loader::set, never from resonate.toml or RESONATE_*. abort_on_panic is pinned off. The server listens on its own listener, 127.0.0.1:8001 (Resonate's SDK default, already reserved in §01 §3), speaks the standard protocol (2026-04-01) so the official SDKs work unchanged, and refuses non-loopback addresses until the unified auth plan (D111) configures authentication. Amends D19 (no gateway role) and §10 §2's example addressOwner direction of 2026-09-27 ("link resonate plugins to loams binary"). The linking spike built and ran it beside an axum 0.8 app with public API only, passed the Python fan-out example and survived kill -9 of the binary in the human-in-the-loop example (§21 §13). An unauthenticated durable API lets any caller schedule work and, through push, make Loam send requests to arbitrary addresses, so the refusal is Live's stricter rule, not D111's warningProposed (owner direction; details proposed); the "on by default" clause is amended by D262 (the durable feature is opt-in)
D1392026-09-27Durable storage backends (§21 §3.3): SQLite (resonate-server-sqlite) for operon dev and operon standalone, one file per tenant under the data directory, refused by operon cluster; TiDB through Resonate's MySQL plugin for Loam cloud now, one database per tenant in a system TiDB pool on keyspace loam_durable_sql, with migrations run as an explicit upgrade step; a TiKV Store for the blob server (upstream PR 2) as the target (D2), in keyspace loam_durable on the Live cluster with tenants as key prefixes and large tenants in their own keyspace. The self-hosted cluster backend without TiDB is open (Q39)Owner direction. TiDB passes Resonate's engine differential, port differential and porcupine on v8.5.8 with no new engine code (research); the TiKV store inherits the blob server's TLA+-checked one-document-per-origin design. D123's one-TiDB-per-keyspace rule is for tenant SQL, where tenants connect; here only Loam connectsProposed (owner direction; layout proposed); the TiDB clause is superseded by D261 (2026-09-29): durable state on the native TiKV backend, mysql:// legacy
D1402026-09-27Resonate comes from the fork dina-kar/resonate at a pinned revision (§21 §11.2): branch loam/0.10.1 = upstream 28dfd01 plus the dependency-hygiene commit (upstream PR 0c: rusqlite 0.32, sqlx 0.8.6, validator 0.20, prometheus 0.14, reqwest on rustls, versions on internal path dependencies), the TiDB fixes (PR 0a, PR 1) and, when D2 needs it, the library-entry change (PR 4); each commit is also proposed upstream. deny.toml gains allow-git = ["https://github.com/dina-kar/resonate"] and ignores RUSTSEC-2023-0071 (rsa via sqlx-mysql, no fix) with the rationale that Loam's TiDB connections use TLS or the cluster network and mysql_native_passwordUpstream 28dfd01 fails Operon's cargo deny (8 advisories, including sqlx 0.8.0's RUSTSEC-2024-0363), links OpenSSL through reqwest defaults, and pins sqlx to 0.8.0 through rusqlite 0.31's links = "sqlite3" conflict, which would also hold M2's operon-meta-postgres at 0.8.0. The hygiene commit builds with no source change and leaves one advisory (spike)Proposed (owner direction: fork at a pinned revision)
D1412026-09-27Transports and Loam's own durable functions (§21 §3.4–§3.5): SDK workers use http-poll (SSE on the durable listener); http-push is linked but off by default (--durable-push), and stays off in cloud until the auth plan and a per-org outbound allowlist exist; Loam's own workflows run on the Resonate Rust SDK (0.6, Apache-2.0) over an in-process Network that calls ResonateServer::process directly, with a Loam worker plugin for the scheme inproc://any@loam/<node>; a loopback-HTTP fallback is allowed if the adapter hits a blocker. The connect-rust loam:// transport is D4, after upstream's axum 0.8 bump (PR 3a)Owner direction (start with http-poll/push; connect later). Push on an unauthenticated server is a server-side request forgery primitive. Using the SDK keeps replay semantics (checkpoints, fan-out, sleep, promises) bought, not rebuiltProposed
D1422026-09-27Tenant isolation: one Resonate instance per namespace (§21 §5): server, workers and routes each, over the namespace's own SQLite file, TiDB database or TiKV prefix/keyspace; built on first use, evicted when idle, kept warm on the router's owner of (ns, "durable") when timers are pending; a dispatcher on the durable listener maps the authenticated credential to the namespace. D1 serves only the default namespace and has no dispatcher. Multi-tenancy depends on the unified auth plan (D111, Q30): namespace-bound tokens, SDK Authorization headers, and the actions durable:invoke, durable:resolve, durable:scheduleResonate has no tenant concept: groups, schedules, searches and timers are global to one server, so id prefixes would share them. The dispatcher needs upstream PR 4 (a public gateway router) or the same change in the forkProposed (owner direction: per-tenant namespace or keyspace); the TiDB-database layout is superseded by D261
D1432026-09-27Resonate patterns become Loam features with owner milestones (§21 §6): D1: the async operations API (d), bulk import with per-file fan-out (b), scheduled incremental import (a), idempotency keys (g). M2 (on D1's engine): GDPR erasure orchestration as a saga and its deadline sweep (c, a); restore and online index backfill (D97) as operations (d). D2: approval gates for destructive operations and agent actions (e); tenant provisioning and deprovisioning sagas (c) with R2. M2.x / D3: shard builds and bulk re-embedding (b), re-embedding operations (a). D3/D4: durable Live actions, the agent runtime, MCP operation tools, and agent traces in Loam (f). M1.3's compaction, merges, hot builds and GC stay lease-fenced worker tasks, and no M1 code is rewritten before M1 exitsDurability pays where one logical job outlives a task run or crosses systems; M1.3's loops are short, state-driven, idempotent and fenced, and a cron would only add latency and a dependency (§21 §6.1)Proposed (owner direction: all useful patterns)
D1442026-09-27Durable conformance (§21 §10): Loam does not re-run Resonate's engine and port differentials (the fork's CI runs them at the pinned revision for SQLite and TiDB); Loam's CI runs upstream's conctrace + conccheck (porcupine) against operon dev with the durable debug flag, per PR for SQLite (path-filtered) and nightly for TiDB (16 × 400, three seeds), plus a nightly SDK example suite (Python hello-world, fan-out with --crash, human-in-the-loop with kill -9 of operon, schedule, money-transfer; TypeScript hello-world and fan-out). Until upstream PR 0b, 503s are counted and treated as not appliedEmbedding changes no engine code; what it adds (configuration, listener, in-process network, workflows) is what a black-box run against the Loam binary testsProposed; the nightly TiDB leg moves to the TiKV backend with D261
D1452026-09-27Track D, beside M1, M2 and R (§21 §14), interleaved on the one-build machine like track R (D127). D1: the embedded server behind durable, SQLite and TiDB backends, the listener, the fork, the in-process runtime, the operations API, bulk import from object storage (the operation) and scheduled incremental import (the schedule), the conformance run and example suite. D2 (tenancy, TiKV backend, approvals, provisioning sagas), D3 (durable agents and Live actions) and D4 (connect transport, §14 Phase B) follow. Resonate Phase A leaves M3 (amends D46), and §20's "durable actions (Resonate), R4" moves to D3Bulk import exists on M1's code (Flight DoPut mapping + CollectionService::write), is the longest thing Loam does today, is every user's first step, and exercises fan-out, idempotency and the operations API at once; scheduled import is the one schedule that a worker loop does not already serve (§21 §7.1)Proposed; D1's TiDB backend is legacy under D261, and the TiKV backend moves from D2 into D1
D1462026-09-27The operations API and idempotency keys (§21 §6.4, §6.7): POST of a long operation answers 202 with Location: /v1/operations/{id} (D90's shape, now general); GET /v1/operations/{id}, GET /v1/namespaces/{ns}/operations, POST /v1/operations/{id}/cancel. The operation id is the root durable promise id; state maps from the promise and task; progress counts tagged branch promises (no second store); cancel settles the root rejected_canceled. An Idempotency-Key header maps to the operation id (op- + 26 hex chars of SHA-256(namespace ‖ key)); the same key with different parameters answers 409 idempotency_key_reused. A step's promise id is the idempotency key of every write it makes: collection upserts, stream producer ids (D72) and Live Mutate keys (D118)One API shape for import, restore, backfill, copy and erasure; the durable promise already is the dedup recordProposed
D1472026-09-27Durable payloads, retention and erasure (§21 §8): Loam's own promise ids and tags carry no personal data (opaque op-… ids and hashes); the docs tell SDK users to pass references, not personal data, in param and value and to use the SDKs' encryptor hook; D1 prunes Loam's own finished operations after 7 days; tenant deletion drops the tenant's durable store; general settled-promise retention is proposed upstream (PR 5) and open (Q40)The protocol has no delete for settled promises and no backend prunes them; pruning a promise that a late retry re-creates re-runs the work, so a general horizon needs careProposed
D1482026-09-29Neon and WeSQL run as separate, unmodified services on Loam's RustFS store; Loam never links either (§23 §3): it integrates through their HTTP admin APIs and wire protocolsOwner direction (sidecar model); WeSQL is GPL-2.0-only, so linking would breach D11; neither engine is a libraryProposed (owner direction)
D1492026-09-29Neon backs the showcase apps' Postgres OLTP (Plane, Zulip, GlitchTip, Keycloak, OpenFGA; §23 §5); Forgejo keeps its database as shipped (no TiDB, D-SC-16; WeSQL fails, D156), Matomo keeps MariaDB (D-SC-14) and OpenPanel keeps its own Postgres until tested on Neon; the Postgres front end over TiKV narrows to Loam Live's reactive, Convex-style API. Amends D-SC-12 and D-SC-16 (TiKV for Live, meta and durable state; Neon for Postgres-wire OLTP; still no TiDB)Neon is Postgres (16.9/17.5) with its durable state in the bucket; OpenFGA's and GlitchTip's migrations ran on it unchanged in the spike; a Postgres-compatible engine over TiKV would take yearsSuperseded 2026-09-29 by D230 (CNPG for the showcase apps, §28) (was: Proposed; needs owner confirmation (Q45))
D1502026-09-29Loam is Neon's control plane (§23 §3.2, §6.1): crate operon-neon behind feature neon drives the pageserver/storage-controller API, writes compute specs and starts computes through a ComputeRuntime trait; one Neon tenant per (namespace, database), branches are timelines, ids derived deterministicallyNeon's own control plane was never open source; a tenant per database keeps branches, quotas and erasure per app; deterministic ids make saga steps idempotentProposed; amended 2026-09-29 by D232 (Loam serves the control-plane API, §28 §5) (was: Proposed)
D1512026-09-29Neon is pinned and owned, with plain Postgres as the exit (§23 §4.1): images pinned by digest to the last public build (2025-08-26); a fork dina-kar/neon only when a fix is needed; routing and bridges address "a Postgres backend", so vanilla Postgres replaces Neon for them by configuration (an engine = postgres record naming an operator-provisioned instance, routed directly, not to a branch compute); its provisioning and a non-Neon runtime are not designed yet, and branching has no plain-Postgres equivalentThe public repository is nearly dormant since August 2025 (≥100 commits a month until July 2025, ~1 a month since; last release 2025-07-29; Postgres 16.9/17.5) after the Databricks acquisitionSuperseded 2026-09-29 by D231 (the fork is owned now, §28 §10) (was: Proposed)
D1522026-09-29Mapping records in operon-meta-tikv (§23 §6.2): x/<ns>/<db> external database (engine, Neon tenant and main timeline or WeSQL instance, state), X/<ns>/<db>/<branch> branch (timeline, parent, ancestor LSN, endpoint, owner, TTL), b/… bridge checkpoints; CAS on version; secrets stay in the auth plan's credential storeThe metastore already holds every other routing record; the prefixes are unused in keys.rsProposed
D1532026-09-29Routing by database name in Loam's wire listeners (§23 §6.3): the pg listener peeks the startup packet and splices mapped databases to the branch's compute (SCRAM end to end), otherwise PG1's analytics; the MySQL listener must terminate the handshake (the database arrives after the server's scramble), then reconnects to WeSQL with a stored credential. Loopback-only until the unified auth plan (D111)One endpoint per tenant; Postgres sends the database before the server speaks, MySQL does notProposed; amended 2026-09-29 by D236 (PgDog routes Postgres OLTP; the pg listener keeps analytics) (was: Proposed); MySQL half: vtgate in front of WeSQL proposed 2026-10-01 by D320 (§31, Q313)
D1542026-09-29Change bridges into collections (§23 §6.4): pgoutput logical replication from Neon and the row binlog from WeSQL become DocOps on a collection's implicit stream with idempotent producer (database, branch, bridge) and sequence (commit LSN, ordinal) or (binlog file, pos); confirm to the source only after the append; lag alarms and resyncSame exactly-once contract as D129; the spike showed pgoutput decoding and a slot surviving compute replacementProposed
D1552026-09-29A Neon branch per agent workspace (§23 §6.5, §15 §9) as a Resonate saga (§21 §6.3): record, create timeline, start compute, return a routed DSN, with compensation; at the end keep (TTL), discard, or promote only if main has not diverged; never mergedBranch creation took 36 ms and isolated both ways in the spike; Postgres cannot merge timelines, so results land as migrationsProposed
D1562026-09-29WeSQL is gated and lower priority (a proposed amendment, D273, applies only after W1 and W2 pass; until then this stands) (§23 §9.2): not for Forgejo (SmartEngine refuses foreign keys: Error 1235); a Matomo candidate only after Q50 and Q51; MariaDB stays (D-SC-14); dev compose only until thenSpike: Forgejo 16.0.5 migrations failed; commits after the last snapshot were lost when the local volume was lost; beta images only (last 2025-01-21), single node since 2026-08-22Proposed
D1572026-09-29Track N (§23 §11): N1 operon-neon client; N2 metastore records; N3 pg routing (after PG1 Tasks 0–5); N4 logical-replication bridge; N5 branch-per-workspace saga; N6 WeSQL deployment, binlog bridge and MySQL routing (after Q50, Q51); each behind feature neon, changing no default pathSmall stacked PRs; N1 and N2 need nothing from PG1Proposed; N3 amended 2026-09-29 by D236 (PgDog config rendering, not the pg listener; §28 P3)
D-PG-12026-09-28Postgres wire access to collections through datafusion-postgres (datafusion-contrib, Apache-2.0; pgwire 0.40, DataFusion 54), placeholder number assigned at merge. Behind the operon feature pgwire (off by default until upstream fixes land), --pg-listen <addr> serves the Flight SQL surface (collections as tables in schema collections, the search table functions) over the Postgres protocol; the Postgres database is the namespace; datafusion-pg-catalog supplies pg_catalog; every query plans through operon_query::sql::plan_read_only. The listener is off unless given and refuses non-loopback addresses (no auth or TLS until the unified auth plan, D111, Q30). Read-only by default: other statements answer SQLSTATE 25006. Writes behind --pg-allow-writes (PG1 Tasks 6–10) map onto collection write paths: INSERT = create-if-absent with 23505 on an existing key, ON CONFLICT (_id) DO UPDATE = upsert (a partial SET = merge patch), DO NOTHING = skip existing; UPDATE … WHERE = patch_by_filter, DELETE … WHERE = delete_by_filter (D87); COPY … FROM STDIN = the Flight DoPut bulk path; CREATE TABLE is not mapped in PG1. Every statement is autocommit: a BEGIN … COMMIT block may hold one write, applied when it runs; a second write, a ROLLBACK after an applied write, and SAVEPOINT answer 0A000. Consistency tokens return as the ParameterStatus operon_consistency_token and a NOTICE; SET operon.consistency_token reads at least a token. The Postgres wire is analytical and ingest access, not OLTP: clients needing multi-statement transactions use Loam Live / TiDB (D123). Plan: docs/plans/2026-09-28-pg1-postgres-wire.md, the first milestone after M1The spike (docs/plans/pgwire-spike.md) served psql (\dt, \d, filters, aggregates, vector_search/text_search/rrf), psycopg 3 and node-postgres through about 300 lines of Operon code, refused every write with 25006, and passed cargo deny; arrow and DataFusion unify. Buying the protocol, encoder and catalog is cheaper than building them. The blockers are upstream and small: arrow-pg panics on FixedSizeList (vector columns) and the three crates enable DataFusion's default features (parquet, compression; +344 s and 261 crates on the first feature build, +74 MB debug). Operon has no multi-statement ACID over collections, so the transaction rules refuse what they cannot honour instead of pretendingProposed (owner approved the spike; writes requested by the owner)
D-SC-1 (number assigned at merge)2026-09-27The showcase suite is named Loam Commons (§22 §1), in its own repository dina-kar/loam-commons, separate from the engine workspaceA suite of third-party apps and glue has a different license profile and release cadence from the engine; a separate repository keeps D11 cleanProposed
D-SC-2 (number assigned at merge)2026-09-27The suite's apps are Plane, Forgejo, Zulip, PostHog and GlitchTip, each run unmodified from official images as a separate service; GlitchTip replaces Sentry (§22 §4.3)Sentry is FSL-1.1-Apache-2.0, not OSI open source; GlitchTip is MIT, takes Sentry SDKs by DSN and runs on Postgres with optional Valkey, against Sentry self-hosted's dozens of containersProposed
D-SC-3 (number assigned at merge)2026-09-27Keycloak is the suite's OIDC provider (§22 §4.4); Loam and every app are clients of realm commons; Loam's gateway stays an OIDC relying party (§19 P7)Apache-2.0 with no enterprise split, full admin API, group claims, SAML brokering; already §19's named broker. Authentik and Ory keep SCIM or SAML in enterprise editions; Dex has no user store; ZITADEL is AGPL-3.0-only (allowed as a separate service, but second choice)Superseded 2026-10-01 by the owner's Authentik ruling (D404, D447; §22 §4.4) (was: Proposed)
D-SC-4 (number assigned at merge)2026-09-27One OpenFGA model spans Loam and every app (§22 §7): a suite project fans out to a Loam namespace, a Plane project, Forgejo repositories, Zulip channels, a PostHog project and a GlitchTip project, in the store Loam shares with Lakekeeper (D67), with §18 §7's tenant fence. OpenFGA is the source of truth; each app's ACL is a projectionOne grant model across apps is the point of the control plane; the apps do not call OpenFGA, so their ACLs must be derivedProposed
D-SC-5 (number assigned at merge)2026-09-27Enforcement by provisioning sagas on Loam Durable (§22 §7.3): outbox tuple changes start idempotent, compensating workflows that call each app's admin API; a durable schedule reconciles drift with audit events; Keycloak groups become group#member tuples; OIDC group claims (Forgejo, Zulip) are a coarse second layerSagas are §21's pattern (c); an outbox avoids orphaned tuples (§18 §7); reconcile bounds drift from changes made in app UIsProposed
D-SC-6 (number assigned at merge)2026-09-27Suite code is Apache-2.0; apps keep their licenses as separate, unmodified services (mere aggregation); the repo ships NOTICE, LICENSES.md and a license-check CI job; any patch goes upstream first, and one that must ship lives in a public fork with its license's source offer (AGPL-3.0 §13 for Plane)D11 governs linked dependencies, not separate services (as D60 did for Alternator); copyleft obligations attach only on modification or redistributionProposed
D-SC-7 (number assigned at merge)2026-09-27What Loam replaces in the suite (§22 §6): every app's object storage on RustFS; Forgejo's issue indexer on Loam's ES API (its queries are in Phase A; the code indexer needs Phase B highlighting and aggregations); all logs over OTLP; cross-app search, activity and assistant on Loam; Forgejo and OpenFGA on TiDB only if their suites pass. Postgres stays for Plane, Zulip, PostHog, GlitchTip, Keycloak and OpenFGA; ClickHouse and Kafka stay for PostHogThe apps' OLTP needs Postgres transactions and Postgres-specific SQL; Loam's Postgres wire is analytical and autocommit, and D2 holds for the engine. Claiming more would not survive a demoProposed
D-SC-8 (number assigned at merge)2026-09-27The showcase features (§22 §8): unified hybrid search over every app with OpenFGA filtering at query time, an activity feed on Loam Live, provisioning sagas with the operations API, OTLP logs, and an MCP assistant on §19's user delegationEach is a Loam feature no single app has, visible in products people knowProposed
D-SC-9 (number assigned at merge)2026-09-27Docker Compose for development, an umbrella Helm chart for Kubernetes (§22 §9), using upstream charts where they exist; PostHog is an optional profile (its hobby compose is its only self-host path)Compose is how most self-hosters start; upstream charts avoid maintaining app deploymentsProposed
D-SC-10 (number assigned at merge)2026-09-27SC1 is a post-release milestone (after v1.0, the implemented unified auth plan Q30, §19's OIDC login, M2's OTLP logs and D2's durable tenancy); it changes no engine code, and engine gaps it finds become engine issuesThe suite is a consumer of Loam; building it before auth exists would force a throwaway scheme (D111)Proposed
D-SC-11 (number assigned at merge)2026-09-28Loam Commons is a separate project (§22 §13a): its own repository, roadmap and releases; SC1 is its first plan, and the engine roadmap lists it only as a consumerOwner direction, 2026-09-28Proposed (owner direction)
D-SC-12 (number assigned at merge)2026-09-28Postgres write compatibility in Loam (§22 §13a): a Postgres wire surface that accepts transactional writes, designed and planned separately in the engine; TiKV (§20) is the candidate store; apps move off Postgres one at a time, only when their own test suites pass. Amends D2 and §22 §2.2 and §6.4Owner direction, 2026-09-28; lets the suite's apps run on Loam end to endProposed (owner direction); design pending (Q-SC-9); amendment proposed by D149 (§23 §5: Neon serves Postgres OLTP, the TiKV front end narrows to Loam Live's API), pending Q45
D-SC-13 (number assigned at merge)2026-09-28Fork PostHog onto Loam's Iceberg analytics (§22 §13a): dina-kar/posthog-loam from posthog-foss (MIT, never ee/); ingestion through the native stream API into Iceberg tables (M4), HogQL compiled to DataFusion SQL instead of ClickHouse SQL; staged scope in its own design docOwner direction, 2026-09-28; the strongest showcase: a known analytics product on Loam with no ClickHouseProposed (owner direction); superseded by D-SC-15 (2026-09-28): PostHog is dropped
D-SC-14 (number assigned at merge)2026-09-28Suite apps: Plane, Forgejo, Zulip, GlitchTip, OpenPanel (AGPL-3.0) and Matomo (GPL-3.0, PHP + MariaDB); PostHog and Sentry dropped (§22 §13b). OpenPanel has no generic OIDC (a gap: upstream a patch or use forward-auth); Matomo uses the third-party LoginOIDC plugin (GPL-3.0)Owner direction, 2026-09-28Proposed (owner direction)
D-SC-15 (number assigned at merge)2026-09-28Fork OpenPanel onto Loam in place of ClickHouse (dina-kar/openpanel-loam, AGPL-3.0 source offer): events through the stream API into Iceberg (M4), its ClickHouse queries rewritten for DataFusion over Flight SQL; its own design doc. Replaces D-SC-13Owner direction; OpenPanel's ClickHouse surface is far smaller than PostHog'sProposed (owner direction)
D-SC-16 (number assigned at merge)2026-09-28No TiDB in the suite; TiKV only, through Loam's tikv-client fork (§22 §13b). Withdraws Q-SC-3 and Q-SC-5 and SC1 Task 7; points D-SC-12 at a Postgres front end over TiKV. The engine-wide effect on D123 and D139 is decided with the Resonate + TiKV + Dapr workstreamOwner direction, 2026-09-28Proposed (owner direction); amendment proposed by D149 (§23 §5: Neon beside TiKV for Postgres-wire OLTP, still no TiDB), pending Q45; applied engine-wide by D260 and D261 (2026-09-29)
D1702026-09-29Loam Functions positioning (§24 §1, §2): a CPU-time serverless runtime for agent and backend functions that spend most of their time waiting, placed next to Loam's data (retrieval, Live on TiKV, durable workflows, Neon branches). Frontends get static hosting plus the workerd framework preset, not a Vercel cloneOwner change 7 to the draft "CPU-Time Serverless Runtime" v2. §24 §9's cost model shows that CPU price alone cannot beat Cloudflare, so placement and suspension are the advantageProposed · approved in conversation 2026-09-29
D1712026-09-29JavaScript runs on workerd (Apache-2.0), not rquickjs plus homemade polyfills. The Cloudflare Workers presets are reused (Hono, Nitro, Astro, SvelteKit, React Router/Remix, OpenNext). workerd's README says it "is not a hardened sandbox", so there is one workerd process per tenant, under gVisor or seccomp (§24 §4.1). D120 (QuickJS for Live) is unchanged; Q35 gains workerd as a candidateOwner change 1. Reuses the frameworks' existing adapters instead of maintaining a Web platformProposed · approved in conversation 2026-09-29
D1722026-09-29T1 is wasmtime (v49), with no WasmEdge benchmark. Before Loam builds a host of its own, F0 evaluates Spin and wasmCloud (wash-runtime), both Apache-2.0, as the T1 host (§24 §4.1, §11)Owner change 2Proposed · approved in conversation 2026-09-29
D1732026-09-29Suspend-on-await scope (§24 §6): short waits keep the instance resident (zero CPU, a few MB); long waits go through Resonate durable promises; only Resonate SDK code (deterministic and replayable) is evicted and resumedOwner change 3. Arbitrary code cannot be resumed without memory snapshots, and T3 defers themProposed · approved in conversation 2026-09-29
D1742026-09-29T3 Firecracker snapshots are deferred (§24 §4.1). In Kata's Firecracker driver, PauseVM, SaveVM and ResumeVM are no-ops (src/runtime/virtcontainers/fc.go), so snapshot restore needs firecracker-containerd or orchestration of our ownOwner change 4Proposed · approved in conversation 2026-09-29
D1752026-09-29Metering sources (§24 §7): Wasm fuel or epochs for T1; per-process CPU (cgroup cpu.stat of the tenant's workerd) for T0; cgroup cpu.stat for gVisor. eBPF comes last, as a cross-checkOwner change 5. A per-switch BPF map update taxes every tenant's hot pathProposed · approved in conversation 2026-09-29
D1762026-09-29Phase-1 protocols: HTTP/1.1, HTTP/2, HTTP/3, gRPC, WebSocket and SSE. Kafka arrives with the M5 Kafka gateway (D74). NATS, MQTT and AMQP go through the shared Dapr runtime (§24 §1)Owner change 6Proposed · approved in conversation 2026-09-29
D1772026-09-29Track F (§24 §11): F0 is spikes. F1 is the manifest, tenant identity in the Rust Dapr server, workerd and wasmtime with Hono, Resonate suspend and resume, and fuel or per-process CPU metering into the WAL. F2 is gVisor (T2). Later: Firecracker, eBPF and more protocolsOwner change 8Proposed · approved in conversation 2026-09-29
D1782026-09-29RustFS is the default object store for self-hosted and GitOps deployments. Cellar and other S3 stores are providers behind an ObjectStoreProvider trait in the forked operator, which refuses any provider without conditional writes (§25 §5)Owner change 9; extends D61Proposed · approved in conversation 2026-09-29
D1792026-09-29The runtime's metastore is TiKV (not "raft or postgres"), deployed by TiDB Operator v2 as Cluster + PDGroup + TiKVGroup with no TiDB (v2.0.0 GA 2025-12-18, v2.0.1 2026-03-26; the CRDs are still core.pingcap.com/v1alpha1). Upstream's examples/pdms has exactly that shape (§25 §6.4). operon dev and standalone keep openraftOwner change 10; follows D-SC-16 and D124Proposed · approved in conversation 2026-09-29
D1802026-09-29Billing is CPU time plus requests, not wall time. Meter events are idempotent records on a _meter system stream, rolled up into Iceberg (§24 §7)The draft's principle "waiting must cost nothing"; the WAL is already the durable spine (D4)Proposed (owner's draft); amended by D190
D1812026-09-29Three app contracts: static, fetch (a Hono-style handler) and http-port. The contract selects the tier (§24 §4.3)The draft, narrowed by D171 and D174Proposed (owner's draft)
D1822026-09-29Tenant identity is checked in the Rust Dapr server (loam-dapr), from mTLS/SPIFFE or a sandbox token. The namespace comes from the credential, never from a request field. OpenFGA decides, and meter events are recorded there (§24 §5). The server does not exist yet: today's operon-stream-grpc is StreamService.Produce and deploy/dapr/edge is a Dapr app, both uncommitted (§24 §3.1)The draft's hardening item; D142's dispatcher ruleProposed; refined by D200 (loam-dapr records no meter events)
D1832026-09-29Go daprd only for the long tail: one shared runtime per cluster, reached after loam-dapr's tenant check. Dapr Workflow stays denied (§24 §5)The draft; Resonate owns durable executionProposed (owner's draft)
D1842026-09-29A Rust gateway with a CloudEvents 1.0 envelope, and Envoy as the edge. Sōzu is rejected (AGPL-3.0, no HTTP/3, no gRPC routes). Pingora is the library if a Rust L7 component is needed; River is stalled (§25 §3)D176 needs HTTP/3 and gRPC in phase 1; the license ruleProposed
D1852026-09-29Fork CleverCloud/clever-kubernetes-operator (MIT, Rust, kube 3.1) as loam-operator: keep crates/core, svc/k8s, svc/http and the Helm chart; drop the Clever add-on CRDs; rename the API group; add the Loam, ObjectStore, Function and RuntimePool CRDs. kaniop (AGPL-3.0) is a design reference only (§25 §5)The draft. A permissive Rust skeleton, even though its reusable core is small (about 1,500 lines)Proposed (owner's draft)
D1862026-09-29GitOps with Argo CD: an app-of-apps over one umbrella chart, deploy/helm/loam-stack. Waves: −2 CRDs; −1 operators and gVisor node setup; 0 RustFS; 1 PD/TiKV; 2 Resonate; 3 loam-dapr, the gateway and Envoy; 4 Loam; 5 runtime tiers. Lua health checks for Application, the TiDB groups, Loam and Function (§25 §6)Sync waves and CRD health gates match the plan; hub and spoke for cloud and BYOC. The chart stays usable from FluxProposed
D1872026-09-29Clever Cloud adoption verdicts (§25 §2). Linked: biscuit-auth (Apache-2.0) and, optionally, clevercloud-sdk (MIT). Fork: clever-kubernetes-operator. Tools and services, on Clever targets only: terraform-provider-clevercloud, karpenter-provider-clever-cloud, biscuit-cli. Rejected: Sōzu, sozu-gateway (links the LGPL sozu-command-lib), the Pulsar, Warp 10 and FoundationDB tooling, and unlicensed repositories. Reference only: clever-tools, clever-components, kawa, nlrsThe owner's request of 2026-09-29 to evaluate every Clever Cloud tool Loam could adopt; the license ruleProposed
D1882026-09-29Biscuit carries sandbox tokens and delegation inside the runtime. The control plane issues node tokens; the supervisor attenuates them per invocation and per call; loam-dapr verifies them offline. JWTs (§19 §5.3) stay for external clients. OpenFGA stays the authority, and token Datalog carries only narrowing facts (§25 §4)Offline, cryptographic attenuation is exactly §19's "vended tokens can only attenuate", without a round tripProposed
D1892026-09-29Tenant secrets go through Dapr's secrets building block (§24 §5.1), covering every store in dapr/components-contrib secretstores/: AWS Secrets Manager and SSM Parameter Store, Azure Key Vault, GCP Secret Manager, HashiCorp Vault/OpenBao, Kubernetes, Alibaba, Tencent and Huawei. The shared Go daprd serves them off the hot path. There is one component per tenant and store, holding the tenant's own credentials. loam-dapr enforces the scoping by mapping the credential's tenant to its component names before calling GetSecret. Bulk reads never reach the shared daprd: loam-dapr answers them with one GetSecret per key that is on the tenant's allow list and granted by OpenFGA, since GetBulkSecret returns every secret in the store. Values are cached per node with a TTL. workerd and gVisor receive secrets over loopback, and wasmtime through a host importOwner decision 2026-09-29; one path to every cloud's secret store, with no client per cloudProposed · owner decision 2026-09-29
D1902026-09-29Billing and metering move to the private loam-platform repository. §24 exposes only open hooks: Prometheus metrics, cgroup labels, Envoy access logs and OTLP spans. Amends D180, and the "into the WAL" part of D177Owner decision 2026-09-29Proposed · owner decision 2026-09-29; refined by D200–D202 (the usage-hooks contract, §27)
D2002026-09-29The Rust Dapr server does not meter (§27 §2): loam-dapr checks tenant identity and authorizes (D182) and records no meter events; it exports only its own call metrics as hooks. Refines D182 and D190The server does not exist yet (§24 §3.1) and sees API calls, not CPU; usage is node-level data, read best from outside the sandboxesProposed
D2012026-09-29The usage-hooks contract (§27 §3): metric families and labels per namespace and function (Prometheus, and OTLP delta metrics for high cardinality); the cgroup layout (loam.slice/tenant-<org>.slice/workerd.scope for T0, loam.slice/wasm-host.scope for T1, the pod cgroup with loam.dev/* labels for T2), removed only after the consumer acknowledges the final read (or a retention window when no consumer is connected); per-invocation host reports (loam.meter.v1.HostReport) on /run/loam/meter.sock, buffered and resent until acknowledged, with cpu_usec flagged as an estimate on T0; Envoy access logs with the gateway-set x-loam-tenant (client values stripped), sink off by default. Versioned like an API. Refines D190 and §24 §7Any metering system can consume standard interfaces; self-hosters get usage visibility without billingProposed
D2022026-09-29The dependency runs one way (§27 §4): metering and billing (D190) are neither in the engine nor in its chart, and the engine never depends on them. Quota enforcement stays in the engine (D65, D98), with limits from configuration or a control planeOpen-core boundary: what a single organisation needs to self-host stays open sourceProposed
D2042026-09-29Loam Jobs: Celery, BullMQ, PySpark and Flink jobs on Loam (§26 §1): Celery through a kombu transport, result backend and beat scheduler (loam-celery); BullMQ through @loam/bullmq; PySpark through Sail run by Loam as the Spark Connect endpoint, with Iceberg on RustFS; Flink SQL on RisingWave (Arroyo optional) and DataStream jobs on unmodified Flink under the Kubernetes operator, reading Loam through the Kafka gateway (M5). One Rust crate, operon-jobs, under all four. Durability comes from Resonate's own SDKs against the embedded server; no new decorator systemOwner: "deploy PySpark, Flink, BullMQ, Celery jobs on Loam … with some modification like adding Resonate decorators … clean Rust API"; buy over buildProposed (direction approved by the owner, 2026-09-29)
D2052026-09-29The Jobs trait (§26 §5): enqueue, enqueue_bulk, lease (long-poll, several queues, batched), extend, complete with a typed Outcome (Ok, Retry, Fail{dead_letter}, Release, Delay, WaitChildren, RateLimited), report, schedule/unschedule, flow, submit_engine_job/control_engine_job, watch (resumable by cursor) and admin calls; every call takes a Ctx whose namespace comes from the credential; leases are LeaseToken{queue, job, epoch} fencing tokens; one JobsError enum mapped one to one onto Connect codesThe owner's sketch, refined for tenancy, fencing, typed outcomes and errors, and the admin surface Celery and BullMQ needProposed
D2062026-09-29loam.jobs.v1 over connect-rust (§26 §5.5): protos under proto/loam/jobs/v1/, operon-jobs-proto generated in build.rs like operon-live-proto; unary Lease with a wait (no push buffer); the jobs listener on 127.0.0.1:7720, loopback until the unified auth plan; Python (connectrpc) and TypeScript (@connectrpc/connect) clients generated with buf; the adapters wrap themD128's one proto toolchain; one surface for SDKs, adapters and the consoleProposed (direction approved by the owner, 2026-09-29)
D2072026-09-29Job semantics (§26 §6): at-least-once delivery; the outcome recorded exactly once through per-job lease epochs checked in the completing transaction (check_fence's rule); no exactly-once execution claim, with idempotency keys, the lease epoch as a fencing token and durable mode as the tools; job_key, idempotency keys (TTL) and BullMQ's four dedup modes as separate mechanisms; BullMQ's priority model (none first, then 1…2^21−1, lower first); delays in a delayed index, never as invisible leases; server-side stalled sweep; queue rate limits and global concurrency in the lease transaction; dead-letter queues; retention; an outbox relayed to a per-queue event stream with an idempotent producer (D72)Matches what Celery on Redis and BullMQ give users today, while removing the ETA-longer-than-visibility-timeout redeliveryProposed
D2082026-09-29loam-celery (§26 §7.1): a kombu virtual transport with a server-side visibility timeout (Loam leases, extended by a transport thread), registered by broker_transport = "loam_celery.transport:Transport" or an import-time TRANSPORT_ALIASES entry (kombu has no transport entry points); a BaseKeyValueStoreBackend result backend with incr for native chord counters (entry point celery.result_backends); LoamScheduler that upserts beat entries as Loam schedules (entry point celery.beat_schedulers), so running beat twice cannot duplicate ticks; fanout for remote control; job_key = (task id, retries); Celery priorities mapped to Redis's order; mingle and gossip off (hard-coded to other driver_types)Celery 5.6.3 / kombu 5.6.2 source; queue-mode canvas runs unchangedProposed
D2092026-09-29@loam/bullmq implements BullMQ v6's IQueueBackend (§26 §7.2): a BackendFactory for the unmodified bullmq package (peer dependency bounded to tested minors, >=6.3.9 <6.4 at J2, widened per minor after BullMQ's suite passes), plus BullMQ's classes re-exported bound to it, so the import change works and so does setDefaultBackendFactory; BullMQ's backend-neutral test suite is the J2 gate; Redis emulation rejected (49 Lua scripts, 5,243 lines, 49 Redis commands); v6 only, no v5 shim (owner, 2026-09-29, Q100); BullMQ Pro features excludedBullMQ 6.3.9 (MIT) has a documented pluggable backend with Redis and Postgres implementations; the API is BullMQ's own and cannot driftProposed; refines the owner's "drop-in package with the same API"; v6-only approved (owner, 2026-09-29)
D2102026-09-29How Resonate is used (§26 §8): schedules are Resonate schedules whose ticks enqueue with the tick's promise id as idempotency key (every intervals as a durable sleep loop); flows are Loam durable functions whose joins are per-node completion promises settled by the queue's outbox; queue mode creates no promises, durable mode is opt-in per task by writing it as a Resonate function called with an id derived from the job, so a redelivery resumes it; the queue is not built on Resonate tasks (no priorities, rate limits, counts or listing); loam helpers only configure Resonate's SDKs (URL, token, namespace, group)Keeps millions of short jobs out of the durable store (§21 Q40) while giving step durability where side effects matterProposed (direction approved by the owner, 2026-09-29)
D2112026-09-29PySpark on Sail (§26 §9): one Sail server per namespace (it runs Python UDFs in its own process), scaled to zero, never linked (D51); Loam's gateway routes authenticated sc:// connections to it; Iceberg through Lakekeeper on RustFS; SparkBatch jobs as durable workflows; Resonate at job level only; RDD, SparkContext, JVM UDFs, MLlib, pandas-on-Spark and Structured Streaming fall back to Apache Spark 4 on Kubernetes. Amends §17 §6.1 ("Loam serves no Spark Connect endpoint"): Loam still implements no Spark Connect, it hosts Sail. Note on D55: Sail's main already has DV reads, so M4 verifies instead of contributingSail 0.7.1: ~91 % of the Spark 3.5 Connect suite, ~71 % on 4.2; the data-processing research of 2026-09-29Proposed (direction approved by the owner, 2026-09-29)
D2122026-09-29Flink (§26 §10): Flink SQL is ported to RisingWave (the D22/D74 companion; Postgres dialect, porting guide, no translator); Arroyo is a documented alternative, not a managed engine; DataStream jobs run unmodified as FlinkDeployments under the Flink Kubernetes Operator (1.16.1) with checkpoints and savepoints on RustFS, reading Loam through the Kafka gateway (M5) or Iceberg (M4); Loam manages only the lifecycle (deploy, FlinkStateSnapshot savepoints, upgrade, rollback with rollback.enabled on, suspend) as durable workflows, never a second checkpoint layerNo Rust replacement for DataStream; Arroyo has had no release since 2025-12 and runs on a DataFusion 48 forkProposed (direction approved by the owner, 2026-09-29)
D2132026-09-29Jobs tenancy and security (§26 §11): everything namespaced, the namespace from the credential; the jobs listener loopback-only until the unified auth plan (D111), which adds jobs:enqueue, jobs:consume, jobs:admin, jobs:schedule, jobs:engine; workers are principals; engines in the namespace's Kubernetes namespace behind network policies; quotas per D65; namespace deletion removes jobs, results, streams, payloads and engine deploymentsSame rule as the durable listener (D138) and Live (D121)Proposed
D2142026-09-29Where job state lives (§26 §3.2, §6.8): a JobStore trait with TikvJobStore (keyspace loam_jobs, tenant prefix, operon-tikv's runner; optimistic enqueue, pessimistic lease) and LocalJobStore (redb, operon dev only; refused by standalone and cluster); payloads and results over 16 KiB on the object store with a Freshness; the event log on Loam streams; schedules and flows on Resonate; jobs require TiKV on every deployment but operon dev, self-hosted included; no Postgres or DynamoDB jobs backend (owner, 2026-09-29, Q97)Dequeue is a small multi-key transaction; TiKV prefers small valuesProposed; the TiKV requirement approved (owner, 2026-09-29)
D2152026-09-29Track J (§26 §15), beside M, R and D, interleaved on the one-build machine: J1 Celery (core, LocalJobStore, listener, schedules, loam-celery, TikvJobStore); J2 BullMQ (@loam/bullmq, flows, events); J3 PySpark on Sail (engine runners, the Spark Connect proxy, the Spark fallback); J4 Flink (RisingWave SQL, the operator lifecycle)Celery and BullMQ first (owner); J3 needs the auth plan and M2/M4, J4 needs M5 for reading streamsProposed (direction approved by the owner, 2026-09-29)
D2162026-09-29Jobs licences (§26 §12): everything linked is Apache-2.0 or MIT/Apache (connect-rust, buffa, redb, tikv-client, the Resonate crates, kube-rs); adapters depend on BSD-3-Clause Celery/kombu and MIT BullMQ; Sail, RisingWave, Arroyo, Flink and Spark run as separate services; BullMQ Pro, Ververica Platform and RisingWave Premium features are excluded; upstream proposals (a kombu transport entry-point group, Celery capability checks) are optional and not postedThe rule: nothing linked AGPL, BSL, SSPL or ELv2Proposed
D2202026-09-29Open-core boundary (open-core.md): everything needed to self-host Loam as a single organisation stays Apache-2.0 in this repository (engine and every wire API, Live on TiKV and the TiKV metastore, change-feed bridges, durable and jobs, the function runtime and gateway, namespaces, OIDC/API-key auth, OpenFGA, quota enforcement, observability and usage hooks, Helm, the operator, backup and restore, SDKs and CLI); what is needed only to run the multi-tenant paid cloud (metering and billing, the control plane and per-plan quota setting, fleet and multi-region operations, hosted Neon/WeSQL fleet automation, BYOC management, abuse, the internal admin console and runbooks) lives in the managed Loam Cloud platform (proprietary, separate repository, loam-platform). This repository never depends on loam-platform; the platform consumes its open hooks and APIs. Borderline: Neon/WeSQL routing, change feed and basic branching are open; a single-cluster admin UI is open; audit and SSO are split by D221. Refines §00 §8Self-hosters get a complete product and reliability or performance is never withheld (§00 §8); the managed cloud's revenue comes from operating many tenants, which a single organisation does not needApproved (owner, 2026-09-29); reconfirmed 2026-10-02 by D403 (D440)
D2212026-09-29Audit and SSO split (open-core.md). Open source: audit events for every admin, auth and data-access action, emitted as OTel logs to a Loam stream (the usage-hook pattern); an audit query API and CLI with a short default retention set by the operator; plain OIDC SSO, with SAML brokered through Keycloak by self-hosters. loam-platform: the hosted audit UI (search, filters, per-user and per-org timelines); long retention (1 year or more), tamper-evident storage and legal hold; continuous SIEM export (Splunk, Datadog) and compliance report packs; SCIM provisioning, org-wide enforced SSO and cross-org admin. Extends D100: its admin and security events, record fields (never document contents) and _audit stream stay, and D221 adds auth and data-access actions to the open-source scope; D100's M2.x console view and SIEM exports become the hosted audit UI and SIEM export above. Replaces D220's undecided SSO item; §00 §8's "SSO/audit UI" becomes "hosted audit UI, long retention, SIEM export, SCIM and enforced SSO"No SSO tax on SAML or OIDC; the paid features are the operational ones at scale or across tenantsApproved (owner, 2026-09-29); amended 2026-10-01 by D404: SAML is brokered through Authentik, not Keycloak (D447)
D2302026-09-29The showcase apps run on CloudNativePG (§28 §4): CNPG 1.30.1 (Apache-2.0), plain Postgres 17.11 (17.11-standard-trixie), backups and PITR to RustFS through the Barman Cloud CNPG-I plugin 0.15.0 (ObjectStore with endpointURL), not the in-tree barmanObjectStore (deprecated since 1.26, removed in 1.31). Supersedes D149; answers Q45 "no"Plain Postgres now, no fork work on the suite's critical path; PITR to the same bucket storeApproved (owner, 2026-09-29)
D2312026-09-29Loam Postgres is a fork of Neon, dina-kar/neon (GitHub fork of neondatabase/neon, created 2026-09-29, Apache-2.0); Loam owns releases, images and the Postgres patch rebases; nothing is posted upstream. Supersedes D151; answers Q46Upstream is dormant since August 2025 and its compute Postgres is 16.9/17.5 (§28 §10)Approved (owner, 2026-09-29)
D2322026-09-29Loam's control plane is Loam Postgres' primary control plane (§28 §5): it calls the storage controller, pageserver and compute_ctl APIs and serves the compute spec endpoint (GET /compute/api/v2/computes/{id}/spec), the storage controller's hooks (PUT /notify-attach, /notify-safekeepers) and, if the proxy is used, its auth API (get_endpoint_access_control, wake_compute, JWKS); state in operon-meta-tikv (x/, X/, new C/ compute records). Amends D150Neon's control plane was never open source; the fork's components call itApproved (owner, 2026-09-29)
D2332026-09-29Loam's WAL replaces Neon's safekeepers, behind feature loam-wal, per database (x/….wal = safekeepers | loam), and becomes the default only when the pgbench gate passes: p99 commit latency ≤ the stock-safekeeper baseline and throughput not worse, same hardware (§28 §7)Owner: safekeepers are the weaker part; the gate keeps the claim honestApproved direction (owner, 2026-09-29), gated on benchmarks
D2342026-09-29Option A on TiKV (§28 §6.2): a WAL record is acknowledged once durable in TiKV (Raft quorum, leader in the compute's AZ); Loam's log group-commits it to the bucket (250 ms); TiKV's copy is trimmed below min(bucket backup_lsn, pageserver remote_consistent_lsn, commit_lsn); target single-digit-ms acks, p99 < 5 msA pure bucket WAL measured 22 ms p50 / 148 ms p99 in the spike, an order of magnitude over safekeepersApproved (owner, 2026-09-29)
D2352026-09-29Loam speaks Neon's safekeeper protocol (§28 §6.8–§6.9): crate operon-safekeeper (codecs v2/v3, acceptor state machine ported from safekeeper.rs, WalStore trait with a TiKV backend, START_REPLICATION readers with the interpreted sender over the fork's wal_decoder, storage-broker publishing); walproposer and the pageserver unmodified. P4 = P4a protocol crate + TiKV backend, P4b benchmark harness, P4c switch-over merged only when the gate passesNo change to the patched compute or the pageserver; like-for-like benchmark; drop-in upgrade and rollbackApproved (owner, 2026-09-29)
D2362026-09-29PgDog routes Loam Postgres connections, as an unmodified separate service only (§28 §8): AGPL-3.0 (v0.1.60), never linked, vendored or patched; Loam renders pgdog.toml/users.toml and sends RELOAD; database names <db> and <db>__<branch>; transaction pooling by default; SCRAM terminated in PgDog (passthrough would force plain). Fallbacks: Neon's proxy (Apache-2.0, scale-to-zero) and PgBouncer (ISC). Amends D153 for Postgres OLTP (the pg listener keeps analytics)Pooling, replica balancing and sharding without Loam code; AGPL §13 bites only on modification, and D11 forbids linkingApproved (owner, 2026-09-29)
D2372026-09-29TxnKV 1PC + async commit, not RawKV, for the WAL hot tier (§28 §6.4): each append is one optimistic transaction that reads and writes the timeline's head key (term, flush_lsn, …) and puts the chunk keys, so a term bump conflicts write-write in either order (check_fence's rule, no timing assumption); votes and truncation are the same shapeRawKV cannot check the term atomically with the write (fenced raw = put + CAS, two Raft writes); spike on one store: 1PC get+put p50 6.55/7.14 ms, raw CAS+put 29.8/13.8 ms, 2PC 25.8/19.0 msProposed (design within D234)
D2382026-09-29WAL key layout (§28 §6.5): keyspace loam_pgwal; per timeline H/<tl> head then W/<tl>/<begin_lsn BE> chunks of ≤ 128 KiB; a pre-split per timeline so head and tail share a region (1PC; async commit when split); one append in flight per timeline, batching the queue up to 1 MiB; load-based split not relied on (it counts reads only)Keeps appends single-region and ordered; batching mirrors the safekeeper's fsync-when-queue-drainsProposed
D2392026-09-29One logical acceptor per timeline on a stateless, shared WAL pool (§28 §6.3): neon.safekeepers names one per-AZ Service (quorum of one); acceptor state lives in TiKV, so any instance serves any timeline and a crash is a reconnect; no membership changes, pull_timeline, peer recovery, eviction or partial uploads; not a compute sidecarTiKV already replicates across AZs; three acceptors on one store would triple writes; a sidecar dies before the pageserver has read the tailProposed
D2402026-09-29Readers and retention (§28 §6.7): the pageserver, replicas and neon_walreader read over the unchanged START_REPLICATION protocols from tail cache → TiKV → Loam log; 'z' feedback updates the head (coalesced) and is relayed to the compute for backpressure; every instance publishes SafekeeperTimelineInfo to the unmodified storage broker; the bucket copy is an operon-log standard stream with idempotent producer (timeline, term); losing TiKV entirely loses at most one group-commit intervalSame consumer contract as safekeepers; RPO on total hot-tier loss 250 ms versus a segment or 15 minutesProposed
D2412026-09-29Fork maintenance cadence (§28 §10): fork neondatabase/postgres as dina-kar/postgres (the submodule URLs are relative), catch up to 16.15/17.11 (2–3 engineer-weeks), enable PG 18 (3–6 engineer-weeks), then a nightly merge-and-build job and a release-day merge of each quarterly minor with images within a week (2–4 engineer-days a quarter for two majors); support 17 and 18, 16 until its EOLNeon's compute is 6 minors behind on 17 and misses at least 39 CVE fixes; minor releases rarely touch Neon's patched areasProposed
D2602026-09-29TiKV only: no TiDB anywhere in the engine or Loam cloud (§20 status note and §10, §18 §2.4, §21 §3.3). TiKV, through operon-tikv and Loam's tikv-client fork, is the transactional store for Loam's own state in clusters, Loam cloud and self-hosted deployments: the metastore (operon-meta-tikv, --meta tikv://<pd-hosts>/<keyspace> behind the tikv feature, D124, D179), the control plane (ControlStore on the Live system app _control, D125), Loam Live, job state (D214) and durable state (D261). PD and TiKV run unmodified: TiDB Operator v2 with no TiDBGroup on Kubernetes (D179), tiup playground without TiDB in dev and CI. No TiDB process, pool or keyspace, no TiCDC, no TiProxy, no TiFlash, no operon-meta-tidb. Supersedes D123, and the TiDB parts of D116, D126, D127 and D131, and of D58's and D71's backend lists. Unchanged: openraft stays the default metastore for operon dev and standalone (D10), the Postgres and DynamoDB MetaStore backends stay (D58), and Neon for the showcase apps' Postgres OLTP stays proposed (D149). MySQL wire access is an open question (Q260), not decidedOwner, 2026-09-29: self-hosted clusters use TiKV (jobs require TiKV; the durable store has a native TiKV backend). Extends D-SC-16 from the suite to the engine: one store and one client stack, with no MySQL dialect as a second data-access styleApproved (owner, 2026-09-29)
D2612026-09-29Resonate durable state on the native TiKV backend (§21 §3.3): --durable-store tikv://<pd-hosts>/<keyspace> (for example tikv://127.0.0.1:2379/loam_durable) selects a Resonate server plugin, resonate-server-tikv: Resonate's blob server over a TiKV Store, one key per object with a random revision, compare-and-set under a pessimistic lock, behind the cargo feature durable-tikv. It is the store for operon cluster, Loam cloud and self-hosted clusters, and moves from D2 into D1 (D145). main does not have it yet: main's --durable-store accepts sqlite:<path> and mysql://… only, and operon cluster refuses SQLite. mysql:// (feature durable-mysql, Resonate's MySQL plugin, D139) is legacy and deprecated: kept for dev and tests until the TiKV backend merges, then removed. SQLite stays for operon dev and operon standalone. Supersedes D139's TiDB clause and D142's TiDB layout; answers Q39Owner, 2026-09-29; follows D260 and the forks-first Resonate + TiKV + Dapr validation (§22 §13b item 4)Approved (owner, 2026-09-29)
D2622026-09-29The durable cargo feature is opt-in (§21 §3.1): on main, crates/operon/Cargo.toml has default = ["es", "flight", "hnsw", "qdrant"], and durable is not in it (owner ruling O1: toggling it rebuilds about 480 crates). Release builds, CI's durable jobs and the Loam cloud build turn it on; without it a --durable-* flag only logs. Amends D138's "on by default"The docs now match the code on mainApproved (owner ruling O1; docs corrected 2026-09-29)
D2632026-09-30compio is the runtime of the Arm A WAL data path (§28 §7.2): per-core shards (one compio runtime per thread) own timelines by hash; a timeline's START_WAL_PUSH connection, journal writes and durable writes run on its shard; custom OpCodes for RWF_DSYNC writes and registered buffers, the raw io-uring crate only where compio cannot reach; tokio stays for the control plane (TiKV metadata, S3 offload, admin HTTP, the feeder), bridged by bounded channels off the commit path; the tokio WAL service stays as the fallback front end. monoio is not usedThread-per-core with a completion driver; compio is MIT and actively released (0.19.2 on 2026-08-18), monoio, glommio and tokio-uring have not released since 2024Approved (owner, 2026-09-30)
D2642026-09-30Arm A: walproposer is the only sequencer over three loam-wal acceptors (§28 §7.2): Paxos as over safekeepers; commit = a majority durable on local disk (1 RTT + 1 durable write); the term fence is in memory behind a durable vote, votes flush the timeline first. Supersedes D239 for Arm A onlyTakes TiKV's Raft hop, apply and three client RPCs off the commit path (§7.1's causes); losing one acceptor stalls nothingProposed
D2652026-09-30One shared journal per shard (§28 §7.2): preallocated, pre-zeroed, recycled 64 MiB segments (NOCOW on btrfs); CRC32C-framed Append/Truncate/Progress records seeded by segment sequence; 4 KiB-aligned flush units; group commit of every pending record on the shard with --io-depth units in flight; recovery stops at the torn tailOne durable write covers every timeline on the shard (TiKV raft-engine's precedent); no allocation, renames or fsyncs on the commit pathProposed
D2662026-09-30A tiered I/O layer chosen at startup (§28 §7.2): uring (compio io_uring, registered files and buffers, O_DIRECT + RWF_DSYNC/FUA, or write → fsync where FUA is absent; SQPOLL optional; IOPOLL only with poll queues), pwritev2 (O_DIRECT + RWF_DSYNC thread pool, for seccomp-blocked io_uring), buffered (pwrite + back-to-back fdatasync); probed and logged, --io overrides; NVMe passthrough documented as a future tier, no SPDKOwner request, 2026-09-30; Docker's default seccomp profile and GKE Autopilot's RuntimeDefault profile block io_uring (GKE Standard does not apply it automatically)Proposed (owner-requested)
D2672026-09-30Pipelined appends (§28 §7.2): writes issued as AppendRequests arrive, AppendResponse carries the highest durable flush_lsn; per-call-durable stores keep the old behaviour through the trait default. Answers Q115Single-flight appends capped group commit in §7.1Proposed: this is the Q115 direction, not adopted until the P4b results (§7.3) show it helps
D2682026-09-30TiKV holds only acceptor metadata in Arm A (§28 §7.2): per (node, timeline) term, history, membership, server info, start and trimmed LSNs, written by 2PC on create, vote, elected and trim; a local control-file backend for developmentElections are rare; the commit path needs no TiKV writeProposed
D2692026-09-30Offload and trim in Arm A (§28 §7.2): one offloader per timeline by a TiKV lease (B/<tl>), committed WAL to pgwal/<tenant>/<timeline>/<begin>-<end>.lwal (header, zstd body, CRC32C trailer, operon-log conventions), backup_lsn in the lease record; segments recycle once every timeline passed min(backup_lsn, remote_consistent_lsn, commit_lsn)Bounds local disk; RPO on node loss is the other acceptors, on total loss the offload intervalProposed
D2702026-09-30Streams speak CloudEvents 1.0 (§02 §7.4), approved in conversation 2026-09-30. An event is one record laid out as the CloudEvents Kafka binding's binary mode: ce_<attr> headers (content-type for datacontenttype), the key from partitionkey else subject, the timestamp from time, the value from data/data_base64, and a loam_ce_types header for non-string extension types. Ingest: HTTP binary, structured and batched modes on POST …/streams/{stream}/events, gRPC ProduceCloudEvents in the protobuf format, and the Kafka binary-mode layout (ce_ headers), which plain produce stores as sent without deduplication (the M5 gateway later). Consume: GET …/partitions/{p}/events, structured batch or binary; CloudEvents-ingested and ce_-headed records round-trip byte for byte on attributes, other records get a synthesized envelope. Required attributes are validated (id, source, type, specversion = 1.0). Ingest is idempotent by SHA-256(source, id) per stream through a metastore ledger (claim, append, complete; CommitUnknown keeps the claims), for a 1-hour default window (--cloudevents-dedup-window, at most 24 h). A hand-written codec (operon-cloudevents), not cloudevents-sdk. deploy/dapr/edge passes CloudEvents through and relies on this dedup; CI builds and tests itD184 makes CloudEvents the runtime gateway's envelope, so streams should store and serve it natively instead of each adapter normalizing events its own way. The Kafka binding's layout keeps streams valid Kafka topics of CloudEvents for M5 (D74) and keeps the log unaware of events. Idempotent producers (D72) dedupe by producer sequence and do not fit events from many independent sources; source + id is the spec's own identity. The ledger closes #140's duplicate-trigger gap: a timed-out Produce retried by Dapr no longer appends twice, except when a node dies between append and complete. cloudevents-sdk 0.9 (Apache-2.0) normalizes time and typed extensions on write-back, which breaks the round trip, has no protobuf format, and pulls in chrono, url, uuid and hostnameApproved (owner, in conversation)
D2712026-09-30The pageserver feed stays the feeder for Arm A (§28 §7.2): one designated acceptor feeds min(commit_lsn, its durable flush_lsn) to a --no-sync stock safekeeper on a different filesystem; in-process wal_decoder (Q112) deferredwal_decoder needs postgres_ffi bindgen against Postgres headers and Neon's workspace pinsProposed
D2722026-09-30The compio data path as built, and its first numbers (§28 §7.2, PRs #154–#165): compio 0.19 on per-core shards behind the compio feature (loam-wal --runtime compio --shards N; --runtime tokio stays the default and is the fallback tier). The tokio front end reads the startup packet and authenticates, then hands a START_WAL_PUSH socket to the owning shard (hash of the timeline id, shard count fixed in <data>/SHARDS), which owns its own journal; readers stay on tokio through ShardedStore. Durable writes are compio positional writes on O_DIRECT + O_DSYNC descriptors (FUA where the device has it), or write then fdatasync where it has not (--uring-sync), up to --io-depth in flight; SQPOLL is --uring-sqpoll. This differs from D263/D266 in what it leaves out: no custom OpCodes, no registered buffers and no IO_LINK, because the workspace forbids unsafe and compio 0.19's safe API exposes none of them; O_DSYNC descriptors replace per-write RWF_DSYNC, and IOPOLL is not used (the test drive has no poll queues). Custom opcodes and registered buffers come back only if a benchmark shows they matter. First laptop numbers (3 acceptors, 3 runs, p50 / p99 ms; safekeepers, tokio pwritev2, compio, compio + SQPOLL): commit-1 8.48 / 34.02, 9.23 / 39.05, 6.52 / 27.83, 6.03 / 26.83; commit-16 13.86 / 89.82, 8.57 / 57.68, 5.74 / 68.45, 5.63 / 57.84; tpcb-16 32.35 / 304.22, 18.99 / 231.68, 14.70 / 205.74, 12.70 / 118.14. The commit-latency gate holds for both compio tiers, but the gate fails for every tier on bulk (11 to 14 MB/s against 118 MB/s), a cause above the I/O tier that is still open. SQPOLL costs about 18 times the CPU per commit. A shared laptop with a client NVMe and run-to-run p99 noise of 39% to 97%, so these are direction, not the gateD263 chose compio; the first numbers say whether it earns its complexity over the pwritev2 pool: it does on latency, not yet on throughputProposed (the departures from D263/D266 are Loam's; compio itself is the owner's D263). Replacing the safekeepers stays gated on P4b on server hardware, and does not hold today
D2732026-10-01WeSQL (the fork ostrium-labs/wesql) is Loam's candidate MySQL OLTP engine with storage entirely on the bucket, built as milestones W1–W4 (§29 §1, §11), a separate process (D148). Proposes to amend D156: WeSQL leaves "dev compose only" only once W1 and W2 pass their acceptance tests (D156 and Q50 stand until then)Owner-approved direction (2026-10-01): verify transactions, close the gaps in four milestones, keep analytics on Iceberg, do not port StarRocksProposed
D2742026-10-01Claim only what the source verifies about SmartEngine transactions (§29 §4): single-node ACID, READ COMMITTED and REPEATABLE READ (snapshot isolation, first-committer-wins), point locks and no gap or next-key locks, no SERIALIZABLE or READ UNCOMMITTED, ROLLBACK TO SAVEPOINT fails after a write, user XA untested (its tests are disabled). Gaps are backlog T1–T4, not milestonesRead from wesql/storage/smartengine and its MTR results, 2026-10-01Proposed
D2752026-10-01W1: foreign keys enforced in the fork's SQL layer (§29 §5), in ha_write_row/ha_update_row/ha_delete_row, behind wesql_enforce_foreign_keys (default OFF). Current-read checks with a shared lock on the parent's primary-key row; referential actions as nested operations that are written to the row binlog; appliers do not re-cascade. TiDB's FK design is the referenceMySQL 8.0 already prelocks FK tables; point locks suffice for the child-insert vs parent-delete race; InnoDB's unlogged cascades would break the binlog bridgeProposed
D2762026-10-01W2: the binlog, not SmartEngine's redo, is made durable in the Loam WAL quorum before a commit is acknowledged (§29 §6), at the sync stage of ordered commit, one round trip per group commit; the engine WAL stays local. Recovery replays from an applied-offset marker written in each transaction's own SmartEngine write batch, rolls back recovered prepares, and does not use mysql.gtid_executed (persisted only on binlog rotation or shutdown for non-InnoDB engines) (§29 §6.3). Closes Q50's remaining window after fork PR 3The binlog is the commit order, the recovery replay input (fork PR 3) and what replicas and the bridge consumeProposed
D2772026-10-01Protocol boundary and licensing for the WeSQL WAL client (§29 §6.4, §10): a GPL-2.0-only C++ client in the fork speaks the safekeeper v3 message set with a mysql-binlog timeline kind to the Apache-2.0 loam-wal acceptors over TLS. No Loam source is copied into the fork. Keeps D11 and D148Apache-2.0 cannot be combined into a GPL-2.0-only work; a separate process over a socket is the standard boundaryProposed
D2782026-10-01W3: failover rebuilt on the WAL quorum (§29 §7): acceptor terms fence (metadata in TiKV, D268), primary and term in the x/<ns>/<db> record by compare-and-set, a lease that makes a deposed primary stop, a replica that catches up from snapshot, archive and WAL tail, promotion as a Resonate saga. Single writerUpstream removed Raft HA on 2026-08-22; Paxos majorities intersect, so no acknowledged commit is lostProposed
D2792026-10-01W4: analytics through Iceberg; porting StarRocks into WeSQL is rejected (§29 §8): binlog bridge (extends D154) to a changelog stream per table and a keyed Iceberg table in Lakekeeper on RustFS; StarRocks or Loam analytics read it in place as sidecarsStarRocks is another executor over another storage format, Apache-2.0 into a GPL-2.0-only tree, and an unbounded divergence; Iceberg is the contractProposed
D2802026-10-01TiDB on Loam's TiKV is the stated alternative; the rule is: WeSQL only if OLTP storage must live entirely on the bucket (§29 §9). TiDB has distributed transactions, foreign keys and HA, and keeps data on TiKV disks. Choosing it reverses D260, which this document does not do (Q273)Honest comparison requested by the ownerProposed
D2812026-10-01One loams binary for the server, the client CLI and a stdio MCP server (§30 §4). The server commands (dev, standalone, cluster, warm, durable) are unchanged. The client groups live in a new library crate, operon-cli (renamed by D33), with one module per group and no server dependencies, flattened into the binary's clap tree. It is not one crate per group. Built under the working name operon until D33's renameOne artifact to ship and one name on PATH. Keeping the server commands unchanged keeps CI, the crash gate and every plan working. Thirteen per-group crates would add link units on a link-bound machine and share all their state anywayProposed (§30); the binary name loams decided by D401
D2822026-10-01The command tree is loams <group> <verb> in the AWS CLI's shape: init, configure, stack create/describe/list/start/stop/restart/run/logs/delete, storage inspect/prepare, env export, keys create/list/rotate/revoke, login/logout, mcp serve/install/uninstall/tools, pkg add, docs search/snippet, self-update, version, completions. One set of flag names (--name for the stack within stack, --stack elsewhere, --yes, --dry-run, --output) (§30 §5)Models already know the aws <service> <verb> shape; consistent flags let agents chain commandsProposed (§30)
D2832026-10-01The output and exit-code contract (§30 §6): --output table|json|text, default table, set by the flag, LOAM_OUTPUT or the profile and never inferred from a TTY. In json mode stdout holds one JSON document, errors go to stderr as {"error": {code, message, hint, exit_code, details}}, and output schemas are snapshotted (output_schema: 1). Exit codes: 0 ok, 1 internal, 2 usage, 3 action required, 4 not found, 5 conflict, 6 unsupported, 7 unavailable, 8 permission, 9 integrity, 130 interrupted. There are no prompts without a TTY, and destructive commands need --yesAgents need a machine-readable result and a classifiable failure for every command; scripts must get the format they asked forProposed (§30)
D2842026-10-01Local state in LOAM_HOME (default ~/.loam, mode 0700): bin/, config.toml (profiles, no secrets), credentials.toml (0600), stacks/<name>/, variants/, receipt.json, mcp-installs.json. Writes are atomic. No telemetry. The only automatic network call is a daily update check, which can be turned off and never runs in json mode, under mcp serve or with CI=true (§30 §7)One directory to inspect, back up or delete; self-hosters and air-gapped users get no phone-homeProposed (§30)
D2852026-10-01Local stacks (§30 §8): stack.toml (schema 1) holds the resolved spec, and stack run turns it into the arguments of loams dev or loams standalone. An engine registry maps each engine to its cargo feature, server flags, port and .env.loam variables. Engines that were not asked for are switched off with the server's --no-* flags. Every address is loopback (D111). Ports are fixed at create time. A detached supervisor (its own process group, pid file, /ready wait, log rotation). pg is D-PG-1's Postgres wire over collections (LOAM_PG_URL). postgres is OLTP Postgres as a companion (D299), and the only engine that sets DATABASE_URLA prebuilt variant serves any subset of its engines, so stacks never need a build. DATABASE_URL means an OLTP store to every ORM, and the pg listener is analytics (PG1)Proposed (§30)
D2862026-10-01Prebuilt variants instead of builds on the user's machine (§30 §9): standard (es, flight, hnsw, qdrant, mcp, pgwire, durable; the default install), full (adds tikv, durable-tikv, stream-grpc, mysql-wire, plus live and jobs once they merge), and cli (no server, after D297). Targets: Linux x86_64 and aarch64 (glibc ≥ 2.35) and macOS aarch64. failpoints, cluster-tests and durable-mysql are never shipped (D29, D260). stack create downloads the smallest covering variant with consent. --from-source (a cargo install of the needed features) is an opt-in escape hatchA server build takes tens of minutes and GB of RAM (D262: durable alone rebuilds about 480 crates), which "paste one command" cannot include; the chat dump also chose prebuilt variantsProposed (§30)
D2872026-10-01NVMe for the cache without a privileged helper (§30 §10). loams storage inspect is read-only. sudo loams storage prepare --device D --yes refuses mounted, system, member, read-only, small, rotational, virtual or signed devices (nine rules, a typed confirmation), formats ext4 (or XFS), mounts by UUID with noatime,nofail, and hands the mount to $SUDO_UID. stack create --storage nvme:D never formats: it exits 3 and prints the exact sudo step. New server flags --cache-dir, --cache-disk-bytes and --cache-ram-bytes put foyer's H1 disk tier on the mount, beside the hot tier (--hot-dir); there was no H1 disk flag beforeMounting needs root, and a setuid helper is attack surface; disk formatting must be explicit and testable. §04's H1 NVMe tier had no flag, so operon dev cached in RAM onlyProposed (§30)
D2882026-10-01Secrets never pass through MCP (§30 §11). Credentials go to .env.loam (0600, rewritten atomically, appended to .gitignore if it is not ignored). --merge .env updates a marker block in place. MCP tools and JSON output return names and redacted values. Secret has no Serialize or Display, Secret::expose is a disallowed-methods entry, and a canary test covers every tool and command. LOAM_API_KEY holds the whole loam_<key_id>_<secret> token (D65), replacing the chat dump's LOAM_KEY_SECRET; LOAM_KEY_ID stays as a non-secret label. For Claude Code, mcp install --protect-env can add Read(./.env.loam) to permissions.denyAgent transcripts are logged and shared; the token is what SDKs and gateways need; an agent with file access could still read the file, hence the deny ruleProposed (§30)
D2892026-10-01loams mcp serve is a stdio bootstrap MCP server (rmcp, the protocol versions of M1.6 Ruling 7) with a fixed tool list: search_docs, get_sdk_snippet, loam_info, stack_status, stack_create, stack_start, env_export and add_package (CLI2). No tool deletes, stops, formats, mints keys, logs in or updates. stack_create from MCP means a local object store, the embedded metastore, no nvme:, at most 8 stacks, and downloads only with allow_download. env_export paths stay inside the project directory. add_package installs only allow-listed Loam packages. The data tools stay on the stack's HTTP MCP endpoint (M1.6, 127.0.0.1:8083/mcp, D111) and are not proxied (§30 §12.1)Destructive tools are omitted rather than confirmed, because elicitation under MCP 2026-07-28 varies by client. Not proxying keeps the tool list fixed, avoids duplicated schemas and lets the cli variant serve MCP without the engineProposed (§30)
D2902026-10-01loams mcp install --agent claude-code|codex|cursor|windsurf (§30 §12.2) registers two entries: loams (stdio, absolute binary path) and loams-<stack> (HTTP, the stack's data tools). It runs claude mcp add or codex mcp add when they are on PATH, and otherwise edits .mcp.json, ~/.codex/config.toml (toml_edit), ~/.cursor/mcp.json or ~/.codeium/windsurf/mcp_config.json in place. Edits are idempotent and preserve other entries. A backup is made before the first edit, an unparseable file is never written, writes are atomic, installs are tracked in mcp-installs.json for uninstall, and --dry-run is availableFormats verified 2026-10-01 against each agent's docs (Windsurf from third-party guides, to re-verify in CLI1 Task 10). GUI agents may not inherit PATHProposed (§30)
D2912026-10-01Docs and SDK snippets are embedded in the binary (§30 §13): docs/guides/** and a new docs/snippets/<engine>/<language>/<task>, packed by build.rs into a zstd tarball and searched with an in-memory Tantivy index. CI runs every snippet against loams dev, and snippets read endpoints from environment variables only. The MCP server reports the binary's version in serverInfo and loam_infoNo docs MCP server exists (the chat dump assumed one; checked 2026-10-01). Embedding keeps the docs matched to the installed binary and works offline; tested snippets do not rotProposed (§30)
D2922026-10-01The release pipeline is cargo-dist 0.33.0 (MIT OR Apache-2.0, released 2026-09-11, verified 2026-10-01): GitHub Releases on ostrium-labs/loams, SHA-256, GitHub artifact attestations (SLSA provenance), and loams-release.json signed with minisign (two embedded public keys for rotation; the secret key lives in the release environment). Archives are loams-<variant>-<target>.tar.xz. Variants come from thin packages or a Loam matrix job, depending on a spike (Q293). Nothing is published before D33's rename and the repository transfer (§30 §17.1)Buy, not build: cargo-dist gives the builds, checksums, attestations and release plumbing; one signed manifest gives in-band verification without gh or cosignProposed (§30)
D2932026-10-01curl -fsSL https://loams.dev/install.sh | sh (§30 §17.2): a POSIX sh script Loam owns (about 250 lines, shellcheck-clean), published with each release; loams.dev redirects (302) to the GitHub release asset. It detects the platform, picks the variant (default standard), verifies the archive's SHA-256 against the manifest (and the minisign signature when minisign is present; --require-signature makes that mandatory), installs to ~/.loam/bin, writes receipt.json, edits PATH once, and offers loams init through /dev/tty. cargo-dist's own installer is published as a fallbackcargo-dist's installer has no variant choice, no minisign check and no init prompt. The trust model (TLS and GitHub on the first install; signatures afterwards) is stated in the guideProposed (§30)
D2942026-10-01loams self-update (§30 §17.3): verify the manifest signature with the embedded keys, then the archive's SHA-256; smoke-test the new binary (version and features must match); keep bin/loams.prev, swap with self-replace 1.5, and offer --rollback; never restart stacks (stack describe reports binary_outdated). It refuses installs it did not make (exit 6 managed_install; the hint is cargo install loams (D400), since crates.io's loam is an unrelated crate)axoupdater (cargo-dist's updater) supports neither Loam's variants nor the signed manifest; self-replace plus minisign-verify is about 150 linesProposed (§30)
D2952026-10-01Keys, login and agent tokens follow §19, after the unified auth plan (D111, Q30), in CLI3 (§30 §15). keys create makes an environment-scoped API key for apps (90-day default), shown once or written to .env.loam. login uses OAuth 2.1 with PKCE and a loopback redirect. Agents never receive API keys (§19 §5.5): mcp install registers an agent principal, and mcp serve uses vended, 15-minute agent tokens (§19 §5.2, flow 3). Until then keys and login exit 6 (auth_not_available) and .env.loam holds endpoints onlyResolves the chat dump's "creates the first key" against D111 and §19Proposed (§30)
D2962026-10-01loams pkg add (§30 §14, CLI2) maps logical names (sdk, bullmq, celery, durable, live) through release/packages.toml to the registry names of D400 (loams on npm, PyPI and crates.io, the @loams npm scope; D33's loamdb is superseded). It pins the CLI's minor version, detects pnpm, npm, yarn, bun, uv, poetry or cargo from the project's files, runs the package manager as a child process, and refuses non-allow-listed names when called from MCP (add_package)The SDK matches the server the CLI runs; an agent cannot use the tool to install arbitrary packagesProposed (§30)
D2972026-10-01The server split (§30 §4.2, CLI2 Task 1): the engine library moves from crates/operon (renamed loams) into loams-server. crates/loams keeps the binary, with loams-server behind a default feature server and every engine feature forwarded, and re-exports it so tests are unchanged. It lands right after D33's rename PRThe cli variant needs a binary without the engine; doing it in the PR after the rename avoids invalidating in-flight branches twiceProposed (§30)
D2982026-10-01Track CLI (§30 §19). CLI1: operon-cli, the output contract, configure, the H1 cache flags, stacks, storage, env export, embedded docs and snippets, mcp serve and mcp install, init, under the working name, beside M1. CLI2: the server split, variants, cargo-dist, the signed manifest, install.sh, self-update, variant downloads, pkg add and service units, after the rename; publishing also needs the transfer, Q281 and Q282. CLI3: keys, login, agent tokens, companions and cloud stacks, after the unified auth planCLI1 depends only on what is on main; distribution depends on the rename; credentials depend on authProposed (§30)
D2992026-10-01Companion services for local stacks (§30 §16, CLI3): postgres (postgres:17 now; Loam Postgres images once track P publishes them, D231), tikv (the pinned tiup playground of R1, needs full), and rustfs (D61). They run through Docker or Podman, pinned by digest in release/companions.toml, follow their stack's lifecycle, and are never linked or bundledThe chat dump's --engines postgres needs real Postgres; separate processes keep the binary's license and size unaffectedProposed (§30)
D3002026-10-01"Loam Router" is Loam's sharding control plane over unmodified, bought routers (§31 §1, §5 ADR-1/ADR-4): PgDog for Postgres (D236), Vitess for MySQL (D302). Loam builds the shard map, config rendering, cutover and failover orchestration, the in-doubt monitor and the verification program; no SQL parser, planner or router data path in RT0–RT5. Replaces the chat dump's "Rust rewrite" and "clean Postgres router" rowsPgDog and Vitess already parse, plan, scatter, merge, aggregate, pool and reshard; the owner prefers buying; DST, the chat dump's reason for Rust, applies to the code Loam writes (D313)Proposed (Q300)
D3012026-10-01The MySQL shard is WeSQL (§29, D273), not a new "InnoDB semantics on TiKV" engine; "Loam SQL" names the product (Vitess in front of WeSQL shards) (§31 §3)InnoDB semantics over TiKV is TiDB rebuilt, which D260 rules out at TiDB's cost; WeSQL is a real mysqld with a row binlog and GTIDsProposed
D3022026-10-01Vitess v24.x (Apache-2.0) vtgate and vttablet in unmanaged mode front WeSQL primaries, unmodified, as separate services, pinned to v24 (the last release supporting MySQL 8.0); adoption gated by the RT3 compatibility gate (§31 §9, §15). Replaces the chat dump's "implement the tablet contract natively"WeSQL is a real mysqld, so the reason for a native tablet disappears; unmanaged tablets exist for externally managed MySQL (vitess.io/docs/24.0)Proposed
D3032026-10-01Shard keys use each router's native scheme (§31 §5 ADR-5, §6.1): Vitess vindexes (hash = null-key DES into keyspace-ID ranges; xxhash) for MySQL; PgDog's Postgres-compatible hash (hashint8extended, hash_bytes_extended, mod n) and range/list mappings for PostgresVitess tooling transfers only for Vitess keyspaces; PgDog matches Postgres's own hash partitioningProposed
D3042026-10-01The shard map record (§31 §6.1): one postcard record per (namespace, database) in the TiKV metastore (proposed prefix xs/<ns>/<db>), CAS on version, monotonic routing generation; the source of truth for PgDog-routed databases, a mirror of Vitess's topology for Vitess keyspacesOne control plane for both engines without overriding Vitess's own workflowsProposed
D3052026-10-01Multi-instance PgDog cutover is Loam's, with a backend fence (§31 §6.4): copy and catch up on a designated instance (RESHARD), PAUSE everywhere, fence the source shards (ALTER ROLE <app> NOLOGIN and terminate), CUTOVER on the designated instance, publish the next generation and RELOAD the others, RESUME; a Resonate saga. Safety never depends on reaching every instancePgDog's open-source RESHARD cuts over one instance only; coordinated cutover is its closed Enterprise Edition (pgdog/docs/RESHARDING.md)Proposed (Q306, Q307)
D3062026-10-01Cross-shard atomic commit only with a durable coordinator log (§31 §10): PgDog 2PC off by default; on only with PgDog as a StatefulSet (NODE_ID = ordinal, DEPLOYMENT_ID, PGDOG_TWO_PHASE_COMMIT_WAL_DIR on a volume) and after RT2's gate. MySQL cross-shard writes are non-atomic (transaction_mode = multi) until Q303. Never across enginesPgDog's 2PC log is local, unchecksummed, and absent without the variable; Vitess 2PC refuses without semi-sync (dt_executor.go)Proposed
D3072026-10-01Loam Postgres becomes a shard by configuration and tests, not engine work (§31 §8): max_prepared_transactions in the compute spec of 2PC databases; durability of prepared transactions, logical slots and the exported-snapshot boundary proven by RT2's testsThe pageserver stores two-phase state (TWOPHASEDIR_KEY) and Neon tests it (test_twophase.py); slots survive compute replacement (§23 §9.1)Proposed
D3082026-10-01Verification by layer (§31 §11–§14): TLA+ for protocols Loam builds or orchestrates; Lean 4 for pure kernels and as the cross-shard oracle; DST for Loam's control plane; contract, differential and nemesis tests for engines and routers. Bought routers are verified as black boxesProving third-party code is out of reach; their behaviour is testableProposed
D3092026-10-01The compatibility inventory method (§31 §15): static extraction, dynamic capture, replay with classification (same, differs, error, unsupported, pending-target), suite pass rates, blessed TSVs in conformance/router/, re-run on every pin bumpMakes "what the router needs from the engine" a versioned specProposed
D3102026-10-01TLA+ specs in spec/tla/router/, TLC v1.7.4 and Apalache v0.62.3 pinned by SHA-256, CI job tla; variants marked expect = "violation:<Inv>" keep unsafe configurations documented and tested (§31 §11.1)Reproducible model checking; the reason for each rule stays executableProposed
D3112026-10-01Trace validation links specs to code (§31 §11.3): machines emit SpecEvents (tracing target loam::spec); <Spec>Trace.tla checks JSON-lines traces from simulation and real runs; an action-coverage test fails on an action with no emitterSpecs drift from code otherwiseProposed (Q309)
D3122026-10-01Lean 4 kernels in spec/lean/ (Lake package LoamRouter, Lean v4.34.1, Plausible): range partition, shard lookup, k-way merge, LIMIT/OFFSET pushdown, aggregate decomposition; compiled to the loam-router-oracle executable for differential tests and mirrored by proptest; hash functions checked by reference vectors only; plan rewrites out of scope (§31 §12)Proves the semantics the routers must meet; the differential tests connect the model to codeProposed (Q308)
D3132026-10-01Bit-exact DST for sans-I/O control-plane code (§31 §7.1, §13): machines read time and randomness only from Ctx; operon-detsim (a new light crate) drives them with models under one seeded ChaCha8Rng, with trace hashes, shrinking, fault points and swarm runs. Amends D28's scope: D28's seeded simulation stays for the engine; madsim and turmoil are not usedOwned, I/O-free code is deterministic without porting a runtime; madsim (last release 2025-10-11) needs patched tokio cratesProposed
D3142026-10-01Shared checkers (§31 §13.4): the WGL checker from operon-meta-conformance; new bank, unique-key, Elle-style list-append, split and liveness checkers in operon-detsim::checkers, the one implementation §20 §14 item 4 also uses (re-exported by operon-sim); Elle (EPL-2.0) never a dependencyOne checker per property across tracksProposed
D3152026-10-01A Rust nemesis harness (operon-nemesis, test-only) for real-process fault tests with toxiproxy, kill/pause, tc netem and clock skew, running the DST workloads and checkers; Jepsen (EPL-1.0) optional and external; RT5 (§31 §14.3)Same checkers in simulation and reality; no Clojure stack to ownProposed (Q311)
D3162026-10-01Contract suites against model and reality (shard_backend_conformance!, router_fleet_conformance!, shard_map_store_conformance!), the metastore_conformance! pattern (§31 §14.1)Simulation models cannot drift silentlyProposed
D3172026-10-01A Loam-built Rust router is the recorded fallback (§31 §5 ADR-1, §7.3), triggered by: a PgDog change config cannot replace and upstream refuses; Vitess failing the RT3 gate on WeSQL with no fork fix; or an unfixed correctness bug in a bought router. Apache-2.0, no PgDog code, Vitess code only with notices; the chat dump's Frontend/Dialect/Planner/Executor/BackendPool seams specified, not builtKeeps the build option open without paying for itProposed
D3182026-10-01Licenses (§31 §16): Apache-2.0 for everything Loam links (D11); Vitess a service, forkable only by a recorded decision; PgDog an unmodified service read only as a reference, no code, text or tests copied into code or specs; PostgreSQL's hash functions ported under the PostgreSQL License, never from PgDog's copy. Answers the chat dump's "router license" row: Apache-2.0AGPL §13 and D11Proposed
D3192026-10-01Track RT (RT0–RT5) replaces the chat dump's M0–M5 (§31 §17); runs beside M2, R, D and P on the one-build machine and changes no M-track codeAvoids a clash with Loam's M0–M6Proposed
D3202026-10-01vtgate is the MySQL front end for WeSQL, sharded or not, replacing §23 §6.3's Loam-built handshake-and-splice proxy (N6); Loam renders vtgate's static auth file and VSchema. Proposes to amend D153's MySQL half; the splice is the fallback for unsharded WeSQL only, and sharded MySQL waits for D317 if RT3's gate fails (§31 §9.1)Buy over build; one MySQL front end for both casesProposed (Q313)
D3212026-10-01The §18 namespace router is unchanged and separate from the SQL routers; SQL resharding copies rows, the retrieval engine's moves stay metadata-only (§31 §1)Different products, different data movementProposed
D3222026-10-01Isolation is promised per shard, never across shards (§31 §10): Postgres semantics per Loam Postgres shard, D274's per WeSQL shard; cross-shard reads may be fractured; no global snapshot; checkers check exactly thisHonest promises the tests can holdProposed
D3302026-10-01Loam Flow, Loam Event Fabric and Loam House are three layers named by role (§32 §1): Flow (routes, connector registry, route compiler), the Event Fabric (event ingestion on Apache Iggy and Apache Fluss), the House (ClickHouse-dialect query and serving on chDB, plus Loam's existing vector and graph surfaces). The draft's "Trigger Engine / Operon Cells" are Loam Functions (§24) and durable functions (§21); its "Loam Gateway" is the House coordinator, so it does not clash with the runtime gateway (D184)The owner's draft "Loam: Final Plan" (2026-10-01); reuse of existing primitives instead of new enginesProposed (owner direction)
D3312026-10-01Two logs, chosen by workload (§32 §5.1): the Loam WAL and streams stay the log for databases and retrieval (Postgres §28, MySQL §29, collections, graphs, Loam-written tables, OTLP logs, trigger-rate CloudEvents, D270); event ingestion goes to the Event Fabric (Iggy + Fluss). Every record has one home log; bridges (D336) copy across, never dual writesOwner, 2026-10-01: "our loam wal is optimized for postgres, mysql, text search, vector search, graph workloads, iggy and fluss for event ingestion so separate is good, and we see more contributions to iggy and fluss"Proposed (owner direction)
D3322026-10-01Apache Iggy is the Fabric's ingest log and protocol edge (§32 §5.2, option A): QUIC/TCP/HTTP/WebSocket transports, SDKs, consumer groups, the Rust connectors runtime and the MCP server, run unmodified as a separate service pinned to server 0.9.0 (SDK 0.11.0), three replicas with persisted durability and bounded retention (72 h proposed, Q346); Loam contributes S3 tiered storage (apache/iggy discussion #3312), the fluss_sink and loam_sink connectors and Kafka-gateway work upstream. Kafka clients of the Fabric use Iggy's Kafka gateway; this proposes to defer D74's Loam-side gateway (Q331). Rejected: Iggy only as an edge drained into the Loam WAL (double write), Loam speaking Iggy's pre-1.0 protocol, Iggy replacing the Loam WALIggy 0.9.0 (2026-09-18, read 2026-10-01; an ASF top-level project since 2026-08-19): Rust, thread-per-core on io_uring/compio, VSR clustering (pre-production), six SDKs, a connectors runtime with 16 sinks and 5 sources, camel-iggy in Camel 4.17+; gaps: local disk only (tiered storage not implemented), pre-1.0 wire, Kafka gateway unreleased (no offset commit), connector SDK unpublishedProposed (owner direction); Iggy as the ingest layer confirmed by D405 (Q330); Q331 open
D3332026-10-01Apache Fluss is the Fabric's real-time table layer (§32 §5.3): Log tables and PK tables (LastRow, FirstRow, Versioned, Aggregation), lookups, remote storage on RustFS, tiering to Iceberg through Lakekeeper; unmodified 1.0.0 with ZooKeeper; v1 accepts Fluss's Flink tiering job, run under the Flink Kubernetes Operator that §26 plans (D212); a Rust tiering service on fluss-rs and iceberg-rust is a contribution candidate (Q332)Fluss graduated to an ASF top-level project on 2026-07-16 and released 1.0.0 on 2026-09-22 with a Rust Gateway (preview) and fluss-rs 1.0.0; its merge engines are what ClickHouse's merge-tree engines need; the tiering service is still a Flink job in 1.0Proposed (owner direction; Q332)
D3342026-10-01One envelope: CloudEvents 1.0 laid out as in D270 (§32 §5.4): on Iggy, ce_<attr> headers plus content-type, the data as payload, and the Iggy message id = the first 16 bytes of D270's dedup key; on Fluss, attribute columns _ce_*; the tenant comes from the credential. loam-fabric ingest deduplicates by source + id with D270's claim/append/complete protocol and a ledger in the Fluss PK table _fabric.ce_dedupOne identity per event across Loam streams, Iggy and Fluss; the draft's envelope fields map onto CloudEvents attributesProposed
D3352026-10-01Fabric schemas and dead letters (§32 §5.6): a Fluss table's schema is authoritative; Iggy topics name schemas through dataschema, with documents under ns/<id>/fabric/schemas/ and versions in _fabric.flow_objects; a Confluent Schema Registry subset waits for Kafka clients (Q344); each route has a DLQ topic <topic>.dlq with ce_loamerrorThe draft puts the schema registry and DLQ in the Fabric; no second registry service in v1Proposed
D3362026-10-01Bridges, not dual writes (§32 §5.7): Iggy → Fluss through the Iggy plugin fluss_sink; Iggy → Loam through loam_sink (ProduceCloudEvents, D270, or collection writes keyed by _ce_source/_ce_id); Fluss → Iceberg by tiering; Loam → Fabric through a link target iggy (FL3); exactly-once where the target dedupes, at-least-once elsewhereKeeps each log's guarantees; consistency tokens (D76) do not cross the bridge, and the docs say soProposed
D3372026-10-01A Flow route is a declarative spec, loam.flow.v1.Route (§32 §6.1): from, steps, to, delivery; served by connect-rust (D128); routes, connector instances and schema versions are rows of the Fabric system table _fabric.flow_objects (Fluss PK); task leases come from the Loam metastore's network API in FL3loam-fabric stays free of engine crates; routes validated against connector capabilities (D353)Proposed
D3382026-10-01The route compiler emits only existing primitives (§32 §6.2): links (§09) between Loam objects; Fluss tables and fluss_sink; House materialized views as per-insert-block transforms; Loam Functions and durable functions for per-event reactions; Iggy plugins and Camel routes for external sinksFlow adds no execution engine; ClickHouse MVs are per-insert-block transforms, which is link semanticsProposed
D3392026-10-01Flow and the House are open source in a separate Cargo workspace fabric/ (binary loam-fabric, roles ingest, house, flow, later coordinator); the Camel runner loam-connect is open source too; the managed connector fleet, per-plan connector limits, hosted secrets UX and connector billing are loam-platform (D220)A single organisation self-hosting Loam needs connectors and the House (open-core.md's rule)Proposed
D3402026-10-01The Flow designer is a console page (§19) editing loam.flow.v1.Route, possibly embedding Apache Camel Karavan's designer (Apache-2.0, 4.18.1) for Camel-backed steps (Q340); not in FL1–FL2The draft's "Karavan-like UI" without a second UI stackProposed
D3412026-10-01No Spark or Flink targets in Flow (§32 §6.2): stateful stream SQL (windowed joins with watermarks) goes to the RisingWave companion (D22) reading Fluss or Iggy; §26 keeps hosting users' Spark (Sail) and Flink jobs (D211, D212); Fluss's tiering job is infrastructureThe draft's goal 1 ("without managed Spark or Flink as the default path"); §00 §7 and §09 stay out of stateful stream processingProposed
D3422026-10-01The House SQL engine is chDB (chdb-core 26.9.x, libchdb, Apache-2.0) through a Loam-owned FFI (loam-chdb-sys, loam-chdb) over the pinned chdb.h, chdb-rust 2.0 as the reference (Q337). Rejected: unmodified clickhouse-server as the engine (cannot merge Fluss's tail or honour consistency tokens; its own disks), a ClickHouse dialect on DataFusion, DuckDBClickHouse semantics without running ClickHouse storage; chdb-rust 2.0.0 (2026-09-20) still says experimentalProposed
D3432026-10-01The House runs in loam-fabric, never in the engine binary; libchdb (~180 MB compressed, process-global) is linked dynamically, fetched by SHA-256; fluss-rs (arrow 59) lives only in fabric/. Amends D51: the engine binary stays free of other engines; a Loam-owned service binary may link an engine libraryArrow lockstep (Lance 12 pins arrow 58 / DataFusion 54) and binary sizeProposed
D3442026-10-01ClickHouse engines map onto Fluss and Iceberg (§32 §7.4): MergeTree → Fluss Log table; ReplacingMergeTree(ver[, is_deleted]) → PK table, Versioned or LastRow, deletes for is_deleted = 1; SummingMergeTree → PK table, Aggregation sum; Replicated* and ON CLUSTER accepted and ignored; ORDER BY/PRIMARY KEY → PK or sort order; PARTITION BY limited to toYYYYMM, toYYYYMMDD, toDate, toStartOfHour and identity; the DDL and ClickHouse types kept in Fluss table properties loam.ch.*Fluss's merge engines are the draft's Tier 1 mapping; no House catalog of its ownProposed
D3452026-10-01House reads are a union of the pinned Iceberg snapshot and the Fluss tail (§32 §7.5), merged by key for PK tables; FINAL accepted as a no-op; INSERT returns X-Loam-Consistency-Token (Fluss bucket offsets, loamfx1. prefix) and a read with it, or in the same session, sees the writeFreshness without waiting for tiering; read-your-writes on the ClickHouse surfaceProposed
D3462026-10-01House writes and DDL go to Fluss (§32 §7.6): bodies in any declared format parsed by chDB into Arrow and written with fluss-rs; INSERT … SELECT runs in chDB; chDB never holds MergeTree storageStateless workers; one write path for the FabricProposed
D3472026-10-01A declared ClickHouse surface, chsurface-1, on the House (§32 §8): the HTTP interface on 127.0.0.1:8123 first (D111), the native protocol in FL5, and a versioned list of statements, engines, formats, settings, functions, table functions (deny by default: file, url, remote, s3, mysql, postgresql, executable, queue engines…), system tables and error codes; "100 % compatible" = 100 % of the declared surface's tests pass, deviations allowlisted. Reverses D45 for the House and amends D42, §00 §7 and §08; Loam tables stay Iceberg-onlyThe draft's compatibility goal, made testable; OpenPanel and ClickHouse-native apps become runnable unchanged (Q334)Proposed (Q333)
D3482026-10-01The allowlist and the release gate (§32 §8.10): conformance/clickhouse/allowlist.toml entries with id, tests, category (merge-timing, semantic, error-text, type-display, unsupported, performance), reason, owner, surface version and, for semantic, the owner's approval; strict xfail (an allowlisted test that passes fails CI); a regression on the declared surface blocks a release; a per-build report of pass ratesThe draft's §7.4 gate, with ownership and expiry so the list cannot grow silentlyProposed
D3492026-10-01ClickHouse compatibility is proven differentially (§32 §10) against a CI-only reference clickhouse-server pinned to the ClickHouse version chdb-core carries: the declared-surface corpus (~400 cases at FL2), the upstream stateless tests inside the surface (compared with their .reference files), engine-semantics cases (the reference run after OPTIMIZE … FINAL), ClickBench and TPC-H correctness, the Rust and Python drivers (Go and Java in FL3), byte equality of formats, S3 faults through toxiproxy, worker and Fluss kills, and a reference cluster for Distributed in FL5D13's rule (compatibility defined by suites) applied to ClickHouseProposed
D3502026-10-01Graph stays native (§07, D44); Grafeo is not the House graph engine (§32 §7.7): Grafeo 0.5.43 (Apache-2.0, released 2026-09-27, four days before this review) comes from a repository created on 2026-01-26 (eight months old), has one maintainer in practice, depends on arrow 60, stores locally without S3 or distribution and brings its own HNSW; grafeo-server (Bolt v5, GQL over gRPC) may be a Flow sink companion if Cypher/GQL is demanded (Q339). One vector source of truth per collection stays Lance + qdrant-edge (§06)Verified 2026-10-01 (release 2026-09-27; repository created 2026-01-26; 90 of the last 100 commits by one author; grafeo-server v0.5.40 of 2026-04-20 has Bolt)Proposed
D3512026-10-01Track FL (§32 §11): FL1 Fabric foundation; FL2 House SQL phase 1 (chDB, HTTP 8123, MergeTree/Replacing/Summing, the differential harness); FL3 Flow routes, MVs and sinks; FL4 graph alignment with M3; FL5 distribution and the native protocol; FL6 hard engines; CN1 the ★ connectors (§33). Interleaved with M, R, D and J on the one-build machine (D127), built in fabric/ so it never rebuilds the engine graphShortest path to a working Fabric and a testable ClickHouse surface; FL1 needs only M1.2 and D270Proposed
D3522026-10-01A connector registry in loam-flow (§33 §1): one manifest per connector (connectors/registry/<id>.yaml, loam.flow.v1.ConnectorSpec), generated for the 200-row catalog from connectors/registry/catalog.csv and hand-written for the ★ set; instances are namespace objects referenced by routes; FlowService lists, describes and validatesThe précis' "Camel + Kestra as one catalog; count unique connectors"Proposed
D3532026-10-01Every connector declares its capabilities (§33 §4): source/sink, streaming, batch, CDC, webhook, resumable position, transactional, upsert/delete, idempotent sink, delivery, ordering, formats, schema, bulk Arrow, backpressure, auth, config schema, secret fields, limits, conformance suites; a route using an undeclared capability is refused; each declared capability of a ★ connector has a contract testThe précis' change 1: Camel producer/consumer and Kestra task/trigger support differ per connectorProposed
D3542026-10-01Connector runtimes, buy first (§33 §5): native Rust in loam-flow for the hot path; Iggy's connectors runtime where Iggy has the plugin; Apache Camel 4.22 unmodified in the JVM service loam-connect (routes ending in camel-iggy) for the long tail and JDBC; Debezium Server 3.7 unmodified for CDC of external databases. Kestra is a companion, not a runtime (Resonate orchestrates, D210); Kafka Connect only through Iggy's Kafka gatewayCamel 4.22.1 has 319 component modules including camel-iggy; Debezium Server has HTTP, Kafka, Fluss and other sinks; Kestra's scheduler and state duplicate Resonate and need their own databaseProposed
D3552026-10-01One envelope and delivery contract for connectors (§33 §6): sources emit CloudEvents with type io.loams.dev.flow.<connector>.<event>.v1, source /connectors/<instance>/<resource> and an id derived from the source position so re-reads are recognisable; sinks deliver at least once with ce_id as the idempotency key; positions commit after the Fabric (or the sink) acknowledgesThe précis' change 3, on D334's envelopeProposed
D3562026-10-01Bulk moves as Arrow (§33 §5–§6, §8): batches of 65 536 rows partitioned by key and written by parallel tasks; ADBC (adbc_driver_manager 0.24 with the Apache-2.0 Snowflake, BigQuery, Postgres, SQLite, DuckDB and Flight SQL drivers) for warehouse reads and loads; native drivers for OLTP upserts; JDBC only in loam-connect; Avro only at Kafka and schema-registry boundariesThe précis' control-plane/data-plane split and ADBC/Arrow guidanceProposed
D3572026-10-01CDC is observed with Debezium, not reinvented (§33 §7): external Postgres, MySQL, MariaDB, SQL Server, Oracle, Db2 and MongoDB through Debezium Server's HTTP sink posting CloudEvents to loam-fabric ingest (Iggy's postgres_source as the lighter Postgres option, Q351); Loam Postgres and WeSQL through Loam's own bridges (D154); current state in Fluss PK tables Versioned on the source LSN, history in Log tables, SCD2 as a House view; no triggers; slot-lag monitoringThe précis' CDC guidance ("20M users is no reason to write your own"); Debezium is Apache-2.0Proposed
D3582026-10-01The ★ set is 21 Loam-owned hot-path connectors (§33 §8): Kafka, Kinesis, Webhooks, PostgreSQL, MySQL, Redis, Elasticsearch, Snowflake, BigQuery, ClickHouse, Iceberg, S3, Parquet, Avro, Arrow IPC/Flight, Debezium-Postgres, Debezium-MySQL, OpenTelemetry, HTTP/REST, JDBC, ADBC, each with a fixed runtime; CN1 ships them in the précis' rollout order; CN2 the P2 rows through stock Camel and Iggy plugins; CN3 the long tail through OpenAPI-generated connectorsThe précis' starred list and rolloutProposed
D3592026-10-01A licence gate per connector (§33 D359): every manifest names its runtime's and drivers' SPDX licences; CI refuses AGPL, BSL, SSPL, ELv2, NOASSERTION and unknown licences for anything shipped or run by default; Airbyte (ELv2 platform) and Redpanda Connect (Redpanda Community License on part of its connectors) are not runtimes; JDBC drivers are mounted by users, not shippedD11 applied to connectorsProposed
D3602026-10-01The standards charter (§34 §1): every external standard Loam speaks is pinned to a spec version behind one crate (HTTP/2 and mTLS, Connect/gRPC/gRPC-Web, CloudEvents 1.0.2, Arrow/Parquet/Iceberg, protobuf with Avro only at schema-registry sinks, OTel and W3C Trace Context, the Resonate protocol, loam.meter.v1). Longevity is replaceability: one crate per standard, a decision-log row per standards choice, deprecation as announce → dual-run (at least one minor release) → sunset, open formats only in stored data and no cloud-specific ids, key and algorithm ids on every signature, cargo deny/cargo vet/SBOMs/reproducible releases/MSRV, restore drills, a dependency review every 3–5 yearsThe owner's draft (§1, §11): "built for 50 years" means the wire contracts, data formats and semantics outlive every implementationProposed
D3612026-10-01Transports (§34 §1): HTTP/2 with mTLS (ALPN h2) between Loam components, h2c only on loopback; HTTP/3 only at the Envoy edge (D176, D184) until an internal-mesh flag (Q372); external HTTP/1.1 traffic is re-originated as h2 by EnvoyDapr and Rust's h3 are behind h2; edge providers already terminate h3. Consistent with D176, whose HTTP/3 is at EnvoyProposed
D3622026-10-01Connect-RPC through connect-rust for every new service (§34 §1): one handler serves Connect, gRPC and gRPC-Web. Checked 2026-10-01: connectrpc 0.9.1 (2026-09-21, Apache-2.0) is the official Connect project's Rust implementation (Connect RFC 007, now connectrpc/connect-rust) and passes the full Connect conformance suite (3,600 server tests). The draft's tonic + hand-written-codec fallback is not needed. Extends D128 and D206The draft asked to verify maturity; the workspace already uses it for loam.live.v1Proposed
D3632026-10-01The narrow waist is loam.stream.v1.StreamService (Produce and D270's ProduceCloudEvents, PR #171): Dapr pub/sub, HTTP push, runner hosts and runtime hooks are adapters into it (commercial adapters in loam-platform use the same service). buf breaking (FILE) is enforced in CI on proto/loam/{stream,events,meter} and crates/operon-stream-grpc/proto; loam.live.v1 joins when R2 freezes it; evolution is compatible-only (§34 §1)One owned internal contract; buf.yaml already names FILE, unenforced since R1Proposed
D3642026-10-01The Loam CloudEvents profile (§34 §2): CloudEvents 1.0 plus required tenantid (<org>/<namespace>, stamped by the gateway or runner host from the credential; a client value is overwritten and counted) and traceparent (W3C, generated at the first hop if missing; tracestate optional). id is the idempotency key (D270's source + id), so no idempotencykey extension; versions are the type suffix .v<major> plus dataschema = urn:loam:proto:<message full name>, so no schemaversion extension; types are io.loams.dev.<domain>.<name>.v<major> (owner ruling 2026-10-01, answering Q361)Two idempotency keys would disagree; CloudEvents already has dataschema; §02 §7.4's dev.loam.stream.record moves to the owner's prefix with the rename PRProposed; the prefix decided by D402
D3652026-10-01High-rate events bypass the dedup ledger (§34 §3): usage events and other per-request events are batched plain produce in D270's Kafka binary-mode record layout (ce_ headers; flush every 50 ms or 1 MiB; bounded queue of 65 536 events or 64 MiB that drops the oldest and counts loam_events_dropped_total); deduplication by (source, id) happens in the keyed event table, not at ingestD270's ledger costs two metastore proposals per request; §02 §7.4 already says bulk ingest should use plain produceProposed
D366–D371, D373, D377, D3792026-10-02Moved to loam-platform (private), 2026-10-02: §34's protocol gateway (D403, D440); not OSS decisions—Moved
D3722026-10-01Events → Arrow (§34 §4): one Arrow schema per event type, derived from the data message's protobuf descriptor (buffa-descriptor); attribute columns, time as Timestamp(Nanosecond, "UTC"), ext as a string map, data as a typed struct, data_raw as the original bytes (on by default); proto path and tag in field metadata; Iceberg field ids assigned by the catalog and matched by proto path; additive-only evolution checked in CI; event tables keyed on (source, id), partitioned by day(time), sorted by (type, time); ns timestamps need Iceberg v3, else a time_ns column (Q368)No official CloudEvents Arrow format exists (v1.0.2 defines JSON, protobuf and Avro; drafts add Avro compact, CBOR and XML)Proposed
D3742026-10-01State rules (§34 §1): the tenant scope on every stored key in Loam's existing forms (ns/<ns>/ in the bucket, keyspace plus prefix in TiKV), not a literal tenant/{id}/; money as int64 micros plus ISO 4217 currency; int64 ns UTC in new protobuf contracts and Arrow (existing versioned contracts keep their units); TiKV only through transactions or atomic CASThe draft's rules, mapped onto Loam's layoutProposed
D3752026-10-01One Runner trait (§24 §16; amends §24 §3): kind, capabilities, deploy (idempotent by digest), invoke (returns the response and, except for the supervisor, a Usage), undeploy, health. SupervisorRunner (Loam's nodes, §24's tiers, built with F1) is the default and the only runner with D170's placement advantage; ProcessRunner (dev and tests) and LambdaRunner (Rust, arm64, provided.al2023, Loam's bootstrap) in RN1; KnativeRunner in MT2 (§38 D441); runners outside this repository as RunnerKind::External (the commercial Cloudflare runner, loam-platform); Cloud Run and Container Apps on demand (Q367). External runners run no Dapr; D183's one shared daprd is unchangedThe draft's four co-equal targets would give up placement next to the data, which is the product's advantage (§24 §9)Proposed
D3762026-10-01Usage from every runner reaches §27's contract (§27 §3.6): exactly one reporter per invocation (the supervisor for its tiers, RunnerHost for external runners); Lambda CPU from the bootstrap's getrusage(RUSAGE_SELF) delta in a bootstrap-owned x-loam-usage header, capped by the REPORT line's billed duration × the fractional CPU share memory_mb / 1 769 (Q366, which gates RN1 Task 5); the Tail Worker join for the commercial Workers runner is specified in §27 §3.6 and built in loam-platform; loam.meter.v1.Invocation gains runner (12), region (13), provider_billed_ms (14), compile_usec (15), overhead_usec (16); the socket gets a framing envelope (HostMessage/ConsumerMessage, RN1 Ruling 1); the CloudEvent form (dev.loam.meter.usage.v1 in §27 today, io.loams.dev. after the rename, D402) is specified but built in loam-platform (§38 D440). Platform CPU (supervisor, loam-dapr, daprd, Envoy) never lands in tenant cgroups; compile CPU is reported apart. Rating, the ledger and reconciliation against provider invoices stay loam-platform (D190, D202). Refines D201The draft's metering section, reconciled with D190/D202 and the owner's 2026-10-02 ruling: hooks only in this repository, one contract for every runnerProposed
D3782026-10-01Languages and embedding (§34 §1): Rust for the data plane; Java (Quarkus) and Go through buf-generated Connect clients; in-process embedding only after profiling (Java FFM, final since JDK 22, with jextract over a cbindgen header; Go cgo, whose baseline overhead Go 1.26 cut by ~30%; the Arrow C Data Interface; Wasm under Chicory or wazero). No core path depends on Go or Java SIMD: Go 1.27 (August 2026) adds a portable simd package and arm64/Wasm simd/archsimd, both still behind GOEXPERIMENT=simd; the Java Vector API is in its eleventh incubator in JDK 26 (JEP 529), waiting on ValhallaChecked 2026-10-01 against the Go 1.26 and 1.27 release notes and JEP 529; JIT and GC CPU work against CPU-billed tenantsProposed
D380, D383–D3872026-10-02Moved to loam-platform (private), 2026-10-02: the former §35 Cloudflare target (D403, D440); D381 and D382 stay, in §36 §17—Moved
D3812026-10-01No Loam-built S3 service (§36 §17): the draft's "Rust S3-compatible service" is RustFS where Loam runs its own storage (D61, D178) and R2 on Cloudflare; operon-store stays the client; tenants are isolated by vended, prefix-scoped credentials (ObjectStoreProvider::issue_credentials, §25 §5), on R2 by temporary credentials (POST /accounts/{id}/r2/temp-access-credentials with prefixes and object-read-only/object-read-write, ttlSeconds ≤ 604800). D178's provider list gains r2. Extends D178R2 already speaks the S3 API; a facade would duplicate SigV4, multipart and conditional writes for no gain; buy over buildProposed
D3822026-10-01The Fs trait (§36 §17, crate operon-fs): read, read_range, write (Overwrite/CreateNew), append, rename, remove, stat, list, caps; ?Send on wasm32. Backends NativeFs, MemFs, DoSqliteFs (Durable Object SQLite, 1 MiB chunks, one transaction per mutation) and R2Fs (immutable blobs), with one conformance suite. Fs holds warm, rebuildable state and scratch only: Durable Object storage is a cache and a lease, never the source of truthD1; code that runs in Workers and natively needs one file abstraction; thin mode has no std::fsProposed
D3882026-10-01Loam Git extends §15 §3 (§36 §1, §4): a repository's refs move from one CAS'd refs document to a per-repository WAL of create-only segments plus checkpoints under ns/<ns>/repos/<repo_id>/; packs stay immutable and content-addressed; forks stay O(1) (checkpoint 0 names the parent and seq). Repos are a service with their own bucket WAL, not a sixth object kind and not a Loam stream (proposed answer to Q15). Amends §15 §3.1–§3.2. Track GT (GT1–GT5) carries W1's repository scope and parts of W2A single document rewritten per push costs O(refs) and contends on one key; agent fleets create thousands of branchesProposed (amends an approved doc; needs the owner)
D3892026-10-01The commit point is the create-only PUT of the next segment, wal/<seq:020>.lgw with If-None-Match: * (§36 §4.2); no CAS'd head pointer; head is a hint written ≤ 1/s; readers stop at the first missing number; a linearizable read is one GET past the reader's state. Holds on S3, R2, GCS (ifGenerationMatch=0), Azure and RustFS (atomic ≤ 1 MiB, so segments are capped at 1 MiB)R2 allows one write per second per key (R2 limits, 2026-06-08), which would cap a CAS'd head at one commit a second; one fewer PUT per batch; Delta Lake's log uses the same protocolProposed
D3902026-10-01Git formats (§36 §4.3): a WAL record is a CloudEvent 1.0 in protobuf (D270) with tenantid, idempotencykey, traceparent, schemaversion, loamseq, types io.loams.dev.git.{reftxn,packset,config}.v1, data loam.git.v1.{RefTransaction,PackSetChange,ConfigChange}; segments frame one CloudEventBatch (LGITWAL\0, version 1, seq, batch id, CRC32C, LGITWEND), ≤ 1 MiB; checkpoints (LGITCKPT) hold refs, symrefs, protections, the pack set, the fork parent and the idempotency window; a push is one object packs/<checksum>.lpk (pack ‖ idx ‖ 32-byte footer). Amends §15 §3.1's packs/<ulid>.pack + .idxThe draft's CloudEvents envelope; §03 §6's magic-and-version rule; one data PUT per push and idempotent retriesProposed; not yet reconciled with D364, which drops the idempotencykey and schemaversion extensions in favour of id and the type suffix plus dataschema (to settle before GT1 Task 1)
D3912026-10-01One sequencer per repository with group commit (§36 §4.4): the rendezvous owner of (ns, repo, repo_id) (D75) under the lease task/git-seq/<ns>/<repo_id> elsewhere; one segment PUT in flight, arrivals form the next group (≤ 64 transactions, ≤ 1 MiB); validation of object ids, protections and scopes in memory; pack upload, fast-forward and connectivity checks before queueing. Correctness never depends on one sequencer: the segment name fences every writer, including direct-to-bucket helpers. Target 30 pushes/s per hot repository (draft)Continuity's batching and primary-only writes, without gossip or replicasProposed
D3922026-10-01Four traits in operon-git (§36 §5): WalStore (fenced, monotonic, contiguous; idempotent on the batch id), RefLog (commit atomic and linearizable, cas_ref, snapshot(Latest/AtLeast/Exactly), watch; idempotent per key within a 1 h window, IdempotencyMismatch on a reused key with another digest), BlobStore (create-only content-named put, put_stream, get_range, head; deletion only by GC) and Materializer (touch(scope, agent), read; cone-mode scopes; coalesced range reads, never a GET per blob)The draft §14.4, with idempotency split between the layer that can check it: batches in the store, keys in the ref logProposed
D3932026-10-01Git, build-cache and mirror usage through §27's hooks only (§36 §11): metric families loam_git_*, loam_buildcache_*, loam_packages_* per org and namespace; a CloudEvent per committed transaction mirrored to the namespace stream _git (at least once, repair sweep). No meter events or ledger in this repository. Amends the draft §14.7D190, D200–D202Proposed
D3942026-10-01Smart HTTP (§36 §6.1): upload-pack speaks protocol v2 (ls-refs, fetch with filter/shallow/wait-for-done, object-info); v0/v1 upload-pack only if a W1 matrix client lacks v2 (Q387); receive-pack speaks v0/v1 (v2 has no push) with report-status, report-status-v2, atomic, delete-refs, side-band-64k, ofs-delta, push-options, quiet. The server loop is Loam's on gitoxide primitives (gix-pack builds for wasm32-unknown-unknown); stock git (GPL-2.0) runs only as an unmodified separate process: as the test oracle, the repack worker, and the git pack-objects/index-pack steps of git-remote-loam's pushes on the client (D395). Answers §15 Q1gitoxide's server-side upload-pack/receive-pack plumbing, delta compression and bitmap writing are unchecked in its crate-status.md (read 2026-10-01); the serving path must run in WorkersProposed
D3952026-10-01git-remote-loam (§36 §6.2): GT1 is a serverless helper over the bucket (fetch, push, option; loam::<store-url>/ns/<ns>/repos/<repo_id> addresses; whole-pack fetch with stored indexes; git pack-objects/index-pack for pushes; each push process is its own fenced sequencer). GT2 adds stateless-connect (v2 to an in-process upload-pack or a server) and loam://<host>/<ns>/<repo> addresses, for partial clone and lazy fetchPhase 1 of the draft needs no server; stateless-connect is how v2 features reach a helperProposed
D3962026-10-01Compaction off the push path (§36 §7): checkpoints every 256 segments or 8 MiB of replay; git repack --geometric=2 -d --write-midx --write-bitmap-index and git commit-graph write on a worker's NVMe mirror under the lease task/git-compact/<ns>/<repo_id>, committed as a PackSetChange through the sequencer; segment GC after 24 h Exactly retention; pack GC by fork-family reachability after the §03 §7 grace with GC claims. On Cloudflare a Durable Object alarm schedules it in a ContainerContinuity reports compaction as its bottleneck; gitoxide cannot yet write deltas or bitmapsProposed
D3972026-10-01loam-vfs is §15 §5.1's /workspace lower layer (§36 §6.4, GT4): cone-mode scopes in the vended token; fetch on first read; server-side write admission refuses commits outside the scope (OutOfScope); a commit is one WAL record whatever folders it touches; no per-scope WAL partitions until GT4 measures the sequencer as the bottleneck (Q391)§15 principle 2 (one writer per workspace); per-scope partitions would break atomic cross-folder commits to fix contention group commit already removesProposed
D3982026-10-01Build cache: sccache (Apache-2.0, v0.18.0) (§36 §8): a direct path (sccache's S3 backend on R2, RustFS or S3 with vended prefix credentials; no Loam code) and a gateway path (sccache's WebDAV backend against operon-buildcache: hit/miss/put metering, refresh-on-hit approximate LRU, TTL and quota sweeper, server-enforced trust). Keys ns/<ns>/cache/sccache/<repo>/<class>/; trusted classes (protected-branch CI) write trusted/; untrusted (forks, PRs, agent sandboxes) read only, with an optional private scratch/<principal>/. BuildCache (zlib) only on demand; sccache-action only on GitHub-hosted runners. Amends §15 §7Cache poisoning across trust boundaries; R2 temporary credentials scope by prefix and read-only permission; sccache has multi-level caches and read-only modesProposed
D3992026-10-01Crates mirror (§36 §9, operon-registry): a crates.io sparse-index read-through in the gateway role; config.json with dl using {sha256-checksum} and auth-required: true; index files cached with ETag revalidation (60 s TTL) and filtered per namespace at serve time; .crate files verified by cksum and stored once in ns/_public/packages/crates/sha256/; §15 §6 policy (allow/deny, pins, quarantine) and an audit CloudEvent per download (io.loams.dev.packages.download.v1) to _packages. The index lives in the object store, not a Durable Object or D1D1; three routes are smaller than adapting kellnr, panamax or ktra, none of which runs in Workers or stores into a Loam namespaceProposed
D4002026-10-01The package namespace is loams on crates.io, PyPI and npm (scope @loams), superseding D33's loamdb (M1 overview A32). Crates publish as loams-* in the rename PR (the working names operon-*, and loam-* in fabric/, stay until then); Go modules use the path loams.dev/... (a go-import vanity responder on loams.dev); Java SDKs are deferred (no Loam-written Java; §33 D354, Q356). Answers Q283loamdb was never published; one namespace across the registries, the domain and the binary (D401)Approved (owner, 2026-10-01; recorded 2026-10-02); amends D33 and A32
D4012026-10-02The binary is loams and the domain is loams.dev, bought on Cloudflare. Every CLI command is loams <group> <verb> (§30, D282) and the installer is https://loams.dev/install.sh (D293); hosting the site and its redirects is loam-cloud's concern. Answers Q284 and the registration half of Q281loam collides with other tools (crates.io's loam-cli); one name for the binary, the packages (D400) and the domainApproved (owner, 2026-10-02)
D4022026-10-01CloudEvents types Loam defines use io.loams.dev.<domain>.<name>.v<major> (§34 D364, §32 D355, §36 D390, §37 D436). §02 §7.4 (D270) builds dev.loam.stream.record and §27 specifies dev.loam.meter.usage.v1 today; the rename PR changes that code and those docs. Answers Q361One prefix, derived from the owned domain (D401)Approved (owner, 2026-10-01)
D4032026-10-02The open-core split of D220 is reconfirmed, for adoption and for VC funding. Commercial designs (the protocol gateway of §34, the former §35 Cloudflare target, the hosted Loams Cloud, metering and billing ledgers) live in the private loam-platform; this repository keeps only §27's usage hooks. Knative is open source here, with no metering (§38, D441–D446). Detailed by D440Owner, 2026-10-02: "keep loam cloud and loams-cloud private; may add Knative in OSS but no metering; I want adoption and also to raise money from VCs; move Cloudflare, OpenRTB etc. commercial to private repos"Approved (owner, 2026-10-02); reconfirms D220
D4042026-10-01Authentik's open-source edition replaces Keycloak and Clerk wherever this repository needs an identity provider (the Kubernetes distribution, the showcase, the apps' sign-in): no paid plan or licence key (§38 D447, D458). Supersedes D-SC-3 and amends D221's SAML broker; the hosted cloud's identity is decided in loam-platform (Q442)Owner, 2026-10-01: "for auth Authentik so no paid plan"; scoped to OSS by D403Approved (owner, 2026-10-01; scoped 2026-10-02)
D4052026-10-01Apache Iggy is the event-ingest layer beside the Loam WAL (§32 §5.2 option A, D331, D332): the Loam WAL stays the log for databases and retrieval, and event ingestion goes to the Event Fabric. Answers Q330; D332's pins, retention and upstream contributions stay proposals (Q331, Q346)The owner's direction of 2026-10-01, quoted in D331Approved (owner direction, 2026-10-01; confirmed 2026-10-02)
D4062026-10-02The repository is renamed loams and moves to the ostrium-labs GitHub organisation (ostrium-labs/loams) after the rename PR (D33, D400). Release URLs and attestations (D292), container images (Q289) and the mobile apps' proto ref (D439) use that pathOne name for the repository, the binary and the packagesApproved (owner, 2026-10-02)
D4202026-10-01Three apps, one contract (§37 §1, §2): Loams Desktop (Tauri 2; macOS, Linux, Windows), Loams for iOS (SwiftUI) and Loams for Android (Jetpack Compose). Every application call from an app to Loam is Connect-RPC generated from the protos connect-rust serves (connect-es, connect-swift, connect-kotlin); the named exceptions are the console's OpenAPI /api/v1 (REST until Q423), sign-in at Authentik and the Loam gateway's OAuth token endpoint (OIDC and OAuth over HTTP), the instance-to-gateway push API and APNs/FCM, the desktop's local CLI JSON contract (D283), and the updater's manifests. Personas: a developer with a local stack on the desktop, an operator or approver on the phone, a platform engineer extending the consoleOwner direction 2026-10-01 ("Connect-RPC everywhere", Tauri desktop, native phones); one contract keeps three clients and the server in stepProposed (§37)
D4212026-10-01Borrow the DeepSeek Harness repos' patterns; fork neither (§37 §3). Both are MIT (desktop: © 2026 DeepSeek; mobile: © 2026 DSH Mobile contributors). Nothing is copied by default, so no attribution is owed; a copied file keeps its MIT notice in THIRD_PARTY_NOTICES.md. The desktop repo's VRM avatar (VRM Public License) and native/landlock-run (BSD-3-Clause) are never taken. cordis is a direct MIT dependency; @koishijs/plugin-console (AGPL-3.0 per npm metadata) is never usedThe desktop host is about 160 lines whose value is the pattern; the mobile app speaks another protocol to another server; Loam's fixes (auth, supervision, pinning, push) touch most of it anywayProposed (§37)
D4222026-10-01The console is a cordis v4 application (§37 §5) in the browser (/ui, §19 P1) and inside Tauri. A small host (@loams/console-host) boots a cordis Context and the stock loader from a boot manifest; layout, pages, engine views, connector forms, approval renderers and the RPC clients are plugins. cordis's browser half is used; the harness's Node host half is replaced by the Rust engine and the Tauri host, and its Typert RPC by ConnectOwner direction (2026-10-01 correction): any code loadable as a plugin. cordis gives injected services, activation gated on dependencies, and scoped disposal; the harness proves the browser half with about 40 client pluginsProposed (§37)
D4232026-10-01Plugin manifest and catalog (§37 §5.3): a loams.plugin block in package.json (kind, entry, tier, inject, provides, slots, permissions, requires.console, requires.api, editions, config, server); the catalog loams.yml in cordis v4 loader format (id, name, config, group, disabled, inject), composed from a base list and patch lists (- id: replaces config, - insert: adds rows) as the harness composes bundles and profiles. !!js is refused in every catalogOne format the stock loader reads; editions become patch files; !!js needs unsafe-eval and is code execution through configProposed (§37)
D4242026-10-01Service contracts (§37 §5.4): transport, one rpc.<service> per proto service (provided only when GetInstance.api_versions lists the package, so dependents activate by themselves), api (the OpenAPI client), session, platform (web or desktop, with desktop-only platform.stacks, platform.auth, platform.deeplink, platform.updates), slots, router, settings, flags, i18n. A value import between plugins is a build error; plugins share types through @loams/proto and declaration mergingPages appear on instances that serve their API; swapping the transport (switching environments) reloads dependents for free; the harness rule keeps plugins independently loadableProposed (§37)
D4252026-10-01Typed slots (§37 §5.5), adapted from the harness's ui-slots: kinds single, list, keyed; ctx.slots.register inside an effect; one error boundary per entry; components receive props, never ctx. Initial catalog: root, console.nav, console.page, console.settings.section, environment.overview.card, collection.tab, engine.view, connector.config, flow.step.editor, approval.renderer, agent.tool.view, operation.detail, palette.command, shell.overlay. The §26, §30, §32, §33, §34 and §21 integrations map onto them (§37 §5.5 table)Keyed slots fit engine adapters, connector forms and approval kinds; a generated slot catalog and a props-change check keep plugins compatibleProposed (§37)
D4262026-10-01Trust tiers and isolation (§37 §5.6): core and first-party run in the console realm behind a guard proxy exposing only injected services; third-party always runs in a sandboxed iframe (allow-scripts only, opaque origin, connect-src 'none') behind a permission-checked postMessage bridge, and its calls carry a vended, attenuated token (§19 §5.2 flow 3: scopes = manifest permissions ∩ the user's rights, act = plugin:<id>@<version>, 15-minute TTL) that the server enforces and audits. The host decides the tier (bundled, or @loams/* with npm provenance from ostrium-labs), never the manifest. Third-party plugins are off until the unified auth plancordis has no sandbox and the harness treats plugins like shell access; an iframe is the only robust browser boundary; server-side attenuation makes client checks a convenience rather than the boundaryProposed (§37)
D4272026-10-01Plugin sources and reload (§37 §5.7): bundled; npm @loams/* at build time with provenance; a private registry configured at build time (scoped .npmrc entries plus the build's trustedPublishers), which is how the private hosted plugins of D220 arrive and how self-hosters ship internal plugins; installed on an instance by an org owner through loam.console.v1.PluginService (stored in the bucket, served with integrity hashes and immutable caching; an approval gate in protected environments); a local path in development. Development reload is Vite HMR plus a fiber refresh; production applies WatchManifest changes as fiber add, replace or dispose without a page reload. requires.console is semver against the host; slot prop changes need a host version bumpSelf-hosters extend without forking; the open/private split needs an extension point, not knowledge of private packages; scoped disposal makes live enable and disable safeProposed (§37)
D4282026-10-01Editions are plugin sets; D220 stands (§37 §5.8): oss (web/apps/console/catalog/base.yml) and desktop (a patch adding platform-tauri, stacks, MCP install, notifications, updates) in this repository. The console's multi-tenant, hosted and billing parts are private plugins in loam-cloud and loam-platform, built into the hosted console from a private registry with a private catalog patch; §37 does not design them. The open host offers extension points only: a registry source, trustedPublishers, catalog patches, the open slot catalog, flags.edition/flags.features, and a replaceable platform service. This repository never depends on the private setD220's "basic single-cluster admin UI" is the oss set; the owner's 2026-10-02 retraction keeps the splitProposed (§37)
D4292026-10-01The desktop shell and its sidecar (§37 §6.1–§6.2): one Tauri 2 window with the bundled console (desktop set); a Rust host loams-desktop in its own Cargo workspace under web/apps/desktop/src-tauri. The sidecar is the loams CLI binary (standard variant, externalBin); every stack action is loams stack <verb> --output json under D283's contract, with LOAM_HOME shared with the terminal. Stacks outlive the app; the app restarts keep_running stacks (1 s … 30 s backoff, 5 in 10 minutes); create and delete stay terminal commands in AP1. "Add loams to PATH" writes install_method: "desktop", which self-update refuses (amends D294). Windows is remote-onlyOne supervisor of record (the CLI, §30 §8.3); terminal and app see the same stacks; the harness had no supervision after readinessProposed (§37)
D4302026-10-01Desktop lockdown and the network bridge (§37 §6.3–§6.4): one local capability with an explicit command allowlist (no shell, fs, http, process or Stronghold permission; a CI check), a CSP with no remote source and no unsafe-eval, no remote content in the window. Every console request goes through net_fetch: Rust performs it, streams head, chunks and end over a Tauri Channel into a standard Response, adds the bearer (and DPoP after Q438), strips JavaScript-set credentials, pins TLS for paired environments, and reaches only the active environment's originsTokens never reach JavaScript; Connect server streams work (a custom URI scheme cannot stream); servers need no CORS for app originsProposed (§37)
D4312026-10-01Desktop sign-in against Authentik (§37 §6.5): OIDC authorization code + PKCE at the instance's Authentik (open-source edition; owner ruling) in the system browser with a loopback redirect (RFC 8252 §7.3), public client loams-desktop (an Authentik application per instance); the Authentik token is exchanged at the Loam gateway (RFC 8693) for Loam's access and refresh tokens, and Authentik's are discarded. The refresh token in the OS keychain through keyring 4.2, rotated on use; the access token in Rust memory only; Linux without a Secret Service signs in for the session only. Approvals on the desktop use step-up SESSION (re-authentication at Authentik with max_age=0 within 5 minutes)RFC 8252's recommendation; Authentik authenticates, Loam authorizes (§19 §5.3); tauri-plugin-stronghold is deprecated and absent from Tauri v3 (plugins-workspace#3494)Proposed (§37)
D4322026-10-01Desktop updates, signing, platforms and deep links (§37 §6.6–§6.7): tauri-plugin-updater 2.13 (mandatory signatures) with static latest.json per channel (stable, beta) from tauri-action on GitHub Releases behind a loams.dev/desktop/… redirect, signed with its own key, separate from the CLI's; macOS Developer ID + notarization; Windows Authenticode (Q421); Linux AppImage, deb, rpm; local stacks on macOS aarch64 and Linux x86_64/aarch64; loams:// deep links (open, approvals, stacks) are navigation only, parsed in Rust against an allowlist and forwarded by tauri-plugin-single-instanceSeparate keys limit a leak's blast radius; deep links are attacker-reachable, so they never actProposed (§37)
D4332026-10-01Native mobile, no Kotlin Multiplatform (§37 §7.1, §7.8): SwiftUI (iOS 17+) with connect-swift 1.2 over URLSession, and Compose with connect-kotlin 0.9 over OkHttp; Connect protocol, binary codec. Shared: the protos, golden fixtures (canonical decision bytes, pairing payloads, sealed notifications) and the conformance scenarios; LoamsCore also builds on Linux. Revisit if connect-kotlin gains KMP support and shared logic exceeds about a third of either appconnect-kotlin does not support KMP (connect-kotlin#140, open); Swift export is Alpha; the shared logic is small and pureProposed (§37)
D4342026-10-01Pairing on §19's identity model, with Authentik as the identity provider (§37 §7.2): a phone is a device credential of a user principal, not a new principal kind. A signed-in user calls DeviceService.CreatePairing; the phone scans a QR v1 (issuer = the Loam gateway, instance_id, spki pin set, jkt = the instance's token-signing key thumbprint, a single-use 128-bit code, user_code, exp 5 minutes), pins before the first byte, and redeems at the gateway's token endpoint with the extension grant urn:loams:params:oauth:grant-type:pairing and a DPoP proof, receiving DPoP-bound access and rotating refresh tokens (dev and cnf.jkt claims). Alternatives sign in at Authentik: PKCE in the system browser, or Authentik's device-code flow (RFC 8628) with a six-word fingerprint check; both are exchanged at the gateway (RFC 8693) for the same tokens. The anchor is jkt; TLS pins rotate under pin sets signed by it. No relay in track APMaps the harness's relay token and pinned TLS key onto Loam's existing tokens, JWKS and revocation instead of a parallel scheme; Authentik authenticates and Loam binds; the harness re-paired on every key change and its relay saw plaintextProposed (§37)
D4352026-10-01Approvals as a shared service (§37 §7.3): loam.approvals.v1 (List, Get, Watch, Decide) over §21 §6.5's approval promises, with server-rendered summary and detail_lines. A decision carries a proof: a compact JWS (ES256 from Secure Enclave or StrongBox, EdDSA for software keys) over {approval_id, revision, decision, iat, jti} by a key that needs biometrics or the passcode, or comes from a session younger than 5 minutes, or the policy says step_up: none. The requester (or the user an agent acts for) cannot approve by default (Q432); DESTRUCTIVE needs the target typed; reject needs a reason; no "always allow", no lock-screen approve, no queued decisionsA stolen bearer token cannot approve; a stale card cannot be approved; three platforms and push show the same wordsProposed (§37)
D4362026-10-01Push is a sealed wake-up through a gateway (§37 §7.4): the engine's notifier projects io.loams.dev.* CloudEvents into a per-user inbox (loam.notifications.v1), seals each notification with HPKE (RFC 9180, X25519, ChaCha20-Poly1305; AAD = instance id ‖ notification id) to each device's key, and posts it to loams-push, which holds the APNs and FCM credentials and sees only ciphertext. loams-push is Apache-2.0 in this repository; Loam operates the instance the store apps use (push.loams.dev); self-hosters with their own app builds run their own. Android's unifiedpush flavor receives the same payload with no gateway. The desktop holds WatchApprovals instead of pushOnly the store apps' publisher can push to them (the Sygnal, Mattermost, ntfy and Home Assistant pattern); sealing keeps Apple, Google and the gateway blindProposed (§37)
D4372026-10-01Offline and background (§37 §7.5): phones cache approvals, operations, jobs and the inbox with freshness timestamps and show stale data read-only; no background sockets (push, WorkManager, BGAppRefreshTask); failed sends show "not sent" with a retry that reuses the idempotency key; sign-out or revocation deletes the instance's cache, keys and tokensPlatform background limits; a late or duplicated decision is worse than noneProposed (§37)
D4382026-10-01The app proto surface (§37 §8, AP0): loam.instance.v1 (GetInstance with api_versions, issuer, JWKS, tls_pins, push config, minimum app versions; WhoAmI), loam.devices.v1 (pairing, devices, push targets, preferences), loam.approvals.v1, loam.operations.v1 (the Connect face of D146), loam.notifications.v1 (the inbox), loam.errors.v1 (ErrorInfo.reason). Rules: unary and server-streaming only; Watch* sends snapshot, changes, a heartbeat every 15 s, and resumes from a cursor (or snapshot_reset); reads marked NO_SIDE_EFFECTS (HTTP GET); every mutation takes idempotency_key. Served first by operon-apps-mock (stateful, YAML scenarios, shared acceptance module). Package names stay loam.* until Q422Browsers and URLSession cannot stream requests; AWS ALB ignores HTTP/2 PING and idles out after 60 s, Cloudflare reads time out at 125 s; resumable streams survive mobile networksProposed (§37)
D4392026-10-01Repository layout and track AP (§37 §9, §13): the console (web/packages/*, web/plugins/*, web/apps/console) and the desktop (web/apps/desktop with its own Cargo workspace) in this repository; both phone apps in one new repository, ostrium-labs/loams-mobile (android/, ios/, conformance/proto-ref.lock), generating from a pinned git ref of this repository's proto/. Plans AP0 (protos, mock), AP1a (cordis console), AP1 (desktop), AP2 (Android), AP3 (iOS); AP4 (server side: AP0's services, the pairing grant and DPoP, proof verification, the notifier, loams-push, PluginService) after the unified auth planThe desktop shares the console and the loams release; mobile toolchains, macOS runners and store secrets stay out of the engine's CI; one mobile repository keeps scenarios and fixtures singleProposed (§37)
D4402026-10-02The open-core boundary stands (§38 §1, §7; D220 reconfirmed). Knative, Authentik and the GitOps layout are self-hosting features, so open; no metering in this repository, only §27's hooks. The protocol gateway (§34: D366–D371, D373, D377, D379, plans GW1–GW4) and the Cloudflare target (§35 on PR #179, plan CF1) move to loam-platform; §34 becomes a stub keeping D360–D365, D372, D374–D376, D378; RN1 loses its usage-event taskThe owner's ruling of 2026-10-02: adoption from open source, the hosted cloud and commercial components private for fundraisingProposed · owner ruling 2026-10-02 (recorded as D403)
D4412026-10-02Knative Serving runs the http-port contract (§38 §3.1–§3.2): KnativeRunner implements D375 (one Knative Service per function, one revision per version, scale to zero, runtimeClassName: gvisor); T0 and T1 stay on the node supervisor. Refines §24 §11's F2 (Knative schedules T2 pods when enabled)A server on $PORT is exactly a Knative Service; the supervisor's many-tenants-per-process model is notProposed
D4422026-10-02Kourier is Knative's ingress, internal only (§38 §3.2): Envoy stays the edge (D184); the gateway authenticates and authorizes, then calls Kourier's internal service; every Knative Service is cluster-localNo function is reachable around the gatewayProposed
D4432026-10-02One Kubernetes namespace per Loam namespace for Knative (§38 §3.3): loam-ns-<namespace>, default-deny NetworkPolicy, ResourceQuota and LimitRange from the namespace's limits, service accounts with no API token; created by loam-operator. Quotas are enforced here; setting them per plan stays loam-platform (D220)Kubernetes' own isolation primitives, driven by the limits D65 already storesProposed
D4442026-10-02No meter on Knative (§38 §3.4): KnativeRunner returns usage: None and writes no host reports; Knative pods are T2 sandboxes under §27 §3.2's labels; queue-proxy and activator metrics and edge access logs are further open hooks. Adds §27 §3.5aThe owner's "no metering" (2026-10-02); D190, D202Proposed · owner ruling 2026-10-02
D4452026-10-02Knative Eventing is an adapter (§38 §3.5): Loam streams (D270) and the Event Fabric (§32 D331) stay the logs; loam-knative-source delivers stream records to any Knative sink as binary-mode CloudEvents, at least once; Loam's POST …/streams/{stream}/events is documented as a Knative sink; InMemoryChannel for development; the production broker is Q443Eventing is a delivery layer; a third log would split the event modelProposed
D4462026-10-02Knative via the Knative Operator (§38 §3.2, §9): KnativeServing and KnativeEventing 1.23 with kubernetes.podspec-runtimeclassname, kubernetes.podspec-securitycontext and kubernetes.podspec-volumes-emptydir enabled; knative.enabled: false by defaultgVisor needs the runtime-class flag (disabled by default at 1.23); the Operator's CRs give Argo CD a health signalProposed
D4472026-10-02Authentik's open-source edition is the default IdP of the Kubernetes distribution and the showcase (§38 §4): an unmodified separate service (2026.8.3, Postgres only), only code outside authentik/enterprise/, no licence key ever. Supersedes D-SC-3 (Keycloak); amends D221 (SAML brokered through Authentik)The owner's "Authentik so no paid plan"; MIT core with blueprints, passkeys and outpostsProposed · owner ruling 2026-10-01 (recorded as D404)
D4482026-10-02Authentik's usable features (§38 §4.2): OAuth2/OIDC provider (code + PKCE, client credentials with JWT federation, device code, refresh, RFC 8693 token exchange since 2026.8.0, RFC 7591 registration); SAML, SCIM (static token), LDAP, RADIUS (PAP), Proxy, RAC providers; OAuth, SAML, LDAP, Kerberos, SCIM sources; flows and stages incl. TOTP, WebAuthn and passkeys; RBAC; brands; blueprints; outposts. Excluded (Enterprise, 2026-10-02): multi-tenancy, Google Workspace and Entra providers, SSF, WS-Federation, agent accounts, SCIM OAuth auth, RADIUS EAP-TLS, the source stage, mTLS stage, account lockdown, password history, enhanced audit, reports and CSV exports, lifecycle, PAM, device connectorsChecked against authentik/enterprise/* at 2026.8.3 and the Enterprise features pageProposed
D4492026-10-02Loam's gateway stays the authority for Loam tokens (§38 §4.3, §5): people sign in through Authentik (code + PKCE; device code for loam login); the gateway exchanges the Authentik token (RFC 8693) for a Loam access token; listeners verify only Loam tokens; agents stay Loam principals (§19 P5) with federation, delegation and vending unchanged; Biscuit (D188) and OpenFGA (D66, D67) unchanged; groups become team#member tuples at sign-inOne verifier on every listener; agent features do not depend on EnterpriseProposed
D4502026-10-02The single binary keeps built-in sign-in (§19 P7); Authentik is the default where Kubernetes is; any OIDC IdP works (§38 §4.5)Air-gapped and laptop installs; IdP-agnostic gatewayProposed
D4512026-10-02D111 narrowed (§38 §1): identity, MFA, SSO and SAML brokering are Authentik's; the remaining unified-auth work (verification on every listener, API keys, agent tokens, TLS, leaving loopback) is MT1 for the native API, console API and MCP, and each other listener's plan adopts MT1's verifierSplits a plan that blocked every listener into one plan and per-listener adoptionProposed
D4522026-10-02Authentik configured by blueprints, installed from its upstream chart (§38 §4.4): Loam's blueprints (Apache-2.0) under deploy/authentik/blueprints/; the chart goauthentik/helm is GPL-3.0, so it is referenced by an Argo CD Application, never vendoredConfiguration as code; no copyleft files in this repositoryProposed
D4532026-10-02"GitOps from Clever Cloud" = Clever's open-source operator and infrastructure tooling (§38 §6.1): the clever-kubernetes-operator fork (D185), terraform-provider-clevercloud and karpenter-provider-clever-cloud on CKE, clever-tools as CLI reference. Clever publishes no GitOps engine (checked 2026-10-02), so Argo CD stays (D186)The owner's "for gitops clever cloud", made concrete against what Clever actually publishesProposed
D4542026-10-02A Flux layout with the same order for the single-node profile (§38 §6.1); Argo CD stays the default and tested pathArgo CD's footprint on small clusters (Q-RT-11)Proposed
D4552026-10-02New sync waves (§38 §6.3; amends §25 §6.3): CNPG and Knative Operator CRDs at −2 and operators at −1; authentik-db at 1; Authentik at 2; KnativeServing/KnativeEventing at 3; loam-knative-source at 5; Lua health checks for the new kindsHealth-gated ordering, as D186Proposed
D4562026-10-02The reference small cluster is k3s or k3d with --disable traefik (§38 §6.4)Envoy and Kourier own ingress; matches CIProposed
D4572026-10-02Licences (§38 §9): Knative (Serving, Eventing, Operator, Kourier, func) Apache-2.0; Authentik MIT outside authentik/enterprise/, unmodified; its chart GPL-3.0, referenced only; CNPG, Argo CD, Flux, k3s Apache-2.0; the operator fork MITD11; nothing enterprise or copyleft linked or vendoredProposed
D4582026-10-02A CI guard keeps Authentik free of Enterprise (§38 §1; MT1 Ruling 7): no licence in values, no AUTHENTIK_TENANTS__ENABLED, a blueprint lint against enterprise app labels derived from the image, and an e2e check that the licence summary is emptyAn open-core dependency needs a mechanical check, not a promiseProposed
D4592026-10-02Track MT (§38 §10): MT1 Authentik identity, MT2 Knative, MT3 GitOps; MT1 and MT2 independent, MT3 wires bothSmall stacked PRs on the one-build machineProposed

Open questions

#QuestionOwnerNeeded by
Q1Final project name Resolved 2026-09-25: Loam, package loamdb, after M1 (D33); package namespace loams since 2026-10-01 (D400)FounderResolved
Q2Can Lakekeeper's catalog backend be implemented on Operon meta, or must we bundle Postgres? Lakekeeper with its own Postgres brought forward for FL1 (§32); production open as Q335.EngM4 design
Q3express class semantics on GCS Rapid and Azure (conditional writes, append, zone redundancy)EngM5
Q4Depth of Kafka transactions required by target users. Reopened 2026-09-26: the Kafka gateway is in M5 without transactions (D74); they are added on demandProductAfter M5
Q5Exact Neo4j procedures/APOC functions used by Graphiti, LangChain, LlamaIndex, LightRAG Moot: no Neo4j surface (D44)EngResolved
Q6Lance multivector support depth vs. implementing multivector in the hot tierEngM1 Phase B
Q7Qdrant code boundary: qdrant-edge crate vs. forking lib/segmentEngM1
Q8Governance path (company-led → LF AI & Data / CNCF) and commercial modelFounderBefore 1.0
Q9Relationship with HelixDB (compete vs. collaborate on shared SlateDB/graph pieces)FounderM3
Q10Elastic REST YAML spec test license compatibility for conformance useEngM1
Q11Resonate SDKs: per-namespace auth headers vs. base paths; S3 Express/GCS Rapid/Azure If-Match support for low-latency durable namespaces. Narrowed 2026-09-27: every SDK has a token option, so tenants are chosen by Authorization headers through the unified auth plan (§21 §5.3, D142; verify for Go and Java); the If-Match part applies only if the blob backend is chosen (Q39)EngD2 plan
Q12Contribute the Operon Resonate server plugin upstream, or keep it in-tree. Narrowed 2026-09-27: fixes go upstream first, one concern per PR, and the fork carries them until they merge (§21 §12: PRs 0a, 0b, 0c, 1, 2, 3a, 3, 4, 5; D140); the TiKV Store is offered upstream with an in-tree fallbackFounderD2 plan
Q13arrow encoding layout (IPC per chunk vs. per-segment file with column index) and whether the WAL writes it directlyEngM0.3 (field), M4 (layout)
Q14Approve §15 agent workspaces Approved 2026-09-24; W0 (MCP) with M1, W1 after M3 (D46)FounderResolved
Q15Repos as a sixth object kind (with an implicit stream) or a service like durable execution Proposed answer 2026-10-01 (D388, §36): a service with its own bucket WAL; events mirrored into the _git stream.EngW1 design
Q16Benchmark for the 100-agent demo (§16, approved 2026-09-24 as the agent-track launch demo)FounderBefore W1 plan
Q17Postgres metastore schema and throughput ceiling (partition-head row locks, WAL-commit batching, whether LISTEN/NOTIFY serializes commits at high commit rates) before operon-meta-postgres is planned. Narrowed 2026-09-26: Lakekeeper's patterns are adopted and there is no single clock row (D58, D59)EngM2 plan
Q18Which LightRAG and LlamaIndex storage-backend test suites are store-agnostic enough to gate our native graph adaptersEngM3 plan
Q19A tensor column type for collections (Arrow fixed_shape_tensor; idea from TileDB) and how Lance stores itEngM4 design
Q20Can pylance open a detached Lance version by id, as D53's external readers need, or must Loam expose a manifest path / custom opener? Resolved 2026-09-26 (M1.2 Task 0): yes, pylance 12.0.0 opens it with lance.dataset(uri, version=<detached id>) and the Rust lance crate 12.0.0 with DatasetBuilder::from_uri(uri).with_version(id).load(), both with Lance's default session and registry; no custom opener is needed (§17 §3.5)EngResolved
Q21Can Loam's and Lakekeeper's OpenFGA model-migration managers share one OpenFGA store (D67), or does each product need its own store?EngM2.x plan
Q22Is floci's TransactGetItems isolated from concurrent writes? If not, which conformance cases need DynamoDB Local or real DynamoDB (D60)?EngM2 plan
Q23How openraft state scales past one snapshot: M2's fix for snapshots over 5 GiB (a multipart Store put or a chunked snapshot) and M6's (applied state in redb with incremental snapshots, or one openraft group per shard) (D63)EngM2 plan (5 GiB), M6 plan
Q24TiDB 8.5 details for operon-meta-tidb: RETURNING support, the errno a client receives for an undetermined commit, and the throttling codes (D58) Moot 2026-09-29 (D260: no TiDB)EngM6 plan
Q25Owner confirmation of the defaults: OpenFGA in M2.x with a store shared with Lakekeeper (D67); the erasure policy (D69; its encryption timing was decided on 2026-09-26 by D96); unsharded ids with namespaced calls (D70) Resolved 2026-09-27: all three defaults accepted as they stand (D114)FounderResolved
Q26How Kafka topic names map to Loam namespaces and streams (a <namespace>.<stream> topic name, or a namespace bound to the credential), and which SASL mechanisms carry Loam API keys (D74)EngM5 plan
Q27A metadata restore to a point before a completed erasure (§10 §6): does it replay the completed erasures from the erasure log (which must then survive the restore), or rebuild the affected manifests without the erased objects, whose versions are gone (D68, D69)? Resolved 2026-09-27: a restore replays the completed erasures from the erasure log before serving, and the log is kept while any snapshot, backup or time-travel version older than the erasure exists, plus 30 days (D115)EngResolved
Q28Where a branch's Lance commits live: detached versions in the source's dataset directory (D34), or a directory of its own that references the source's data files; and how lineage GC finds every member (D90)EngM2 plan
Q29Whether Tantivy's block-max WAND bounds, computed with k1 = 1.2 and b = 0.75, still cover per-field BM25 parameters with D78's slack, or whether pruning must be off for them (D93)EngM2 plan
Q30The unified auth plan (D111): API keys and TLS on the native API, the Qdrant and Elasticsearch gateways, MCP and the console; multi-tenant keys and quotas; OpenFGA; and when the non-loopback warning becomes a refusal Narrowed 2026-10-02 (D451, §38): MT1 is the identity half; listeners adopt MT1's verifier in their own plans.FounderAfter M1, before the M2 plan
Q31Serializable range reads in Loam Live mutations: guard keys per index equality-prefix bucket, or validation against the commit journal after prewrite (§20 §5.2, D118)EngR2 plan
Q32MVCC GC for txn-API keyspaces: does TiKV honour keyspace-level safe points set through PD's GC-state RPCs, or must a unified GC TiDB run per cluster; will tikv-client accept a patch exposing keyspace GC (§20 §9.3, D122) Answered by R1 Task 0 (2026-09-27, docs/plans/r1-dependency-spike.md): no. PD v8.5.8 has no GC-state RPCs (AdvanceTxnSafePoint/GetGCState answer Unimplemented), and TiKV v8.5.8 reads only the cluster safe point: a keyspace safe point set with UpdateGCSafePointV2 was ignored for 90 s, while the cluster path (UpdateServiceGCSafePoint + UpdateGCSafePoint) dropped old versions within 10 s. TiKV serves reads below the safe point without an error. Loam's GC loop therefore acts as the cluster's GC worker for every keyspace, with barriers as PD service safe points, and operon-tikv refuses reads older than the GC window (R1 plan rows R5–R7). Per-keyspace GC returns when a pinned PD release ships the GC-state API. Still open: whether client-rust accepts a patch exposing GC safe points (R1 exit report)EngAnswered (R1 Task 0); upstream patch: R1 exit report
Q33Do released classic TiDB binaries (v8.5.x) support keyspace-name? Verified on playground v8.5.8 (2026-09-27 spike): yes, with the keyspace pre-allocated in PD and TiKV on API v2 (§20 §10.1, D123). Moot 2026-09-29 (D260: no TiDB); the follow-up on tidb-operator v2 keyspaces for classic TiDB is withdrawnEngMoot
Q34Can TiCDC capture a keyspace-mode TiDB's tables into a Kafka or storage sink on a classic cluster, for SQL tables → collections (§20 §10.3) Moot 2026-09-29 (D260: no TiDB)EngR3 plan
Q35The long-term function engine: QuickJS only, V8 (deno_core) for CPU-bound functions and npm compatibility, or wasmtime for Rust and Go functions (§20 §6.3, D120)EngR2 plan
Q36Keyspaces per TiKV cluster before region overhead dominates, and whether TiKV request units can be attributed to a txn-API keyspace; they set the size-class thresholds and per-app quotas (§20 §9.2, D122)EngR2 plan
Q37Do BR's log backup and PITR restore cover a txn-API keyspace (not only TiDB tables), per keyspace, on a classic cluster; and what recovery point does the default log flush interval give (§20 §10.4, D131)EngR2 plan
Q38What tiup playground --mode tidb-x / tidb-cse (next-gen, S3-backed TiDB, in tiup 1.17.1) runs: where its TiKV binaries come from, their license, and whether they can be self-hosted; no public next-gen TiKV source was found. If open, S3 could become Live's source of truth (§20 §10.4, D130, D131)EngR2 plan
Q39The durable backend for self-hosted clusters without TiDB: Resonate's Postgres plugin (D58 brings Postgres in M2) or the blob server over operon-store (keeps D1, needs no new service) (§21 §3.3, D139) Resolved 2026-09-29 by D261: self-hosted clusters use the native TiKV backendEngD2 plan
Q40Retention of settled promises: an upstream server setting (PR 5) or Loam-side deletes per backend, and the safe horizon given that a late retry re-creates a pruned id and re-runs the work (§21 §8, D147)EngD2 plan
Q41The TiKV durable layout: one loam_durable keyspace with tenant prefixes, or keyspaces per size class; depends on Q36 (§21 §5.1, D142)EngD2 plan
Q42The largest origin the blob/TiKV backend should hold (one document per origin; TiKV prefers values under 1 MiB) and the split rule for large fan-outs (§21 §6.2)EngD2 plan
Q43OTLP traces ingest (D73 is logs only) for agent traces and durable step spans, and the link from spans to an execution graph (§21 §6.6)EngD3 plan
Q44Durable Live actions: a Resonate context bound into QuickJS over the Rust SDK, or actions run by the TypeScript SDK in a Node worker; ties to Q35 (§21 §6.6)EngD3 plan
Q45Owner confirmation of D149: Neon for the showcase apps' Postgres OLTP, D-SC-12's TiKV front end narrowed to Loam Live's API, D-SC-16 amended (§23 §5) Answered 2026-09-29 by the owner: no, CloudNativePG serves the showcase apps (D230, §28 §4)FounderBefore the N3 plan
Q46Given Neon's dormant upstream (§23 §4.1), does Loam accept owning a Neon fork (Postgres minor rebases, security fixes), or default the suite to plain Postgres 17 and keep Neon for workspace branching only (D151) Answered 2026-09-29 by the owner: yes, Loam forks and owns Neon (D231, §28 §10)FounderBefore the N1 plan
Q47Scale-to-zero and SNI: implement the Neon proxy's control-plane API in Loam (get_endpoint_access_control, wake_compute, JWKS) and run the proxy, or keep Loam's listener as the only router and start computes on first connect (§23 §3.1) Superseded by Q113 (§28 §8)EngNeon production plan
Q48The Neon storage controller's own Postgres database: where it lives, and whether a single-pageserver install needs the controller at all (§23 §3.1) Answered 2026-09-29: a small CNPG Cluster; TiKV later (§28 §5.2, §9)EngNeon production plan
Q49Branch selection on the Postgres wire (options=-c loam.branch=… or <db>@<branch>), and whether bridges run on branches (inherited slots) (§23 §6.3) Answered 2026-09-29 for the syntax: <db>__<branch> database names in PgDog (D236); how the principal is presented so Loam authorizes the branch before routing, and bridges on branches, stay openEngN3 plan
Q50WeSQL durability: does archive recovery replay binlog slices from the bucket when configured, so losing the local volume loses no committed transaction (§23 §9.2)EngBefore N6
Q51Matomo on WeSQL: do its installer and schema run on SmartEngine (no foreign keys, verify), and does its archiver fit a single node (§23 §13)EngBefore N6
Q110Fork neondatabase/postgres as dina-kar/postgres (required by dina-kar/neon's relative submodule URLs); base the catch-up on the REL_1x_STABLE_neon heads (16.12/17.8, possibly needing unpublished extension changes) or on Neon main's pins (16.9/17.5) (§28 §10)Founder (the fork), EngP2a
Q111PGroonga for Zulip on CNPG: a custom image on 17.11-standard-trixie, or a separate Cluster (§28 §4)EngP1
Q112Where the interpreted WAL sender builds: in Operon behind a feature with the fork's Postgres headers cached in CI (postgres_ffi runs bindgen against them), or as a small binary in the fork's workspace linking operon-safekeeper (§28 §6.8)EngP4a plan
Q113Scale-to-zero for Loam Postgres: Neon's proxy with Loam's wake_compute in front of PgDog, or always-on computes first (§28 §8). Supersedes Q47FounderP3 plan
Q114Re-measure RawKV and TxnKV 1PC write latency on a three-node, three-AZ TiKV cluster with PLP NVMe; published TiKV p99 figures for WAL-sized values (§28 §6.4)EngP4a
Q115More than one in-flight append per timeline, if single-flight group commit limits throughput (§28 §6.5)EngP4b results
Q116The WAL on the metastore's TiKV cluster (own stores through placement rules) or on a dedicated cluster (§28 §6.6)EngP4a plan
Q117Leader placement granularity (per timeline or per AZ group) and shorter Raft ticks on WAL stores, given the ~10 s default election timeout stalls a failed leader's timelines (§28 §6.6, §12 row 10)Eng, Founder (acceptable failover time)P4b
Q118BatchCommands RPC batching in Loam's client-rust fork (client-go's max-batch-wait-time), if P4b shows per-RPC overhead (§28 §6.6)EngP4b
Q119The bucket copy's log layout: a stream per tenant with a partition per timeline, or a stream per timeline (§28 §6.7)EngP4a
Q-SC-1Showcase SSO gaps (§22 §5): is chained SSO through Forgejo's OAuth2 provider enough for Plane (generic OIDC is paid), or do we upstream generic OIDC; for PostHog (SAML/OIDC in ee/), forward-auth plus provisioned accounts or an upstream OIDC backend outside ee/EngSC1 Task 0
Q-SC-2PostHog FOSS: is there a maintained image built from posthog-foss, or must the suite build one; what does the FOSS build lose beyond SSO and RBAC; is PostHog worth its weight in the showcase (§22 §4.2)EngSC1 Task 0
Q-SC-3Withdrawn 2026-09-28 (D-SC-16). OpenFGA's MySQL datastore on TiDB: do its migrations and tests pass (§22 §6)EngSC1 Task 7
Q-SC-4The language of commons-control and the portal: TypeScript or Python on the Resonate SDKs, or Rust (§22 §3.1)EngSC1 Task 0
Q-SC-5Withdrawn 2026-09-28 (D-SC-16). Forgejo on TiDB (not in Forgejo's supported list): does make test-mysql pass against TiDB (§22 §6.3)EngSC1 Task 7
Q-SC-6PostHog's hobby deploy on Apache Kafka instead of Redpanda (BSL-1.1), or Redpanda documented as a separately licensed service (§22 §6.4)EngSC1 Task 0
Q-SC-7Plane's public API makes workspace membership read-only (changes by invite): enough for the projection, or upstream a write API (§22 §7.3)EngSC1 Task 4
Q-SC-8Drift policy for changes made in an app's own UI: revert to OpenFGA's state, or adopt them as tuples (§22 §7.3)FounderSC1 Task 5
Q-SC-9Postgres write compatibility (D-SC-12): which store backs it (a Postgres front end over TiKV transactions, or a layer in front of TiDB), and which Postgres features are in scope first. Proposed answer 2026-09-29: Neon behind Loam's pg listener (D149, §23 §5), pending Q45. Updated 2026-09-29 (owner, D230): the showcase apps use plain Postgres 17 on CloudNativePG; serverless Postgres is Loam Postgres, a Neon fork (D231, §28); the TiKV front end stays long-termFounderThe engine's Postgres-write design doc
Q-SC-10The PostHog fork Withdrawn 2026-09-28 (D-SC-15): the OpenPanel fork's scope goes in its design doc. Was: the PostHog fork (D-SC-13): the staged scope, keeping HogQL's semantics on DataFusion, and how far the fork may drift from upstream before rebases stop being practicalEngThe PostHog-on-Loam design doc
Q95Arroyo: a documented alternative only (proposed), or a managed engine beside RisingWave (§26 §10.1, D212)EngJ4 plan
Q96Engine tenancy: Sail per namespace with scale-to-zero (proposed); RisingWave per namespace or a shared cluster with a database per small namespace (§26 §9.2, D211)EngJ3 plan
Q97Jobs on self-hosted clusters without TiKV: require TiKV, or a JobStore over the Postgres or DynamoDB MetaStore backends Resolved 2026-09-29 by the owner: jobs require TiKV; redb for operon dev only; no Postgres or DynamoDB jobs backend (§26 §6.8, D214)EngResolved
Q98A server-streamed LeaseStream beside the long-polled unary Lease (§26 §5.5, D206)EngJ2 plan
Q99Strict per-key ordering (FIFO groups): out of scope (proposed) or a later queue kind (§26 §6.5)FounderAfter J2
Q100BullMQ v5 apps: require the v6 upgrade or ship a v5 shim Resolved 2026-09-29 by the owner: BullMQ v6 only, no v5 shim (§26 §7.2.5, D209)FounderResolved
Q101Celery beat's solar and custom schedule classes: keep the standard scheduler for them (proposed) or support them in Loam (§26 §7.1.6, D208)EngJ1 plan
Q102Whether Loam hosts Celery and BullMQ workers, with the CPU-time runtime of doc 24 (§26 §17)FounderDoc 24
Q103The jobs listener on its own port (7720, proposed) or mounted on the native API listener (§26 §5.5, D206)EngJ1 plan
Q104The Spark fallback: Loam-managed (Spark Connect server and the Kubeflow Spark Operator per namespace) or documented only (§26 §9.4, D211)FounderJ3 plan
Q-RT-1Name and home of the Rust Dapr server (loam-dapr), and whether operon-stream-grpc and deploy/dapr are merged to main first (§24 §3.1, D182)FounderF1 plan
Q-RT-2The resident threshold for short waits and the wall limit for non-SDK code; whether resident memory beyond an allowance is billed (§24 §6, D173)FounderF1 plan
Q-RT-3What the Cloudflare adapters' output needs that open-source workerd lacks; whether capnp is generated from wrangler.toml or by a Loam builder (§24 §4.3)EngF0
Q-RT-4Durable Objects in clusters: out of scope, or on a TiKV keyspace with router placement (§24 §4.2)EngF2
Q-RT-5Does the Resonate TypeScript SDK run inside workerd unchanged (§24 §6)EngF0
Q-RT-6Per-invocation CPU attribution inside a shared workerd process (§24 §7, D175)EngF1 plan
Q-RT-7The request fee and CPU price, given §24 §9's cost modelFounderBefore cloud beta
Q-RT-8The operator's API group and domain, and where the fork lives (§25 §5, D185)FounderOperator fork
Q-RT-9Does Cellar honour If-None-Match/If-Match on PUT, so that it can hold the WAL, or is it for static assets only (§25 §5, D178)EngCellar provider
Q-RT-10Biscuit for §19's vending flow too, or only inside the runtime; the cargo deny result for biscuit-auth 6.0 (§25 §4, D188)FounderF1 plan
Q-RT-11Argo CD's footprint on k3d and in small BYOC clusters; whether BYOC-local-meta needs a Flux option (§25 §6.1, D186) See Q449 and MT3 Task 6 (2026-10-02).EngF1
Q-RT-12Envoy Gateway, or plain Envoy with xDS from loam-gateway (§25 §8)EngF1 plan
Q-RT-13Clever Kubernetes Engine as a first-class target (CI and a clever-cke env), or documented only (§25 §8)FounderBefore cloud beta
Q-RT-14Contribute crates/core improvements back to clever-kubernetes-operator instead of diverging (§25 §8)FounderAfter the fork
Q-RT-15Does daprd v1.18 hot-reload per-tenant secret-store components, and which tenants get a dedicated daprd (§24 §5.1, D189)EngF1 plan
Q-UH-1Per-namespace metric cardinality: OTLP delta metrics only, or a Prometheus endpoint limited to the namespaces active on a node (§27 §3.1)EngF1 plan
Q-UH-2How the usage-hooks contract (metric names, cgroup layout, labels, loam.meter.v1) is versioned, and where its conformance tests live (§27 §6)EngF1 plan
Q260MySQL wire access after D260 (no TiDB): Loam's own MySQL wire front end over DataFusion (read-only analytics over collections and tables, the MySQL counterpart of D-PG-1's pgwire listener), or no MySQL surface at all. Not decided (§20 §10) Narrowing to analytics proposed 2026-10-01 by Q314 (§31: OLTP MySQL wire access is vtgate in front of WeSQL, D320).FounderBefore any MySQL-protocol work
Q261Arm A membership changes: reuse Neon's generations and pull_timeline, or a control-plane member set in TiKV with re-sync from the bucket plus a peer (§28 §7.2)?EngP4c
Q262Rewrite (purge) idle timelines' live journal records forward so they stop pinning old segments (§28 §7.2)?EngBefore multi-tenant production
Q263Arm A defaults on PLP NVMe: --io-depth, SQPOLL, IOPOLL with poll queues (§28 §7.2)EngThe three-node gate run
Q264Multishot recv with provided buffer rings on the shards' sockets, and SO_BUSY_POLL by default (§28 §7.2)?EngAfter Q263
Q270CloudEvents dedup (D270): a window per stream instead of the server-wide --cloudevents-dedup-window, and a cap on ledger entries per namespace, if the ledger's memory matters beyond webhook and trigger event rates (§02 §7.4)EngM2 plan
Q271Legal review of the WeSQL WAL client's licensing boundary: GPL-2.0-only client speaking a documented protocol to an Apache-2.0 server, written from the specification without copying Loam code (§29 §6.4, D277)Owner and counselBefore the W2 plan
Q272One archive or two: keep the fork's binlog_archive slices beside the acceptors' offload, or replace them (§29 §6.5)EngW2 plan
Q273Do W1–W4 beat reversing D260 for the bucket-only requirement (§29 §9, D280)OwnerBefore W2 starts
Q274Forgejo's minimum MySQL version versus WeSQL's 8.0.x base: rebase to 8.4 or accept a documented deviation (§29 §11 W1)EngW1 plan
Q275W1 in the handler wrappers or in the engine's handler; decided by the spike on a current-read probe from ha_smartengine (§29 §4.3, §5)EngThe W1 spike
Q276ROLLBACK TO SAVEPOINT with writes (T2) before or after W1 (§29 §4.3)EngAfter the W1 spike
Q277Consistency tokens for WeSQL replica reads: carry the binlog offset in D76's token (§29 §7.4)EngW3 plan
Q278mysql-binlog timeline kind in the safekeeper v3 greeting, or a separate listener for non-Postgres kinds (§29 §6.4)EngW2 plan
Q279Stream-scoped authentication for loam-wal's mysql-binlog kind: client certificates with the stream in the SAN, or control-plane tokens (§29 §6.4)EngW2 plan
Q281loams.dev has no DNS record (checked 2026-10-01): who registers and hosts it? Answered 2026-10-02 by the owner: loams.dev is bought on Cloudflare (D401). Still open: the proposal is a redirect route in loam-cloud for /install.sh and /releases/* to GitHub Releases. Should the product docs' source move into this repository's docs/guides so that the website and the embedded bundle share it (§30 §13, §17.2)?FounderCLI2 publish (Task 9)
Q282Release signing: who generates and holds the minisign key (offline, the owner), and which reviewers gate the release environment? Or use Sigstore keyless only, which gives no in-band verification without gh or cosign (§30 §17.1, D292)?FounderCLI2 Task 4
Q283The @loam npm scope (@loam/sdk in the chat dump, @loam/bullmq and @loam/durable in §26): register the scope, or rename to loamdb-* (§30 §14, D296)? Answered 2026-10-02 by the owner: the namespace is loams with the npm scope @loams (D400)FounderResolved
Q284The binary name loam collides with other tools that may install a loam binary (crates.io's loam-cli, "Loam CLI for building smart contracts"). Keep loam, or ship loamdb with a loam alias (§30 §20 risk 1)? Answered 2026-10-02 by the owner: the binary is loams (D401)FounderResolved
Q285Windows: ship the cli variant (x86_64-pc-windows-msvc), or support WSL2 only (§30 §2.2)?FounderCLI2 Task 2
Q286Intel macOS server variants: skip them (proposed), or cross-build (§30 §9.2)?EngCLI2 Task 2
Q287The default install variant: standard (proposed: one command gives a working stack) or cli (small, with a server downloaded on the first stack create) (§30 §9.2, CLI2 Ruling 4)?FounderCLI2 Task 5
Q288Local stacks as detached processes (proposed default) or as systemd user units and launchd agents by default, which survive logout and reboot? CLI2 Task 9 adds them as an opt-in (§30 §8.3)EngCLI2 Task 9
Q289Container images per variant on GHCR (ghcr.io/ostrium-labs/loams:<version>-<variant>): in CLI2 or later?FounderCLI2
Q290A Homebrew tap, an npm shim (npx loams) or a PyPI shim: which, and when?FounderAfter CLI2
Q291Cloud stacks: the loam-platform public API that stack create --target cloud calls, and whether the cloud client lives in this repository (D220 says the CLI is open) (§30 §2.2)FounderCLI3 plan
Q292May the MCP tool stack_create start a local process without a human confirming it? The proposal is yes, within its limits: loopback, a local object store, no disks, no downloads without allow_download, at most 8 stacks (§30 §12.1, CLI1 Ruling 8)FounderCLI1 Task 9
Q293The cargo-dist spike: does cargo-dist build several variant packages that each produce a binary named loams, or does Loam own a matrix job for the extra variants (§30 §17.1)?EngCLI2 Task 3
Q294Stack default ports for the pg wire (15432) and MySQL wire (13306), which avoid local Postgres and MySQL. D-PG-1 gives the server no default port. Keep these, or use 5433 and 3307 (§30 §8.2)?EngCLI1 Task 4
Q300Accept D300 (buy PgDog and Vitess; build only the control plane and the evidence) over the chat dump's Loam-built Rust router with MySQL and Postgres frontends (§31 §5)FounderRT1 plan start
Q301Vitess topology on a dedicated etcd (proposed) or on PD's embedded etcd (support for external etcd v3 clients not verified) (§31 §4.1)EngRT3 plan
Q302WeSQL's rebase to MySQL 8.4 before Vitess v24's support ends (estimate: about 2027-04); joint with §29 Q274 (§31 §9.2 C-3)Founder, EngRT3 plan
Q303Atomic MySQL cross-shard commits: a semi-sync replica so Vitess 2PC is allowed, or no MySQL cross-shard atomicity (§31 §9.2 C-2, §10)Eng, FounderRT4 plan
Q304Vitess _vt sidecar tables on SmartEngine or InnoDB (serverless_honor_innodb_engine), and whether vttablet's sidecar diff loops on the engine (§31 §9.2 C-1)EngRT3 Task 0
Q305When sharded Loam Postgres through PgDog leaves beta (§31 §18, risk 1)FounderRT2 exit
Q306Does CUTOVER on one PgDog plus RELOAD of the rendered swap on the others give identical routing, without PgDog's proprietary Enterprise Edition (§31 §6.4)EngRT4 plan
Q307The Postgres fence: ALTER ROLE <app> NOLOGIN plus termination (proposed), or per-table write revokes (§31 §6.4)EngRT1 Task 0
Q308The Lean CI job on every relevant PR (proposed) or nightly only (§31 §12.3)EngRT0 Task 8
Q309Trace validation with TLC and the trace-spec method (proposed), Apalache, or a Rust port of each spec's Next (§31 §11.3)EngRT1 Task 8
Q310The nightly DST seed budget (1 000 000 proposed) against runner minutes (§31 §13.5)EngRT1 exit
Q311The nemesis harness in Rust (proposed) or Jepsen as an external tool (§31 §14.3)EngRT5 plan
Q312Keep RouterSession (a black-box contract of bought routers) or drop it (§31 §11.2)EngRT4 plan
Q313vtgate in front of unsharded WeSQL too (D320), retiring §23's N6 splice (§31 §9.1)FounderRT3 plan
Q314Narrow Q260 to analytics: OLTP MySQL wire access is vtgate in front of WeSQL (§29, D320); whether Loam also serves read-only MySQL wire over DataFusion stays open (§31 §19)FounderWith Q273
Q330Confirm §32 §5.2 option A: Iggy as the event stream engine beside the Loam WAL, with bounded retention until Iggy's tiered storage lands (D332) Answered 2026-10-02 by the owner: Iggy is the event-ingest layer beside the Loam WAL (D405); retention stays Q346FounderResolved
Q331D74's Kafka gateway: defer it and contribute to Iggy's Kafka gateway (proposed), keep D74 in M5 as well, or cancel it (D332)FounderBefore the M5 plan
Q332Fluss tiering in v1: the Flink job under the operator (proposed) or a Rust tiering service first (D333)FounderFL1 Task 0
Q333Approve reversing D45 for the House, and chsurface-1's scope (§32 §8, D347)FounderBefore FL2 Task 2
Q334D-SC-15: run OpenPanel's ClickHouse queries unchanged on the House (proposed) instead of rewriting them for DataFusionFounderAfter FL2
Q335Lakekeeper's catalog database in production: its own Postgres, or a backend on Loam's metastore (Q2)EngM4 plan
Q336House tenancy: one house process per node with per-namespace sessions and quotas (proposed), or a process per namespaceEngFL2 Task 0
Q337chDB binding: Loam's own loam-chdb-sys (proposed) or chdb-rust 2.0 directly (D342)EngFL2 Task 0
Q338The native ClickHouse protocol (port 9000) in FL5 (proposed), or earlier if driver suites show HTTP is not enoughFounderAfter FL2
Q339Demand for Cypher, GQL or Bolt (a grafeo-server companion), or §07's native surface only (D350)FounderFL4
Q340Flow UI: embed Karavan's designer for Camel-backed steps, or Loam's own editor only (D340)FounderFL3
Q341Fluss 1.0's Iceberg layout for PK-table tiering (equality deletes or merge-on-read) and whether chDB's Iceberg reader reads it correctlyEngFL2 Task 0
Q342fluss-rs APIs for lake-snapshot offsets, log scans from offsets and write offsets (§32 §7.5): present, or contributionsEngFL2 Task 0
Q343Iggy credentials: a user per namespace with personal access tokens per API key, or a user per API keyEngUnified auth plan (Q30)
Q344A Confluent Schema Registry REST subset for the Fabric's Kafka clients: in loam-fabric, or upstream in Iggy's gatewayEngWith Q331
Q345Fluss and Iggy footprint on small and self-hosted clusters (JVM heap, ZooKeeper, Flink) and a single-node profileEngFL1 Task 0
Q346Iggy retention before tiering (72 h proposed) and the replay story when Fluss is down longerFounderFL1
Q347After D331, do D72's explicit streams (named consumers, subscribe) stay in M2 as planned, or narrow to trigger-rate and internal useFounderM2 plan
Q348rdkafka (librdkafka, C build) or a pure-Rust client for the Kafka ★ connector (rskafka 0.6 lacks consumer groups)EngCN1 Task 0
Q349Shipping ADBC drivers (Snowflake, BigQuery) inside the loam-fabric image: licences, NOTICE, platform builds, the in-process trust modelEngCN1 Task 0
Q350Debezium Server's offset and schema-history stores: a file on a volume, Redis, or a Fluss/Iggy-backed store contributed upstreamEngCN1 Task 8
Q351Iggy postgres_source (CDC mode) as the default for Postgres sources with Debezium for the rest, or Debezium everywhereEngCN1 Task 8
Q352Camel Main or Camel Quarkus (JVM or native) for loam-connectEngCN2 Task 0
Q353loam-connect's home: connect/ in this repository, or its own repositoryFounderCN1 Task 12
Q354Connector credentials: Dapr secret components per namespace (D189), or Loam-vended short-lived credentials (AWS STS, GCP workload identity) where providers support themEngUnified auth plan
Q355Which P2 connectors move into CN1 if a launch customer needs themFounderCN1 start
Q356Publish a Kestra plugin for Loam (House queries, Fabric produce, route control); deferred with all Loam-written Java (owner, 2026-10-01)FounderWhen Java is un-deferred
Q357Contribute Loam's native connectors (Kafka, ADBC, Kinesis) to Iggy's connectors runtime as plugins, so one Rust connector set serves bothFounderAfter CN1
Q358The OpenAPI connector generator for CN3 (progenitor 0.15 or openapi-generator) and how cursors and webhooks are declared beside a specEngCN3 plan
Q359Connector metering hooks (§27): which counters (events, bytes, API calls) the platform needs per instanceFounderBefore the cloud beta
Q360, Q363–Q365, Q370, Q371, Q373, Q374Moved to loam-platform (private), 2026-10-02: §34's protocol gateway (D440)—Moved
Q361Event type prefix: dev.loam. (as §02 §7.4 builds) or dev.loams. (the registered domain), for §02's synthesized type and D364's types alike (§34 §2) Answered 2026-10-01 by the owner: io.loams.dev.<domain>.<name>.v1 (D402); §02 currently builds dev.loam., so the rename PR changes the codeFounderResolved
Q362Showback and single-organisation billing in the open repository or loam-platform only Answered 2026-10-02 by the owner: no metering in OSS (D403, D440, D444)FounderResolved
Q366Lambda CPU attribution: the bootstrap's getrusage delta capped by billed duration × memory_mb / 1 769, or billed duration as the meter on Lambda (§24 §16, §27 §3.6); gates RN1 Task 5FounderRN1 Task 5
Q367Cloud Run and Container Apps runners: build or document only (§24 §16)FounderAfter RN1
Q368Iceberg v3 timestamptz_ns on the pinned iceberg-rust and Lakekeeper by M4, or a time_ns column (§34 §4)EngEvent-table plan
Q369Move operon-stream-grpc from tonic/prost to connect-rust/buffa (D128), and when (§34 §5)EngM2 stream API plan
Q372Internal HTTP/3: the condition that enables it (§34 §3)EngLater
Q380–Q383, Q396–Q399Moved to loam-platform (private), 2026-10-02: the former §35 Cloudflare target (D440)—Moved
Q384Segment PUT latency on R2, S3 Standard, S3 Express One Zone and RustFS, and whether S3 Express directory buckets honour If-None-Match: * (§36 §4.4)EngGT1 Task 10
Q385RustFS conditional PUT above 1 MiB: can a reader ever see a partial large object (§36 §4.2)EngGT1 Task 3
Q386Does current git drive partial clone and lazy fetch through a remote helper's stateless-connect (§36 §6.2)EngGT2 Task 0
Q387Do libgit2 and JGit fetch over protocol v2, or does W1's client matrix need a v0/v1 upload-pack (§36 §6.1, D394)EngGT2 Task 0
Q388When to support SHA-256 repositories, given gitoxide's open SHA-256 parity item (§36 §2.2)EngAfter GT2
Q389Git authentication before the unified auth plan (D111): loopback only in the open-source gateway until MT1 (§38 D451, PR #182) (§36 §6.1)FounderGT2 Task 0
Q390Is age-since-write eviction enough for the direct cache path, or must quotas require the gateway path (§36 §8.2)EngGT3 results
Q391Per-scope WAL partitions, if GT4 measures sequencer contention on a hot monorepo (§36 §6.4, D397)EngGT4 plan
Q392A server-side merge queue that rebases disjoint-path agent commits onto a shared branch: GT4, later, or never (§36 §6.4)FounderGT4 plan
Q393Should the WAL carry pack bytes for very small pushes, saving the separate .lpk PUT (§36 §16)EngGT1 results
Q394The WebDAV subset sccache's backend needs (PROPFIND, MKCOL, HEAD, GET, PUT) (§36 §8.2)EngGT3 Task 0
Q395Start GT1–GT3 now as track GT beside M, R, D and J, or keep §15's W1 after M3 (§36 §13)FounderBefore GT1
Q420Store and signing accounts for the ostrium-labs entity: Apple Developer Program (D-U-N-S, legal name) and Google Play; who holds the Developer ID identity and the upload keys (§37 §6.6, §7)FounderAP1 Task 11, AP2 Task 10, AP3 Task 10
Q421Windows code signing: Azure Artifact Signing (organisation validation), a Key Vault certificate, or unsigned betas (§37 §6.6)FounderAP1 Task 11
Q422Rename every proto package loam.* → loams.* in D33's rename PR, before any app is published? Connect URL paths contain the package (§37 §8.2)FounderBefore the AP2/AP3 store releases
Q423Keep the console's OpenAPI /api/v1 (§19 P9) for identity administration (proposed for M2), or move it to Connect (loam.console.v1) so every console call is an rpc.* service (§37 §8.1, §16)?FounderAP1a Task 4
Q424The official push gateway push.loams.dev: free for every self-hosted instance using the store apps? Limits, a privacy policy, instance registration (§37 §7.4)FounderAP4
Q425Reaching private instances from phones (a laptop stack, a cluster behind a firewall): an end-to-end relay (TLS passthrough by SNI), or VPNs and tunnels only (proposed for track AP) (§37 §7.2.4)?FounderAfter AP2/AP3
Q426Third-party console plugins in OSS: allowed, off by default, installed by org owners (proposed)? Is npm provenance from ostrium-labs enough for first-party, or only an explicit allowlist (§37 §5.6)?FounderAP1a Task 6
Q427cordis: pinned npm cordis@4.0.0-rc.10 behind @loams/cordis with pnpm patch (proposed), vendor now like the harness, or wait for 4.0? (§37 §14 risk 1)EngAP1a Task 1
Q428Bundle the standard loams binary in the desktop app (proposed; tens of MB, estimate) or download it on first run through the CLI's variants (§37 §6.2)?EngAP1 Task 2
Q429On quit, leave stacks running (proposed) or stop the ones the app started (§37 §6.2)?FounderAP1 Task 3
Q430Linux packages: AppImage, deb and rpm (proposed), plus Flatpak or Snap (§37 §6.6)?EngAP1 Task 11
Q431Minimum OS versions: iOS 17, Android 10 (API 29), macOS 13 (proposed) (§37 §7)FounderAP2/AP3 Task 0
Q432May a user approve an operation an agent requested on their behalf? Proposed: no by default; an org policy may allow it outside protected environments (§37 §7.3)FounderAP0 Task 3
Q433Crash reporting in the apps: none (proposed, D284) or opt-in Sentry, which loam-cloud uses (§37 §16)?FounderAP1/AP2/AP3 release tasks
Q434Store names ("Loams") and a trademark checkFounderStore releases
Q435App review: a bundled demo mode with seed data (proposed) or a hosted demo instance and account (§37 §14 risk 5)?FounderAP2/AP3 Task 10
Q436Does Loam Cloud's console (today a Next.js app in loam-cloud with Clerk, which the Authentik ruling retires) move onto the cordis host as private plugins from the private registry (§19 P1; proposed), or stay separate (§37 §5.8)?FounderThe cloud console's next phase
Q437Windows desktop: remote-only (proposed), stacks through WSL2, or a Windows server variant (ties to §30's Q285) (§37 §6.2)?FounderAP1 Task 11
Q438Does the unified auth plan add, on the Loam gateway, the RFC 8693 exchange of Authentik tokens for Loam tokens, DPoP (RFC 9449) for user tokens issued to devices, device-bound rotating refresh tokens and the pairing extension grant (§37 §7.2, §16)? Which Authentik release provides the device-code flow and max_age step-up the apps rely on? §19 §5.3 lists DPoP as an agent follow-up onlyFounder, EngThe auth plan; AP4
Q439Ship the Android unifiedpush flavor on F-Droid (reproducible builds), and when (§37 §7.4)?FounderAP2 Task 10
Q440Authentik's SCIM provider is free; keep SCIM provisioning into Loam in loam-platform (D221), or move it to OSS for adoption (§38 §4.3)FounderMT1 Task 5
Q441Keep the single binary's built-in password and TOTP (§19 P7), or require an external OIDC IdP once MT1 landsFounderMT1 Task 0
Q442Hosted Loams Cloud identity: Clerk (loam-cloud today) or Authentik Moved to loam-platform (private), 2026-10-02: the hosted cloud's identity is decided there; OSS uses Authentik (D404)FounderMoved
Q443Knative Eventing's production broker: the Kafka broker over Loam's Kafka gateway (M5), a Loam broker class over streams, or none (§38 §3.5)EngMT2 Task 6
Q444gVisor mandatory for every KnativeRunner function, or optional for trusted single-org code (§38 §3.2)FounderMT2 Task 2
Q445Should KnativeRunner emit request and wall-time reports (no CPU) for showback, or stay at usage: None (D444)FounderMT2 Task 4
Q446Kourier, or net-gateway-api on Envoy Gateway so Envoy is the only proxy (§38 §3.2)EngMT2 Task 3
Q447Authentik's Postgres: CloudNativePG now (D230), Loam Postgres (§28) laterEngMT3 Task 3
Q448Does Authentik send OIDC back-channel logout, so a removed user's session ends before its refresh (§38 §4.3)EngMT1 Task 4
Q449Flux as the default for the single-node profile if Argo CD measures too heavy (merges Q-RT-11)EngMT3 Task 6
Q450Give §34's retained vendor-neutral decisions their own document, and split GW1's vendor-neutral tasks (buf breaking, the CloudEvents profile, the event Arrow mapping) into an open planFounderBefore GW1 starts in loam-platform
Q451The loam.dev/* pod labels under the loams rename: keep, or loams.dev/* with the rename PREngRename PR
Q452loam-knative-source for Iggy topics (§32): in MT2 or with FL1EngAfter FL1
Q453Authentik upgrade cadence and security-patch policy for self-hostersEngMT3 Task 2

On this page