On this page
An AI application is still an application. It has users, sessions, chat threads, settings and permissions, and its interface wants to update the moment something changes. Convex showed how pleasant that can be: documents in tables, queries that stay subscribed and push new results, mutations that are transactions and retry themselves.
Loam Live brings that developer model into Loam, on infrastructure we can run and sell, under open-source licenses. It adds one thing Convex does not have: a table can be declared searchable, and Loam keeps a collection in step with it, so vector, full-text and hybrid search run over application data without an ETL job.
Live is the one part of Loam that does not keep its data in the bucket. It runs on TiKV, because an application database needs millisecond transactions over mutable rows. This post explains the design and how far it is built. Live borrows concepts from Convex, not its wire protocol, function names or code; Convex's backend is under the FSL, and we read it for ideas only.
The shape
Generated Connect, gRPC and gRPC-Web stubs plus a thin reactive layer per platform.
One keyspace per app (or a shared keyspace for small apps), one for Loam metadata.
Data model
An app is one Live database. It belongs to a namespace, beside that namespace's collections and streams. An app has tables of documents: maps from field names to values (null, int64, float64, bool, string, bytes, array, object). Every document has an _id and a _creationTime.
Document ids are 16 random bytes inside the table's key range. Their text form encodes the table id and a checksum, so a client can check that an id belongs to the table it claims. Random ids spread inserts across TiKV regions; a time-ordered id would put every insert on one region.
Indexes list up to 16 fields, and _creationTime and _id are appended to every index so entries are unique and stably ordered. A query on an index is a range: equality on a prefix of the fields, then optional bounds on the next field, in either direction, with a limit.
All keys are encoded with an order-preserving tuple codec, so a TiKV range scan returns index entries in value order. Integers flip the sign bit, floats use the standard total-order trick, strings escape zero bytes and end with a double zero. The codec is fuzzed against a reference comparator.
Transactions
A mutation is one TiKV optimistic transaction. The function reads at the transaction's start timestamp, buffers its writes and commits. TiKV's Percolator-style two-phase commit makes the commit atomic across regions. On a write conflict the runner reruns the whole function at a new timestamp, with jittered backoff, up to eight times. That is safe because functions are deterministic: Date.now() returns the transaction's timestamp, Math.random() is seeded from it, and there is no network, filesystem or timer access. Cryptographic randomness throws inside queries and mutations rather than returning something predictable.
Timestamps are the version numbers. PD's timestamp oracle hands out globally ordered timestamps. A query result is "valid at ts", a mutation returns its commit timestamp, and a client keeps an optimistic update on screen until its session has caught up to that timestamp.
Isolation, honestly. TiKV gives snapshot isolation. Convex promises serializability. The gap is write skew: two mutations read overlapping data, write disjoint keys, and both commit. Live closes the common case by promoting every document read by id into the transaction's lock set, so a read-modify-write of one document is serializable. Range reads get snapshot isolation only. Two candidate fixes for serializable ranges are designed (guard keys per index bucket, or validating read ranges against the journal after prewrite), and the choice waits on measurements.
Idempotency. A mutation can carry an idempotency key. The runner stores the key's record inside the same transaction, so a retried call after a lost acknowledgement returns the recorded result instead of applying twice.
Reactivity from read sets
Every query records a read set: the point keys it read and the index ranges it scanned, as encoded key ranges. A filter applied after the scan does not narrow the set, so an insert that falls into a scanned range is always caught. A range that stopped at its limit is recorded up to the last key returned, so inserts past a full page do not invalidate it.
The invalidation signal comes from a commit journal that every mutation writes inside its own transaction. The journal is sharded; a mutation picks a shard, reads its head h, and writes its entry at h + 1 and the head in the same transaction. So each shard's sequence is dense and ordered by commit, and an entry is visible at timestamp T exactly when its transaction committed by T. Each entry lists the document keys and the old and new index keys the transaction touched.
- 01TickTake a timestamp T, read the shard heads at T and scan the new journal entries.
- 02MatchLook up each changed key in an interval tree of read sets, per table and index.
- 03RerunRerun each invalidated query at T. Identical subscriptions share one rerun.
- 04PushSend each session one transition to T with only the queries whose results changed.
The argument that no update is missed is short. Every write to a key in a query's read set belongs to a transaction that also wrote a journal entry naming that key. The entries visible at T but not at the query's last tick t are exactly those between the two heads. If none touches the read set, the old result is still the result at T. Every query in a session is evaluated at the same tick, so a client never shows two results from different moments. A safety net reruns every subscription every five minutes and alerts on any difference.
Why not TiKV's own change feed? We tested it from Rust: TiKV's CDC service does stream a transactional keyspace's changes, with resolved timestamps. But ordered delivery waits about a second for resolved timestamps to advance, the subscriber must track region splits and leader moves itself, and old index keys need extra configuration. The in-transaction journal invalidates within milliseconds and needs no region tracking. CDC stays a validated secondary path for consumers that can tolerate a second of lag.
Server functions
| Kind | Transaction | Side effects | Status |
|---|---|---|---|
| Query | Reads one snapshot | None; rerun on invalidation | Built-in queries available |
| Mutation | One TiKV transaction | None; retried on conflict | Built-in mutations available |
| Query and mutation functions in TypeScript | As above | As above | In progress |
| Action | Each inner query or mutation is its own transaction | fetch to allowed hosts, AI-gateway calls | Planned |
| Scheduled function | Scheduled transactionally by a mutation | As its kind | Planned |
User functions are TypeScript, bundled with esbuild, and run in QuickJS through rquickjs. Convex runs V8. We chose QuickJS for now because Live functions are mostly small and I/O-bound, QuickJS builds from C in seconds, and it gives CPU and memory limits out of the box. Each context serves exactly one invocation and is then discarded, so module-level state can never leak between calls. The function API is engine-neutral, so moving to V8 or workerd later does not change user code. Actions that must survive a crash will run as durable functions on the embedded Resonate server (part 4).
The sync API
The client protocol is protobuf over connect-rust, which serves Connect, gRPC and gRPC-Web from one handler:
service LiveService {
rpc Watch(WatchRequest) returns (stream Transition); // a session: results, then changes
rpc ModifyQuerySet(ModifyQuerySetRequest) returns (ModifyQuerySetResponse);
rpc Query(QueryRequest) returns (QueryResponse);
rpc Mutate(MutateRequest) returns (MutateResponse); // returns the commit timestamp
rpc Deploy(DeployRequest) returns (DeployResponse);
}It is a server stream plus unary calls, not a bidirectional stream, because browsers cannot stream full duplex over fetch and HTTP/1.1 cannot either. A server stream works on every transport, so one design serves every client. A session's state is a version triple (query-set version, identity version, timestamp); a client applies a transition only if it starts at its current version, and a gap means it resumes. Clients for other platforms are generated with Buf (protobuf-es, connect-go, connect-swift, connect-kotlin); each platform adds a small hand-written layer for the session state machine and optimistic updates.
Tenancy: keyspaces
TiKV's API v2 prefixes every key with a mode byte and a 3-byte keyspace id, so one cluster can hold up to 224 isolated keyspaces. Live uses them for tenancy: large apps get their own keyspace, and small apps share a pooled keyspace under an app prefix, because each keyspace costs at least one region. A router maps each app to its keyspace; moving an app from shared to dedicated copies its key range under a write fence. Nobody advances the MVCC garbage-collection safe point for a non-TiDB keyspace, so Loam runs its own GC loop per keyspace.
The same cluster holds Loam's own metadata. A MetaStore backend on TiKV implements the engine's metadata contract (WAL commits, leases, fenced pointer swaps) and passes the same 49-case conformance suite, with linearizability checks, as the default openraft backend. The TiKV deep dive covers that, and the client fork it needs.
The bridge into search
This is the part that makes Live more than a Convex clone. A table in the deployed schema can say searchable: { collection, fields, vectors, text }. A bridge task tails the journal, reads the changed documents at each tick, and appends upserts and deletes to the collection's implicit stream with an idempotent producer id per journal shard and the journal sequence as the producer sequence. A crash between append and checkpoint re-appends, and the producer sequence drops the duplicate: exactly once, end to end. Embeddings come from the collection's own ingest path, so a mutation never waits on a model call.
A search from a Live function can pass after_ts, the commit timestamp of a mutation, and the service waits until the bridge has appended past it. So "write a message, then search for it" works in one function. The bridge is designed, not built.
Durability
TiKV replicates each region three ways with Raft, so a lost store costs a leader election, not data. For losing a whole cluster, the design makes continuous log backup to object storage mandatory: TiKV's backup-stream component writes change logs to the tenant's bucket, periodic full snapshots are taken with BR, and a cluster does not serve traffic until its backup is running and a restore has been tested. Whether BR's point-in-time restore covers a non-TiDB keyspace is an open question the next phase answers first; if it does not, the commit journal already records every write and can be exported instead.
Where this stands
| Part | Status |
|---|---|
| TiKV client layer: keyspace bootstrap, transaction runner with fault hooks, tuple codec, GC loop | Available |
| TiKV metastore, passing the metastore conformance suite (behind a build flag) | Available |
Documents, indexes, LiveTxn, the commit journal, read sets and subscriptions | Available |
loam.live.v1 protos and the generated TypeScript client | Available |
| The served sync API, TypeScript functions in QuickJS, the reactive and transaction checkers | In progress |
| Router and shared keyspaces, actions, scheduled functions, log backup and restore | Planned |
| The collections bridge, auth on the Live API, React hooks, Swift and Kotlin clients | Planned |
| Kubernetes packaging and BYOC for Live | Planned |
Live has no authentication yet, so its listener binds 127.0.0.1 and refuses any other address at startup. The next post covers the other half of "stateful but not in the bucket": durable execution.