Loam is pre-alpha: the engine core runs today; Live and Durable are in progress. See the roadmap

Blog/The Loam platform

Self-hosting Loam with GitOps

How the whole stack is meant to deploy from one Git repository, from a laptop cluster to your own data centre, on RustFS, the TiKV operator, a Loam operator and Argo CD. And the line between what is open source and what only the managed cloud runs.

Self-hosting Loam with GitOps
On this page
  1. The pieces
  2. The Loam operator
  3. Sync waves
  4. What stays true without Kubernetes
  5. Open source and the managed cloud
  6. Where this stands

Proposal The deployment design in this post is under review. Today the engine runs as a single binary (operon dev, operon standalone, operon cluster) with TiKV from tiup playground in development. The open-source boundary at the end is approved.

A platform that lives in your bucket should also be deployable in your cluster, by you, without calling us. The goal is one deployment path: the same Git repository and charts deploy a laptop k3d cluster, a self-hosted cluster on your own hardware, and the managed Loam Cloud.

The pieces

LayerComponentLicenseWhy this one
Object storeRustFSApache-2.0S3-compatible, written in Rust, with the conditional writes Loam's commit protocol needs. It replaced MinIO, whose community edition went into maintenance mode and was archived. Any S3, GCS or Azure store works as a provider instead
Metadata and LiveTiKV and PD, deployed by TiDB Operator v2Apache-2.0v2 models a cluster as separate component groups, so a cluster of PD and TiKV only, with no TiDB, is a supported shape
Durable executionResonate serverApache-2.0Embedded in dev; its own Deployment on the TiKV store in clusters
EdgeEnvoyApache-2.0TLS, HTTP/3, gRPC routing, rate limits, external authorization
LoamThe engine roles, the Live role, the runtime tiersApache-2.0—
OperatorA Loam Kubernetes operatorMIT (fork) and Apache-2.0See below
GitOpsArgo CDApache-2.0Sync waves, custom health checks for CRDs, and one control plane for many clusters

The Loam operator

The proposal forks Clever Cloud's Kubernetes operator (MIT, Rust, kube-rs) as the skeleton of a Loam operator, keeping its MIT notice. What carries over is modest and we say so: a controller registry, finalizers, event recording, a metrics server and a hardened Helm chart, about 1,300 lines. Its own custom resources provision Clever Cloud add-ons through Clever's API, so they are dropped. The operator adds four resources:

ResourceReconciles
LoamA Loam cluster: its roles as StatefulSets or Deployments, the metastore's PD endpoints, the bucket, the Resonate endpoint
ObjectStoreA bucket and credentials from a provider: RustFS by default, any S3, or Clever Cloud's Cellar
FunctionA function version for the runtime: contract, tier, bundle digest, routes, limits
RuntimePoolPer-node capacity for the runtime tiers and their RuntimeClass

Object stores sit behind a provider trait that reports capabilities, and Loam refuses a provider without conditional writes, because the write-ahead log's commit protocol depends on them.

We also evaluated Clever Cloud's edge proxy, Sōzu. It is well engineered and reconfigures without restarts, but it is AGPL-3.0 and, as far as its documentation shows, has neither HTTP/3 nor gRPC routes, both of which the runtime needs. Envoy stays the edge. The same evaluation proposes Biscuit (Apache-2.0, now at the Eclipse Foundation, started at Clever Cloud) for sandbox tokens inside the runtime (part 5).

Sync waves

Argo CD would apply an app-of-apps in waves. Each wave waits for the previous one to be healthy, which needs custom health checks for the child Application resources (Argo CD has none built in), the TiDB Operator's resources and Loam's own.

  1. 01CRDs and operatorsTiDB Operator v2, the Loam operator, the gVisor RuntimeClass
  2. 02Object storeRustFS and the system buckets
  3. 03PD and TiKVAPI v2 on, no TiDB; health from the operator
  4. 04ResonateThe durable server on the TiKV store
  5. 05EdgeEnvoy, the Rust gateway, the Dapr API server
  6. 06LoamThe Loam resource and its roles
  7. 07RuntimeThe node supervisor, workerd and wasmtime tiers
The proposed order. The umbrella chart has no Argo-specific templates, so Flux can consume it too.
deploy/
  helm/loam-stack/          # umbrella chart; every component can be switched off
  gitops/
    bootstrap/argocd/       # Argo CD and its health checks
    root.yaml               # the app-of-apps
    apps/                   # one Application per wave
    envs/{dev-k3d,selfhosted,cloud}/values.yaml

The costs, plainly: Argo CD is heavier than Flux (an application controller, a repo server, Redis and a server), which matters on a laptop and in small clusters, and TiDB Operator v2's resources are still v1alpha1, so each Loam release pins an operator version. A local k3d profile runs one replica of everything; PD, TiKV and TiDB together peaked at about 3.2 GB of memory in our development playground.

What stays true without Kubernetes

The engine is still one binary. operon dev runs every role in one process with the openraft metastore and a local directory or any S3 bucket, and needs no TiKV. TiKV becomes necessary for Loam Live, for jobs and for clusters that use the TiKV metastore.

Open source and the managed cloud

The line is one rule: anything needed to self-host Loam as a single organisation is open source, under Apache-2.0. Anything needed only to run Loam as a multi-tenant paid cloud is part of the managed Loam Cloud. The open repository must never depend on the managed cloud; the cloud consumes the same hooks and APIs any self-hoster can use.

Open sourceManaged Loam Cloud only
The engine: retrieval, streams, Iceberg analytics and every wire APIMetering and billing
Loam Live, the TiKV metastore, the change-feed bridgesThe multi-tenant control plane: provisioning, plans, entitlements
The embedded Resonate server, durable patterns, the jobs API and adaptersSetting quotas per plan
The runtime: the Dapr API server, workerd and wasmtime hosting, gVisor, secretsFleet and multi-region operations, autoscaling policy, pre-warming
Namespaces, OIDC and API-key auth, OpenFGA checks, enforcing quotasHosted Postgres fleet automation, BYOC management
Metrics, OTel, cgroup labels, access logs: every usage hookAbuse, trust and safety; support tooling
The Helm chart, the operator, the Argo CD layout, RustFS defaults, backup and restoreThe multi-tenant console
SDKs, the CLI, generated clients, the design docs

Quotas show the split well: the engine enforces whatever limits it is given; the managed cloud decides what those limits are for each plan. And we do not hold back reliability or performance features from open source. The hot tier is in the open engine today, and the planned replicated write-ahead log will be too, not behind a license key. A few calls are still open, such as enterprise SSO and SCIM.

Where this stands

PartStatus
Single binary: dev, standalone and cluster modesAvailable
TiKV development playground configsAvailable
RustFS as the default store, with the store conformance suite and S3 fault matrix run against itPlanned
The umbrella chart, the Loam operator, Argo CD layout, TiKV through TiDB Operator v2Planned
Continuous backup of TiKV keyspaces to the bucket, restorePlanned
The open-source and managed-cloud boundaryApproved policy

That is the end of the platform series. The other series on this blog, What Loam is built on, takes each open-source project under Loam in turn: what it is, how it works inside, why we chose it, exactly how we use it, its license and its current limits.

More from the blog