Loam is pre-alpha: the engine core runs today; Live and Durable are in progress. See the roadmap

Blog/The Loam platform

Functions billed on CPU time

An agent that computes for 10 ms and waits 20 s on a model should pay for 10 ms. Our proposal for Loam Functions, with workerd, wasmtime and gVisor tiers and Resonate for long waits.

Functions billed on CPU time
On this page
  1. How serverless platforms bill today
  2. Three kinds of waiting
  3. Tiers
  4. The app contracts
  5. A Dapr API, in Rust
  6. Metering through open hooks
  7. Where this stands

Proposal Everything in this post is a design under review. No code for Loam Functions exists yet.

Agent code spends most of its life waiting. It sends a prompt and waits seconds for tokens. It calls a tool and waits. It asks a human for approval and waits hours. Serverless platforms that bill wall-clock time charge for all of that waiting. The design goal of Loam Functions is one sentence: waiting costs no CPU, so it should cost nothing.

The second goal is placement. A function that runs next to Loam's data reaches retrieval, Live, durable promises and streams over a loopback or node-local connection, instead of across the internet with egress on every call.

How serverless platforms bill today

The public price lists show the two models side by side (read on 2026-09-29):

PlatformWhat you pay forWaiting on I/O
Cloudflare Workers (Standard)Requests and CPU milliseconds; "no charge or limit for duration"Free
Vercel Fluid computeActive CPU, plus provisioned memory for the instance's whole life, plus invocationsCPU free, memory billed
AWS LambdaDuration × memory (GB-seconds), plus requestsBilled

Cloudflare's model is the one we want. It works because a V8 isolate that is waiting on a socket costs almost nothing to keep around. The catch is that it only covers code that can run in an isolate, and only for short waits. An agent that sleeps for an hour or waits a day for a human still holds an isolate somewhere. Loam's design adds durable promises for that case.

Three kinds of waiting

WaitExamplesWhat happensCPU billedMemory held
Short (under a threshold, 30 s proposed)a model call, a database query, setTimeoutThe instance stays resident with its event loop parkedNone while idleA few MB
Long, in Resonate SDK codectx.sleep("1h"), waiting for an approval, awaiting another functionThe SDK records a durable promise and returns; the instance is evicted; when the promise settles, the function replays and skips its finished stepsOnly the replay's CPUNone
Long, in ordinary codea plain await on a slow socketStays resident up to the invocation's wall-clock limit, then is cancelledNone while idleA few MB, bounded

Only SDK code is evicted, because only SDK code is replayable: the Resonate SDKs record each step's result as a durable promise, so a re-invocation skips completed steps. Arbitrary JavaScript that holds a socket or a closure cannot be resumed without snapshotting its memory, and we have deferred VM snapshots. We would rather be exact about what is free than promise transparent suspension of anything.

Tiers

Edge
Envoy: TLS, HTTP/1.1, HTTP/2, HTTP/3, gRPC, WebSocket, SSE
Routing
loam-gateway (Rust): route, CloudEvents 1.0, admission, quotas
Node runtime
T0: workerd, one process per tenantT1: wasmtime componentsT2: gVisor sandboxes (later phase)
Platform APIs
Dapr API in Rust: identity, secrets, state, pub/subResonateLoam retrieval, Live, streams
Storage
TiKV (function metadata)Object store: bundles and static assets
The proposed runtime. Envoy terminates TLS and HTTP/3; a Rust gateway routes to a tenant, function and version; a supervisor per node runs the tiers.
TierRuntimeIsolationForMeter
T0workerd, one process per tenant, one isolate per function versiongVisor around the process, or seccomp, a user namespace and a cgroup where gVisor is unavailableJavaScript and TypeScript: Hono, and the Cloudflare adapters of Astro, SvelteKit, Nuxt, React Router and Next.js (OpenNext)Per-process CPU time from the tenant's cgroup
T1wasmtime with the pooling allocator, a component per requestWasm memory safety, inside the host's process and cgroupRust, Go (TinyGo) and Python (componentize-py) components on wasi:httpFuel, a deterministic count of Wasm work rather than CPU time, or epoch interruption (cheaper, coarse)
T2Bun, Node, Python or native binaries in gVisorgVisor's user-space kernelExpress, Fastify, Django, Rails: anything that listens on $PORTThe sandbox's cgroup
T3Firecracker microVMs from snapshotsKVMDeferred—

Why workerd and not our own JavaScript runtime. With a bare JavaScript engine, we would have to write and maintain the web platform ourselves: fetch, streams, crypto.subtle, URL, the Node compatibility surface. The framework adapters for Cloudflare already target workerd's API, so it runs their output unchanged. workerd's own README is explicit about the cost: "workerd on its own does not contain suitable defense-in-depth against the possibility of implementation bugs". So isolates are never shared across tenants. Each tenant gets its own workerd process inside its own sandbox, and isolates only separate one tenant's functions from each other.

Why wasmtime first for Wasm. It is the reference implementation of the component model and WASI 0.2, and its fuel metering gives an exact, deterministic CPU count. Before we write a host, we plan to evaluate Spin and wasmCloud, both Apache-2.0 and both built on wasmtime, as the T1 host.

Why Firecracker snapshots are deferred. We wanted microsecond cold starts from VM snapshots. Kata Containers' Firecracker driver turned out not to restore snapshots (its pause, save and resume calls return without acting), so that path needs firecracker-containerd or orchestration of our own. It waits.

The sandbox deep dive covers the three runtimes in detail.

The app contracts

A manifest picks one of three contracts, and the contract picks the tier:

  • static: files served from the object store and the edge. Vite SPAs, Docusaurus, Hugo.
  • fetch: a fetch(Request) → Response handler in the style of Hono. The reference contract. T0, or T1 for Wasm components.
  • http-port: a server that listens on $PORT. T2.

This is not a Vercel or Netlify clone: no preview deployments, image optimization or CDN product. Frontends get static hosting and the workerd framework presets.

A Dapr API, in Rust

Functions reach the platform through a subset of Dapr's API, served by a new Rust server over a loopback gRPC connection: service invocation, pub/sub on Loam streams, state on Live or a TiKV keyspace, bindings, configuration and secrets. Dapr's workflow API is denied, because Resonate owns durable execution. The long tail of Dapr bindings and brokers (NATS, MQTT, AMQP) goes through one shared Go daprd per cluster, not a sidecar per function.

Tenant secrets come from AWS Secrets Manager, Azure Key Vault, GCP Secret Manager, Vault, OpenBao or Kubernetes through Dapr's secrets building block, with one component per tenant and store, carrying the tenant's own credentials. Loam enforces the scoping from the caller's credential, never from a request field. The Dapr deep dive explains the details and the limits.

Identity inside the sandbox. Every call to the platform API carries either an mTLS certificate with a SPIFFE ID or a sandbox token: a Biscuit the node supervisor mints for each invocation and attenuates for each call. Biscuit tokens can be narrowed offline by any holder, which is exactly what a function calling a sub-function needs. OpenFGA stays the authority for every decision on data; the token only carries narrowing facts.

Metering through open hooks

Metering sources follow the tiers: wasmtime fuel or epochs for T1, the tenant's cgroup cpu.stat for workerd (cross-checked with getrusage), the sandbox cgroup for gVisor. eBPF comes last, as a cross-check only, because updating a BPF map on every context switch taxes every tenant's hot path.

The open-source runtime does not bill anyone. It exposes usage through documented, stable hooks that any metering system can read:

  • Prometheus and OTLP metrics per tenant, function, version and tier: invocations, CPU seconds, wall seconds, resident bytes. For T0, CPU seconds are exact per tenant process and estimated per function and version.
  • A documented cgroup layout and pod labels for every sandbox, kept until a final reading is possible.
  • Per-invocation reports on a node-local socket, where many tenants share one process.
  • Envoy access logs with a tenant header the gateway sets and the client cannot.

Turning those into invoices is part of the managed Loam Cloud. A self-hosting organisation sees the same usage for dashboards, capacity and quotas; quota enforcement stays in the open engine. Part 8 draws that line in full.

Where this stands

PartStatus
Manifest, Rust Dapr API server with tenant identity, workerd and wasmtime with Hono, Resonate suspend and resume, CPU hooks, secrets through DaprPlanned
gVisor tier for http-port appsPlanned
Firecracker snapshots, eBPF cross-check, Kafka triggersPlanned

The first step is a set of spikes: workerd inside gVisor on a local Kubernetes cluster (cold start, memory per tenant, config generation from a Wrangler bundle), Spin against wasmCloud against raw wasmtime, and whether Resonate's TypeScript SDK runs inside workerd unchanged. Functions also depend on the auth plan and on the stream API. The next post covers the other workloads teams already run beside their apps: jobs.

More from the blog