On this page
Proposal Everything in this post is a design under review. No code for Loam Functions exists yet.
Agent code spends most of its life waiting. It sends a prompt and waits seconds for tokens. It calls a tool and waits. It asks a human for approval and waits hours. Serverless platforms that bill wall-clock time charge for all of that waiting. The design goal of Loam Functions is one sentence: waiting costs no CPU, so it should cost nothing.
The second goal is placement. A function that runs next to Loam's data reaches retrieval, Live, durable promises and streams over a loopback or node-local connection, instead of across the internet with egress on every call.
How serverless platforms bill today
The public price lists show the two models side by side (read on 2026-09-29):
| Platform | What you pay for | Waiting on I/O |
|---|---|---|
| Cloudflare Workers (Standard) | Requests and CPU milliseconds; "no charge or limit for duration" | Free |
| Vercel Fluid compute | Active CPU, plus provisioned memory for the instance's whole life, plus invocations | CPU free, memory billed |
| AWS Lambda | Duration × memory (GB-seconds), plus requests | Billed |
Cloudflare's model is the one we want. It works because a V8 isolate that is waiting on a socket costs almost nothing to keep around. The catch is that it only covers code that can run in an isolate, and only for short waits. An agent that sleeps for an hour or waits a day for a human still holds an isolate somewhere. Loam's design adds durable promises for that case.
Three kinds of waiting
| Wait | Examples | What happens | CPU billed | Memory held |
|---|---|---|---|---|
| Short (under a threshold, 30 s proposed) | a model call, a database query, setTimeout | The instance stays resident with its event loop parked | None while idle | A few MB |
| Long, in Resonate SDK code | ctx.sleep("1h"), waiting for an approval, awaiting another function | The SDK records a durable promise and returns; the instance is evicted; when the promise settles, the function replays and skips its finished steps | Only the replay's CPU | None |
| Long, in ordinary code | a plain await on a slow socket | Stays resident up to the invocation's wall-clock limit, then is cancelled | None while idle | A few MB, bounded |
Only SDK code is evicted, because only SDK code is replayable: the Resonate SDKs record each step's result as a durable promise, so a re-invocation skips completed steps. Arbitrary JavaScript that holds a socket or a closure cannot be resumed without snapshotting its memory, and we have deferred VM snapshots. We would rather be exact about what is free than promise transparent suspension of anything.
Tiers
| Tier | Runtime | Isolation | For | Meter |
|---|---|---|---|---|
| T0 | workerd, one process per tenant, one isolate per function version | gVisor around the process, or seccomp, a user namespace and a cgroup where gVisor is unavailable | JavaScript and TypeScript: Hono, and the Cloudflare adapters of Astro, SvelteKit, Nuxt, React Router and Next.js (OpenNext) | Per-process CPU time from the tenant's cgroup |
| T1 | wasmtime with the pooling allocator, a component per request | Wasm memory safety, inside the host's process and cgroup | Rust, Go (TinyGo) and Python (componentize-py) components on wasi:http | Fuel, a deterministic count of Wasm work rather than CPU time, or epoch interruption (cheaper, coarse) |
| T2 | Bun, Node, Python or native binaries in gVisor | gVisor's user-space kernel | Express, Fastify, Django, Rails: anything that listens on $PORT | The sandbox's cgroup |
| T3 | Firecracker microVMs from snapshots | KVM | Deferred | — |
Why workerd and not our own JavaScript runtime. With a bare JavaScript engine, we would have to write and maintain the web platform ourselves: fetch, streams, crypto.subtle, URL, the Node compatibility surface. The framework adapters for Cloudflare already target workerd's API, so it runs their output unchanged. workerd's own README is explicit about the cost: "workerd on its own does not contain suitable defense-in-depth against the possibility of implementation bugs". So isolates are never shared across tenants. Each tenant gets its own workerd process inside its own sandbox, and isolates only separate one tenant's functions from each other.
Why wasmtime first for Wasm. It is the reference implementation of the component model and WASI 0.2, and its fuel metering gives an exact, deterministic CPU count. Before we write a host, we plan to evaluate Spin and wasmCloud, both Apache-2.0 and both built on wasmtime, as the T1 host.
Why Firecracker snapshots are deferred. We wanted microsecond cold starts from VM snapshots. Kata Containers' Firecracker driver turned out not to restore snapshots (its pause, save and resume calls return without acting), so that path needs firecracker-containerd or orchestration of our own. It waits.
The sandbox deep dive covers the three runtimes in detail.
The app contracts
A manifest picks one of three contracts, and the contract picks the tier:
static: files served from the object store and the edge. Vite SPAs, Docusaurus, Hugo.fetch: afetch(Request) → Responsehandler in the style of Hono. The reference contract. T0, or T1 for Wasm components.http-port: a server that listens on$PORT. T2.
This is not a Vercel or Netlify clone: no preview deployments, image optimization or CDN product. Frontends get static hosting and the workerd framework presets.
A Dapr API, in Rust
Functions reach the platform through a subset of Dapr's API, served by a new Rust server over a loopback gRPC connection: service invocation, pub/sub on Loam streams, state on Live or a TiKV keyspace, bindings, configuration and secrets. Dapr's workflow API is denied, because Resonate owns durable execution. The long tail of Dapr bindings and brokers (NATS, MQTT, AMQP) goes through one shared Go daprd per cluster, not a sidecar per function.
Tenant secrets come from AWS Secrets Manager, Azure Key Vault, GCP Secret Manager, Vault, OpenBao or Kubernetes through Dapr's secrets building block, with one component per tenant and store, carrying the tenant's own credentials. Loam enforces the scoping from the caller's credential, never from a request field. The Dapr deep dive explains the details and the limits.
Identity inside the sandbox. Every call to the platform API carries either an mTLS certificate with a SPIFFE ID or a sandbox token: a Biscuit the node supervisor mints for each invocation and attenuates for each call. Biscuit tokens can be narrowed offline by any holder, which is exactly what a function calling a sub-function needs. OpenFGA stays the authority for every decision on data; the token only carries narrowing facts.
Metering through open hooks
Metering sources follow the tiers: wasmtime fuel or epochs for T1, the tenant's cgroup cpu.stat for workerd (cross-checked with getrusage), the sandbox cgroup for gVisor. eBPF comes last, as a cross-check only, because updating a BPF map on every context switch taxes every tenant's hot path.
The open-source runtime does not bill anyone. It exposes usage through documented, stable hooks that any metering system can read:
- Prometheus and OTLP metrics per tenant, function, version and tier: invocations, CPU seconds, wall seconds, resident bytes. For T0, CPU seconds are exact per tenant process and estimated per function and version.
- A documented cgroup layout and pod labels for every sandbox, kept until a final reading is possible.
- Per-invocation reports on a node-local socket, where many tenants share one process.
- Envoy access logs with a tenant header the gateway sets and the client cannot.
Turning those into invoices is part of the managed Loam Cloud. A self-hosting organisation sees the same usage for dashboards, capacity and quotas; quota enforcement stays in the open engine. Part 8 draws that line in full.
Where this stands
| Part | Status |
|---|---|
| Manifest, Rust Dapr API server with tenant identity, workerd and wasmtime with Hono, Resonate suspend and resume, CPU hooks, secrets through Dapr | Planned |
gVisor tier for http-port apps | Planned |
| Firecracker snapshots, eBPF cross-check, Kafka triggers | Planned |
The first step is a set of spikes: workerd inside gVisor on a local Kubernetes cluster (cold start, memory per tenant, config generation from a Wrangler bundle), Spin against wasmCloud against raw wasmtime, and whether Resonate's TypeScript SDK runs inside workerd unchanged. Functions also depend on the auth plan and on the stream API. The next post covers the other workloads teams already run beside their apps: jobs.