Skip to content

Transports – service-plane

Goal: decide which transport to use between two services — by environment, call shape, performance, and cost.

Service Plane offers four transports. They are interchangeable at the ability level (same tokens, same validation, same handlers), so this is purely a routing decision — and it can differ per caller/service pair.

TransportShapeStreamsWhere
cloudflareNativeRpc(binding)session (Workers RPC)yesCloudflare, same account
cloudflareServiceBindingRpc(binding) / httpBatchRpc(url)one HTTP request per call (with pipelining)noeverywhere
websocketRpc(url, { createWebSocket? })long-lived sessionyeseverywhere both ends can hold a socket
customRpcTransport(transport)whatever you bringyestests, message ports, exotic links

Apply in order; the first match wins.

  1. Same Cloudflare account → native binding RPC. Always. No public egress, streams work natively, and billing is favorable: requests through service bindings do not incur additional request fees — CPU time is billed once across the chain. This holds for streaming too.
  2. Request/response between services → HTTP-batch. The default. Stateless, retryable, observable, no connection lifecycle to manage, and promise pipelining resolves chained calls in one round trip. The price: the capability token is verified on every call (~100 µs of ES256). Cap’n Web 0.11 removed Node and Bun’s former ~1 ms batch-scheduling floor.
  3. Chatty pair or streaming → WebSocket, but only if both ends can hold the socket. A session authenticates once, then calls cost ~10 µs; streams require a session transport anyway. Which brings us to the question that decides most real cases:
End of the connectionCan hold a long-lived WebSocket?
Long-running Node / Bun / Deno processYes — and it is essentially free (one TCP socket and some memory; no per-message platform cost)
Stateless Cloudflare Worker as the callerNo. A Worker can open outbound WebSockets, but cannot persist a connection across invocations — each request would pay a fresh upgrade handshake, which is strictly worse than one HTTP-batch POST
Durable ObjectYes, but it bills duration for the whole connection. Normally the WebSocket Hibernation API would make accepted sockets cheap — but Cap’n Web sessions cannot hibernate yet (capnweb#36, open feature request): the session’s in-memory state must survive between messages, so a DO holding a Cap’n Web session — inbound or outbound — bills wall-clock duration for the entire connection
Browser / external clientYes

If either end answers “no”, use HTTP-batch (or restructure so a Durable Object owns the session).

SituationUseWhy
CF worker → CF worker, same account — any shape, including streamingcloudflareNativeRpcRule 1: free through bindings, streams natively, no public surface
CF worker → CF worker, different account, request/responsehttpBatchRpc over the public URLNo bindings across accounts; neither stateless worker can hold a socket, so per-request WebSocket = handshake + teardown every call. Both accounts bill their own requests either way
CF worker → CF worker, different account, streamingWebSocket, with a Durable Object as the caller holding the sessionSomeone must own the socket; only a DO can — and it bills duration for the whole connection (no hibernation for Cap’n Web sessions, capnweb#36). If the traffic doesn’t justify that, reconsider: same-account placement (rule 1), or request/response with batched results
Node service ↔ Node service, both long-running, frequent calls or streamingwebsocketRpcSockets are free on Node; auth amortizes to once per session (~10 µs/call) instead of verifying a token for every batch
Node ↔ Node, occasional calls (webhooks, cron fan-out)httpBatchRpcReconnect/heartbeat upkeep isn’t worth it below a few calls per second
Many stateless CF workers → one Node servicehttpBatchRpcThe callers can’t hold sockets, so a WebSocket server on the Node side gains nothing
Browser or AI session → control plane broker (interactive, streaming tools)WebSocket to /rpcLong-lived by nature. On Cloudflare, serving the socket from a plain Worker costs no duration (only CPU per message); a Durable Object adds cross-connection coordination but bills duration for the whole connection — Cap’n Web can’t hibernate (capnweb#36) — so keep sessions purposeful and close them when idle. Stock AI clients use the MCP endpoint instead
LLM token streamingany session transport + the batching recipeStreams need sessions; message count dominates cost

Doc-backed facts that drive the rules above (see Workers pricing and Durable Objects pricing for current numbers):

  • Service bindings: no additional request fees; CPU time across the chain is billed once. This is why rule 1 has no exceptions.
  • Stateless Workers: no duration billing at all. A WebSocket upgrade counts as one request; incoming WebSocket messages are billed at a favorable 20:1 ratio; outgoing messages are free. So serving WebSockets on a plain Worker is cheap — the constraint is never cost, it’s that a stateless caller can’t keep the socket.
  • Durable Objects: accept()ing a WebSocket bills duration for the entire connection unless the Hibernation API lets the object sleep between messages — and Cap’n Web sessions cannot hibernate today (capnweb#36 is an open feature request): the RPC session keeps in-memory state between messages. Until that lands, treat any DO-held Service Plane session as duration-billed for its whole lifetime, inbound or outbound. Mitigations: idle timeouts that close sessions, plain-Worker WS serving where no cross-connection state is needed, or HTTP-batch.
  • Node (self-hosted): no platform billing dimension; a WebSocket costs a file descriptor, some memory, and your reconnect/heartbeat logic.

From npm run bench (in-memory, network excluded — see Streaming for the streaming numbers):

  • Persistent session: ~10 µs per call after the one-time authenticate; within ~25% of a raw Cap’n Web session.
  • HTTP-batch: per-call token verify (~100 µs) plus batch framing. Cap’n Web 0.11 uses setImmediate on Node and Bun, removing the former ~1 ms timer floor; Service Plane mirrors that scheduler for service-binding batches.
  • Native binding: no serialization at all in-process; on Cloudflare it is also the only transport with zero public egress.
  • A reused brokered session (plane in the data path) benchmarks ~2× faster than a hand-rolled two-hop Hono chain with bearer middleware — connection reuse pays for the real crypto.
flowchart TD
A["Call another service"] --> B{"Same Cloudflare account?"}
B -- yes --> NB["cloudflareNativeRpc"]
B -- no --> C{"Streaming, or sustained chatty pair?"}
C -- no --> HB["httpBatchRpc (default)"]
C -- yes --> D{"Can BOTH ends hold a socket?<br/>(long-running process, DO, browser)"}
D -- yes --> WS["websocketRpc"]
D -- no --> E{"Worth giving the caller a Durable Object?"}
E -- yes --> WS
E -- no --> HB

Next: Streaming, Cloudflare, Node.js, and the reference.

  • TypeScript100%