Reference – service-plane
Goal: quickly look up the main Service Plane API pieces and wire shapes.
For a guided walkthrough, start with Create A Service and Create A Control Plane.
Ability Definition
Section titled “Ability Definition”defineAbility({ id: 'asana.tasks', title: 'Asana Tasks', description: 'Task operations for Asana', exposure: 'private' | 'published', access: 'plane' | 'service', scopes: ['asana.tasks.write'], methods: { createTask: abilityMethod({ input, output, scopes: ['asana.tasks.write'], rest: { method: 'post', path: '/asana/tasks' }, mcp: { name: 'asana_create_task', description: 'Create a task in Asana' }, }), }, rpc: { path: '/rpc/asana.tasks', transports: ['http-batch', 'websocket'], }, handler: ({ context, identity }) => new AsanaTasksHandler(context.env, identity),});Defaults:
exposure: 'private'access: 'plane'rpc.path: /rpc/<abilityId>rpc.transports: ['http-batch']
access: 'plane' means the control plane or gateway owns any upstream product auth decision before calling the service. access: 'service' restricts the ability to authenticated service callers, and is enforced at both ends: the broker refuses it for a non-service caller from the discovered catalog, and the service refuses it from its own definition using the token’s spa claim. Tightening an ability therefore takes effect when the service deploys, not when the plane’s discovery cache catches up.
Ability Method
Section titled “Ability Method”abilityMethod({ input: z.object({ name: z.string() }), output: z.object({ id: z.string() }), scopes: ['asana.tasks.write'],});Each method accepts one input object and returns one output value. The wrapper validates both.
input and output accept any Standard Schema value that also implements Standard JSON Schema — ArkType 2.1.28+, Valibot 1.2+ via @valibot/to-json-schema, VineJS 4.3+, Zod 4.2+, or anything else meeting both contracts. The choice is per schema: one method may take its input from one library and its output from another. See Choosing A Validation Library.
Validation failures raise AbilityValidationError carrying the issues the schema library reported. See Errors.
Optional rest metadata projects the method into the generated OpenAPI 3.2 document and mounts the
live control-plane route; rest.method accepts get, post, put, patch, delete, and query
(HTTP QUERY per RFC 10008 — request parameters in the body, safe and idempotent). rest.status
declares the successful 2xx response and defaults to 200. See OpenAPI and
MCP.
Optional idempotent: true declares that calling the method again with the same input cannot double its effect, so a caller may safely retry an ambiguous failure. See Idempotency.
Streaming Methods
Section titled “Streaming Methods”Some methods produce many results over time — large file transfers, long exports. Declare them with stream: true; the output schema then validates each streamed item, and the handler returns an async iterable (usually an async generator), a sync iterable, or a ReadableStream:
abilityMethod({ input: z.object({ path: z.string() }), output: z.object({ chunk: z.string() }), // validates each streamed item scopes: ['hub.files.read'], stream: true,});There is no custom wire protocol: the wrapper returns the items as a native Cap’n Web ReadableStream with built-in flow control, so callers receive them exactly like any other RPC value:
const api = await abilitySession<AbilityRpc<typeof hubFiles>>({ ... });const stream = await api.readFile({ path: '/big.bin' }); // ReadableStream<{ chunk: string }>for await (const item of stream) { // ...}Cap’n Web streams ride the ongoing session, so streaming methods require a session transport: WebSocket (websocketRpc), the Cloudflare native binding (cloudflareNativeRpc), or a custom bidirectional transport. The one-round-trip HTTP-batch transport cannot carry them — calling a streaming method over HTTP-batch fails with a 405, and an ability that declares streaming methods must enable websocket or cloudflare-binding-rpc in rpc.transports (checked at setup). Unary methods on the same ability keep working over HTTP-batch.
Through the broker, streams proxy transparently: connect to /rpc over WebSocket, and the plane reaches the service over its own session transport — preferring the endpoint’s native ability RPC binding (ServiceEndpoint.abilityRpc, set explicitly via cloudflareServiceBinding({ abilityRpc }) — a Workers stub answers any property with a callable proxy, so it cannot be detected), then WebSocket. When the caller’s own leg cannot carry a stream (HTTP-batch), the ability’s streaming methods are rejected with a 405 and the plane leg stays on HTTP-batch — no socket is opened for a stream that could never be returned. Streaming methods cannot project MCP prompts, resources, or REST operations (single-response surfaces); MCP tools are supported.
For high-frequency streams (LLM token deltas), batch deltas in the handler and declare the batch as the item (output: z.array(...)) — see the coalescing recipe in Streaming.
Full guide, including per-runtime WebSocket wiring and performance guidance: Streaming.
Discovery Document
Section titled “Discovery Document”type ServiceDiscoveryDocument = { id: string; title: string; version: string; capabilities?: CapabilityCatalog; abilities: ServiceAbilityDiscovery[];};Ability discovery includes exposure, access, scopes, RPC path, transports, method names, method scopes, JSON Schemas, optional REST metadata, and optional MCP metadata. Streaming methods carry stream: true, with their outputSchema describing one streamed item.
The control-plane registry accepts a discovery document only when its id, and its optional capabilities.serviceId, match the configured endpoint id. The endpoint configuration is the identity authority; a service cannot publish metadata for another configured service.
Service
Section titled “Service”new ServicePlaneService({ id, title, version, auth, ingress, capabilities, abilities,});Mounted routes:
GET /.well-known/service-plane/service.jsonALL /rpc/<abilityId>ingress is optional. When configured, ability RPC routes require a capability token with a signed broker claim from the configured control-plane service id. Non-brokered tokens are rejected before handler execution.
httpCache is optional. When set (true or { maxAgeSeconds, staleWhileRevalidateSeconds, tags }), the discovery route emits Cache-Control and Cache-Tag headers so an edge cache (e.g. Cloudflare Workers Cache) can serve it without executing the Worker. See Cloudflare.
Control Plane
Section titled “Control Plane”new ServicePlaneControlPlane({ signingKeys, authenticateCaller, invocationMiddleware, services, openapi, rpc, mcp,});Mounted routes:
POST /.well-known/service-plane/capability-tokenGET /.well-known/service-plane/jwks.jsonGET /openapi.jsonPOST /mcp (when published MCP projections exist)ALL /rpc (only when `rpc` is configured)* <published rest.path> (REST facade)The plane serves the OpenAPI document and mounts every published non-streaming REST projection as a
live route. Mount a documentation UI yourself on plane.app (e.g. @hono/swagger-ui or
@scalar/hono-api-reference) pointed at /openapi.json.
REST input is assembled as query, then JSON body, then path parameters. The generated operation
removes path fields from its body schema, represents them as required path parameters, and exposes
top-level string/string-array fields as optional query fallbacks. The service validates the combined
value. Every {name} in the route must match a top-level input field; inconsistent definitions or
discovery documents are rejected, and empty request segments do not match variables.
openapi.security and openapi.securitySchemes describe the public authentication performed
by invocationMiddleware; neither is invented by default because middleware may use any scheme or
an explicit anonymous caller.
The top-level rpc option controls the control plane’s public Cap’n Web broker route. rpc: {}
mounts it at /rpc; rpc: { path: '/custom' } overrides that path, and omitting rpc (or setting
rpc: false) leaves it unmounted. This is separate from an ability’s rpc metadata, which describes
how the control plane reaches that ability on its owning service.
The JWKS route is served from the signing authority (signingKeys, issuer) and never
resolves services or fetches discovery documents, so key publication survives a service-discovery
outage. The capability-token, REST, broker, and MCP routes additionally need the authorization catalog
(discovered capabilities and grants) and fail closed when it cannot be built. See
auth.md.
httpCache is optional and mirrors the service option: when set, the OpenAPI and JWKS routes emit Cache-Control and Cache-Tag headers. The capability-token endpoint always responds with Cache-Control: no-store and Pragma: no-cache. Broker and MCP RPC responses are never cache-eligible.
discoveryCache caches the discovered service catalog for every route that needs it — token issuance, REST, broker, MCP, and OpenAPI. It defaults to a process-local cache; pass a RegistryCache to share one across a fleet, false to resolve fresh every time, or an object keyed by token (issuance, REST, broker and MCP), openapi, and default to give either path its own store. openapi.cache is separate and caches the generated document rather than the catalog behind it; set its TTL with openapi.cacheTtlSeconds.
services(context) resolves runtime bindings and deployment configuration for one logical service
catalog. Its endpoint set and discovery metadata must not vary by caller or organization. Services
own organization-specific data scoping behind their stable ability definitions; applications that
need genuinely different catalogs should use separate control-plane instances and discovery caches.
The shared top-level invocationMiddleware is a Hono MiddlewareHandler that runs only for a
matched published REST route, an available MCP endpoint, or the enabled broker. Before it calls
next(), it must set servicePlaneCaller to a BrokerCaller; it may also set
servicePlaneConnInfo. It can short-circuit with an application-owned response, preserving a 401
and its WWW-Authenticate challenge. Calling next() without a caller is a configuration error and
returns 500.
Global middleware on a supplied Hono app may set the same variables instead.
servicePlaneInvocation is available for policy and auditing: REST sets resolved service, ability,
method, scopes, and path before invocation middleware runs; MCP enriches the surface during protocol
dispatch; broker HTTP sessions expose the broker surface because later RPC calls can select multiple
abilities.
The plane can open a typed, disposable ability session for trusted code running in the same process:
type ControlPlaneAbilitySessionOptions = { abilityId: string; caller?: { id: string; kind: 'service' | 'user'; orgId?: string; principalKind?: string }; connInfo?: ConnInfo; idempotencyKey?: string; requestId?: string; scopes: string[]; targetServiceId: string; timeoutMs?: number;};
const api = await plane.abilitySession<AsanaTasksApi>(options, bindings);caller: { kind: 'user', ... } creates an RFC 8693 delegated subject and stamps
callerAccess: 'plane'; kind: 'service' preserves the service id and stamps
callerAccess: 'service'; omitting caller
creates a plane-class call under controlPlaneServiceId without a subject. This is a trusted API,
not an authentication boundary: application code must authenticate a user or service before passing
that identity. The method reuses the plane catalog, grants, signing material, discovery cache,
broker authorization, and transport selection. It automatically mints a brokered token for a target
that advertises required ingress. timeoutMs includes catalog resolution and token issuance.
The returned AbilitySession<Scoped> supports streaming and must be disposed with using or
disposeAbilitySession(). Caller-facing capability-token endpoints still reject subject.
mcp.streamLimits accepts maxItems and maxBytes for streaming tools (defaults: 10,000 items and 1 MiB). maxBytes independently caps serialized item aggregation and cumulative optional progress-notification bytes. Exhausting the item aggregation budget fails the tool call in-band; exhausting only the progress budget stops further notifications while the bounded final result continues.
The MCP endpoint accepts protocol revisions 2025-11-25, 2025-06-18, and 2025-03-26; missing
MCP-Protocol-Version means 2025-03-26, while unsupported values return 400. Incoming browser
Origin headers must match the endpoint origin. mcp.allowedOrigins adds exact trusted origins for
intentional cross-origin clients; other origins return 403 before invocation middleware.
Caller
Section titled “Caller”const api = await abilitySession<AbilityRpc<typeof asanaTasks>>({ abilityId: 'asana.tasks', callerServiceId: 'workflow-runner', targetServiceId: 'asana', scopes: ['asana.tasks.write'], requestToken, transport,});Persistent WebSocket/custom sessions and native binding targets are disposable. Prefer using so
the transport closes at the end of the block:
{ using api = await abilitySession<AbilityRpc<typeof asanaTasks>>({ ... }); await api.createTask(input);}Otherwise, call await disposeAbilitySession(api) from finally. Disposal is idempotent. For a
stateless transport there is no connection to release, but disposal still permanently closes the
session object so accidental reuse fails consistently.
Transports:
cloudflareServiceBindingRpc(binding)cloudflareNativeRpc(binding)httpBatchRpc(url)websocketRpc(url, { createWebSocket? })— the optional factory receives the final URL afterrequest_idpropagation, allowing Node runtimes without a globalWebSocketto inject a standards-compatible client without requiring the application to install a persistent global; the compatibility path uses a temporary synchronousWebSocket.CONNECTINGshim that is restored immediately.customRpcTransport(transport)
cloudflareNativeRpc(...) can call ingress-protected services only with brokered capability tokens. Normal direct caller tokens are rejected.
Control-plane endpoints may additionally provide ServiceEndpoint.createWebSocket. Configure it
with httpsService({ createWebSocket }) (or cloudflareServiceBinding({ createWebSocket })) so
broker and MCP calls can reach WebSocket-only abilities on runtimes without a global client.
ServiceEndpoint.abilityRpc is likewise explicit: pass cloudflareServiceBinding({ abilityRpc: env.ASANA })
for a binding whose target forwards connectAbility(...). It is never inferred from the binding,
because a Workers service-binding stub returns a callable RPC proxy for every property name.
Which transport fits which pair of services — by environment, performance, and cost — is covered in Choosing A Transport.
Token requesters:
controlPlaneRpcTokenRequester(...)controlPlaneJwkTokenRequester(...)controlPlaneHmacTokenRequester(...)
Capability Token Claims
Section titled “Capability Token Claims”Capability tokens are ES256 JWS tokens with a closed claim set. Unknown claims are dropped at verification.
Tokens come in two shapes, and sub always answers the same question: who is this token about. A plain service-to-service token is about the calling service. A delegated token uses RFC 8693’s act actor-claim semantics: it is about the plane-class principal, while the calling service moves into act.sub. The presence of act is what switches the interpretation, and the verifier resolves it for you: identity.serviceId is always the calling service, and identity.subject is set only when a principal is delegated.
Plain service token:
{ "iss": "control-plane", "sub": "workflow-runner", "aud": "asana", "scp": ["asana.tasks.write"], "spa": "service" }→ identity.serviceId = 'workflow-runner', no identity.subject.
Delegated (plane-principal) token:
{ "iss": "control-plane", "sub": "key-123", "act": { "sub": "control-plane" }, "spk": "api-key", "spo": "org-42", "aud": "asana", "scp": ["asana.tasks.write"], "spa": "plane" }→ identity.serviceId = 'control-plane' (from act.sub), identity.subject = { id: 'key-123', kind: 'api-key', orgId: 'org-42' }.
| Claim | Plain service token | Delegated token (act present) |
|---|---|---|
sub | calling service → identity.serviceId | delegated principal → identity.subject.id |
act | absent | acting service, { sub } → identity.serviceId |
spk | rejected at verification | optional principal kind → identity.subject.kind |
spo | rejected at verification | subject’s org → identity.subject.orgId |
iss | control-plane issuer → identity.issuer | same |
aud | target service id → identity.audience | same |
scp | granted scopes → identity.scopes | same |
spa | caller access class → identity.callerAccess; 'service' in the example above, 'plane' when the plane calls without a service caller (e.g. an anonymous broker) | always 'plane' — a delegated subject is a fronted caller, and the issuer refuses the other pairing |
spb | broker service id on brokered (ingress) tokens → identity.brokerServiceId | same |
cnf | { jkt } on tokens bound to a caller key (always, for JWK callers) → identity.confirmation, only after a matching proof verified | same |
jti | token id → identity.tokenId | same |
exp | expiry → identity.expiresAt; iat/nbf are also enforced | same |
The act delegation relationship comes from RFC 8693 and cnf from RFC 7800 (with the jkt confirmation method registered by RFC 9449). scp, spa, spk, spo, and spb are Service Plane-specific claims, and /.well-known/service-plane/capability-token is the package’s JSON capability endpoint, not an RFC 8693 token-exchange endpoint. spk is an optional application-owned string; its absence preserves the legacy user-subject shape, and it never influences spa or service access.
spa is the access class the control plane authenticated for the caller. It is service for a caller the plane proved to be another service — the capability-token endpoint, issueCapabilityTokenForCaller, and invocation middleware setting kind: 'service' — and plane for every caller the plane fronts itself: users, API keys, anonymous traffic. Services compare it against the ability’s own access and reject a mismatch with 403 before the handler is created. A token carrying no spa reads as plane, so a control plane that predates the claim can only reach access: 'plane' abilities.
That default dictates the rollout order: upgrade the control plane before any service declares access: 'service'. A service on this version behind an older plane refuses every caller of its service-only abilities — legitimate service callers included — until the plane mints the claim. The reverse mix is the transitional gap, not a hole in the new guarantee: a service still on an older package version never checks spa, so for that service tightening access keeps depending on the plane’s catalog refresh until the service upgrades.
Delegated subjects are minted only by control-plane code — ServicePlaneControlPlane.abilitySession(), invocation middleware setting a BrokerCaller with kind: 'user' and optional orgId / principalKind, or a low-level direct issueCapabilityToken({ subject, ... }) call. The capability-token endpoint and issueCapabilityTokenForCaller reject caller-supplied subjects with 403, and the shipped token requesters fail fast locally instead of transmitting one. Direct issue mints a non-brokered token; abilitySession() and the broker select issueBrokeredCapabilityToken automatically for ingress-required targets. See auth.
Logging And Request Correlation
Section titled “Logging And Request Correlation”Every request that enters a ServicePlaneControlPlane gets an X-Request-Id (incoming header value or a generated UUID, via hono/request-id). The REST, broker, and MCP endpoints forward that id on every outbound call to a service: as the X-Request-Id header for HTTP-batch and service-binding transports, as the request_id query parameter for WebSocket transports (SERVICE_PLANE_REQUEST_ID_QUERY_PARAM), and as the requestId field on connectAbility(...) for Cloudflare native RPC. ServicePlaneService adopts the propagated id into its own requestId context variable and echoes it on responses, so one id correlates plane and service logs end to end.
Connection info about the original client rides the same three channels when middleware sets servicePlaneConnInfo: the X-Service-Plane-Conn-Info header, the conn_info query parameter (SERVICE_PLANE_CONN_INFO_QUERY_PARAM), and the connInfo field on connectAbility(...). Services expose it to handlers as connInfo only for brokered calls with ingress enabled — see Forwarded Connection Info.
Deadlines
Section titled “Deadlines”Two bounds, layered. A service-side ceiling that always exists, and an end-to-end budget a caller may set on top of it. Whichever expires first wins.
The ceiling you get for free
Section titled “The ceiling you get for free”Every unary ability method is bounded at DEFAULT_ABILITY_TIMEOUT_MS — 10 seconds — without anyone configuring anything.
That default is deliberate. gRPC and Connect leave deadlines entirely to the caller, and the standing advice in gRPC’s own guidance is to “always set a deadline” — a rule that only needs stating because the unset case is unbounded. Systems that own a default do not need the reminder: Envoy routes time out at 15s, and Armeria’s server request timeout is 10s. 10s matches the closest analogue — a server bounding its own request handling.
Tune it where it belongs:
new ServicePlaneService({ timeout: { methodMs: 2_500 }, // service-wide ceiling; `false` removes it});
bigExport: abilityMethod({ timeoutMs: 120_000, ... }); // the one slow methodbigMigration: abilityMethod({ timeoutMs: 0, ... }); // opt this one out entirelyMethod values are validated at definition time — a negative, fractional, or absurdly large value refuses the service instead of silently dropping or clamping the ceiling — and a method’s own ceiling is deliberately not clamped to the 10-minute wire limit: that limit bounds what a caller may ask for, not how long a service allows its own export to run.
Raise the exception, not the ceiling. Streaming methods are never bounded this way — for the reason Envoy documents about its own route timeout, a bound that suits a request is wrong for a stream. Session lifetime is untouched either way.
The effective ceiling is advertised per method in the discovery document, so a gateway can size its own wait against it.
The budget a caller sets
Section titled “The budget a caller sets”A caller states how long it is willing to wait; every hop spends from that budget rather than granting a new one.
const api = await abilitySession<AbilityRpc<typeof syncAbility>>({ // ... timeoutMs: 5_000,});The value travels on the same three channels as the request id: the X-Service-Plane-Timeout header, the timeout query parameter (SERVICE_PLANE_TIMEOUT_QUERY_PARAM), and the timeoutMs field on connectAbility(...). It is relative milliseconds remaining, not an absolute timestamp — two clocks that disagree would shift an absolute deadline by the whole skew, and this package already assumes clocks can differ. Each hop measures its own elapsed time on its own clock and forwards what is left, which is the trade grpc-timeout makes for the same reason.
What each participant does with it:
- The caller bounds its own wait per method call and rejects with
ServicePlaneTimeoutErrorwhen the budget elapses. Cap’n Web has no cancel message, so this frees the caller, not the callee. - The control plane reads an inbound
X-Service-Plane-Timeouton REST, broker, and MCP requests and forwards what is left after its own work — resolving the catalog, minting a token. If nothing is left, the invocation fails before a service session is opened. - The service turns it into the
signalits ability handlers receive, and the validating wrapper fails the method if the handler outlives it. A handler that ignoressignaltherefore loses the work, not correctness.
handler: ({ signal }) => new MyApi(signal), // pass it to outbound fetch, long loops, DB callsChains
Section titled “Chains”A budget only survives a chain if each service passes on what is left of its own. Handlers get remainingTimeoutMs() for exactly that:
handler: ({ remainingTimeoutMs, signal }) => ({ async run(input) { const downstream = await abilitySession({ ...opts, timeoutMs: remainingTimeoutMs?.() }); return downstream.doWork(input); },});Skip it and the next hop starts a fresh budget: A(5s) → B where B calls C with its own 5s means the end-to-end bound A asked for is gone. Nothing enforces this for you — a service that calls onward has to opt in.
At exhaustion the pattern stays safe: remainingTimeoutMs() returns 0 once the budget is gone, and a session opened with timeoutMs: 0 fails every call immediately with a timeout error instead of running unbounded — the same fail-fast the broker applies before opening a service leg.
What This Does And Does Not Bound
Section titled “What This Does And Does Not Bound”Three of the four mechanisms are plain timers over a duration, so they do not read a clock and cannot drift:
- the caller’s own wait (
setTimeout), - the
signalhanded to handlers (AbortSignal.timeout), - the wrapper’s refusal to resolve a method past the deadline.
Only the plane’s decrement does clock arithmetic — Date.now() at request entry versus at the moment it opens the service leg. Both readings are on the same machine, so there is no cross-host skew to worry about.
It does not:
- cancel the peer. Cap’n Web has no cancel message. A caller-side timeout frees the caller; the service keeps running until its own budget expires. The forwarded budget is what actually stops work.
- close a session. The deadline fails a method. A WebSocket session stays open, so on Cloudflare a Durable Object holding one keeps billing duration — see Transports. Use an idle timeout to bound that, not a deadline.
- bound a stream’s lifetime. It bounds the call that returns the stream, not consumption of its items.
On Cloudflare
Section titled “On Cloudflare”Workers freeze Date.now() during synchronous execution and advance it on I/O (a Spectre mitigation). That suits this design rather than breaking it: the plane’s decrement measures waiting — the discovery fan-out, the token mint — and waiting is I/O, which is exactly when the clock moves. What stays invisible is pure CPU time, which the Workers CPU limit already bounds and which is small next to a network hop. The effect is that a plane’s decrement can slightly under-count, never over-count, so a service is handed a budget that is generous rather than short.
Two caveats worth stating:
- The runtime matrix in #11 does not run yet, so the above reflects documented workerd behavior, not a test result on workerd.
- If Cap’n Web ever gains WebSocket Hibernation (capnweb#36), a hibernating Durable Object would lose the in-memory
AbortSignal.timeoutbehind a session-scoped deadline, and it would silently never fire on wake. Today Cap’n Web sessions cannot hibernate, so this is not reachable — but a deadline set before hibernation is not something to assume survives it.
Values are clamped to MAX_SERVICE_PLANE_TIMEOUT_MS (10 minutes) and anything that is not a positive integer count of milliseconds is ignored. A caller that sends nothing forwards nothing — the service-side ceiling above is what still bounds the call.
Both shells take a policy for what they will accept:
new ServicePlaneControlPlane({ timeout: { defaultMs: 10_000, maxMs: 60_000 } });new ServicePlaneService({ timeout: { defaultMs: 5_000, maxMs: 30_000 } });defaultMs supplies a budget when the caller sent none — on per-call transports (HTTP-batch) only. A session transport (WebSocket, native binding) resolves its budget once at session open, so a manufactured default would become a death timer for long-lived sessions whose callers never asked for one; an explicit caller budget on a session transport still applies. maxMs clamps any budget — explicit or defaulted — that asks for more than you are willing to hold a connection for. Invalid policy values (0, negatives, fractions) are refused at construction rather than silently loosening at runtime. The plane has no built-in default on purpose — it forwards a budget rather than doing the work, so the bound that must always exist lives at the service. Set defaultMs when you want the plane to be the policy point, the role Envoy’s route timeout plays.
A caller’s own local wait is set slightly above the budget it forwards (SERVICE_PLANE_TIMEOUT_GRACE_MS, 250ms). Armeria does the same thing — its client response timeout of 15s sits above its 10s server request timeout — so that the service’s own enforcement fires first and the caller gets the error the service actually raised instead of a bare local abort that says nothing about what happened downstream.
When a deadline fires
Section titled “When a deadline fires”| This package | gRPC | Envoy | Armeria | |
|---|---|---|---|---|
| Caller sees | code: 'timeout', status: 504 inside the RPC payload — the HTTP response is 200 | DEADLINE_EXCEEDED (maps to 504) | 504 Gateway Timeout | ResponseTimeoutException |
| Service sees | The method rejects; signal is aborted | Context cancelled (CANCELLED) | Upstream stream reset | RequestTimeoutException, work cancelled |
| Peer is told | No | Yes | Yes | Yes (RST_STREAM / close) |
The last row is the honest gap: Cap’n Web has no cancel message, so a caller giving up cannot tell the service. That is why the budget is forwarded rather than relied on locally — the service’s own copy is what stops the work. Everyone else in that table can signal the peer; we compensate by making the service-side bound the one that always exists.
retryable is true for a timeout, matching Envoy’s treatment of 504 as a gateway-error worth retrying — but only retry when the method is also idempotent. See Idempotency.
status is a classification, not an HTTP status code. A method’s failure is a value inside the Cap’n Web batch, so the HTTP response is 200 and the error travels in its body. Read the classification with servicePlaneErrorInfo; do not expect to see 504 on the wire. The number matters when a gateway maps the failure onto its own response. On the MCP surface even an exhausted forwarding budget stays inside the protocol: the refusal is a JSON-RPC-framed tool failure the client can correlate, never a bare HTTP error body.
Unlike forwarded connection info, a deadline is honoured from any caller without requiring ingress. It is not an authorization input: a caller shortening its own budget can only cut itself off, and a long one is clamped.
One limit worth knowing: the budget rides the transport, and a session transport is established once. Over HTTP-batch and native bindings a session is one call, so the budget is per call. Over WebSocket it is fixed when the socket opens and therefore bounds every call on that session.
Idempotency
Section titled “Idempotency”Deadlines create ambiguous failures — a call that timed out may or may not have run — so a caller needs two things to retry correctly: whether the method is safe to call again, and a way for the service to recognize the retry.
The method says whether it is safe. Mark it in the ability definition, the same way stream is marked:
lookupTask: abilityMethod({ idempotent: true, input: TaskQuery, output: Task, scopes: ['asana.tasks.read'],});It is projected into the discovery document so callers and gateways can read it. An unmarked method is absent from the projection rather than false: it makes no claim, which is the safe reading. Note that this package never retries on its own — retry policy is the caller’s, and mesh-level retry belongs to your platform.
Combined with retryable from the error taxonomy, the decision is: retry only when the failure was transient and the method is idempotent.
The caller says which attempt this is. Pass a key and it travels the same three channels as the request id — X-Service-Plane-Idempotency-Key, the idempotency_key query parameter, and the idempotencyKey field on connectAbility(...) — reaching the handler as idempotencyKey:
const api = await abilitySession({ /* ... */ idempotencyKey: 'attempt-7f3a' });
// service sidehandler: ({ idempotencyKey }) => new TaskApi(idempotencyKey);The package forwards the key and nothing else. Deduplicating means storing a result and expiring it, which needs a store and a retention policy — the same reason discovery snapshots and token caches are yours to supply.
Two things to get right when you build that store:
- Scope the key by method name. The key identifies the caller’s attempt, not one method call, because it rides the transport rather than the RPC payload. Two different methods on one session would otherwise collide. Store under
${idempotencyKey}:${methodName}. - Keys are validated on both send and receive: word characters,
-, and=only, up to 255 characters. Anything else is dropped rather than forwarded, so a key can never smuggle a separator into a log line or a store key.
Both shells log structured JSON events to the console by default. Every event carries event, level, and (when known) requestId.
Service events (ServicePlaneLogEvent):
service_plane.discovery.servedservice_plane.request.completedservice_plane.request.failedservice_plane.ability.handler_failed— a handler throw the wrapper replaced with an opaque error; carries the original name and message
Control-plane events:
service_plane.broker.connect.completed/service_plane.broker.connect.failed(ServicePlaneBrokerLogEvent)service_plane.mcp.tool.completed/service_plane.mcp.tool.failed(ServicePlaneBrokerLogEvent)service_plane.mcp.resource.completed/service_plane.mcp.resource.failed(ServicePlaneBrokerLogEvent)service_plane.mcp.prompt.completed/service_plane.mcp.prompt.failed(ServicePlaneBrokerLogEvent)service_plane.rest.completed/service_plane.rest.failed(ServicePlaneBrokerLogEvent)service_plane.caller_auth.not_configured(ServicePlaneControlPlaneLogEvent)service_plane.caller_auth.hmac_unauthorized/service_plane.caller_auth.jwk_unauthorized(caller-auth middleware, ownlogoption). Thereasonfield names the check that failed.
Where the events go is up to the app. Each surface takes a log callback that is invoked once per event; when it is omitted, the package writes the event as one JSON line to the console. The package never talks to a logging framework itself — you forward events to whatever logger the app uses:
new ServicePlaneService({ // ... logger: { log: (event, context) => appLogger.info(event) }, // or false to disable request logging requestId: { generator: myIdGenerator }, // customize hono/request-id; the middleware itself is always on});
new ServicePlaneControlPlane({ // ... log: (event, context) => appLogger.info(event), // or false to silence REST/broker/MCP/config events});The log callback receives the Hono Context as a second argument when the event was emitted inside a request, so a request-scoped logger stored on the context by your own Hono middleware (e.g. c.set('logger', child)) is reachable from it. On the service, middleware mounted via the middleware option can also read the emitted events after await next() with servicePlaneLogEvents(context) — useful when you prefer to do all log shipping in one place in your own middleware.
Caches
Section titled “Caches”Use separate caches for:
- service discovery snapshots
- generated OpenAPI document
- control-plane JWKS fetched by services
- caller capability tokens
Token cache keys include caller id, target service id, ability id, normalized scopes, optional TTL, and the complete delegated subject when present — including principal kind — so tokens cannot collide across principals or principal categories.
Errors
Section titled “Errors”- Missing or invalid token:
CapabilityAuthErrorwith 401-style status. - Missing scope:
CapabilityAuthErrorwith 403-style status. - Invalid caller input:
AbilityValidationErrorwith 422-style status. - Invalid service output or streamed item:
AbilityValidationErrorwith 500-style status — the handler broke its own declared contract. - Deadline elapsed:
ServicePlaneTimeoutErrorwith 504-style status, thrown by whichever hop notices first. See Deadlines.
AbilityValidationError.issues carries the schema library’s issues as { message, path? } entries, so a gateway can build a field-level response without parsing the joined message:
import { AbilityValidationError } from 'service-plane/service';
try { await asana.createTask(input);} catch (error) { if (error instanceof AbilityValidationError) { return Response.json({ errors: error.issues }, { status: error.status }); } throw error;}Reading An Error A Caller Received
Section titled “Reading An Error A Caller Received”instanceof works in-process, but not on an error that arrived over RPC. Cap’n Web rebuilds a received error as a plain Error: its class table holds only built-in error types, and the sent class name is used to choose from that table rather than restored onto the result. Own enumerable properties do survive, which is why the taxonomy lives in code, status, and retryable. Read them with servicePlaneErrorInfo, which works for both a local instance and a received one:
import { servicePlaneErrorInfo } from 'service-plane/service';
const info = servicePlaneErrorInfo(error);if (info?.retryable) return retryLater();if (info?.code === 'capability_auth') return refreshTokenAndRetry();code | Meaning |
|---|---|
capability_auth | Token, scope, ingress, or proof-of-possession check refused the call |
ability_validation | Input or output did not satisfy the method’s schema |
timeout | The caller’s deadline elapsed |
handler | The handler failed deliberately and chose what the caller sees |
internal | Anything else, including a handler failure the service did not shape |
retryable means the failure is transient — the same call may succeed later. It does not mean retrying is safe: for a non-idempotent method a retry can still double an effect. It defaults from the status (408, 429, 502, 503, 504) and can be set explicitly. A 500 is deliberately not retryable by default: a handler that broke once usually breaks again, and saying otherwise invites a retry storm against a service already failing.
Every field is re-validated when read, so a hostile or buggy peer cannot make a refusal look retryable.
What An Ability Handler May Throw
Section titled “What An Ability Handler May Throw”Errors this package raises are already shaped for callers and pass through untouched. Everything else a handler throws is replaced with an opaque 500 before it leaves the service:
Service-Plane ability handler failed: <methodName>That is deliberate. A database driver error or a TypeError was written for an operator, not a caller, and routinely carries connection strings, internal hostnames, SQL, or row data. The same replacement applies to a streaming method that fails mid-stream.
The replacement also holds across chains: an error that already carries the taxonomy — thrown by a downstream service and rebuilt as a plain Error on the way through — passes intermediate hops untouched instead of being re-replaced, so the original code/status/retryable/reason reach the first caller.
Every replacement is logged service-side as a service_plane.ability.handler_failed event carrying the original error’s name and message (the RPC response is a 200 batch, so request.failed never fires for it). The original object also stays reachable in-process via handlerFailureCause(error).
To choose what the caller sees, throw AbilityHandlerError:
import { AbilityHandlerError } from 'service-plane/service';
throw new AbilityHandlerError('Monthly export quota is used up', { reason: 'quota_exhausted', // your own discriminator, carried alongside code: 'handler' retryable: false, status: 429,});The original failure is not lost — it is held beside the replacement, reachable in-process with handlerFailureCause(error) so a service can log it. It is deliberately not attached as cause: Cap’n Web serializes cause unconditionally, which would defeat the replacement.
A schema that deviates from the Standard Schema contract fails closed: a validator that throws, or returns neither a value nor issues, raises AbilityValidationError rather than letting the value through. A schema missing ~standard.validate or ~standard.jsonSchema is rejected when the service is defined, not on the first call.
Next: auth, OpenAPI and MCP, and Cloudflare.