Trust & Risk · Resilience
The interesting question is what happens when it breaks
Each surface has a written degradation contract: what it stops doing, what it keeps doing and what it tells the member. Drills are run monthly and the recovery time is published, not estimated.
Drills passed
5/5
Worst measured RTO
18m
full trace-store restore
Kill-switch time
2.1s
every tool token, org-wide
Duplicate sends after failover
0
idempotency by run id
Drills
Run one now — it exercises production paths in a controlled window
Passed — queued work replayed with no duplicate sends.
Passed — regulated surfaces queued instead of downgrading.
Passed — agents degraded to batch scores and said so in the UI.
Passed — audit continuity kept for the whole window.
Passed — one control revoked every tool token org-wide.
Degradation contracts
What a member sees when a dependency is down
Stoxo research
Answers from cache only, clearly timestamped. New questions queue rather than answer without sources.
Voice concierge
Inbound falls back to a scripted flow and an offer to call back. No improvised turns.
Support desk
Agent keeps drafting, sends nothing. A human queue takes over with the drafts attached.
Compliance guard
Publishing stops completely. Nothing member-facing ships unchecked, ever, for any reason.
Lifecycle messaging
Sends pause and resume in order. Members never receive a burst of stale messages at once.
Ad spend
Allocations freeze at their last approved state. No agent changes money during an outage.

