StockGro6

Trust & Risk · Resilience

The interesting question is what happens when it breaks

Each surface has a written degradation contract: what it stops doing, what it keeps doing and what it tells the member. Drills are run monthly and the recovery time is published, not estimated.

Drills passed

5/5

Worst measured RTO

18m

full trace-store restore

Kill-switch time

2.1s

every tool token, org-wide

Duplicate sends after failover

0

idempotency by run id

Drills

Run one now — it exercises production paths in a controlled window

Gateway region failoverRTO 42slast 9 days ago

Passed — queued work replayed with no duplicate sends.

Model provider outageRTO 8slast 16 days ago

Passed — regulated surfaces queued instead of downgrading.

Feature store cold startRTO 3m 10slast 23 days ago

Passed — agents degraded to batch scores and said so in the UI.

Full trace-store restoreRTO 18mlast 31 days ago

Passed — audit continuity kept for the whole window.

Rogue-agent kill switchRTO 2slast 4 days ago

Passed — one control revoked every tool token org-wide.

Degradation contracts

What a member sees when a dependency is down

Stoxo research

Answers from cache only, clearly timestamped. New questions queue rather than answer without sources.

Voice concierge

Inbound falls back to a scripted flow and an offer to call back. No improvised turns.

Support desk

Agent keeps drafting, sends nothing. A human queue takes over with the drafts attached.

Compliance guard

Publishing stops completely. Nothing member-facing ships unchecked, ever, for any reason.

Lifecycle messaging

Sends pause and resume in order. Members never receive a burst of stale messages at once.

Ad spend

Allocations freeze at their last approved state. No agent changes money during an outage.