AI Platform · Experiment Lab
Agent behaviour changes are experiments, not opinions
Every behaviour change ships against a holdout with a stated guardrail. An experiment that wins on its primary metric but moves a guardrail the wrong way is stopped, and we keep the record of it.
Running
3
Shipped this quarter
2
Stopped on a guardrail
1
Members under test
3,55,240
Experiments
Tap a row to read its arms and its stop rule
| Experiment | Surface | Arms | Exposure | Lift | Confidence | Status |
|---|---|---|---|---|---|---|
| Risk sentence before the answer, not after | Stoxo Desk | control / risk-first | 1,84,000 | +6.2% | 97% | running |
| Voice concierge instead of push for stalled KYC | Onboarding | push / voice / both | 41,000 | +18.4% | 99% | shipped |
| Coaching drill assigned same evening | Advisor Coaching | next-day / same-evening | 240 | +11% | 88% | running |
| Discount-free save play | Retention | discount / value-reframe | 22,000 | +4.1% | 94% | shipped |
| Vernacular creative in 6 languages | Marketing Studio | english / vernacular | 96,000 | +22.6% | 96% | running |
| Agent-written research titles | Research Factory | human / agent / hybrid | 12,000 | -2.3% | 91% | stopped |
Risk sentence before the answer, not after
Stoxo Desk · control / risk-first
- Exposure
- 1,84,000
- Lift on primary metric
- +6.2%
- Confidence
- 97%
- Guardrail
- Complaint rate flat
- Status
- running
Rules of the lab
Why the numbers here can be trusted
- A primary metric and a guardrail are declared before launch, never chosen afterwards.
- Randomisation is at member level and sticky, so a member never sees two voices.
- A permanent 2% holdout measures the whole agent programme, not just each change.
- Compliance-relevant arms are reviewed before exposure; an arm cannot test a claim we would not publish.
- Negative results stay visible. The retired discount-first save play is on this board on purpose.

