AI Platform · Eval Harness
Graded cases, not demos
Every suite is a set of real StockGro cases with a graded expected outcome. Four suites gate releases: a failing gate blocks the promotion regardless of who is asking.
Graded cases
2490
Mean pass rate
96%
Gating suites
6
block promotion when red
Open regressions
1
Suites
Run any suite and the result is written into the release record
Fleet quality
Weighted across suites
96
/100
mean pass rate
7-week trend
Open regressions
- Support resolution without escalation
-1.9 pts on payment-failure intents
How a case is graded
Four grader types, none of them a thumbs-up
Deterministic
Refusals, PII, tool arguments and format rules are checked in code. Pass or fail, no opinion.
Source-grounded
Every claim in a research answer must resolve to a retrieved passage id.
Rubric-scored
A grader model scores against a written rubric, sampled and audited by the owner weekly.
Human-labelled
A rotating panel labels 40 cases a week; disagreement with the grader becomes a new case.

