StockGro6

AI Platform · Eval Harness

Graded cases, not demos

Every suite is a set of real StockGro cases with a graded expected outcome. Four suites gate releases: a failing gate blocks the promotion regardless of who is asking.

Graded cases

2490

Mean pass rate

96%

Gating suites

6

block promotion when red

Open regressions

1

Suites

Run any suite and the result is written into the release record

Research accuracy & citation integritygate94.3%
420 casesowner Ananya Raolast run 42 min ago
Refusal boundary (tips, guarantees, SME calls)gate99.2%
260 casesowner Ritika Shahlast run 42 min ago
SEBI advertisement code adherencegate97.7%
310 casesowner Ritika Shahlast run 2 hours ago
Support resolution without escalation88.1%
500 casesowner Kabir Menonlast run 1 hour ago-1.9 pts on payment-failure intents
Hinglish and regional comprehension91.4%
380 casesowner Dev Sharmalast run 3 hours ago
Voice barge-in and latency budgetgate95.6%
180 casesowner Dev Sharmalast run 26 min ago
Tool-call argument validitygate98.8%
240 casesowner Meera Iyerlast run 18 min ago
PII leakage in generated copygate100%
200 casesowner Meera Iyerlast run 18 min ago

Fleet quality

Weighted across suites

96

/100

mean pass rate

7-week trend

Open regressions

  • Support resolution without escalation
    -1.9 pts on payment-failure intents

How a case is graded

Four grader types, none of them a thumbs-up

Deterministic

Refusals, PII, tool arguments and format rules are checked in code. Pass or fail, no opinion.

Source-grounded

Every claim in a research answer must resolve to a retrieved passage id.

Rubric-scored

A grader model scores against a written rubric, sampled and audited by the owner weekly.

Human-labelled

A rotating panel labels 40 cases a week; disagreement with the grader becomes a new case.