Your classifier got measurably better at the share of harmful messages flagged, holding the false-positive rate fixed: the governed lift is +0.174, certified to at least +0.120 on data the optimizer never saw.
harm caught at a fixed false-positive rate moved 0.302 → 0.476 on the held-out set; the governed lift is +0.174 (95% lower bound +0.120, measured on held-out data).
One guardrail can't be cleared yet: on short the evidence is too thin to rule out a regression beyond your tolerance. Review it before shipping.
Computed on held-out labels the optimizer and the selection never saw. One-sided 95% lower bound. Note: the committed 1% false-positive rate realizes ≈3.4% on this score axis — the Δ is measured at the frozen thresholds shown in Operating point.
harm caught at a fixed false-positive rate · base vs optimized, same rows as the certified Δ
95% intervals: base 21–48% · optimized 36–65%
Decimal 0.302 → 0.476.
verified labels (minimum 20) · 115 positive (minimum 10)
committed 1% · realized 3.4% on the candidate's negatives · τ pinned by 3 row(s)
These diagnostics do not change the verdict — review them before acting on the Δ.
Labeled-only Δ interval (no correction): [-9.5, +56.3] pts
The same held-out labels, read two ways: gold only — your verified labels on their own — and governed, which reads those same labels alongside the rest of your data. Same rows, same quantity; the governed interval is simply narrower.
Δ recall at 1% FPR · the dashed line is no change · 37 held-out labels + 220 unlabeled. Both intervals are the headline quantity on the headline slice — only λ differs.
Deploying will replace the live prompt behind your key with the certified candidate, with the review below recorded against the run. Until that ships, copy the prompt above into your own deployment.
Rollback will be one API call — POST /v1/deployments/rollback — restoring the prior prompt in under a minute. Nothing is routed or shipped in this build.
Losses shown first. A slice is flagged when the evidence can't rule out a regression worse than your tolerance (−5.0 pt). At 95% confidence, with small slices, flags are expected and are the honest state.
Interval withheld: too little data to show as evidence.
Interval withheld: too little data to show as evidence.
Why flagged: the interval reaches -19.2 pts, past tolerance. With 11 positive anchors the slice is underpowered — a data problem before it is a verdict.
What would clear it: more verified short labels would narrow the interval. Request labels — or acknowledge and ship with staged monitoring.
Search-slice estimates — the certified numbers live on the Overview. Each dot is one candidate. Across — its recall on your raw labels with a wide interval. Up — the governed estimate on those same rows plus your unlabeled pool, with a tight interval. Points below the diagonal are prompts the raw labels flatter; the winner is chosen on the governed estimate, never the raw one.
▶ Replay the search — animates the run as it happened. Useful for a walkthrough; the plot above is the same data, already settled.
Certified bound. A one-sided 95% lower bound on the improvement, anchored by your verified labels and corrected for the fact that this candidate was chosen from many. Selection happens on one split; certification is computed only on a held-out split the optimizer never touched, so there is exactly one certified number per run.
What certification requires. A run is only certified when the coverage floor is met — enough verified labels and enough positive anchors — and the bound clears zero on the held-out split. A run that misses either returns CANNOT CERTIFY YET and tells you how many more labels would settle it, rather than reporting a number it can't stand behind.
Guardrail slices. Each slice carries its own 95% interval. A slice is flagged when that interval reaches past your tolerance — small slices are expected to flag, and that is the honest state, not a verdict about the model.
Error budget. At 95%, about 1 in 20 certified bounds will be too optimistic. We audit certified runs against later ground truth and will publish the realized rate. Staged deploys exist because the other 19 don't announce themselves.
Known limits of this run. Single optimizer seed — multi-seed replication is scheduled. Underpowered slices are flagged as a measurement statement, not a verdict about the model.
The estimator itself is documented separately, in reviewed form — read the methodology docs.
| Run | e2b33ec4-519d-449a-8d7e-b975c74c0d5e · immutable · digest 67f8…d319 |
| Tokens | 5,371,512 metered |