Methodology
We benchmark quarterly against a 3,200-image test set balanced across the 10-point Monk Skin Tone (MST) scale, three lighting conditions (bright, mixed, low), and three age brackets. Test sources: FairFace, CelebA-Dialog, and 640 consented Autoproctor session frames used with explicit re-use consent. We report false-non-match rate (FNMR — a real face is missed), false-match rate (FMR — the wrong face is accepted), and mean absolute yaw error for gaze estimation.
Latest results
| Skin tone bucket | Samples | FNMR | FMR | Gaze error |
|---|---|---|---|---|
| MST 1–2 (lightest) | 640 | 0.70% | 0.09% | ±3.1° |
| MST 3–4 | 640 | 0.60% | 0.11% | ±3.3° |
| MST 5–6 | 640 | 0.80% | 0.12% | ±3.5° |
| MST 7–8 | 640 | 1.10% | 0.13% | ±3.7° |
| MST 9–10 (darkest) | 640 | 1.40% | 0.14% | ±4.1° |
FNMR gap between lightest and darkest buckets: 0.7 pp. Our published target is < 1.5 pp; when a quarterly run exceeds that gap we retrain before releasing to production and mark that model version as failed in our changelog.
What we do about it
- Any single-signal flag (face lost, gaze off) is never auto-terminating — it always triggers human review.
- Students can opt into Low-bandwidth mode when their connection is unstable, so poor video quality doesn't become a false violation.
- Instructors can enable Accessibility mode per exam, which uses dyslexia-friendly typography and a keyboard-first flow.
- Browser lockdown (copy/paste blocking) is off by default — it's an opt-in per exam, not a mandate.
- Every flagged attempt is appealable in-product; the appeal ships with the tamper-evident event log.
Report an accuracy issue
If our system flagged you unfairly, email fairness@autoproctor.com with your attempt reference. We respond within 2 business days and include any resulting model changes in the next quarterly report.
