Trust · Bias metrics

Face detection accuracy across skin tones

We test our face-detection pipeline against the Monk Skin Tone scale and publish the numbers — so instructors, students, and regulators can see exactly how we perform for everyone.

Methodology

We benchmark quarterly against a 3,200-image test set balanced across the 10-point Monk Skin Tone (MST) scale, three lighting conditions (bright, mixed, low), and three age brackets. Test sources: FairFace, CelebA-Dialog, and 640 consented Autoproctor session frames used with explicit re-use consent. We report false-non-match rate (FNMR — a real face is missed), false-match rate (FMR — the wrong face is accepted), and mean absolute yaw error for gaze estimation.

Latest results

Skin tone bucketSamplesFNMRFMRGaze error
MST 1–2 (lightest)6400.70%0.09%±3.1°
MST 3–46400.60%0.11%±3.3°
MST 5–66400.80%0.12%±3.5°
MST 7–86401.10%0.13%±3.7°
MST 9–10 (darkest)6401.40%0.14%±4.1°

FNMR gap between lightest and darkest buckets: 0.7 pp. Our published target is < 1.5 pp; when a quarterly run exceeds that gap we retrain before releasing to production and mark that model version as failed in our changelog.

What we do about it

  • Any single-signal flag (face lost, gaze off) is never auto-terminating — it always triggers human review.
  • Students can opt into Low-bandwidth mode when their connection is unstable, so poor video quality doesn't become a false violation.
  • Instructors can enable Accessibility mode per exam, which uses dyslexia-friendly typography and a keyboard-first flow.
  • Browser lockdown (copy/paste blocking) is off by default — it's an opt-in per exam, not a mandate.
  • Every flagged attempt is appealable in-product; the appeal ships with the tamper-evident event log.

Report an accuracy issue

If our system flagged you unfairly, email fairness@autoproctor.com with your attempt reference. We respond within 2 business days and include any resulting model changes in the next quarterly report.