Facial-Analysis Benchmark Methodology
A reproducible, protocol-first way to evaluate landmarks, Action Units, expression-category scores, temporal events, and failure handling on consented data.
Protocol v1.0 · Published September 7, 2026 · Results pending a locked run
What this benchmark measures
The benchmark evaluates whether a system can locate and describe observable facial movement under stated capture conditions, and whether it reports uncertainty or a usable failure state when it cannot. It does not establish what a participant feels, thinks, intends, or remembers.
Dataset scope
The v1 target is 50 consenting adults with a participant-disjoint 60/20/20 development, validation, and test split fixed before evaluation. Each participant records a neutral reference segment and one short clip in four conditions: frontal diffuse light, turned pose, low light, and partial occlusion. The final report must replace targets with observed counts and disclose exclusions.
Reference annotations and metrics
- Landmarks: Frame-level reference points, visibility flags, normalized mean error, visibility precision/recall, and failure rate.
- Action Units: Independent presence, intensity, and onset/apex/offset annotations; AU precision, recall, F1, intensity error, and inter-rater agreement.
- Expression categories: Reviewer labels of visible expression patterns—not reports of felt emotion—scored with macro-F1, balanced accuracy, AUROC, Brier score, calibration error, and abstention rate.
- Temporal events: Onset, apex, offset, duration, and identity scored with a locked tolerance, event F1, timing error, and duplicate/missed-event counts.
Uncertainty and failure handling
Use participant-level bootstrap resampling for 95% confidence intervals and report valid and failed items beside every metric. Count no-face, multiple-face, blur, occlusion, decode, timeout, malformed, and untraceable outputs separately. Never turn a documented quality failure into a forced class label.
Consent, privacy, and limitations
Recruit only adults who provide informed consent for the stated purpose. Store source media and consent keys separately with access controls and deletion procedures. Keep raw face video restricted and publish only aggregate annotations, metrics, code, configuration, and hashes where release cannot re-identify participants.
A consented staged capture set is not representative of all people, cultures, devices, expressions, or real-world video. A benchmark score describes system performance on the tested data; it does not certify safety, fairness, legal compliance, clinical validity, or generalization.
Results status
No performance result is claimed yet. A completed report must name the exact dataset, model and API version, configuration, environment, split, counts, exclusions, per-condition metrics, uncertainty, inter-rater agreement, latency, failure rates, and deviations from this protocol.
Read the Facial Analysis Guide | Review the Developer API | Download the machine-readable artifact bundle | Cite Protocol v1.0