Scale-aware multi-path deep neural networks for unconstrained face detection
Bibliographic record
Abstract
Unconstrained face detection is the task of robustly finding and locating faces in an image subject to possible variations in facial scale, blur, pose, illumination, occlusion, and facial expression.It is a critical first step towards a host of modern surveillance applications, including but not limited to face verification, face recognition, face tracking, and human-computer interaction.Though much progress has been made in unconstrained face detection during the past decade, the majority of work focuses on improving the detection robustness on variations caused by blur, pose, illumination, occlusion and facial expression.Facial scale, despite its immense influence on face detection accuracy, has received much less attention than have the above factors.This is partially due to the fact that most traditional face detection benchmark datasets tend to collect faces of relatively large size and with modest scale variation.Nonetheless, in real-world applications, such as surveillance systems, it is imperative to possess an equal ability to detect both big faces (close to camera) and tiny ones (far away from the camera) at the same time.To the best of our knowledge, no published face detection algorithm can detect a face as large as 1000 1000 pixels while simultaneously detecting another one as small as 10 10 pixels within a single image with similarly high accuracy.We introduce a Multi-Path Face Detection Network (MP-FDN) to filter an image for simultaneously proposing and verifying different sized faces in parallel paths.This is the first time that faces across a large span of scales are detected by a single network with forked detection paths.More importantly, the division of the paths are not handcrafted, but totally based on the scale sensitivity inherent in the convolutional networks that was also discovered in this thesis for the first time.MP-FDN consists of two stages.The first stage is a Multi-Path Face Proposal Network (MP-FPN) that suggests faces at three different scale ranges.This design is based on our observation that the hierarchical multiscale layers of deep convolutional networks (ConvNet) can inherently represent face patterns at multiple scales.In particular, low-level ConvNet layers are more sensitive to tiny faces, while high-level ConvNet layers are more discriminative to big faces.To this end, MP-FPN utilizes three parallel outputs of the convolutional feature maps to simultaneously predict small, medium and large candidate face regions, respectively.The second stage is a Multi-Path Face Verification Network (MP-FVN) that further eliminates false positives while including false negatives.MP-FVN utilizes the same three parallel paths as MP-FPN.First and foremost, I would like to express my sincere gratitude to my supervisor, Professor Martin D. Levine,
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".