Ensemble Stacking for Grading Facial Paralysis Through Statistical Analysis of Facial Features
Bibliographic record
Abstract
In medical diagnostics, the accurate assessment of facial paralysis (FP) represents a significant challenge, necessitating intricate analysis of facial spatial information, notably asymmetry.This condition, characterized by the inability to regulate facial muscles effectively during specific actions, often demands the discernment of clinicians, which lacks a quantitative foundation.In response to this challenge, the present study introduces two innovative models aimed at enhancing the diagnostic process for FP.The first model employs a binary classification framework to differentiate between affected individuals and those without the condition.The second, more complex model, utilizes an ensemble stacking technique to categorize the severity of FP into four distinct grades: normal, mild, moderate, and severe.Data for this analysis was sourced from a collection comprising 21 individuals diagnosed with FP and 20 healthy counterparts, extracted from publicly accessible datasets.Utilizing the OpenFace 2.0 toolkit, three categories of facial features were analyzed: landmarks, facial action units, and eye movement metrics.A comprehensive evaluation was conducted to determine the optimal model through a series of tests that integrated individual and combined facial feature sets alongside dimension reduction techniques.The findings revealed that the Support Vector Machine (SVM) method, applied to the binary classification of FP, attained an accuracy of 97.7%.Conversely, the ensemble stacking approach, incorporating Logistic Regression (LR) and SVM, demonstrated an 88.2% accuracy rate in the grading of FP severity.These outcomes suggest significant potential for the application of such models in telemedicine, facilitating early detection and ongoing remote monitoring of facial nerve functionality, thereby reducing the need for direct patientclinician encounters.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".