La critique d’un outil d’intelligence artificielle dans l’évaluation des paralysies faciales périphériques
Bibliographic record
Abstract
La paralysie faciale périphérique (PFP) est une altération du fonctionnement de certains muscles du visage, suite à une lésion du nerf facial. Cette pathologie entraîne des conséquences fonctionnelles et esthétiques qui impactent la qualité de vie des individus. Leur prise en soin est donc cruciale et débute par une évaluation précise. Actuellement, on utilise principalement des échelles de notation telles que Sunnybrook Facial Grading System (SFGS) ou House-Brackmann Grading System (HBGS), basées sur le jugement du clinicien. Cependant, ces méthodes d’évaluation laissent place à une certaine subjectivité. Grâce aux récentes avancées technologiques, on s’intéresse davantage à l’intelligence artificielle (IA). L’IA pourrait permettre de développer un outil d’évaluation objectif, automatisé et quantitatif, applicable en milieu clinique. Cette approche vise à diminuer la subjectivité induite par les évaluations actuelles. Nous avons mené une étude rétrospective auprès de 38 patients présentant une PFP modérée-sévère à totale. L’objectif de l’étude était de recenser les bénéfices et les limites d’Emotrics+, un logiciel de métriques faciales basé sur l’IA, afin de déterminer si cet outil est applicable en clinique. Ce protocole s’est déroulé à deux périodes différentes (14 jours et 1 an post-PFP) en utilisant l’échelle SFGS et le logiciel Emotrics+. Nous avons évalué les fiabilités inter-juges et intra-juge afin de déterminer la fiabilité et la reproductibilité des deux outils. Puis, nous avons établi une corrélation entre les deux outils pour déterminer si Emotrics+ suivait la tendance de SFGS. Nos résultats actuels ne justifient pas l’applicabilité immédiate de cet outil. Cependant, avec des ajustements appropriés, Emotrics+ présente un potentiel certain. Peripheral facial palsy (PFP) is an alteration in the functioning of some facial muscles following an injury to the facial nerve. This pathology has functional and aesthetic consequences that impact the quality of life of patients. Their care is essential and begins with an accurate assessment. Currently, scoring scales such as Sunnybrook Facial Grading System (SFGS) or House-Brackmann Grading System (HBGS) are used, based on clinician judgment. However, these evaluation methods can be subject to a certain degree of subjectivity. Recent advances in technology have led to increased interest in artificial intelligence (AI). AI could make it possible to develop an objective, automated and quantitative assessment tool, applicable in a clinical setting. This approach aims to reduce the subjectivity induced by current evaluation. We conducted a retrospective study of 38 patients with moderate-severe to total PFPs. The objective of the study is to identify the benefits and limitations of Emotrics+, a facial metrics tool based on AI, in order to determine whether the tool is applicable in the clinic. This protocol took place at two different time periods (14 days and 1 year post-PFP) using the SFGS scale and the Emotrics+ software. We evaluated the inter-rater and intra-rater reliability in order to determine the reliability and the reproducibility of the two tools. Then, we established a correlation between the two tools to determine if Emotrics+ followed SFGS's trend. Our currents results do not support the immediate applicability of this software. However, with appropriates adjustments, Emotrics+ has a certain potential.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".