Artificial Intelligence for Evaluation of Emotions behind Face Masks
Bibliographic record
Abstract
Boonipat, Yan, and Uldis1 present a study using machine learning to assess whether face masks hinder the ability to identify emotional states by means of facial expression. Unsurprisingly, they found that “covering the face with a mask leads to a significant loss of emotional information conveyed.” Furthermore, they found that “happy” faces were frequently misconstrued by their artificial intelligence tool as “neutral,” and masked faces were frequently interpreted as “angry or sad.” This seems intuitive if we believe that emotional expression is universal, yet decades of research in psychology and neuroscience have proven otherwise.2 In 1872, Charles Darwin wrote The Expression of the Emotions in Man and Animals, priming the scientific stage for the belief that mental states produce external behaviors including a set of facial movements known as “expressions.” Darwin’s findings were elaborated on by evolutionary psychiatrists who argued that accurately interpreting these expressions conferred a greater advantage to various tribal animals, including humans, who relied on such social cues to survive. Since that time, psychologists and neuroscientists have argued both with and against this theory through a dizzying amount of literature that consistently reinforces one fact: there is little consensus regarding whether emotions can be interpreted by facial expression alone.2 A study of the hunter-gatherer Hadza tribe of Northern Tanzania published in Nature in 2020 was specifically designed to investigate whether the evolutionary psychology hypothesis was plausible: through two studies, scientists demonstrated an absence of universal emotion perception and that individuals (from both the United States and the Hadza tribe) infer emotional states from facial expression as a result of knowledge that is steeped in cultural context, not universal meaning.3 Simply, context and culture influence the way emotions are expressed and interpreted, and no two contexts and cultures are the same. Similarly, the machine learning program used for the study by Boonipat et al. likely learned to interpret facial expression as a function of the cultural context of its creators, not a superior ability to identify universal facial expressions of emotion. Still, this study is not without its merits: it highlights the need for receivers (those interpreting emotional expression) to confirm how the sender (the individual displaying the emotion) is feeling, as opposed to making assumptions based on facial expression alone, especially in an era of facial mask coverings. However, it is time for the surgical subspecialties to keep pace with other scientific fields. The area of emotions research is centuries old and stands to advance our current practices, including in areas of surgical performance and mental skills training.4 However, to do so, we must take the time and interest to learn from this extensive pool of experimentation and literature. Although artificial intelligence has potential to make great advances in our field, not all science falls at the knees of the Cartesian model that logic stands free from emotion. Instead, the balance is likely more reflective of our current neuroanatomical states: a constant communication of emotion and cognition5 that results in our higher ability to communicate and collaborate beyond initial impressions. DISCLOSURE The author has no financial interest to declare in relation to the content of this communication.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.013 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".