Low sensitivities but surprisingly high efficiencies for face-gender discrimination from interattribute distances
Bibliographic record
Abstract
According to an influential view, relational cues such as the distances between the main internal features of a face (i.e. mouth, eyes, eyebrows and nose) play a predominant role in face processing (Maurer et al., 2002; but see Taschereau-Dumouchel et al., 2010). Studies on face-gender perception are no exception (see Campbell et al.,1999). Nevertheless, the use of real-world interattribute distances for face-gender discrimination in humans has never been examined. This was the aim of the present study. In Exp. 1 we tested whether observers can discriminate the gender of faces based solely on real-world interattribute distances. Participants had to discriminate the gender of two androgynous faces of the same identity that were presented simultaneously on the screen: one had real-world interattribute distances of a woman and, the other, of a man. Despite relatively low sensitivities (average d’= 0.40 +/- 0.31, ranging from 0.82 to 0.05), 9 out of 11 observers performed significantly above chance (p<0.05). Surprisingly, statistical efficiencies were relatively high (M=13.91%, SD=12.85%, ranging from 37.35% to 0%). This is because real-world interattribute distances contain little face-gender information: A linear classifier trained on the interattribute distances of 250 faces (125 men) and tested on 250 novel faces (125 men) obtained a d’=1.34. Would real-world interattribute distances still contribute to gender discrimination when more informative cues such as attribute shapes and skin properties are available? In Exp. 2, we tested this by manipulating realistically the interattribute distances of 500 faces to make them more or less congruent with the gender of the face. Results showed that indeed congruency had a significant positive effect on the sensitivity (F(2.6, 44.4)=16.75, p<0.0001). Meeting abstract presented at VSS 2012
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.014 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".