The Impact of Gender on High‐Stakes Dental Evaluations
Bibliographic record
Abstract
The purpose of this study was to determine whether gender affects high-stakestest performance among dental students. Our sample consisted of 128 women and 323 men from six consecutive dental classes for which we recorded AADSAS overall and science predental GPAs; Dental Admission Test (DAT) scores; National Board Dental Examination (NBDE) I and II scores and pass/fail status; North East Regional Board of Dental Examiners (NERB) pass/fail status; and cumulative GPAs following the spring quarter of year two and summer quarter of year four of dental school. DAT scores, when controlled for previous academic performance, revealed that men significantly outperformed women in all areas except reading comprehension and biology, where the women's scores significantly exceeded the men's and were comparable, respectively. NBDE I results favored men and approached significance (p = 0.066), while for Part II men significantly outscored women. NBDE I and II and NERB pass rates showed no significant differences. These board results were also controlled for previous academic performance. Although we found that differences existed between genders, which appear to be the ramification of the classic high-stakes dilemma (women do as well as men in the classroom and on course-related tests, but less well on gatekeeper board exams), the context mitigates their operational effects. DAT differences are likely reduced by most admissions processes, but may be problematic when selected predictive algorithms are used. Practically, the NBDE I and II results are unlikely to meaningfully influence women's academic progress in dental school or postgraduate education admissions due to their magnitude and timing.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.017 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".