Bibliographic record
Abstract
Better Science by Beating Back Bias T he human mind takes shortcuts by using past experiences to fill in missing information.This special talent helped our forebears avoid unfamiliar dangers and facilitated the development of modern civilization.Today, it allows us to quickly size up new social situations and connect with complete strangers.As researchers, it helps us see patterns in nature that explain how the world works.But this inherently human characteristic has its flaws.In its most benign manifestation, our reliance on shortcuts makes us susceptible to optical illusions or a magician's slight of hand.More troubling is our tendency to fill in missing facts by making broad generalizations that can lead us to draw erroneous conclusions about our fellow researchers and the quality of their work.The idea that our lazy minds and the ways that we are socialized can cause us to draw unjustified conclusions, a concept known as implicit bias, calls into question the integrity of the peer review process.After all, if one of the main pillars of modern science is affected by preconceived notions, how can we be sure that we are publishing the most reliable and important research?Upon learning about implicit bias in the peer review process, most of us assume that we are not the culprits.But just as we can be tricked by a skilled magician, all of us are susceptible to implicit bias in the peer review process.Implicit bias can creep into every stage of the review process, causing us to misjudge research abilities and quality due to assumptions associated with gender, country of origin, and the academic reputation (earned or presumed) of our authors and reviewers.Through our experiences as faculty members at institutions that take diversity seriously, our years as members of diverse research teams, and our personal commitments to diversity, we thought that we were truly objective when we served as authors, peer reviewers, and editors.But a simple exercise that forces you to confront your implicit biases about students and peers, coupled with statistics about the review process in a journal in a closely related field, leads us to question this notion.In 2017, Lerbeck and Hanson analyzed the gender of reviewers of the 20 peer-reviewed journals published by the American Geophysical Union (AGU).They found that both men and women authors suggested fewer women reviewers than expected on the basis of AGU membership or prior authorship (i.e., 28% of AGU members and 27% of first authors are women compared to 21% and 15% of the reviewers suggested by women and men, respectively).AGU editors also invited fewer women to serve as peer reviewers than expected (22% and 17% of the invited reviewers by female and male editors, respectively, were women).In addition, even though AGU-accepted authors (both female and male) reside in roughly equal parts North America, Europe, and Asia, AGU reviewers came primarily from the United States, Canada, and Europe, suggesting geographic bias.The existence of the AGU data set was fortuitous because the computer system that their journals used made it feasible to assess potential bias.Although we have not repeated this
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".