MétaCan
Menu
Back to cohort
Record W2981403469 · doi:10.1177/2059799119884273

A new innovative method to measure the demographic representation of scientists via Google Scholar

2019· article· en· W2981403469 on OpenAlexaff
Ehsan Jozaghi

Bibliographic record

VenueMethodological Innovations · 2019
Typearticle
Languageen
FieldDecision Sciences
Topicscientometrics and bibliometrics research
Canadian institutionsBC Centre for Disease ControlUniversity of British Columbia
Fundersnot available
KeywordsRanking (information retrieval)CensusGlobeCitationUnderrepresented MinorityHigher educationPopulationCategorizationRepresentation (politics)Diversity (politics)InstitutionPolitical sciencePublic relationsMedical educationGeographyLibrary scienceSociologySocial scienceComputer sciencePsychologyDemographyMedicineInformation retrievalLaw

Abstract

fetched live from OpenAlex

Many countries around the globe have seen increases in the enrollment of female and visible minorities in postsecondary education. Therefore, it is critical to evaluate whether recent demographic changes at the postsecondary institution have translated to employment opportunities in scientific fields for women and previously underrepresented groups. Instead of relying on algorithm indices, surveys, or anonymous census data, this study is the first research to utilize an innovative approach to report the demographic representation of top-ranking scientists from around the world. The recently developed Google Scholar profile platform, university ranking system, and the search engine are the main methods that allowed this study to identify and categorize the top scientists from countries in which English is one of the official languages, or where English is used as the language of instruction in higher education. Overall, findings reveal that at top-ranking universities in which the majority of the population is Caucasian, women and minorities are severely underrepresented in all areas of science, capturing 7.3% and 6.4% of the total citations, respectively. Each country’s highest concentration of scientists in each field, based on citation and percentage of researchers, is highlighted. There are recommendations offered to help make scientific advancement more favorable to underrepresented groups, and also to encourage institutions of higher education to adapt and build new capacities.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.007
metaresearch head score (Gemma)0.029
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Bibliometrics
Consensus categoriesnone
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.993
Threshold uncertainty score0.036

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0070.029
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0330.048
Science and technology studies0.0010.001
Scholarly communication0.0030.003
Open science0.0010.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0070.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.848
GPT teacher head0.657
Teacher spread0.191 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designSimulation or modeling
DomainEvaluation
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations7
Published2019
Admission routes1
Has abstractyes

Explore more

Same venueMethodological InnovationsSame topicscientometrics and bibliometrics researchFrench-language works237,207