Contributions of International Medical Graduates to US Biomedical Research: The Experience of US Medical Schools
Bibliographic record
Abstract
International medical graduates (IMGs) constitute an appreciable fraction of full-time faculty at US medical schools and of principal investigators (PIs) on National Institutes of Health (NIH) research project grants. Information from the Faculty Roster of the Association of American Medical Colleges (AAMC) and from the NIH Consolidated Grant Applicant File (CGAF) was examined to assess IMGs' contribution to US medical school faculty and research. The study found that the number of IMG full-time faculty more than doubled over two decades-from 7,866 individuals in 1984 to 17,085 individuals in 2004, but that IMGs remained relatively stable as a share of physician full-time faculty (from 18.8 to 19.4%); the share is somewhat higher (20.0% of full-time physician faculty in 1984 to 23.7% in 2004) if faculty with degrees of unknown provenance are included. From 1984 to 2004, IMGs increased as a share of full-time physician faculty who are principal investigators on NIH research grants from 16.5% (540) to 21.3% (1,143). Including faculty with incomplete data on degree provenance, the corresponding IMG share increases to 18.0 and 24.0% respectively. Thus, IMGs comprise at least one-fifth and more likely one-fourth of all full-time faculty physicians who are PIs on NIH research project grants. The proportion of IMG full-time physician faculty who are in basic science departments is about twice that of their US/Canadian counterparts, as is the proportion of IMG physician PIs. Slightly fewer than half (48%) of full-time IMG faculty PIs pursue human subjects research (as coded by the NIH), while the majority of US/Canadian counterparts pursue human subjects research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.023 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.001 | 0.005 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".