In Reply to Rosenkranz and Hu and to Wolfson and Arora
Bibliographic record
Abstract
We appreciate the thoughtful commentary on our review, provided by Rosenkranz and Hu and by Wolfson and Arora. Here, we will attempt to clarify a few points from our review1 that were discussed in these letters. Rosenkranz and Hu caution that some items (including self-reported medical student research participation and publication rates) discussed in our review may reflect selection bias, and that our recommendation to increase curricular time devoted to medical student research may not be suitable for every institution, given resource constraints. Selection bias should be considered when interpreting voluntary survey data, particularly regarding extracurricular initiatives. Students choosing to participate in extracurricular research activity may value research differently than the entire medical student population. We did not necessarily suggest that medical education programs invest more (of already limited) formal, curricular time in medical student research. Rather, we suggested that giving students the option of extending their (often extracurricular) research project timeline could provide a more fulfilling, robust scholarship experience to the select students who voluntarily choose this option. Allowing for greater time would also address factors (such as mentor interaction and time available to devote to the project) frequently cited in student feedback regarding the research programs reviewed as well as our own 10-week Summer Studentship Program at the University of Ottawa.2 More time for research may also lead to more mature research products, facilitating increased dissemination rates, thereby addressing student desire to be recognized for their scholarly contributions. We do agree that research program duration should be consistent with an institution’s available resources, including dedicated research mentors. Ultimately, local participant feedback should also be considered when optimizing program time duration. Wolfson and Arora suggested that using publication metrics to assess medical student research programs may be inconvenient and proposed relying on more immediate measures, including whether research programs positively impacted student attitudes towards future research activity. However, many reviewed studies have already assessed this; our review and our local program evaluation study both described that students generally perceive their research experience positively and self-report increased interest in future research activity.1,2 What is not clear is whether this self-reported, perceived interest in future research is typically realized as future research contributions. Given that the average time to publication is less than 16 months in our local research program,2 tracking dissemination outcomes may be an achievable short-term metric to evaluate these programs, complementing the student feedback that is typically collected. Long-term, tracking actual future research activity may provide more tangible information than the “intent” to conduct research, for the purposes of program evaluation. Christopher J. Ramnanan, PhD Assistant professor, Department of Innovation in Medical Education, University of Ottawa Faculty of Medicine, Ottawa, Ontario, Canada; [email protected] Youjin Chang, MD Resident, Department of Surgery, University of Ottawa Faculty of Medicine, Ottawa, Ontario, Canada.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.029 | 0.236 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.003 |
| Bibliometrics | 0.004 | 0.005 |
| Science and technology studies | 0.004 | 0.006 |
| Scholarly communication | 0.008 | 0.011 |
| Open science | 0.007 | 0.006 |
| Research integrity | 0.035 | 0.041 |
| Insufficient payload (model declined to judge) | 0.009 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".