Bibliographic record
Abstract
A review of: Patel, Manesh R., Connie M. Schardt, Linda L. Sanders, and Sheri A. Keitz. “Randomized Trial for Answers to Clinical Questions: Evaluating a Pre-Appraised Versus a MEDLINE Search Protocol.” Journal of the Medical Library Association 94.4 (2006): 382-6. Objective – To determine the success rate of electronic resources for answering clinical questions by comparing speed, validity, and applicability of two different protocols for searching the medical literature. Design – Randomized trial with results judged by blinded panel. Setting – Duke University Medical Center in Durham, North Carolina, United States of America. Subjects – Thirty-two 2nd and 3rd year internal medicine residents on an eight-week general medicine rotation at the Duke University Medical Center. Methods – Two search protocols were developed: Protocol A: Participants searched MEDLINE first, and then searched pre-appraised resources if needed. Protocol B: Participants searched pre-appraised resources first, which included UpToDate, ACP Journal Club, Cochrane Database of Systematic Reviews, and DARE. The residents then searched MEDLINE if an answer could not be found in the initial group of pre-appraised resources. Residents were randomised by computer-assisted block order into four blocks of eight residents each. Two blocks were assigned to Protocol A, and two to Protocol B. Each day, residents developed at least one clinical question related to caring for patients. The questions were transcribed onto pocket-sized cards, with the answer sought later using the assigned protocol. If answers weren’t found using either protocol, searches were permitted in other available resources. When an article that answered a question was found, the resident recorded basic information about the question and the answer as well as the time required to find the answer (less than five minutes; between five and ten minutes; or more than ten minutes). Residents were to select answers that were “methodologically sound and clinically important” (384). Ten faculty members formally trained in evidence-based medicine (EBM) reviewed a subset of therapy-related questions and answers. The reviewers, who were blinded to the search protocols, judged the applicability and internal validity of the answers. Results – In total, 120 questions were searched using protocol A and 133 using protocol B; 104 answers were found by the protocol A group and 117 by the protocol B group. In protocol A, 97 answers were found in MEDLINE (80.8%) and six answers were found in pre-appraised resources (5.0%). In protocol B, 85 answers were found in pre-appraised resources (64.6%) and 31 were found in MEDLINE (23.3%). UpToDate was the major resource for answers in protocol B. A statistically greater number of answers were found in less than five minutes in protocol B (p
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.310 | 0.709 |
| Meta-epidemiology (narrow) | 0.008 | 0.007 |
| Meta-epidemiology (broad) | 0.022 | 0.014 |
| Bibliometrics | 0.064 | 0.041 |
| Science and technology studies | 0.005 | 0.007 |
| Scholarly communication | 0.015 | 0.019 |
| Open science | 0.008 | 0.016 |
| Research integrity | 0.013 | 0.006 |
| Insufficient payload (model declined to judge) | 0.082 | 0.022 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".