Bibliographic record
Abstract
The process results of the 8 December 2025 searches of Google, OVID, PubMed, Scopus, and Web of Science for the keywords “tumor” and “time to discovery” are in Figure 1. The initial number of records identified was 168: Google Scholar (n = 13), OVID (n = 8), PubMed (n = 32), Scopus (n = 62), and Web of Science (n = 53). The Duplicate records removed in total were (n = 10). By databases, they were OVID—duplicated with itself—(n = 3), Scopus—duplicated with PubMed—(n = 2), and Web of Science—duplicated with Scopus—(n = 5). Those that were not an empirical study were (n = 5). They are from Google Scholar ( n = 2), OVID (n = 1), and Web of Science (n = 1). Those not peer-reviewed were from Google Scholar (n = 3), leaving the “Records screened” (n = 150). The excluded records were then those studies of non-human subjects (n = 18). By database: PubMed (n = 2), Scopus (n = 4), and Web of Science (n = 12), leaving the “Reports sought for retrieval” (n = 132). The “Reports not retrieved” were (n = 7). PubMed had (n = 5) unretrievable, and Scopus (n = 2), resulting in the “Reports assessed for eligibility” (n = 125). The “Reports excluded” regarded no tumor or no time to discovery. “No tumor” was (n =10). The breakdown by database was Google Scholar (n = 2), PubMed (n = 7), and Scopus (n = 1). Those lacking time to discovery were (n = 19). By database, they were OVID (n = 1), PubMed (n = 5), Scopus (n = 11), and Web of Science (n = 2), leaving the Studies included in review (n = 96). Of these, Google Scholar represents (n = 6), OVID (n = 3), PubMed (n = 13), Scopus (n = 42), and Web of Science (n = 33). All the included studies are in the References. There were (n = 117) “Reports of included studies”. Those studies with more than one report were (n = 2) from OVID, (n = 4), (n = 8) from PubMed, (n = 4), and (n = 8) from Scopus. Counting the one study itself, the total of reports that were in addition to the 96 studies was 21.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.039 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.041 | 0.039 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.007 | 0.003 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.134 | 0.084 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".