Dr Singer et al reply
Bibliographic record
Abstract
To the Editor: We appreciate the comments shared by Wang and colleagues.1 Here we provide responses to the key comments that were raised and highlight relevant aspects in our original publication2 and from our research. The first point highlighted by Wang et al was regarding the limitations associated with using diagnosis codes for the identification of rheumatoid arthritis (RA) and herpes zoster (HZ) cases.1 As mentioned in the Discussion of our original publication,2 although the large sample sizes provided by the administrative claims database is a benefit of using such a data source, it is important to acknowledge there are limitations associated with this.2 One of these limitations, which was noted in the original publication, is that the study relied on “diagnosis codes that were used previously by other researchers to define HZ cases, [although] it is possible that some HZ cases may not have been identified or some cases may not have been true cases … Address correspondence to Dr. D. Singer, GSK, Philadelphia, PA 19103, USA. Email: david.a.singer{at}gsk.com.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".