Raise the MAST (DOI) – Assessing Preprints for Data DOIs
Bibliographic record
Abstract
Since launching JWST, the Space Telescope Science Institute (STScI) Metrics Office has partnered with the Mikulski Archive for Space Telescopes (MAST) to identify arXiv preprints that use JWST data. Preprints are then evaluated to determine whether the data are cited using a durable data DOI. Authors are contacted with instructions for including data DOIs before the final paper goes to print. While this process does not catch every paper, it identifies enough to make preprint review a worthwhile endeavor. The Metrics Office reviews preprints throughout the year. While the overall paper count is not reported until the first quarter of the following year, we are able to refer to these identified preprints as a preliminary count well before annual reporting has been finalized. Additionally, contacting authors increases the chance of a data DOI being included in the published version of the paper. For JWST preprints in 2024, the rate of data DOIs in published papers was significantly higher when we were able to contact the author during pre-publication (43%) than when we were unable (25%). This increase justifies this review effort. This poster will provide an overview of our process and its inspiration from ALMA; an explanation of the relationship between preprints and final bibliographic entries; statistics that demonstrate success compared to HST data DOIs; our plans to incorporate an LLM in this process; and steps for incorporating preprint review into bibliographic work.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.243 | 0.668 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.033 | 0.022 |
| Science and technology studies | 0.009 | 0.006 |
| Scholarly communication | 0.055 | 0.032 |
| Open science | 0.005 | 0.021 |
| Research integrity | 0.007 | 0.009 |
| Insufficient payload (model declined to judge) | 0.091 | 0.167 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".