Comparison of administrative/billing data to expected protocol‐mandated chemotherapy exposure in children with acute myeloid leukemia: A report from the Children's Oncology Group
Bibliographic record
Abstract
BACKGROUND: Recently investigators have used analysis of administrative/billing datasets to answer clinical and pharmacoepidemiology questions in pediatric oncology. However, the accuracy of pharmacy data from administrative/billing datasets have not yet been evaluated. The primary objective of this study was to determine the concordance of Pediatric Health Information System (PHIS) administrative/billing chemotherapy data with Children's Oncology Group (COG) protocol-mandated chemotherapy and to assess the implications of this level of concordance for further PHIS research. PROCEDURE: Data from 384 pediatric patients (1,060 courses of chemotherapy) with acute myeloid leukemia treated on COG clinical trial AAML0531 were previously merged with PHIS data. PHIS chemotherapy administrative/billing data were reviewed for the first three courses of chemotherapy. Accuracy was assessed using three metrics: recognizability of chemotherapy pattern by course, chemotherapy administration pattern by individual medication, and concordance with the number of days of protocol-defined chemotherapy. RESULTS: The chemotherapy pattern was recognizable in 87.3% of courses when course-wide accuracy was assessed. Chemotherapy administration pattern varied by medication. Cytarabine had perfect concordance 70.9% of the time, daunorubicin had perfect concordance 77.4% of the time, and etoposide had perfect concordance 67.8% of the time. CONCLUSIONS: The accuracy of chemotherapy administrative/billing data supports the continued use of PHIS data for epidemiology studies as long as investigators perform data quality control checks and evaluate each specific medication prior to undertaking definitive analyses.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".