P02.17. Building a database of validated pediatric outcomes: an investigation of compliance with established reporting standards
Bibliographic record
Abstract
Ten journals with the highest impact factors were searched for pediatric RCTs published between 2000-2010. Two independent reviewers conducted screening and data extraction on 20% of randomly selected included studies. Variables extracted included: journal, sample size, participant age, condition under study, intervention, control, and details of primary outcome and outcome measurement tools. Searches identified 2229 unique references. Screening of a random sample of 2.5% determined that most (97%) were RCTs, thus full text for all references were obtained. Inclusion screening was carried out simultaneously with data extraction. Of the 446 articles screened to date, 66% were included. Participant age ranged from 20 weeks gestation to 20 years. Most (65%) were of treatment rather than prevention. Commonly used controls included placebo (35%) and another intervention (33%). With respect to primary outcome reporting, 34% of trials did not identify a primary outcome. Half (53%) reported at least one primary outcome; of these, 55% described one outcome as primary and 38% identified more than one outcome as primary. One quarter of the trials that included only one primary outcome used a questionnaire or scale-based tool and of these, only 26% presented information on tool clinometrics. This project will help identify gaps in the quality of outcome reporting in pediatric trials published in top journals over the past 10 years, leading to recommendations for improvements in reporting standards.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".