Bibliographic record
Abstract
Most big scientific meetings have sessions reserved for important late breaking trials. Between 1999 and 2002, 86 trials made it into these sessions at meetings of the American College of Cardiology. Unsurprisingly they were bigger, better, and more likely to be subsequently published than trials presented in other sessions, a recent study has found. Despite their high quality and impact, reports of late breaking trials were just as likely as other trials to change between the meeting stage and full publication up to three years later. Overall, 41% of trials reported different effect sizes, and a quarter reported different sample sizes in the initial report and the full paper. In one trial in seven, the statistical significance of the effect also changed. Effect sizes changed by more than one standard deviation, on average. Credit: JAMA The authors don't explore in detail why discrepancies are so common, but at least some of the sample sizes changed because preliminary results presented at the meetings were later extended. The authors do question, however, whether early reports of important papers should be included in the evidence base for treatments, since so many of them seem unstable. JAMA 2006;295: 1281–1 [OpenUrl][1][CrossRef][2][PubMed][3][Web of Science][4] Doctors have been treating croup with humidified air for well over 100 years. That's probably long enough, say researchers from Canada, after their carefully controlled and blinded trial failed to find any evidence of benefit. Children given 40% oxygen at 100% humidity did no better than children who were effectively given enriched room air to breathe. One control group had the standard humidification from flexible tubing directed towards the face by a parent. The other control group had 40% oxygen at 40% humidity via a face mask. The researchers did not include an untreated group because humidification is the standard treatment for croup. Although the trial was small … [1]: {openurl}?query=rft.jtitle%253DJAMA%26rft.stitle%253DJAMA%26rft.issn%253D0002-9955%26rft.aulast%253DToma%26rft.auinit1%253DM.%26rft.volume%253D295%26rft.issue%253D11%26rft.spage%253D1281%26rft.epage%253D1287%26rft.atitle%253DTransition%2BFrom%2BMeeting%2BAbstract%2Bto%2BFull-length%2BJournal%2BArticle%2Bfor%2BRandomized%2BControlled%2BTrials%26rft_id%253Dinfo%253Adoi%252F10.1001%252Fjama.295.11.1281%26rft_id%253Dinfo%253Apmid%252F16537738%26rft.genre%253Darticle%26rft_val_fmt%253Dinfo%253Aofi%252Ffmt%253Akev%253Amtx%253Ajournal%26ctx_ver%253DZ39.88-2004%26url_ver%253DZ39.88-2004%26url_ctx_fmt%253Dinfo%253Aofi%252Ffmt%253Akev%253Amtx%253Actx [2]: /lookup/external-ref?access_num=10.1001/jama.295.11.1281&link_type=DOI [3]: /lookup/external-ref?access_num=16537738&link_type=MED&atom=%2Fbmj%2F332%2F7543%2F713.atom [4]: /lookup/external-ref?access_num=000235972600026&link_type=ISI
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.045 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.006 | 0.006 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.013 | 0.008 |
| Open science | 0.004 | 0.002 |
| Research integrity | 0.008 | 0.008 |
| Insufficient payload (model declined to judge) | 0.249 | 0.155 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".