Comparison of evidence of treatment effects in randomized and nonrandomized studies on allergen immunotherapy
Bibliographic record
Abstract
Abstract Nonrandomized studies (NRS) on allergen immunotherapy (AIT) particularly lend themselves to evaluate outcomes that are insufficiently addressed in randomized controlled studies (RCTs). However, NRS are prone to several sources of bias, which limit their validity. We aimed at comparing AIT effects between RCTs and NRS and evaluate the reasons for discrepancies in study results. In this analysis, we compared NRS on AIT (including subcutaneous and sublingual immunotherapy, SCIT and SLIT, respectively) with SLIT and SCIT RCTs from published meta‐analyses, assessing the Risk of Bias (RoB) for each study and the certainty of evidence from NRS and RCTs using the GRADE approach. We found: (1) very serious RoB in the 7 NRS included in the meta‐analysis showing a large difference between AIT and controls (standardized mean difference [SMD] for symptom score [SS], −1.77; 95% CI, −2.30, −1.24; p < .001; I 2 = 95%) with very low certainty evidence; (2) serious RoB in the 13 SCIT‐RCTs reporting a moderate‐to‐high difference between SCIT and controls (SMD for SS, −0.81; 95% CI, −1.12, −0.49; p < .001; I 2 = 88%) with moderate certainty evidence; (3) low RoB in the 13 SLIT‐RCTs reporting a small benefit (SMD for SS, −0.28; 95% CI, −0.37, −0.19; p < .001; I 2 = 54.2%) with high certainty evidence. Similar results were reported for medication score. Our evidence is sufficient to conclude that the magnitude of effect estimates of NRS and RCTs directly correlate with the degree of RoB and inversely with the overall evidence certainty. NRS, which are more affected than RCTs by bias resulting in low certainty evidence, showed the largest effect size. Sound NRS are needed to complement RCTs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.012 | 0.002 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".