Effect-Size Discrepancies in Literature Versus Raw Datasets from Experimental Spinal Cord Injury Studies: A CLIMBER Meta-Analysis
Bibliographic record
Abstract
Translation of spinal cord injury (SCI) therapeutics from pre-clinical animal studies into human studies is challenged by effect size variability, irreproducibility, and misalignment of evidence used by pre-clinical versus clinical literature. Clinical literature values reproducibility, with the highest grade evidence (class 1) consisting of meta-analysis demonstrating large therapeutic efficacy replicating across multiple studies. Conversely, pre-clinical literature values novelty over replication and lacks rigorous meta-analyses to assess reproducibility of effect sizes across multiple articles. Here, we applied modified clinical meta-analysis methods to pre-clinical studies, comparing effect sizes extracted from published literature to raw data on individual animals from these same studies. Literature-extracted data (LED) from numerical and graphical outcomes reported in publications were compared with individual animal data (IAD) deposited in a federally supported repository of SCI data. The animal groups from the IAD were matched with the same cohorts in the LED for a direct comparison. We applied random-effects meta-analysis to evaluate predictors of neuroconversion in LED versus IAD. We included publications with common injury models (contusive injuries) and standardized end-points (open field assessments). The extraction of data from 25 published articles yielded n = 1841 subjects, whereas IAD from these same articles included n = 2441 subjects. We observed differences in the number of experimental groups and animals per group, insufficient reporting of dropout animals, and missing information on experimental details. Meta-analysis revealed differences in effect sizes across LED versus IAD stratifications, for instance, severe injuries had the largest effect size in LED (standardized mean difference [SMD = 4.92]), but mild injuries had the largest effect size in IAD (SMD = 6.06). Publications with smaller sample sizes yielded larger effect sizes, while studies with larger sample sizes had smaller effects. The results demonstrate the feasibility of combining IAD analysis with traditional LED meta-analysis to assess effect size reproducibility in SCI.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".