What Is a Meaningful Difference When Using Infarct Volume as the Primary Outcome?: Results From the HERMES Database
Bibliographic record
Abstract
BACKGROUND: Ischemic stroke lesion volume at follow-up is an important surrogate outcome for acute stroke trials. We aimed to assess which differences in 48-hour lesion volume translate into meaningful clinical differences. METHODS: We used pooled data from 7 trials investigating the efficacy of endovascular treatment for anterior circulation large vessel occlusion in acute ischemic stroke. We assessed 48-hour lesion volume follow-up computed tomography or magnetic resonance imaging. The primary outcome was a good functional outcome, defined as modified Rankin Scale (mRS) scores of 0 to 2. We performed multivariable logistic regression to predict the probability of achieving mRS scores of 0 to 2 and determined the differences in 48-hour lesion volume that correspond to a change of 1%, 5%, and 10% in the adjusted probability of achieving mRS scores of 0 to 2. RESULTS: In total, 1665/1766 (94.2%) patients (median age, 68 [interquartile range, 57-76] years, 781 [46.9%] female) had information on follow-up ischemic lesion volume. Computed tomography was used for follow-up imaging in 83% of patients. The median 48-hour lesion volume was 41 (interquartile range, 14-120) mL. We observed a linear relationship between 48-hour lesion volume and mRS scores of 0 to 2 for adjusted probabilities between 65% and 20%/volumes <80 mL, although the curve sloped off for lower mRS scores of 0-2 probabilities/higher volumes. The median differences in 48-hour lesion volume associated with a 1%, 5%, and 10% increase in the probability of mRS scores of 0 to 2 for volumes <80 mL were 2 (interquartile range, 2-3), 10 (9-11), and 20 (18-23) mL, respectively. We found comparable associations when assessing computed tomography and magnetic resonance imaging separately. CONCLUSIONS: A difference of 2, 10, and 20 mL in 48-hour lesion volume, respectively, is associated with a 1%, 5%, and 10% absolute increase in the probability of achieving good functional outcome. These results can inform the design of future stroke trials that use 48-hour lesion volume as the primary outcome.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.037 | 0.087 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.004 | 0.008 |
| Bibliometrics | 0.003 | 0.005 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".