COSMOS: Interrater and Intrarater Reliability Study of a Novel Outcome Measure
Bibliographic record
Abstract
The vast majority of patients with minor stroke achieve what are considered good or excellent outcomes on the modified Rankin Scale (0-1/0-2), yet many are dissatisfied with their outcomes. There is a need for a functional outcome measure tailored for minor stroke that better reflects the spectrum of clinical outcomes within this population. We developed the Canadian Outcome Scale for Minor Stroke (COSMOS) and performed an interrater and intrarater reliability study. COSMOS is a 7-point scale ranging from 0 (no symptoms) to 6 (loss of independence for an instrumental or basic activity of daily living or worse), which accounts for performance limitations and losses of a person's hobbies or passions and of their employment, educational, service, or caregiving pursuits, besides just activities of daily living. One hundred test case vignettes were developed. Stroke physicians, fellows, and research nurses/staff were invited to review training materials and provide the COSMOS grade for 20 cases representing all COSMOS grades (0-6). After a minimum 2 weeks' wash-out period, participants were asked to grade the same 20 cases again. Interrater and intrarater agreement were assessed using Cohen κ, weighted κ, percentage agreement, and intraclass correlation coefficient. Among 33 participants (18 attending physicians, 9 stroke fellows, and 6 research staff/nurses; median 12.5 years of experience), COSMOS had substantial interrater reliability (80.5% agreement [95% CI, 75.7%-85.3%]; Cohen κ, 0.77 [95% CI, 0.72-0.84]) and almost-perfect intrarater reliability overall (87.1% agreement [95% CI, 84.4%-89.7%]; Cohen κ, 0.85 [95% CI, 0.82-0.88]); weighted κ showed almost perfect agreement for both interrater (0.88 [95% CI, 0.85-0.92]) and intrarater reliability (0.92 [95% CI, 0.90-0.94]). The overall chance-adjusted simultaneous intrarater/interrater agreement using intraclass correlation coefficient was 0.95 (95% CI, 0.94-0.97). Results were similar with substantial to almost-perfect agreement when considering key subgroups based on position (attendings, fellows, research nurses/staff) and years of experience. In conclusion, the newly proposed COSMOS scale demonstrated substantial interrater and intrarater reliability. The scale merits further study in cohort studies and clinical trials of minor stroke.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.105 | 0.146 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".