Inter-Rater Reliability of EEG-Based Encephalopathy Grading
Bibliographic record
Abstract
PURPOSE: Visual EEG Confusion Assessment Method-Severity (VE-CAM-S) quantifies encephalopathy severity based on electroencephalography features. This study evaluated inter-rater reliability among experts using the VE-CAM-S scale. METHODS: Nine experts from six institutions independently reviewed 32 15-second electroencephalography samples in an online test, assessing 29 features (16 in the VE-CAM-S and 13 additional, or "VE-CAM-S+"). A consensus of three experts served as the gold standard. Performance was measured by the median Matthews correlation coefficient between expert and gold-standard VE-CAM-S+ scores, along with average sensitivity and specificity. Qualitative analysis identified common feature-recognition errors affecting scores. RESULTS: Experts achieved a median Matthews correlation coefficient of 0.82 [95% CI: 0.74-0.99]. Specificity exceeded 90% for most features except background β (87%) and generalized delta (71%). Sensitivity was ≥65% except for burst suppression with epileptiform activity (61%), extreme delta brush (EDB; 61%), posterior dominant rhythm (50%), background α (59%) and β (42%). Common errors included missing subtle findings, confusing features, and misidentifying extreme delta brush. CONCLUSIONS: This pilot study offers some initial support for the reliability of VE-CAM-S+ scoring. The largest errors occurred when experts missed or falsely identified features with higher weight in the VE-CAM-S. Encephalopathy grading through VE-CAM-S may be improved by breaking high-stakes features into smaller parts, creating a "cheat sheet" with scored examples, and designing teaching materials.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | no category Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | low |
| gpt | no category Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | medium |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.044 | 0.093 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".