Good interobserver and intraobserver agreement in the evaluation of the new ILAE classification of focal cortical dysplasias
Bibliographic record
Abstract
PURPOSE: An International League Against Epilepsy (ILAE) consensus classification system for focal cortical dysplasias (FCDs) has been published in 2011 specifying clinicopathologic FCD variants. The aim of the present work was to microscopically assess interobserver agreement and intraobserver reproducibility for FCD categories among an international group of neuropathologists with different levels of experience and access to epilepsy surgery tissue. METHODS: Surgical FCD specimens covering a broad histopathology spectrum were retrieved from 22 patients with epilepsy. Three surgical nonepilepsy specimens served as controls. A total of 188 slides with routine or immunohistochemical stainings were digitalized with a slide scanner to allow Internet-based microscopy review. Nine experienced neuropathologists were invited to review these cases twice at a time gap of 3 months and different orders of case presentation. The 2011 ILAE FCD consensus classification served as instruction. Kappa analysis was calculated to estimate interobserver and intraobserver agreement levels. In a third evaluation round, 21 additional neuropathologists with different experience and access to epilepsy surgery reviewed the same case series. KEY FINDINGS: Interobserver agreement was good (κ = 0.6360), with 84% consensus of diagnoses during the first evaluation (21 of 25 cases). Kappa values increased to 0.6532 after reevaluation, and consensus was obtained in 24 (96%) of 25 cases. Overall intraobserver reproducibility was also good (κ = 0.7824, ranging from 0.4991 to 1.000). Fewest changes in the classification were made in the FCD type II group (2.2% of 225 original diagnoses), whereas the majority of changes occurred in FCD type III (13.7% of 225 original diagnoses). In the third evaluation round, interobserver agreement was reflected by the level of experience of each neuropathologist, with κ values ranging from moderate (0.5056; high level of experience >40 cases/year) to low (0.3265; low level of experience <10 cases/year). SIGNIFICANCE: Our study achieved a good and reliable interobserver agreement among the group of expert neuropathologists originally involved in the ILAE FCD consensus classification system. Intraobserver reproducibility in this group was even more robust. These results showed considerable improvement compared to a previous study evaluating the 2004 Palmini FCD classification. Agreement levels were lower in our second group of neuropathologists and were related to their level of access and experience with epilepsy surgery specimens. These results suggested that the more precise ILAE definition of FCD histopathology patterns improves operational procedures in the diagnosis of FCDs. On the other hand, microscopic assessment of FCD is a challenge and requires sustained experience and teaching. The virtual slide review system allowed testing of this hypothesis and reached a widespread group of participating colleagues from different centers all over the world. We propose to further use this tool as a teaching device and also to address other epilepsy-associated entities still difficult to classify such as hippocampal sclerosis, long-term epilepsy-associated tumors, or mild malformations of cortical development (mMCDs), which were not yet covered by current ILAE classification systems.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".