Improvements in Neoplasm Classification in the International Classification of Diseases, Eleventh Revision: Systematic Comparative Study With the Chinese Clinical Modification of the International Classification of Diseases, Tenth Revision
Bibliographic record
Abstract
Background The International Classification of Diseases, Eleventh Revision (ICD-11) improved neoplasm classification. Objective We aimed to study the alterations in the ICD-11 compared to the Chinese Clinical Modification of the International Classification of Diseases, Tenth Revision (ICD-10-CCM) for neoplasm classification and to provide evidence supporting the transition to the ICD-11. Methods We downloaded public data files from the World Health Organization and the National Health Commission of the People’s Republic of China. The ICD-10-CCM neoplasm codes were manually recoded with the ICD-11 coding tool, and an ICD-10-CCM/ICD-11 mapping table was generated. The existing files and the ICD-10-CCM/ICD-11 mapping table were used to compare the coding, classification, and expression features of neoplasms between the ICD-10-CCM and ICD-11. Results The ICD-11 coding structure for neoplasms has dramatically changed. It provides advantages in coding granularity, coding capacity, and expression flexibility. In total, 27.4% (207/755) of ICD-10 codes and 38% (1359/3576) of ICD-10-CCM codes underwent grouping changes, which was a significantly different change (χ21=30.3; P<.001). Notably, 67.8% (2424/3576) of ICD-10-CCM codes could be fully represented by ICD-11 codes. Another 7% (252/3576) could be fully described by uniform resource identifiers. The ICD-11 had a significant difference in expression ability among the 4 ICD-10-CCM groups (χ23=93.7; P<.001), as well as a considerable difference between the changed and unchanged groups (χ21=74.7; P<.001). Expression ability negatively correlated with grouping changes (r=–.144; P<.001). In the ICD-10-CCM/ICD-11 mapping table, 60.5% (2164/3576) of codes were postcoordinated. The top 3 postcoordinated results were specific anatomy (1907/3576, 53.3%), histopathology (201/3576, 5.6%), and alternative severity 2 (70/3576, 2%). The expression ability of postcoordination was not fully reflected. Conclusions The ICD-11 includes many improvements in neoplasm classification, especially the new coding system, improved expression ability, and good semantic interoperability. The transition to the ICD-11 will inevitably bring challenges for clinicians, coders, policy makers and IT technicians, and many preparations will be necessary.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.023 | 0.032 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.007 | 0.011 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".