Revisited: Walch Classification of the Glenoid in Glenohumeral Osteoarthritis
Bibliographic record
Abstract
Background Shoulder osteoarthritis is characterized by progressive wear of the joint. To grade the degree of joint deformity, the Walch classification of glenohumeral arthritis has been proposed. This classification is based on five categories (A1, A2, B1, B2, C), although its validity has been questioned. Methods The present study proposed a new classification in three categories and compared it in terms of inter- and intra-observer reliability with the complete Walch classification and regroup classification (A, B, C). Results One hundred and sixteen computed tomography scans of patients with shoulder arthritis were revised by three independent evaluators and were classified according to the three classifications. The kappa statistics were identical between the new classification and the complete Walch classification (0.87 and 0.874). The regroup Walch classification (A, B, C) demonstrated higher reliability (kappa = 0.92). Most of the disagreement between observers was observed between glenoid B1 and B2. Discussion We report the first study on the Walch classification to use a large number of patients and challenge its reliability. According to the results obtained, there is no advantage in changing the classification. Therefore, surgeons must be aware of higher risk of mistake for glenoid type B. The superiority of a classification in terms of the prediction of surgical decisions and outcome has to be determined.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.045 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.007 | 0.004 |
| Science and technology studies | 0.001 | 0.005 |
| Scholarly communication | 0.002 | 0.004 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".