Reliability of the Sigmoid Notch Classification of the Distal Radioulnar Joint
Bibliographic record
Abstract
Abstract Background The Tolat sigmoid notch classification is a commonly used classification to characterize the distal radioulnar joint (DRUJ). This classification was based on a limited assessment of the entire joint, which may lead to inaccuracies in sigmoid notch evaluation. Questions/Purposes The purpose of this study is to assess the reliability of the Tolat classification for sigmoid notch characterization. Methods The sigmoid notch of 52 models of cadaveric forearms was assessed by applying the Tolat classification to the three-dimensional (3D) modeled notch and then slices at the start of the notch (0 mm) and 4 mm more proximal. The inter- and intrarater agreement was assessed using Cohen's and Fleiss' kappa statistic. Results Agreement between iterations regardless of slices or surgeons/radiologists was moderate. Intrarater agreement between pairs of slices (0 vs 4 mm, 0 mm vs 3D, 4 mm vs 3D) was moderate, whereas agreement between all slices was slight. Agreement between surgeons and between radiologists was moderate, while agreement across all raters and slices was fair. Models described as “other” were more consistent in 3D classifications and were commonly classified as a reverse ski slope. Conclusions Classification using the Tolat scheme is fair to moderate at best. Classification of the sigmoid notch using an axial view of the distal radius may not accurately reflect the anatomy throughout the notch. Clinical Relevance The Tolat classification supplies a limited analysis of the sigmoid notch, and does not represent a comprehensive evaluation of the entire joint. Future classification systems should characterize the entire sigmoid notch.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".