MétaCan
Menu
Back to cohort
Record W3046179501 · doi:10.1186/s12891-020-03520-x

Binary Tönnis classification: simplified modification demonstrates better inter- and intra-observer reliability as well as agreement in surgical management of hip pathology

2020· article· en· W3046179501 on OpenAlexfundno aff
Jacob Shapira, Jeffrey Chen, Rishika Bheem, Philip J. Rosinsky, David R. Maldonado, Ajay C. Lall, Benjamin G. Domb

Bibliographic record

VenueBMC Musculoskeletal Disorders · 2020
Typearticle
Languageen
FieldMedicine
TopicOrthopaedic implants and arthroplasty
Canadian institutionsnot available
FundersPacira BioSciencesEwing Marion Kauffman FoundationPacira PharmaceuticalsMAKO Surgical CorporationStrykerGraymontArthrex
KeywordsMedicineReliability (semiconductor)Cohen's kappaKappaPhysical therapyMachine learningComputer scienceMathematics

Abstract

fetched live from OpenAlex

BACKGROUND: The traditional Tönnis Classification System has inherent drawbacks as it is vulnerable to the subjectivity of a four-grade system. A two-grade classification could potentially be more reliable. The purpose of this study is to (1) compare the inter-observer and intra-observer reliability of the traditional Tönnis Classification System and a simplified Binary Tönnis Classification System for hip osteoarthritis and to (2) evaluate the clinical applicability of both systems. Our hypothesis is that the proposed Binary Tönnis Classification System will have better reliability and agreement for surgical decision-making. METHODS: Forty consecutive patients were selected to participate in this study. Patients were included in this study if they were between 35 and 60 years old. Patients were excluded if they had prior hip surgeries or conditions. All radiographs were randomized and blinded by a non-observer. Five fellowship-trained hip surgeons from a single center, in a fully crossed design, analyzed and graded all the radiographs utilizing the traditional Tönnis Classification System and the proposed Binary Tönnis Classification System. Intra- and inter-observer reliability values for both the systems were calculated using the Cohen's κ coefficient. A multi-rater κ was calculated using the weighted Fleiss method. RESULTS: The study sample contained 40 anterosuperior hip radiographs. For the traditional Tönnis Classification System, the weighted κ showed a fair inter-observer reliability (κ = 0.474) and excellent intra-observer reliability (κ mean = 0.866). For the proposed Binary Tönnis Classification System, both inter-observer and intra-observer reliability demonstrated excellent values, (κ = 0.858 and 0.928, respectively). On average, the Binary Tönnis Classification System correctly captured 87% of cases. When the traditional Tönnis Classification System was dichotomized, the capture rate was 84%. CONCLUSION: A simplified binary Tönnis Classification System demonstrates better reliability and clinical implementation than the traditional Tönnis Classification System.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.177
Threshold uncertainty score0.780

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.026
GPT teacher head0.288
Teacher spread0.262 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations12
Published2020
Admission routes1
Has abstractyes

Explore more

Same venueBMC Musculoskeletal DisordersSame topicOrthopaedic implants and arthroplastyFrench-language works237,207