Interobserver reliability in Pirani clubfoot severity scoring between a paediatric orthopaedic surgeon and a physiotherapy assistant
Bibliographic record
Abstract
The Ponseti method, now regarded as the standard of care for congenital clubfoot, is equally effective whether provided by orthopaedic surgeons or orthopaedic paramedics. Therefore, it is particularly suitable for under-resourced nations with lack of surgeons and physicians. At the Sudan Clubfoot Clinic, physiotherapy assistants (3-year diploma nurses with additional physiotherapy experience) are part of the Ponseti clubfoot treatment team, with the role of assessing the degree of deformity by the Pirani score to assist the team in providing treatment. However, the reliability of Pirani scores measured by physiotherapy assistants in this context is unknown. After obtaining informed consent, we measured the interobserver reliability between a physiotherapy assistant and an orthopaedic surgeon in measuring Pirani scores in 91 virgin clubfeet in 54 infants (41 males and 13 females) at the Sudan Clubfoot Clinic. Scores were measured independently before the onset of treatment and analysed by the κ statistic for interobserver reliability. The κ statistic was 0.61 for posterior crease, 0.72 for empty heel, 0.51 for rigid equinus, 0.54 for the hid-foot score, 0.57 for medial crease, 0.54 for curved lateral border, 0.56 for lateral head of talus, 0.50 for the midfoot score and 0.50 for the total score. The mean percentage of agreement of both observers for all Pirani components was 83%. We found moderate to substantial interobserver reliability for the Pirani clubfoot severity score and all its subcomponents. Properly trained physiotherapy assistants are efficient in assessing the degree of severity of clubfoot. This is particularly useful in developing countries, where orthopaedic surgeons are few. Clubfoot treatment can be made more affordable by using paramedical healthcare workers such as physiotherapy assistants.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.031 | 0.057 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".