Reliability of Clinical Symptoms in Diagnosing Temporomandibular Joint Arthritis in Juvenile Idiopathic Arthritis
Bibliographic record
Abstract
OBJECTIVE: Temporomandibular joint (TMJ) arthritis, commonly considered oligoarthritic/asymptomatic, occurs frequently in children with juvenile idiopathic arthritis (JIA), and gadolinium-enhanced magnetic resonance imaging (Gd-MRI) has proved to be a sensitive diagnostic tool in this context. We compared the reliability of clinical examinations to Gd-MRI results in diagnosing the condition. METHODS: Patients with JIA (134 consecutive) underwent routine clinical and Gd-MRI examinations. The clinical items examined were clicking, tenderness (TMJ/adjacent muscles), and mouth-opening capacity. Blinded MRI reading focused on inflammation (synovitis/hypertrophy). After statistical power analysis, the clinical findings for 134 healthy controls were included. Contingency analysis was used to determine the sensitivity, specificity, and frequency of clinical symptoms (JIA/healthy controls); Cohen's κ was used to establish the interrater reliability. RESULTS: Statistically significant differences were observed between JIA and healthy control groups with regard to the concise screening items (power analysis > 0.95), whereas no differences in mouth-opening capacity were noted. In 80% of the patients with JIA, Gd-MRI revealed signs of TMJ arthritis, with positive correlations between concise screening items and Gd-MRI results. The average specificity was 0.81, but the sensitivity was low, at 0.42. Combining items led to a marked increase in the sensitivity (0.73). There was a high rate of both false-negative and false-positive results (corresponding to clinical underdiagnosis or overdiagnosis of TMJ arthritis). CONCLUSION: Despite a relatively high specificity, clinical examination alone does not seem sufficiently sensitive to adequately detect TMJ arthritis. Thus, a relatively high number of cases will be missed or overdiagnosed, potentially leading to undertreatment or overtreatment. Gd-MRI may support correct diagnosis, thereby helping to prevent undertreatment or overtreatment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.023 | 0.010 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".