Diagnosing Early or Rheumatoid Arthritis. Which Is Better: Expert Opinion or Evidence?
Bibliographic record
Abstract
All clinicians know about the complexity in diagnosing rheumatoid arthritis (RA), especially early in the disease. Because RA lacks pathognomonic features — that is, there are no clinical, biological, or radiological characteristics specific to RA diagnosis — doubt about the diagnosis may persist for some patients1,2. Examples are patients with “nude” polyarthritis [i.e., without positivity for serum rheumatoid factor (RF), anti-citrullinated peptide antibodies, typical erosion, or all 3], or even elderly people with erosive RF-positive polyarthritis associated with psoriasis or calcium crystal deposition disease features seen on joint radiography. When RA is neither obvious nor completely excluded, the clinician strikes a balance between possible or probable RA, depending on the level of confidence. In this context, in clinical research, RA classification criteria may be of some help because they ensure, at the group level, the diagnosis of RA with minimal error. However, in clinical practice, RA criteria cannot be used as the gold standard, especially in early arthritis (EA), as was previously shown3–5. In this issue of The Journal, Morvan, et al report on a cohort of patients with EA followed for 10 years to investigate discrepancies in RA diagnosed by American College of Rheumatology (ACR) classification criteria and final diagnosis by an office-based rheumatologist6. The authors noted poor agreement at the onset of the disease, as has been shown, but also at 2 years, when ACR criteria are supposed to be more accurate. If one assumes that the rheumatologist is an expert, who is right: the expert or the criteria? Eminence … Address correspondence to Dr. B. Fautrel, Department of Rheumatology, Pitié-Salpêtrière Hospital, 83 boulevard de l’Hôpital, 75651 Paris cedex 13, France. E-mail: bruno.fautrel{at}psl.aphp.fr
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.026 | 0.140 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.005 | 0.002 |
| Bibliometrics | 0.005 | 0.003 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.007 | 0.013 |
| Open science | 0.006 | 0.002 |
| Research integrity | 0.015 | 0.012 |
| Insufficient payload (model declined to judge) | 0.011 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".