Agreement and concurrent validity between telehealth and in-person diagnosis of musculoskeletal conditions: a systematic review
Bibliographic record
Abstract
OBJECTIVES: To assess the concurrent validity and inter-rater agreement of the diagnosis of musculoskeletal (MSK) conditions using synchronous telehealth compared to standard in-person clinical diagnosis. METHODS: We searched five electronic databases for cross-sectional studies published in English in peer-reviewed journals from inception to 28 September 2023. We included studies of participants presenting to a healthcare provider with an undiagnosed MSK complaint. Eligible studies were critically appraised using the QUADAS-2 and QAREL criteria. Studies rated as overall low risk of bias were synthesized descriptively following best-evidence synthesis principles. RESULTS: We retrieved 6835 records and 16 full-text articles. Nine studies and 321 patients were included. Participants had MSK conditions involving the shoulder, elbow, low back, knee, lower limb, ankle, and multiple conditions. Comparing telehealth versus in-person clinical assessments, inter-rater agreement ranged from 40.7% agreement for people with shoulder pain to 100% agreement for people with lower limb MSK disorders. Concurrent validity ranged from 36% agreement for people with elbow pain to 95.1% agreement for people with lower limb MSK conditions. DISCUSSION: In cases when access to in-person care is constrained, our study implies that telehealth might be a feasible approach for the diagnosis of MSK conditions. These conclusions are based on small cross-sectional studies carried out by similar research teams with similar participant demographics. Additional research is required to improve the diagnostic precision of telehealth evaluations across a larger range of patient groups, MSK conditions, and diagnostic accuracy statistics.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.004 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".