Using ROC to Examine the Association between Attendance and Compliance
Notice bibliographique
Résumé
Dear Editor-in-Chief In the January 2014 edition of the journal, Miller et al. (1) outlined a method of assessing compliance with a prescribed exercise program on the basis of measuring HR as a marker of intensity (in addition to duration and frequency). In doing so, they were able to estimate a precise dose–response effect of compliance with prescribed exercise against several different health outcomes. Interestingly, they compared their HR measure of compliance to attendance at exercise sessions (they described the latter as a measure of adherence). Subjects who attended at least 80% of scheduled sessions were classified as adherent. Not surprisingly, there was less than perfect agreement between adherence and compliance; 9.1% of subjects were misclassified. With less than 10% misclassification though, it could be argued that attendance is a reliable proxy for more complicated measures of compliance (e.g., HR) in trials such as this. Because it is not always possible or feasible to directly assess prescribed compliance using HR over time but it is relatively easy (and feasible) to record attendance, knowing the precise ability of the latter to predict the former would be helpful for future intervention research. However, Miller et al. (1) did not report such results. To do so, receiver operator characteristic (ROC) curve analysis (2), which estimates sensitivity, specificity, positive predictive value, and negative predictive value of a proxy measure (i.e., recoded attendance) compared against a criterion or gold standard (i.e., compliance measured using HR), would need to be calculated. My analysis of their reported data using ROC confirmed that sensitivity is excellent, suggesting that when subjects attend 80% or more sessions, there is a conditional probability of meeting compliance on the basis of an HR of 95%. Specificity is also very good. The conditional probability of not meeting compliance for those attending less than 80% of sessions is 78%. On the basis of these data, if we were using attendance records only as a proxy for compliance, the 80% or higher threshold would correctly identify 93% of all subjects who were compliant on the basis of HR (positive predictive value) and 83% would be classified as noncompliant subjects (negative predictive value) on the basis of the same criterion. There are many conceivable circumstances where this level of agreement is acceptable. For example, if we were screening for entry into a follow-up study, we may use the 80% threshold to maximize the chance of selecting participants who were compliant. Of course, it is possible to set a higher threshold (e.g., >85% sessions), which may in fact improve specificity but at the cost of lowering sensitivity. ROC analysis can then be used to define an optimal cut point that balances sensitivity against specificity. Such calculations, however, were not possible on the basis of the data provided by Miller et al. (1). In sum, on the basis of these data, simple reported attendance rate is a reasonable proxy for compliance measured using HR. However, it is important to remember that participants were in a structured and supervised exercise intervention. Therefore, use of attendance as a proxy may be reasonable but only under conditions similar to the ones reported in this study. John Cairney, PhD Departments of Family Medicine Kinesiology, and Clinical Epidemiology and Biostatistics McMaster University Hamilton, Ontario, Canada
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,036 | 0,140 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,003 |
| Bibliométrie | 0,012 | 0,008 |
| Études des sciences et des technologies | 0,000 | 0,001 |
| Communication savante | 0,004 | 0,002 |
| Science ouverte | 0,002 | 0,001 |
| Intégrité de la recherche | 0,003 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,002 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».