Bibliographic record
Abstract
Dear Editor-in-Chief In the January 2014 edition of the journal, Miller et al. (1) outlined a method of assessing compliance with a prescribed exercise program on the basis of measuring HR as a marker of intensity (in addition to duration and frequency). In doing so, they were able to estimate a precise dose–response effect of compliance with prescribed exercise against several different health outcomes. Interestingly, they compared their HR measure of compliance to attendance at exercise sessions (they described the latter as a measure of adherence). Subjects who attended at least 80% of scheduled sessions were classified as adherent. Not surprisingly, there was less than perfect agreement between adherence and compliance; 9.1% of subjects were misclassified. With less than 10% misclassification though, it could be argued that attendance is a reliable proxy for more complicated measures of compliance (e.g., HR) in trials such as this. Because it is not always possible or feasible to directly assess prescribed compliance using HR over time but it is relatively easy (and feasible) to record attendance, knowing the precise ability of the latter to predict the former would be helpful for future intervention research. However, Miller et al. (1) did not report such results. To do so, receiver operator characteristic (ROC) curve analysis (2), which estimates sensitivity, specificity, positive predictive value, and negative predictive value of a proxy measure (i.e., recoded attendance) compared against a criterion or gold standard (i.e., compliance measured using HR), would need to be calculated. My analysis of their reported data using ROC confirmed that sensitivity is excellent, suggesting that when subjects attend 80% or more sessions, there is a conditional probability of meeting compliance on the basis of an HR of 95%. Specificity is also very good. The conditional probability of not meeting compliance for those attending less than 80% of sessions is 78%. On the basis of these data, if we were using attendance records only as a proxy for compliance, the 80% or higher threshold would correctly identify 93% of all subjects who were compliant on the basis of HR (positive predictive value) and 83% would be classified as noncompliant subjects (negative predictive value) on the basis of the same criterion. There are many conceivable circumstances where this level of agreement is acceptable. For example, if we were screening for entry into a follow-up study, we may use the 80% threshold to maximize the chance of selecting participants who were compliant. Of course, it is possible to set a higher threshold (e.g., >85% sessions), which may in fact improve specificity but at the cost of lowering sensitivity. ROC analysis can then be used to define an optimal cut point that balances sensitivity against specificity. Such calculations, however, were not possible on the basis of the data provided by Miller et al. (1). In sum, on the basis of these data, simple reported attendance rate is a reasonable proxy for compliance measured using HR. However, it is important to remember that participants were in a structured and supervised exercise intervention. Therefore, use of attendance as a proxy may be reasonable but only under conditions similar to the ones reported in this study. John Cairney, PhD Departments of Family Medicine Kinesiology, and Clinical Epidemiology and Biostatistics McMaster University Hamilton, Ontario, Canada
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.036 | 0.140 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.012 | 0.008 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".