Retrofit Weight-Loss Outcomes at 6, 12, and 24 Months and Characteristics of 12-Month High Performers: A Retrospective Analysis
Bibliographic record
Abstract
BACKGROUND: Obesity is the leading cause of preventable death costing the health care system billions of dollars. Combining self-monitoring technology with personalized behavior change strategies results in clinically significant weight loss. However, there is a lack of real-world outcomes in commercial weight-loss program research. OBJECTIVE: Retrofit is a personalized weight management and disease-prevention solution. This study aimed to report Retrofit's weight-loss outcomes at 6, 12, and 24 months and characterize behaviors, age, and sex of high-performing participants who achieved weight loss of 10% or greater at 12 months. METHODS: A retrospective analysis was performed from 2011 to 2014 using 2720 participants enrolled in a Retrofit weight-loss program. Participants had a starting body mass index (BMI) of >25 kg/m² and were at least 18 years of age. Weight measurements were assessed at 6, 12, and 24 months in the program to evaluate change in body weight, BMI, and percentage of participants who achieved 5% or greater weight loss. A secondary analysis characterized high-performing participants who lost ≥10% of their starting weight (n=238). Characterized behaviors were evaluated, including self-monitoring through weigh-ins, number of days wearing an activity tracker, daily step count average, and engagement through coaching conversations via Web-based messages, and number of coaching sessions attended. RESULTS: Average weight loss at 6 months was -5.55% for male and -4.86% for female participants. Male and female participants had an average weight loss of -6.28% and -5.37% at 12 months, respectively. Average weight loss at 24 months was -5.03% and -3.15% for males and females, respectively. Behaviors of high-performing participants were assessed at 12 months. Number of weigh-ins were greater in high-performing male (197.3 times vs 165.4 times, P=.001) and female participants (222 times vs 167 times, P<.001) compared with remaining participants. Total activity tracker days and average steps per day were greater in high-performing females (304.7 vs 266.6 days, P<.001; 8380.9 vs 7059.7 steps, P<.001, respectively) and males (297.1 vs 255.3 days, P<.001; 9099.3 vs 8251.4 steps, P=.008, respectively). High-performing female participants had significantly more coaching conversations via Web-based messages than remaining female participants (341.4 vs 301.1, P=.03), as well as more days with at least one such electronic message (118 vs 108 days, P=.03). High-performing male participants displayed similar behavior. CONCLUSIONS: Participants on the Retrofit program lost an average of -5.21% at 6 months, -5.83% at 12 months, and -4.09% at 24 months. High-performing participants show greater adherence to self-monitoring behaviors of weighing in, number of days wearing an activity tracker, and average number of steps per day. Female high performers have higher coaching engagement through conversation days and total number of coaching conversations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".