Evaluating the Mental Health of Physician-Trainees Using an SMS Text Message–Based Assessment Tool: Longitudinal Pilot Study
Bibliographic record
Abstract
BACKGROUND: Physician burnout is a multibillion-dollar issue in the United States. Despite its prevalence, burnout is difficult to accurately measure. Institutions generally rely on periodic surveys that are subject to recall bias. SMS text message-based surveys or assessments have been used in health care and have the advantage of easy accessibility and high response rates. OBJECTIVE: In this pilot project, we evaluated the utility of and participant engagement with a simple, longitudinal, and SMS text message-based mental health assessment system for physician-trainees at the study institution. The goal of the SMS text message-based assessment system was to track stress, burnout, empathy, engagement, and work satisfaction levels faced by users in their normal working conditions. METHODS: Three SMS text message-based questions per week for 5 weeks were sent to each participant. All data received were deidentified. Additionally, each participant had a deidentified personal web page to follow their scores as well as the aggregated scores of all participants over time. A 13-question optional survey was sent at the conclusion of the study to evaluate the usability of the platform. Descriptive statistics were performed. RESULTS: In all, 81 participants were recruited and answered at least six (mean 14; median 14; range 6-16) questions for a total of 1113 responses. Overall, 10 (17%) out of 59 participants responded "Yes" to having experienced a traumatic experience during the study period. Only 3 participants ever answered being "Not at all satisfied" with their job. The highest number of responses indicating that participants were stressed or burnt out came on day 25 in the 34-day study period. There were mixed levels of concern for the privacy of responses. No substantial correlations were noted between responses and having experienced a traumatic experience during the study period. Furthermore, 12 participants responded to the optional feedback survey, and all either agreed or strongly agreed that the SMS text message-based assessment system was easy to use and the number of texts received was reasonable. None of the 12 respondents indicated that using the SMS text message-based assessment system caused stress. CONCLUSIONS: Responses demonstrated that SMS text message-based mental health assessments are potentially useful for recording physician-trainee mental health levels in real time with minimal burden, but further study of SMS text message-based mental health assessments should address limitations such as improving response rates and clarifying participants' sense of privacy when using the SMS text message-based assessment system. The findings of this pilot study can inform the development of institution-wide tools for assessing physician burnout and protecting physicians from occupational stress.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".