Prevalence, associated factors, and machine learning-based prediction of depression, anxiety, and stress among university students: a cross-sectional study from Bangladesh
Bibliographic record
Abstract
BACKGROUND: Mental health challenges are a growing global public health concern, with university students at elevated risk due to academic and social pressures. Although several studies have exmanined mental health among Bangladeshi students, few have integrated conventional statistical analyses with advanced machine learning (ML) approaches. This study aimed to assess the prevalence and factors associated with depression, anxiety, and stress among Bangladeshi university students, and to evaluate the predictive performance of multiple ML models for those outcomes. METHODS: A cross-sectional survey was conducted in February 2024 among 1697 students residing in halls at two public universities in Bangladesh: Jahangirnagar University and Patuakhali Science and Technology University. Data on sociodemographic, health, and behavioral factors were collected via structured questionnaires. Mental health outcomes were measured using the validated Bangla version of the Depression, Anxiety, and Stress Scale-21 (DASS-21). Statistical analyses included chi-square tests and binary logistic regression, while seven ML models including, K-Nearest Neighbors (KNN), Random Forest (RF), Gradient Boosting Machine (GBM), Extreme Gradient Boosting (XGBoost), Categorical Boosting (CatBoost), Logistic Regression (LR), and Support Vector Machine (SVM) were employed to predict mental health outcomes. RESULTS: The prevalence of depression, anxiety, and stress was 56.9%, 69.5%, and 32.2%, respectively. Significant associated factors for depression included unfriendly family relationships, enrollment in commerce, and cigarette smoking. Female gender, unfriendly family relationships, academic year, and cigarette smoking were significant factors for stress. No significant factors were identified for anxiety. Among ML models, SVM achieved the highest accuracy for depression prediction (accuracy = 0.5693; precision = 0.7560; log loss = 0.6847), LR for anxiety (accuracy = 0.6948; precision = 0.7881), and CatBoost for stress (accuracy = 0.6706; precision = 0.6454; F1-score = 0.5777; log loss = 0.6284). Feature importance analyses highlighted faculty of study and relation with family as the top predictors. ROC-AUC values indicated moderate discriminatory performance (all ≥ 0.5). CONCLUSIONS: Integrating machine learning with conventional analyses enhances the identification and prediction of factors associated with depression, anxiety, and stress among university students. These findings support the implementation of campus-based mental health screening, accessible counseling, and peer support programs, and highlight the value of data-driven approaches for developing targeted university mental health policies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".