Assessment of Population Well-being With the Mental Health Quotient: Validation Study
Bibliographic record
Abstract
BACKGROUND: The Mental Health Quotient (MHQ) is an anonymous web-based assessment of mental health and well-being that comprehensively covers symptoms across 10 major psychiatric disorders, as well as positive elements of mental function. It uses a novel life impact scale and provides a score to the individual that places them on a spectrum from Distressed to Thriving along with a personal report that offers self-care recommendations. Since April 2020, the MHQ has been freely deployed as part of the Mental Health Million Project. OBJECTIVE: This paper demonstrates the reliability and validity of the MHQ, including the construct validity of the life impact scale, sample and test-retest reliability of the assessment, and criterion validation of the MHQ with respect to clinical burden and productivity loss. METHODS: Data were taken from the Mental Health Million open-access database (N=179,238) and included responses from English-speaking adults (aged≥18 years) from the United States, Canada, the United Kingdom, Ireland, Australia, New Zealand, South Africa, Singapore, India, and Nigeria collected during 2021. To assess sample reliability, random demographically matched samples (each 11,033/179,238, 6.16%) were compared within the same 6-month period. Test-retest reliability was determined using the subset of individuals who had taken the assessment twice ≥3 days apart (1907/179,238, 1.06%). To assess the construct validity of the life impact scale, additional questions were asked about the frequency and severity of an example symptom (feelings of sadness, distress, or hopelessness; 4247/179,238, 2.37%). To assess criterion validity, elements rated as having a highly negative life impact by a respondent (equivalent to experiencing the symptom ≥5 days a week) were mapped to clinical diagnostic criteria to calculate the clinical burden (174,618/179,238, 97.42%). In addition, MHQ scores were compared with the number of workdays missed or with reduced productivity in the past month (7625/179,238, 4.25%). RESULTS: >0.99). Furthermore, the aggregate MHQ scores were systematically related to both clinical burden and productivity. At one end of the scale, 89.08% (8986/10,087) of those in the Distressed category mapped to one or more disorders and had an average productivity loss of 15.2 (SD 11.2; SEM [standard error of measurement] 0.5) days per month. In contrast, at the other end of the scale, 0% (1/24,365) of those in the Thriving category mapped to any of the 10 disorders and had an average productivity loss of 1.3 (SD 3.6; SEM 0.1) days per month. CONCLUSIONS: The MHQ is a valid and reliable assessment of mental health and well-being when delivered anonymously on the web.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.015 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".