Identification by cluster analysis of patients with asthma and nasal symptoms using the MASK-air® mHealth app
Bibliographic record
Abstract
BACKGROUND: The self-reporting of asthma frequently leads to patient misidentification in epidemiological studies. Strategies combining the triangulation of data sources may help to improve the identification of people with asthma. We aimed to combine information from the self-reporting of asthma, medication use and symptoms to identify asthma patterns in the users of an mHealth app. METHODS: We studied MASK-air® users who reported their daily asthma symptoms (assessed by a 0-100 visual analogue scale - "VAS Asthma") at least three times (either in three different months or in any period). K-means cluster analysis methods were applied to identify asthma patterns based on: (i) whether the user self-reported asthma; (ii) whether the user reported asthma medication use and (iii) VAS asthma. Clusters were compared by the number of medications used, VAS asthma levels and Control of Asthma and Allergic Rhinitis Test (CARAT) levels. FINDINGS: We assessed a total of 8,075 MASK-air® users. The main clustering approach resulted in the identification of seven groups. These groups were interpreted as probable: (i) severe/uncontrolled asthma despite treatment (11.9-16.1% of MASK-air® users); (ii) treated and partly-controlled asthma (6.3-9.7%); (iii) treated and controlled asthma (4.6-5.5%); (iv) untreated uncontrolled asthma (18.2-20.5%); (v) untreated partly-controlled asthma (10.1-10.7%); (vi) untreated controlled asthma (6.7-8.5%) and (vii) no evidence of asthma (33.0-40.2%). This classification was validated in a study of 192 patients enrolled by physicians. INTERPRETATION: We identified seven profiles based on the probability of having asthma and on its level of control. mHealth tools are hypothesis-generating and complement classical epidemiological approaches in identifying patients with asthma.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".