Feature Analysis and Hierarchical Classification of Anxiety Severity during early COVID-19
Bibliographic record
Abstract
Distress, confusion, and anger are common responses to COVID-19. Statistics Canada created the Canadian Perspectives Survey Series (CPSS) to understand social issues and effects of COVID-19 on the Canadian labour force (LF). The evaluation of the health and health-related behaviours were done through surveys collected between April and July. Features are composed of 4600 participants and 62 questions, which include the General Anxiety Disorder (GAD)-7 questionnaire. This work proposes the use of CPSS2 survey data characteristics to identify the level of anxiety within the Canadian population during early stages of COVID-19 and is validated with the use of GAD-7 questionnaire. Minimum redundancy maximum relevance (mRMR) is applied to select the top 20 features to represent user anxiety. During classification, decision tree (DT) and support vector machine (SVM) are used to test the separation of anxiety severity. Hierarchical classification was used which separated the anxiety severity labels into different test sets and classified accordingly. We employ SVM for binary classification with 10-fold cross validation to separate the labels of Minimal and Severe anxiety to achieve an overall accuracy of 94.77±0.05%. After analysis, a subset of the reduced feature set can be represented as pseudo passive (PP) data, which are passive sensors that can augment qualitative data. The accurate classification provides proxy on what gives rise to anxiety, as well as the ability to provide early interventions. Future works can implement passive sensors to augment PP data and further understand why people cope this way.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".