Sentiment and Mobility Analysis on COVID-19 Restrictions with Autoencoder
Bibliographic record
Abstract
In 2020, the world was attacked by a virus known as the COVID-19 virus. Restrictions on people’s activities were conducted in various countries to prevent the spread of the virus. However, since people were vaccinated, restriction levels have been reduced or eliminated, although the new cases of COVID-19 worldwide have not ended. People’s responses to restriction policies vary, including sentiment and human mobility. The possibility of sentiment is either support or resistance, while mobility is staying at home or not. This study analyzes the proportion between the two responses through two types of data: text for sentiment and time series for mobility. Sentiment text data is taken from Twitter and mobility time series data is taken from Google Mobility for February 2020 to April 2022. Twitter and Google Mobility data are collected from several countries using English and implementing restrictions: Australia, Canada, Singapore, the United Kingdom (UK), and the United States (US). The unsupervised Autoencoder model is leveraged to find clusters. Two Autoencoder architectures are proposed for each data type. Before being used in Multilayer Autoencoder, text data is converted to vector data by Word2Vec. On the other hand, LSTM-Autoencoder is used for time series data. Finally, hypothesis tests are performed to determine the mean between the clusters formed is the same or different, out of five countries, only Canada has a null hypothesis is accepted, that means people in Canada tend to be neutral in response to COVID-19 while mobilities are dynamics, it reveals that people in Canada obey the government’s decision on restrictions during the rise of COVID-19 cases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".