Exploring Transit Use during COVID-19 Based on XGB and SHAP Using Smart Card Data
Bibliographic record
Abstract
As the coronavirus (COVID-19) pandemic continues, many protective measures have been taken in Seoul, Korea, and around the world. This situation has drastically changed lifestyle and travel behavior. An important issue concerns understanding the reasons for giving up transit use and the potential impact on travel patterns during the COVID-19 pandemic. To shed light on these issues that are essential for transit policy, this study explores transit use choice, such as whether users have given-up transit use or not, during the COVID-19 pandemic. Two days of smart card data, before and during the COVID-19 pandemic, were used to look at users who gave up transit use during the COVID-19 pandemic. The choice set of the dataset includes two alternatives, for example, transit use and given-up transit use. An extreme gradient boosting (XGB) model was used to estimate the transit use behavior. Shapley additive explanations were performed to interpret the estimation results of the XGB model. The results for the overall specificity, sensitivity, and balanced accuracy of the proposed XGB model were estimated to be 0.909, 0.953, and 0.931, respectively. The feature analysis based on the Shapley value shows that the number of origin-to-destination trip feature substantially impacts transit use. As such, users tend to avoid transit use as travel time increased during the COVID-19 pandemic. The proposed model shows remarkable performance in accuracy and provided an understanding of the estimated results.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".