Application of Machine Learning to Child Mode Choice with a Novel Technique to Optimize Hyperparameters
Bibliographic record
Abstract
Travel mode choice (TMC) prediction is crucial for transportation planning. Most previous studies have focused on TMC in adults, whereas predicting TMC in children has received less attention. On the other hand, previous children's TMC prediction studies have generally focused on home-to-school TMC. Hence, LIGHT GRADIENT BOOSTING MACHINE (LGBM), as a robust machine learning method, is applied to predict children's TMC and detect its determinants since it can present the relative influence of variables on children's TMC. Nonetheless, the use of machine learning introduces its own challenges. First, these methods and their performance are highly dependent on the choice of "hyperparameters". To solve this issue, a novel technique, called multi-objective hyperparameter tuning (MOHPT), is proposed to select hyperparameters using a multi-objective metaheuristic optimization framework. The performance of the proposed technique is compared with conventional hyperparameters tuning methods, including random search, grid search, and "Hyperopt". Second, machine learning methods are black-box tools and hard to interpret. To overcome this deficiency, the most influential parameters on children's TMC are determined by LGBM, and logistic regression is employed to investigate how these parameters influence children's TMC. The results suggest that MOHPT outperforms conventional methods in tuning hyperparameters on the basis of prediction accuracy and computational cost. Trip distance, "walkability" and "bikeability" of the origin location, age, and household income are principal determinants of child mode choice. Furthermore, older children, those who live in walkable and bikeable areas, those belonging low-income groups, and short-distance travelers are more likely to travel by sustainable transportation modes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.008 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".