Identifying the Bias: Evaluating Effectiveness of Automatic Data Collection Methods in Estimating Details of Bus Dwell Time
Bibliographic record
Abstract
Data from automated vehicle location (AVL) systems, automatic passenger counter (APC) systems, and fare box payments have been heavily used to generate dwell time models with the goal of recommending improvements in efficiency and reliability of bus transit systems. However, automatic data collection methods may result in a loss of detail with regard to the dynamics of passenger activity, which may bias the estimates associated with dwell or passenger activity time. The purpose of this study is to understand better any biases that might exist from using data from AVL–APC systems or fare box payments when estimating dwell time. Manually collected data from Montreal, Quebec, Canada, are used to estimate detailed dwell time models. This study compared those estimates to models generated by using data similar to what was reported by AVL–APC systems and fare boxes. The results reveal an overestimation in the passenger activity component of dwell time, which is mainly attributed to excess dwell time that AVL–APC data and fare box payments generally do not capture. While AVL–APC and fare box technologies provide transit agencies with rich data for analysis, adjustments to such data collection methods are warranted to reduce the overestimation of dwell time and to provide a more accurate picture of what is happening on the ground to generate better interventions that can reduce dwell times.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.025 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".