Data-Driven Approach to Quantify and Reduce Error Associated with Assigning Short Duration Counts to Traffic Pattern Groups
Bibliographic record
Abstract
Traffic monitoring agencies collect traffic data samples to estimate annual average daily traffic (AADT) at short duration count sites. The steps to estimate AADT from sample data introduce error that manifests as uncertainty in the AADT statistic and its applications. Past research suggests that the assignment of a short duration count site to a traffic pattern group (TPG), characterized by known traffic periodicities, represents a significant but poorly quantified source of error. This paper presents an approach to quantify the range of errors arising from such assignments and to mitigate these errors using a novel data-driven assignment method. The approach uses simulated 48-hour short duration counts sampled from continuous count sites with known AADT to develop a benchmark of the total error expected when AADT is estimated from such samples. Likewise, the analysis produces a set of AADT estimates using temporal factors from pre-defined TPGs to quantify the range of assignment errors. The data-driven assignment method aims to mitigate these errors by minimizing the absolute mean deviation in AADT estimates produced from multiple short duration counts in a single year. The approach is applied to traffic data collected in Manitoba, Canada, as a case study. The results indicate that the mean absolute error from 48-hour short duration counts is 6.40% of the true AADT and that improper assignment can lead to a range in mean absolute errors of 9%. When applied to previously unassigned sites, the data-driven assignment method reduced mean absolute errors from 10.32%, using a conventional assignment method, to 7.86%.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".