A mixed-effect statistical model for before-after speed studies
Bibliographic record
Abstract
This paper proposes an efficient methodology to conduct observational before-after studies for operating speed data. The employed method has some noteworthy strengths: (i) it can analyze data at a disaggregated level to properly account for variations in the speed profile; (ii) it considers the entire distribution of speed to overcome the bias associated to the traditional approaches that represent the speed distribution with a single point estimate; (iii) it takes advantage of full Bayes methods to avoid the empirical Bayes method limitations. To illustrate the feasibility of the proposed framework, a limited sample of before-after speed dataset from Montreal was used. The effectiveness of a safety countermeasure-a reduction in speed limits-was assessed. The speed data were collected on local urban streets grouped into comparison and treatment sites. For modeling the operating speed, we employed a hierarchical mixed-effect Binomial model using a wide range of site characteristics. This model is capable of dealing with heterogeneity across observations and accounting for site specific effects. The analyses results indicated that lane width, number of lanes, and night hours affect the operating speed positively while presence of parking, peak hours, weekend, one way, and precipitation affect it negatively. Although the speed limit reduction was found to be effective in controlling the operating speed, the analyzed sample may not be representative of the entire urban areas subject to this reduction. This paper also highlights some essential issues in the data collection process and the sensitivity of the outcomes to the collection method used.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.059 | 0.066 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.005 | 0.007 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.010 | 0.004 |
| Research integrity | 0.005 | 0.005 |
| Insufficient payload (model declined to judge) | 0.017 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".