Local Calibration of the MEPDG Distress and Performance Models for Ontario’s Flexible Roads: Overview, Impacts, and Reflection
Bibliographic record
Abstract
Built upon a seven-year local calibration study of Ontario’s flexible pavements, this paper provides a summary of the calibration results and design impact and, more importantly, shares the experience and lessons learned from the process. The best results have been achieved on the local calibration of the rutting, bottom-up fatigue cracking, and international roughness index (IRI) distress models minimizing the residual sum of squares (RSS) while maintaining the average bias at zero. Significant efforts have been made to calibrate the other distress models with limited success. A design impact study found that local calibration of the rutting models was very important, whereas the alligator fatigue cracking did not usually govern the design in Ontario, although the global model was found to under-predict the cracking damage. The performance of the calibrated IRI model in the design of heavy traffic freeways for both reconstructed and rehabilitated sections was unsatisfactory and needs further study. The paper also presents several open questions for future research. These include the handling of section-length effects of observed cracking data, the determination of initial IRI, the updating of standard deviation functions and the overall reliability models, and the prioritization of pavement research under the new paradigm of the Mechanistic–Empirical Pavement Design Guide (MEPDG).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".