Long-Term Performance of Superpave in Specific Pavement Study 9A
Bibliographic record
Abstract
As part of the Long-Term Pavement Performance Program, an experiment identified as Specific Pavement Study 9A (SPS-9A) was established to evaluate the long-term performance of the Superpave® mix design methodology. At each project location, three main test sections were constructed, incorporating the conventional agency mix design, the Superpave Level I mix design with a 98% reliability performance-grade (PG) asphalt binder, and Superpave with an alternative PG binder to evaluate rutting or thermal cracking performance. Further supplementary test sections were constructed at most projects to evaluate the use of other PG asphalt binders with the Superpave mix design. This paper provides a performance comparison and paired T-test analysis of the Superpave Level I mix methodology with conventional agency mix methodology based on the SPS-9A experiment test sections. In addition, the performance impact of PG asphalt binder grades and the use of reclaimed asphalt pavement in Superpave was assessed by statistical analysis of field distresses. Field distresses evaluated in the study include fatigue cracking, thermal cracking, and rutting. Although some trends were identified, overall there was no statistically significant difference in performance between the considered PG asphalt binder grades and mix design types. These results would indicate that asphalt binder and mixture properties probably play a larger role in performance than binder and mix design classification. However, because of the level of asphalt binder and mixture property data in DataPave Release 20.0, a more detailed evaluation of binder and mixture properties on performance was not possible.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".