Multifactorial analysis of AAR development: Integrating laboratory and field data with statistical and probabilistic modelling
Bibliographic record
Abstract
Alkali-aggregate reaction (AAR) is among the most harmful durability issues affecting concrete infrastructure, with prevention being the most effective strategy. Widely used laboratory tests, such as the Accelerated Mortar Bar Test (AMBT) and Concrete Prism Test (CPT), assess aggregate reactivity and often present discrepancies with field performance that remain unquantified. This study introduces a probabilistic, risk-based framework to evaluate the reliability of these tests using a multifactorial analysis that integrates field and laboratory data with Bayesian inference and Beta distribution modelling. The likelihood of AAR occurrence is evaluated considering test outcomes, environmental exposure, and alkali loading. Results show AMBT outperforms in identifying non-reactive cases (41% vs. 61% posterior probability for mixes without supplementary cementitious materials [SCMs]; 16% vs. 30% with SCMs), while both tests perform similarly for reactive cases (i.e., 74% for mixes without SCMs and 50% with SCMs). Moreover, warm climates, high alkali content, and the absence of SCMs increase the risk of the tests’ misclassification, while cold environments with low alkali levels and SCMs enhance their reliability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".