Statistical Analysis and Insight from the Water Research Foundation PCCP Failure Database 1942−2022
Bibliographic record
Abstract
This paper summarizes a statistical analysis of a database collected by the Water Research Foundation (WRF) of prestressed concrete cylinder pipe (PCCP) failures in North America from 1942 to 2022. This database, with almost 50,000 data points collected, is part of two separate research efforts for Projects #4034 and #5069 and is sourced from utilities in the United States and Canada. These WRF projects examine PCCP failures for trends which were co-sponsored and funded by the USEPA and Great Lakes Water Authority. Failures are grouped into one of the three categories, each with descriptive factors, including age, location, size, pipe type, installation date, etc. Via simple yet comprehensive analytics, failures can be graphed as a function of their descriptive factors and subsequent inferences can be made such as: are there locations of PCCP at higher risk than others, which installation dates/periods have the most failures, does size matter, etc.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".