WECC ADS 2034 Hydropower Generation Code (weccadshydro)
Bibliographic record
Abstract
This release contains the raw data, lookup tables, and code used to produce the WECC ADS 2034 hydropower dataset [1]. The core code is contained in Jupyter notebooks, which has both the function definitions for processing the raw data and producting the inputs, and the workflow. The code is also version controlled in a code repository, weccadshydro [2]. Every two years the WECC (Western Electricity Coordinating Council) releases an Anchor Data Set (ADS) to be analyzed with a Production Cost Models (PCM) and which represents the expected loads, resources, and transmission topology 10 years in the future from a given reference year. For hydropower resources, the WECC relies on members to provide data to parameterize the hydropower representation in production cost models. The datasets consist of plant-level hydropower generation, flexibility, ramping, and mode of operations and are tied to the hydropower representation in those production cost models. In 2022, PNNL supported the WECC by developing the WECC ADS 2032 hydropower dataset [3]. The WECC ADS 2032 hydropower dataset (generation and flexibility) included an update of the climate year conditions (2018 calendar year), consistency in representation across the entire US WECC footprint, updated hydropower operations over the core Columbia River, and a higher temporal resolution (weekly instead of monthly)[3] associated with a GridView software update (weekly hydro logic). Proprietary WECC utility hydropower data were used when available to develop the monthly and weekly datasets and were completed with HydroWIRES B1 methods to develop the Hydro 923 plus (now RectifHydPlus weekly hydropower dataset) [4] and the flexibility parameterization [5]. The team worked with Bonneville Power Administration to develop hydropower datasets over the core Columbia River representative of the post-2018 change in environmental regulation (flex spill). Ramping data are considered proprietary, were leveraged from WECC ADS 2030, and were not provided in the release, nor are the WECC-member hydropower data. The generator database was first updated by WECC. Based on a review of hourly generation profiles, 16 facilities were transitioned from fixed schedule to dispatchable (380.5MW). The operations of the core Columbia River were updated based on Bonneville Power Administration's long-term hydro-modeling using 2020-level of modified flows and using fiscal year 2031 expected operations. The update was necessary to reflect the new environmental regulation (EIS2023). The team also included a newly developed extension over Canada [6] that improves upon existing data and synchronizes the US and Canadian data to the same 2018 weather year. Canadian facilities over the Peace River were not updated due to a lack of available flow data. The datasets have been incorporated into the 2034 ADS and are in active use by WECC and the community. WECC ADS 2034 hydropower datasets contain generation at weekly and monthly timesteps, for US hydropower plants, monthly generation for Canadian hydropower plants, and the two merged together. Separate datasets are included for generation by hydropower plant and generation by individual generator units. Only processed data are provided. Original WECC-utility hourly data are under a non-disclosure agreement and for the sole use of developing the dataset. [1] Voisin, N., Broman, D., Abernethy-Cannella, K., Bracken, C., Son, Y., & Harris, K. (2025). WECC ADS 2034 Hydropower Generation Datasets (3.2) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.12617457 [2] https://github.com/HydroWIRES-PNNL/weccadshydro/ [3] Voisin, N., Harris, K. M., Oikonomou, K., Turner, S., Johnson, A., Wallace, S., Racht, P., et al. (2022). WECC ADS 2032 Hydropower Dataset (PNNL-SA-172734). See presentation (Voisin N., K.M. Harris, K. Oikonomou, and S. Turner. 04/05/2022. "WECC 2032 Anchor Dataset - Hydropower." Presented by N. Voisin, K. Oikonomou at WECC Production Cost Model Dataset Subcommittee Meeting, Online, Utah. PNNL-SA-171897.). [4] Turner, S. W. D., Voisin, N., Oikonomou, K., & Bracken, C. (2023). Hydro 923: Monthly and Weekly Hydropower Constraints Based on Disaggregated EIA-923 Data (v1.1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.8212727 [5] Stark, G., Barrows, C., Dalvi, S., Guo, N., Michelettey, P., Trina, E., Watson, A., Voisin, N., Turner, S., Oikonomou, K. and Colotelo, A. 2023 Improving the Representation of Hydropower in Production Cost Models, NREL/TP-5700-86377, United States. https://www.osti.gov/biblio/1993943 [6] Son, Y., Bracken, C., Broman, D., & Voisin, N. (2025). Monthly Hydropower Generation Dataset for Western Canada (1.1.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.14984725 Requests for access to the files in this release containing the code and underlying raw data should be made to: Nathalie Voisin (nathalie.voisin@pnnl.gov). Requests for WECC-member provided hydropower data need to be directed to WECC (https://www.wecc.org/committees/pcds).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.005 | 0.004 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.010 | 0.014 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".