Process and policy insights from an intercomparison of open electricity system capacity expansion models
Bibliographic record
Abstract
Abstract This study performs a detailed intercomparison of four open-source electricity capacity expansion models—Temoa, Switch, GenX, and USENSYS—to evaluate (1) how closely the results of these models align when inputs and configurations are harmonized, and (2) the degree to which varying model configurations affect outputs. We harmonize the inputs to each model using PowerGenome and use clearly defined scenarios (policy conditions) and configurations (model setup choices). This allows us to isolate how differences in model structure affect policy outcomes and investment decisions. Our framework allows each model to be tested on identical assumptions for policy, technology costs, and operational constraints, allowing us to focus on differences that arise from inherent model structures. Key findings highlight that, when harmonized, models produce very similar capacity portfolios under current policies and net-zero scenarios, with less than 1% difference in system costs for most configurations. This agreement among models allows us to focus on how configuration choices affect model results. For instance, configurations with unit commitment constraints or economic retirement yield different investments and system costs compared to simpler configurations. Our findings underscore the importance of aligning input data and transparently defining scenarios and configurations to provide robust policy insights.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".