Data-Driven Inverse Optimization with Applications in Electricity Markets
Bibliographic record
Abstract
Due to the increasing penetration of renewable resources and demand response instruments in the electricity markets, generation planning models have become more complex and require detailed information on the inherent structure of the system, including generator and demand parameters. Demand should be met by cost-effective, adaptable, and efficient power plants to ensure that it is met even in the worst-case scenarios, such as an unanticipated peak or the failure of a critical generating unit. On the other hand, there is a need to consider short-term details in the Planning problems to address the needed system flexibility due to sudden changes in demand and renewables generation. Such short-term details increase the size of the models and their related computations. As a result, there is a trade-off between the complexity of the computation and the level of short-term operational details, which should be considered. \n \nAccessing electricity infrastructure data in North America is often difficult due to the lack of open data standards and the proprietary nature of much of the data. The regulations and policies surrounding the data also vary significantly from province to province, making it difficult to access the data uniformly. Additionally, privacy and security considerations can limit access even further. Despite these limitations, there are indirect methods such as inverse optimization(IO) to derive the market parameters using publicly available data; examples of these parameters include generator costs of generation, their capabilities, etc. The discovery of unobservable information via IO could aid energy models to account for operational details without increasing the complexity of their problem. Furthermore, this information can inform policymakers on potential interventions to improve the efficiency of the electricity market. \n \nIn this research, a MIP model is developed to incorporate capital and operational costs associated with long-term planning problems. The operating costs of each technology are assumed to be approximated by a series of step-wise functions so that model outcomes, such as generation output, are as close as possible to real-world electricity market generation. The proposed method employs a two-stage algorithmic framework using data-driven inverse optimization and regression. In the first stage, constraints are generated based on relationships between cost and electricity prices. In the second stage, these constraints on costs are added to a problem that finds and reconciles the parameters of the cost functions. To evaluate the performances of the proposed IO-based method, it was applied to a DC-OPF model using the IEEE 24-bus system, which helped eliminate power flow constraints. This approach was then applied to a long-term planning model using Ontario's electricity market data. The results indicate that the proposed approach could find a close solution to the conventional models. In the long-term planning model, the IO-based approach showed more moderate investment policies, while the traditional methods tend to over or under-invest.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".