Solar Irradiance Forecasting with Visible Spectrum Sky View Images and Random Forest Regression
Bibliographic record
Abstract
Integrating solar photovoltaic (PV) energy into the electrical grid is challenging due to its inherent production variability stemming from clouds. Solar forecasts and grid inertia from our current electrical infrastructure help maintain grid stability and availability in the wake of PV production fluctuations. However, to transition to future high-renewable-energy-penetration grids, grid operators will require improved forecasting to proactively deploy grid-balancing assets and maintain grid reliability. In this work, we use 3-channel, 8-bit-depth visible spectrum sky-view images from a year-long dataset captured at 15-second intervals, and smart persistence features to predict broadband irradiance in North Cape, Canada over several forecasting horizons. Multiple features are extracted from a single sky-view image to train and test random forest regression models and evaluate feature importance over forecast horizons from 1 minute to 1 hour. The practicality and usefulness of the models are evaluated using testing and training computational times, and the prediction accuracies relative to the smart persistence model, respectively. The models achieve a positive skill score in the range of 0 to 0.1, outperforming the smart persistence model for all forecast horizons.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".