Using a boundary-corrected wavelet transform coupled with machine learning and hybrid deep learning approaches for multi-step water level forecasting in Lakes Michigan and Ontario
Bibliographic record
Abstract
Accurate water level (WL) forecasting is important for water resources management and planning purposes in the Great Lakes. The objectives of this research are two-fold. The first objective is to apply machine learning (ML) (i.e., random forest (RF) and support vector regression (SVR)) and hybrid convolutional neural network(CNN)-long-short term memory (LSTM) deep learning (DL) models for multi-step (i.e., one-, two- and three-monthly step ahead) WL forecasting in the Great Lakes (Michigan and Ontario). The second objective is to integrate the boundary corrected (BC) maximal overlap discrete wavelet transform (MODWT) with SVR, RF, and CNN-LSTM models to improve the performance of the individual models. By employing a BC-wavelet decomposition method, the ‘future data’ issue (i.e., data from the future that is not available), often overlooked in the literature and a major barrier to achieving realistic forecasting performance is overcome. For Lakes Michigan and Ontario, 1212 monthly WL (m) records (spanning Jan 1918–Dec 2018) were used to develop the models. For the non-wavelet-based models (SVR, RF, and CNN-LSTM), candidate model inputs included the WL recorded over the previous 12 months. For the BC-MODWT-based models (BC-MODWT-SVR, BC-MODWT-RF, and BC-MODWT-CNN-LSTM), the lagged input time series were decomposed into BC-wavelet and scaling coefficients by using different mother wavelets (Haar, Daubechies, Symlets, Fejer-Korovkin and Coiflets), filter lengths (from two up to 12) and decomposition levels (from one up to seven). For each method (SVR, RF, and CNN-LSTM), mother wavelet, and decomposition level a model was generated. For both wavelet- and non-wavelet-based models, the particle swarm optimization (PSO) method was used to select the most appropriate inputs to include in the proposed multi-step WL forecasting models. The datasets were partitioned into calibration and validation subsets. After calibrating the models, various performance evaluation metrics, e.g., coefficient of determination (R2), root mean square error (RMSE), mean absolute error (MAE), root mean square percentage error (RMSPE), mean absolute percentage error (MAPE) and the Nash-Sutcliffe efficiency coefficient (NSC) were used to assess model accuracy. Of the ML models, the SVR outperformed RF while the DL models outperformed the ML models for each forecast lead time (one-, two-, and three-step(s) ahead). Results from this case study indicate that not all wavelet families and decomposition levels perform equally and in some cases, the wavelet-based models do not improve performance over the non-wavelet-based models. However, the BC-MODWT-CNN-LSTM using suitable mother wavelets (e.g., Haar) outperforms the individual ML and BC-MODWT-ML-based models. More accurate forecasts were obtained for Lake Michigan although the performance in both Great Lakes was accurate. The outcomes of this research indicate that the BC-MODWT-CNN-LSTM model is a promising tool for generating accurate WL forecasts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".