Soft sensor based on multi‐phase stacking ensemble model with self‐selected primary learner for batch processes
Bibliographic record
Abstract
Abstract In batch processes, existing soft sensing methodologies encounter substantial challenges when confronted with nonlinearity and multi‐phase issues. In response to these challenges, an innovative soft sensing framework known as the multi‐phase stacking ensemble model with self‐selected primary learner is proposed. The main innovation of this framework lies in the solution to the primary learner selection issue within the stacking model. To commence, the batch process is divided into multiple phases employing a Gaussian mixture model, thereby establishing local ensemble models for each phase. Subsequently, the iterative self‐selection of primary learners strategy is proposed, which iteratively selects suitable primary learners for these models, optimizing their combination of primary learners for each local model. This primary learner selection strategy effectively enhances the accuracy of predictions in the stacking ensemble model. To further enhance the performance, Bayesian optimization is utilized to tune the hyperparameters of each local ensemble model. This step guarantees optimal performance of the model across diverse phases. Extensive simulation experiments are conducted on an industrial penicillin fermentation process to validate the effectiveness of the proposed framework. According to the findings, the model demonstrated superior performance compared to existing single‐learner soft sensing methods and commonly utilized ensemble‐based soft sensing methods in terms of both R2 score (0.97584) and RMSE (0.0513). Overall, this framework offers a novel approach for selecting primary learners in stacking ensemble models and enhancing the predictive performance in batch processes for soft sensing.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".