Development of Digital/Visual Twin for Real‐Time Leak Detection in Gas Pipelines Under Multiphase Flow Conditions
Bibliographic record
Abstract
ABSTRACT Leak detection (LD) in gas pipelines (GPs) is critical for ensuring operational safety and environmental protection. This study presents a novel digital/visual twin for detecting single‐ and multiple leaks in GPs under both single‐ and multiphase flow conditions. The framework of the digital twin leverages experimental data from a multiphase flow‐testing loop and synthetic data generated using OLGA software to validate and optimize machine learning (ML) models for leak detection and localization. Several ML models, including random forest (RF), support vector machine (SVM), k ‐nearest neighbors ( k ‐NNs), decision tree regression (DTR), and eXtreme gradient boosting (XGBoost), were tested individually for their ability to classify leak conditions and localize leaks. Initial results showed moderate performance for individual models, with accuracies ranging from 42% to 57%. However, a significant improvement was observed through the use of advanced techniques such as stacking models, feature engineering, and data averaging. The final stacking regressor model, which combined the strengths of RF, k ‐NN, and SVM, outperformed the individual models, achieving R 2 values exceeding 0.96 with an accuracy of 90% in complex multiple leak scenarios. The digital twin system integrates this ML framework with real‐time data visualization, allowing operators to visualize offshore pipeline conditions, detect leaks, and localize leak positions using a virtual twin representation of the physical pipeline. The virtual twin provides an interactive, high‐fidelity interface that enables users to monitor and analyze leak events as they occur, enhancing situational awareness and decision‐making capabilities. The combination of advanced ML techniques and digital twin technology provides a robust and accurate solution for real‐time LD in offshore pipelines. It significantly improves detection performance in multiphase flow conditions. This innovative approach sets a new benchmark for offshore pipeline monitoring systems, offering superior LD capabilities under a range of operational conditions. The system is readily adaptable for integration with SCADA platforms and pipeline monitoring infrastructures, supporting deployment in offshore oil and gas operations, industrial gas distribution networks, and critical energy corridors where early LD is essential.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".