An Intelligent Selection Method of Main Controlling Factors for Tight Gas Reservoirs Productivity Based on Improved Harris Hawk Algorithm
Bibliographic record
Abstract
Identifying the main controlling factors of oil and gas productivity and making accurate forecasts is crucial for efficient development and reservoir reconstruction. Tight gas reservoirs have complex geological conditions and high-dimensional, nonlinear factors that traditional methods struggle to analyze, complicating the identification of main factors and accurate productivity prediction. In the present work, an improved Harris hawk algorithm (TVLHHO), incorporating a nonlinear escape energy strategy and a time-varying leader structure, is proposed for the feature selection of the main controlling factors of tight gas productivity. The algorithm expands the search space of feature subsets, enhances convergence speed, reduces the risk of local optima, and ensures the accuracy of feature selection. Using a certain tight sandstone gas field as a case study, 59-dimensional features, including geological, logging, and fracturing properties, were used as inputs to study the main controlling factors affecting unimpeded flow rate. Initially, Pearson correlation analysis and XGBoost were used for preliminary feature selection, reducing the features to 23 dimensions. The TVLHHO algorithm was then employed to optimize the selection of the main controlling factors. Through an iterative process of updating the feature subsets and validating predictions, the optimal controlling factors identified included displacement fluid, deviation angle, azimuth angle, fracture half-length, gas relative density, perforation thickness, and initial gas saturation. The study shows that compared with six other well-known algorithms, TVLHHO not only demonstrates faster convergence but also achieves an R 2 mean value exceeding 0.9 on the evaluator, resulting in the highest prediction accuracy. Furthermore, the main controlling factors selected by TVLHHO were used to predict the unimpeded flow rate, effectively identifying the distribution of high- and low-capacity wells. This validates the rationality of the TVLHHO feature selection results and demonstrates the algorithm’s feasibility and effectiveness in practical applications. It provides a powerful tool for identifying the main controlling factors in tight gas reservoirs, addressing challenges related to high-dimensional data and complex relationships, and ultimately offering a more precise foundation for productivity prediction and optimization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".