Global Progress in Oil and Gas Well Research Using Bibliometric Analysis Based on VOSviewer and CiteSpace
Bibliographic record
Abstract
Studies related to oil and gas wells have attracted worldwide interest due to the increasing energy shortfall and the requirement of sustainable development and environmental protection. However, the state of oil and gas wells in terms of research characteristics, technological megatrends, article-produced patterns, leading study items, hot topics, and frontiers is unclear. This work is aimed at filling the research gaps by performing a comprehensive bibliometric analysis of 6197 articles related to oil and gas wells published between 1900 and 2021. VOSviewer and CiteSpace software were used as the main data analysis and visualization tools. The analysis shows that the annual variation of article numbers, interdisciplinary numbers, and cumulative citations followed exponential growth. Oil and gas well research has promoted the expansion of research fields such as engineering, energy and fuels, geology, environmental sciences and ecology, materials science, and chemistry. The top 10 influential studies mainly focused on shale gas extraction and its impact on the environment. More studies were produced by larger author teams and inter-institution collaborations. Elkatatny and Guo have greatly contributed to the application of artificial intelligence in oil and gas wells. The two most contributing institutions were the Southwest Petr Univ and China Univ Petr from China. The People’s Republic of China, the US, and Canada were the countries with the most contributions to the development of oil and gas wells. The authoritative journal in engineering technology was J Petrol Sci Eng, in environment technology was Environ Sci Technol, in geology was Aapg Bull, and in materials was Cement Concrete Res. The keyword co-occurrence network cluster analysis indicated that oil well cement, new energy development, machine learning, hydraulic fracturing, and natural gas and oil wells are the predominant research topics. The research frontiers were oil extraction and its harmful components (1992–2016), oil and gas wells (1997–2016), porous media (2007–2016), and hydrogen and shale gas (2012–2021). This paper comprehensively and quantitatively analyzes all aspects of oil and gas well research for the first time and presents valuable information about active and authoritative research entities, cooperation patterns, technology trends, hotspots, and frontiers. Therefore, it can help governments, policymakers, related companies, and the scientific community understand the global progress in oil and gas well research and provide a reference for technology development and application.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.030 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.239 | 0.249 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.007 | 0.006 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".