Taking a Different Approach to Drilling Data Aggregation to Improve Drilling Performance
Bibliographic record
Abstract
Abstract Currently, there is a multitude of commercially available real-time drilling data aggregation and distribution systems, yet the industry remains plagued with issues that limit the usability and effectiveness of data before, during, and after a well is drilled. There are challenges with moving, merging, analyzing, qualifying, and formatting data as well as having access to like-data in sufficient quantity and on a reliable data frequency. This paper discusses a novel, adaptable, and low cost approach to building a system to drive drilling performance and set the stage for future automation. The Operator embarked on a project to develop a powerful, low cost system in order to leverage both high and low frequency data to gain value from real-time data models and algorithms at the rig site. High frequency data is defined as 1 to 100 Hertz data frequency. Low frequency data is defined as longer than once per hour or asynchronous, and is usually contextual - BHA information, mud reports, rig state, etc. Existing commercial systems fail to meet the requirements due to multiple factors. These include an inability to handle and process high frequency data, communicate with different protocols, and work across different proprietary systems. The result leads to higher costs, extra human resources and efforts, and a lack of consistency across a diverse rig fleet. Druing this process severe data quality issues were discovered at the rig site and needed the flexibility to modify, replace, or add sensors and data streams to remedy the problem. After evaluating more than thirty potential process controls and other industry applications, a software solution was selected, prototyped, tested and deployed to seven North American land rigs within a ten month period. This effort employed the agile development methodology which is an incremental, iterative work cadence using empirical feedback for rapid deployment of updated versions. The system was designed to take in all forms of data, file types, and communication protocols for seamless integration. The system includes rig state determination, data quality verification, a real time Bayesian model for analytics and smart alarms, integration to the Daily Drilling Report (DDR) database, real-time visualizations, and an open application layer with a Human Machine Interface (HMI) - all at the rig site. Ultimately the platform can also be used as a building block to assist automated drilling due to it being a Supervisory Control Advisory and Data Acquisition (SCADA) system although this is not the goal for this project.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".