The Application of Artificial Neural Networks With Small Data Sets: An Example for Analysis of Fracture Spacing in the Lisburne Formation, Northeastern Alaska
Bibliographic record
Abstract
Summary Artificial neural networks (ANNs) have been used widely for prediction and classification problems. In particular, many methods for building ANNs have appeared in the last 2 decades. One of the continuing important limitations of using ANNs, however, is their poor ability to analyze small data sets because of overfitting. Several methods have been proposed in the literature to overcome this problem. On the basis of our study, we can conclude that ANNs that use radial basis functions (RBFs) can decrease the error of the prediction effectively when there is an underlying relationship between the variables. We have applied this and other methods to determine the factors controlling and related to fracture spacing in the Lisburne formation, northeastern Alaska. By comparing the RBF results with those from other ANN methods, we find that the former method gives a substantially smaller error than many of the alternative methods. For example, the errors in predicted fracture spacing for the Lisburne formation with conventional ANN methods are approximately 50 to 200% larger than those obtained with RBFs. With a method that predicts fracture spacing more accurately, we were able to identify more reliably the effects on the spacing of such factors as bed thickness, lithology, structural position, and degree of folding. By comparing performances of all the methods we tested, we observed that some methods that performed well in one test did not necessarily do as well in another test. This suggests that, while RBF can be expected to be among the best methods, there is no "best universal method" for all the cases, and testing different methods for each case is required. Nonetheless, through this study, we were able to identify several candidate methods and, thereby, narrow the work required to find a suitable ANN. In petroleum engineering and geosciences, the number of data is limited in many cases because of expense or logistical limitations (e.g., limited core, poor borehole conditions, or restricted logging suites). Thus, the methods used in this study should be attractive in many petroleum-engineering contexts in which complex, nonlinear relationships need to be modeled by use of small data sets.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".