Mathematical Modeling and Analysis of Patterns in Structured Collections of Big Data
Bibliographic record
Abstract
This paper addresses the issue of creating and applying mathematical models and methods for finding generalized solutions when working with structured collections of “big data”. We reviewed the modern methodologies used to solve problems of this class. The mathematical model presented describes an ordered set of all subsets formed from a finite ordered base set of arbitrary size and data type. We explored a set of functional dependencies of five discrete input variables to work with this mathematical model. Some of these functional dependencies are derived for specific solutions with specified boundary conditions. The paper also presents examples of how the derived functional dependencies are applied in the implementation of mathematical methods using this model. This required us to conduct a comparative assessment of the search time for a solution with and without the use of these mathematical methods. Comparative graphs are demonstrated to show the rate of increase in the number of operations depending on the size of the original finite base set with and without the use of these mathematical methods. As a result of this, logical conclusions are drawn regarding the impact of mathematical methods for working with structured collections on minimizing time and computational resources.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.020 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.003 | 0.010 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".