Formulation of Huge Lattice Spatial Adjacency Matrices With Non-rectangular Shape of Socio-economic Grid-Cell Data for the Analysis of Sustainable Economy With High Computational Efficiency
Bibliographic record
Abstract
The advantage of using grid-cell data for socio-economic analysis should be the feasibility to incorporate satellite data that will enrich the regional analysis and has an important role to observe the relationship between socio-economics and nature. This advancement corresponds to the sustainable development goals that balance the socio-economic quality in harmony. In order to perform the analysis, formulation of a spatial adjacency matrix has an important role to project the spatial relationship within regions. However, no precedent research provided a practical formulation for the spatial adjacency matrix in grid-cell data structure (Fitrianto & Tanaka, 2017).The general process that used shapefiles solely, which store geometry and attribute information for the spatial features (ESRI, 1998) to construct the adjacency matrix is not suitable. The problem arises due to the existence of NA cells that represent non-inhabitant areas such as water bodies, yet the shapefile does not contain this information inside the municipal body. The NA cells create a non-rectangular lattice and it is important to exclude them in the analysis to correctly project the real information.This article provides a method to precisely project the real information by using Kronecker product to construct the adjacency matrix and applying a projection matrix to eliminate the NA cells (Tanaka & Nishii, 2009). It showed eminent efficiency compared with commonly used R package called spdep. Experimental results verified that this method, even for huge dimension with a trillion elements, produces more than 2000 times faster elapsed time than the package.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.007 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".