GeoHexViz: A Python package for the visualizinghexagonally binned geospatial data
Bibliographic record
Abstract
Geospatial visualization is often used in military operations research to convey analyses to both analysts and decision makers.For example, it has been used to help commanders coordinate units within a geographic region (Feibush et al., 2000), to depict how terrain impacts vehicle performance (Laskey et al., 2010), and inform training decisions in order to meet mission requirements (Goodrich et al., 2019).When such analyses include a large amount of point-like data, combining geospatial visualization and binning -in particular, hexagonal binning given its properties such as having the same number of neighbours as sides, the centre of each hexagon being equidistant from the centres of its neighbours, and that hexagons tile densely on curved surfaces (Carr et al., 1992;Sinha, 2019) -is an effective way to summarize and communicate the data.Recent examples in the military and public safety domains include assessing the impact of infrastructure on Arctic operations (Hunter et al., 2021) and communicating the spatial distribution of COVID-19 cases (Shaito & Elmasri, 2021) respectively.However, creating such visualizations may be difficult for many since it requires in-depth knowledge of both Geographic Information Systems and analytical techniques, not to mention access to software that may require a paid license, training, and in some cases knowledge of a programming language such as Python or JavaScript.To help reduce these barriers, GeoHexViz -which produces publication-quality geospatial visualizations with hexagonal binning -is a Python package that provides a simple interface, requires minimal in-depth knowledge, and either limited or no programming.The result is an analyst being able to spend more time doing analysis and less time producing visualizations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.083 | 0.031 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".