Geospatial Data Infrastructure Portals: Using the National Atlas as a Metaphor
Bibliographic record
Abstract
The concept of geospatial data infrastructure (GDI) has been put into practice in some countries by providing portals allowing users to search for multiple geospatial data sets. Our review and inquiry activities show that current portals suffer from two potential setbacks: inappropriate navigation tools and a lack of supports for users’ understanding. This article defines a new approach to portal development using the atlas as metaphor. This allows the atlas to be used not only to access assorted thematic maps but also to discover data sets. Within the atlas, an information structure plays an important role in organizing the content. Metadata published by providers are incorporated into this structure as metadata summaries. Based upon the topical relevancy of the data, each metadata summary is linked to a specific map within a particular topic. These summaries can be represented as symbols to support discovery tasks, either loosely or strictly defined. Browsing can be used to deal with the first via navigations and map interfaces. Searching can be used to deal with the second via explorer and search presentation interfaces. A working prototype to enable users to browse and search is built as a Flash-based ArcIMS client. Whether browsing or searching, users are offered interfaces to effectively assess data suitability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.004 | 0.008 |
| Science and technology studies | 0.002 | 0.006 |
| Scholarly communication | 0.010 | 0.022 |
| Open science | 0.001 | 0.006 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.007 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".