Data interoperability across borders : a case study of the Abbotsford-Sumas aquifer (British Columbia-Washington State)
Bibliographic record
Abstract
The ability to integrate data from multiple sources is central to geographic information science (GIS). Although data integration is an active field of research in the GIS community, a number of challenges remain unresolved. Interoperability research addressing data integration challenges experienced by institutions in an international setting also remains sparse. Groundwater is an example of an environmental phenomenon which does not respect political borders, and its management requires data from multiple jurisdictions. The Abbotsford-Sumas aquifer, straddling the Canada US border, is used as a case study to explore integration challenges in an international setting. Development of groundwater management practices to ensure a sustained source of good quality groundwater is dependent, on an understanding of the conceptual model of the aquifer. Due to a lack of geophysical studies, geological information contained in the water well reports, is the chief source of depth-specific lithological information. The use of this information in constructing the conceptual model is constrained by poor data quality and a lack of an integrated and standardized lithological database. To achieve the research goals of exploring integration challenges in an international setting, lithological datasets from BC and Washington State are integrated. The resultant lithological database is used to test the usability of water well reports for constructing the conceptual model. Numerous interoperability challenges such as data availability, lack of metadata, data quality and formats, database structure, semantics, policies and cooperation are identified as inhibitors of data integration. Despite the numerous challenges the lithological database is useful in constructing a generalized conceptual model. This research is important as it presents challenges to data integration that should be considered as a starting point for environmental management projects.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.016 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.007 |
| Science and technology studies | 0.011 | 0.004 |
| Scholarly communication | 0.007 | 0.004 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".