Paskapoo groundwater study part VII: Alberta groundwater wells data dictionary - a view to groundwater data modeling
Bibliographic record
Abstract
The Geological Survey of Canada (GSC) under the Natural Resources Canada, Earth Sciences Sector Groundwater Program has evaluated key aquifers across Canada. The Paskapoo Aquifer System is one of the largest systems identified as a key aquifer in the country. Data requirements to study this 66,768 km2 area went beyond the usual requirements of local scale aquifer studies and a relational database implementation (EarthFX Inc. 2005) of the Alberta Environment (AENV), Groundwater Information Center (GIC) water well database was used to facilitate analysis of the required data. The data dictionaries in this report define the data that was received from AENV "as is" prior to being imported into the EarthFX data model. The Data Dictionary for the April 2003 release of the GIC database defines the data that was subsequently imported into the EarthFX model and provides background information for how the data was interpreted and used in the study. The Data Dictionary for the May 2007 release of the GIC database is an accompanying standalone document that can be used as a guide for evaluating recent changes to the GIC database. The primary purpose of the dictionaries is to help users interpret and understand the GIC data structure and its content for the purpose of extracting meaningful data that can be use for local and regional groundwater studies. Presenting the background information in the context of relational database concepts also provides a view for future database modeling, storage, and management of the data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.007 | 0.017 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.016 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".