Research metadata on the web: selected geospatial data and metadata directories
Bibliographic record
Abstract
This paper will present an overview of the types of research metadata that are available over the Web.It will include a discussion of the electronic tools which are being used to create metadata online and the directories which are being created at the national, federal, and state levels. I The Value of MetadataMetadata allow researchers to identify potentially useful data sets collected by others.Gary Waggoner, Biological Research Division, U.S. Geological Survey, states "Metadata refers to data that are used to describe a database (e-g., describing the extent of the data, coverage, scale, what methods were used to collect the data, by whom and when the data was collected, etc.)With valid and complete metadata, someone can learn enough about a database (without communicating with the "owner" of the data) to determine if the data would be of use or interest to them."I geographic information systems (GIs).In the U.S., the impetus for establishing standards and directories for this type of metadata came from the federal government.Many agencies, e.g., U.S. Geological Survey, Environmental Protection Agency, U.S. Department of Agriculture, share a common base of spatial data needs upon which their unique projects are overlaid.The federal government mandated the development of metadata standards and directories to reduce the costly duplication of geospatial data collection.The Federal Geographic Data Committee (FGDC) was
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.066 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.005 |
| Science and technology studies | 0.034 | 0.012 |
| Scholarly communication | 0.047 | 0.010 |
| Open science | 0.045 | 0.062 |
| Research integrity | 0.000 | 0.005 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".