ISPRS-SHY – OPEN DATA COLLECTOR FOR SUPPORTING GROUND TRUTH REMOTE SENSING ANALYSIS
Bibliographic record
Abstract
Abstract. The 2021 Scientific Initiatives in ISPRS funded this project called ISRS-SHY from “SHare mY ground truth”. It was intended as a collector of geographic data to support image analysis by sharing the necessary ground truth data needed for rigorous analysis. Regression and classification tasks that use remote sensing imagery necessarily require some control on the ground. The rationale behind this project is that often data on the ground is collected during projects, but is not valued by sharing across projects and teams globally. Internet has improved the way that data are shared, but there are still limitations related to discoverability of the data and its integrity. In other words, data are usually kept in local storage or, if in an accessible server, they are not documented and therefore they will not be picked up during search. In this initiative we created a portal using the Geonode environment to provide a hub for sharing data between research groups and openly to the community. The portal was then tested within the framework of three projects, with several participants each. The data that was uploaded and shared covered all types of geographic data formats and sizes. Further sharing was done in the context of teaching activities in higher education.The results show the importance of creating easy means to find data and share it across stakeholders. Qualitative results are discussed, and future steps will focus on quantitative assessment of the portal’s usage, e.g. number of registered users in time, number of visits, and other key performance indicators. The results of this project are to be considered also in light of the effort in the scientific community to make research data available, i.e. FAIR - Findability, Accessibility, Interoperability, and Reuse of digital assets.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.041 | 0.043 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.008 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.005 | 0.009 |
| Open science | 0.003 | 0.012 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.028 | 0.026 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".