MétaCan
Menu
Back to cohort
Record W4290057726

How to improve geospatial data usability: from metadata to quality-aware GIS community

2007· preprint· en· W4290057726 on OpenAlexaffabout
Rodolphe Devillers, Y. Bedard, M. Gervais, Robert Jeansoulin, François Pinet, Michel Schneider, Lotfi Bejaoui, M.A. Levesque, M. Salehi, A. Zargar

Bibliographic record

VenueHAL (Le Centre pour la Communication Scientifique Directe) · 2007
Typepreprint
Languageen
FieldSocial Sciences
TopicGeographic Information Systems Studies
Canadian institutionsUniversité LavalMemorial University of Newfoundland
Fundersnot available
KeywordsGeospatial analysisMetadataUsabilityComputer scienceGeospatial metadataGeospatial PDFGeographic information systemWorld Wide WebMeta Data ServicesQuality (philosophy)Data elementRemote sensingGeographyHuman–computer interaction
DOInot available

Abstract

fetched live from OpenAlex

/ The field of Geomatics/GISciences witnessed major changes since its origins in the 1960s. If developments in this field first looked at ways to transfer and store spatial and semantic information from paper documents into computers (e.g. spatial data structures - raster vs vector; scan; topology), the focus moved later to the design of more advanced ways to store, retrieve and analyze geospatial data. The field grew exponentially and, in the last two decades, organisations started to realise that a large volume of geospatial data were produced, but mostly remained unknown from many potential users (even within a same organisation). In order to have a better Return on Investment (ROI), and to encourage an increased use, organisations and countries started to develop initiatives like Digital Libraries and Spatial Data Infrastructures. This moved the research focus from the systems themselves to data transfers and reuse issues (e.g. interoperability, metadata, ontologies, data fusion). We are nowadays entering the next phase: the widespread usage of these data by people who haven't collected or integrated them and may not have been initially targeted as primary users. With the increasing ease of access to geospatial data and with the user-friendliness of today's Geographic Information Systems (GIS) and related web-based solutions, geospatial information is more than ever reaching the hands of the general public (e.g. Google Earth, Virtual Earth). Similarly, expert users in different fields of applications have also seen their number increasing by orders of magnitudes worldwide. Digital Libraries and Spatial Data Infrastructures are now facing huge downloads. For instance, the Canadian Web portal Geobase that provides free access to geospatial data (e.g. DEM, road network), had an increase from about 210,000 downloads in 2003-2004, to about 600,000 in 2004-2005, and more than 2,200,000 downloads in 2005-2006. In addition, GIS applications are not anymore restricted to traditional land/resources-related uses but now reach most disciplines, ranging from Science/Engineering to Human/Social or Medical Sciences. Consequently, most of the new users have limited or no knowledge of the geospatial field and the underlying nature of spatial referencing (e.g. reference systems/projections, scale, generalisation, accuracy). Furthermore, it appears that problems faced by users from the general public are also emerging more often than ever before. Similarly, users having an expertise in geomatics often just cannot know the key characteristics which are necessary to assess the usefulness of the data being downloaded, nor the added uncertainty resulting from the integration of such datasets. Consequently, an increasingly important research agenda is now to make sure the level of usability of spatial data is better known for contexts that were not always planned when the data were collected. One objective is to try to reduce the risks of misuse of these data and the risks of potential accidents that could result from these misuses.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.046
metaresearch head score (Gemma)0.020
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow), Science and technology studies, Scholarly communication, Open science
Consensus categoriesMetaresearch, Open science
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.812
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0460.020
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.001
Science and technology studies0.0030.001
Scholarly communication0.0030.001
Open science0.0080.014
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.073
GPT teacher head0.327
Teacher spread0.255 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2007
Admission routes2
Has abstractyes

Explore more

Same venueHAL (Le Centre pour la Communication Scientifique Directe)Same topicGeographic Information Systems StudiesFrench-language works237,207