Un modèle UML et des contraintes OCL pour les entrepôts de données spatiales. De la représentation conceptuelle à l'implémentation
Bibliographic record
Abstract
Spatial Data Warehouses (SDW) and Spatial OLAP (SOLAP) systems represent an effec-tive solution to perform spatial analysis on geographical phenomena. However, the quality of such analysis heavily depends on the quality of stored data and how these data are explored: how the different indicators are computed (What aggregate functions are applied to summa-rize the measures and in what order these functions are applied?). In this context, a number of studies have been attempted to address the issues of data quality in SDW by using Integ-rity Constraints (IC). In this paper, motivated by the lack of Model Driven Architecture (MDA)-based implementations, we propose a conceptual framework based on two new clas-sifications to ease identification and implementation of SDW IC. Moreover, following an MDA approach, we propose the MDA-based modeling of most IC categories using the UML (Unified Modeling Language) and OCL (Object Constraint Language) standard languages; and show the automatic implementation of some IC classes using an MDA-based code gen-erator, called Spatial OCL2SQL. / Les Entrepôts de Données Spatiales (EDS) et les systèmes SOLAP représentent une solution efficace pour l'analyse spatiale de phénomènes géographiques. Cependant, la qualité de cette analyse dépend fortement de la qualité des données entreposées et de la manière dont l'exploration de ces données est réalisée : comment les différents indicateurs ou agrégats sont calculés ? (i.e. quels opérateurs d'agrégation sont appliqués aux différentes mesures ? Et dans quel ordre ils sont appliqués ?). Dans ce contexte, quelques travaux essaient de mieux maîtriser la qualité des informations dans les entrepôts de données (spatiales) par exemple par le biais de contraintes d'intégrité. Dans cet article, motivé par le manque d'implémentations basées sur une approche MDA (Model Driven Architecture), nous proposons un «framework» conceptuel basé sur deux nouvelles classifications pour faciliter l'identification et l'implémentation des contraintes d'EDS ; dans le cadre de l'approche MDA, nous proposons la modélisation et l'implémentation de la plupart des types de contraintes en utilisant les standards UML(Unified Modeling Language) et OCL (Object Constraint Language), et le générateur de code Spatial OCL2SQL.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.018 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.007 | 0.007 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.004 | 0.004 |
| Insufficient payload (model declined to judge) | 0.005 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".