An approach for model-driven data reengineering = Un enfoque de reingeniería de datos dirigido por modelos
Bibliographic record
Abstract
Esta tesis se centra principalmente en la aplicacion de tecnicas MDE a un proceso de reingenieria de datos. En concreto, analizamos en que medida el uso de modelos facilita la implementacion de una mejora de la calidad en los datos de un sistema legado mediante la conversion de esquemas, que es un escenario comun de modernizacion. La conversion de esquemas implementada en nuestra solucion aborda la inferencia de restricciones de integridad referencial (declaradas en base de datos como claves ajenas) junto con la comprobacion y correccion de los niveles de normalizacion en un esquema de datos. Se deben proporcionar diferentes tecnicas para el descubrimiento de claves ajenas para obtener resultados mas fiables. Ademas, se ha proporcionado automatizacion al proceso de migracion mediante una herramienta software que soporta la definicion y ejecucion de procesos de migracion y que ha sido validada atraves tomando como caso de estudio nuestro proceso de reingenieria de datos. Por otro lado, las soluciones MDE requieren de la integracion con herramientas de terceros, en nuestro caso para la automatizacion del proceso de normalizacion. Este requisito nos condujo a desarrollar una solucion arquitectonica para facilitar la interoperabilidad de herramientas y poder asi integrar otras herramientas en nuestro proceso MDE. Podemos identificar los siguientes objetivos para la tesis: Una implementacion de un proceso de reingenieria de datos mediante el uso de tecnicas MDE. La herramienta soporta la comprobacion automatica del nivel de normalizacion de la base de datos y su correccion. Uso de diferentes estrategias para la inferencia de claves ajenas en la etapa de restructuracion del proceso de reingenieria. La construccion de una herramienta capaz de automatizar el desarrollo de procesos de reingenieria basados en modelos. Abordar la interoperabilidad basada en modelos mediante la construccion de un puente bidireccional entre herramientas. Metodologia Se ha aplicado la metodologia DSRM (Design Science Research Methodology) que consiste en 6 actividades: (1) identificacion del problema y motivacion, (2) definicion de los objetivos de la solucion, (3) diseno y desarrollo, (4) demostracion, (5) evaluacion y (6) conclusiones y comunicacion. Resultados Describimos a continuacion las contribuciones de la tesis organizadas segun los objetivos identificados. Proceso de Reingenieria de Datos Hasta nuestro conocimiento, este trabajo es una de las primeras contribuciones proporcionando una valoracion del uso de MDE in la reingenieria de datos. El enfoque es validado mediante un sistema legado real, ampliamente usado en la industria sanitaria en Canada: OSCAR. Hemos contrastado ademas nuestro trabajo con enfoques tradicionales de reingenieria de datos, y hemos identificado algunos beneficios e inconvenientes de aplicar MDE, lo que nos permite dar una valoracion de en que medida MDE es aplicable en estos escenario Estrategias de Descubrimiento de Claves Ajenas Se ha abordado el problema de la inferencia de claves ajenas y la combinacion de diferentes tecnicas de reingenieria. Herramienta de Migracion Encontramos tres contribuciones en la herramienta: (1) es la primera propuesta que ejecuta procesos basados en modelos mediante la generacion de tareas automaticas y manuales que son integradas en un entorno de desarrollo; (2) es una de las primeras experiencias mostrando como una solucion MDE puede ser usada para construir herramientas de soporte para la definicion y ejecucion de procesos, asi como la gestion de tareas de migracion; (3) se presenta una solucion para el soporte de procesos de migracion implementados con tecnologias MDE. Interoperabilidad de Herramientas Se ha abordado la implementacion de una arquitectura MDE orientada a conectar herramientas. Se ha contribuido ademas a analizar y discutir a cerca como MDE es capaz de tratar diferentes escenarios de interoperabilidad. This thesis is mainly focused on applying MDE techniques to a data reengineering process. In particular, we analyse to what extent the use of models facilitates the implementation of the data quality improvement of a legacy system by means of a schema conversion, which is a common data modernisation scenario. The schema conversion implemented in our approach addresses the elicitation of implicit referential integrity constraints (declared in database by foreign keys) along with checking and fixing the appropriate normalisation level in a schema. Several techniques for discovering foreign keys should be combined in order to obtain more reliable results. Furthermore, an automation of migration processes is tackled. We have built a tool that supports the definition and enactment of migration processes, which have been validated for the data migration case study. In addition, MDE solutions normally require the integration with a third-party tool which allows an automatic normalisation step. This requirement leads us to develop an architectural solution to ease tool interoperability and then to integrate other useful tools (from the data engineering and requirement areas) to the migration process here proposed. Ee can therefore infer the following objectives of this thesis: An implementation of a data reengineering process by using MDE techniques. An automatic checking of the database normalisation level in the relational schema is supported. Using of different strategies in order to elicit foreign keys for the restructuring stage of the process. Building a tool able to automate the development of model-driven reengineering processes. To tackle the MDE-base tool interoperability through the building of some bidirectional bridge. Methodology We have followed the design science research methodology (DSRM) which consists of six activities: (1) problem identification and motivation, (2) define the objectives of a solution, (3) design and development, (4) demonstration, (5) evaluation and (6) conclusions and communication. Results We shall describe next the contributions of this thesis. They will be categorised according to the goals identified. Data Reengineering Process To the best of our knowledge, this work is one of the first contributions to provide an assessment of the use of MDE in data reengineering. The approach is showcased by means of an information system that is widely used in the healthcare industry in Canada: OSCAR. We have contrasted our work with the tasks usually performed in traditional approaches and have identified some benefits and drawbacks of applying MDE techniques, which could enable us to assess to what extent MDE could be applicable to other problems. Strategies of FK Discovering We devise a process for reengineering legacy information systems with respect to establishing referential integrity constraints and combining existing reengineering methods. Migration Tool There are three contributions in the tool proposed: (1) it is the first proposal that enacts process models by executing automated tasks and generates programming manual tasks which are integrated into a task management tool; (2) our work is one of the first experiences showing how an MDE approach can be used to build a tool supporting software development processes from the definition of software processes to the management of the tasks to be performed by managers and developers; (3) we present a solution to support migration processes that have been implemented with MDE technologies. Tool Interoperability We have devised a model-based architecture aims to bridge the gap between tools. The MDE techniques have proven useful to ease and extend the interoperability capabilities of DB-Main. We contribute to analyse and discuss through this case study how MDE can address several interoperability scenarios.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.014 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.004 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.007 | 0.006 |
| Open science | 0.004 | 0.006 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".