Automatic Conflict Resolution to Integrate Relational Schema
Bibliographic record
Abstract
With the constantly increasing reliance on database systems to store, process, and display data comes the additional problem of ensuring interoperability between these systems. On a wider scale, the World-Wide Web (WWW) provides users with the ability to access a vast number of data sources distributed across the planet. However, a fundamental problem with distributed data access is the determination of semantically equivalent data. Ideally, users should be able to extract data from multiple sites and have it automatically combined and presented to them in a usable form. No system has been able to accomplish these goals due to limitations in expressing and capturing data semantics. Schema integration is required to provide database interoperability and involves the resolution of naming, structural, and semantic conflicts. To this point, automatic schema integration has not been possible. This thesis demonstrates that integration may be increasingly automated by capturing data semantics using a standard dictionary. This thesis proposes an architecture for automatically constructing an integrated view by combining local views that are defined by independently expressing database semantics in XML documents (X-Specs) using only a pre-defined dictionary as a binding between integration sites. The dictionary eliminates naming conflicts and reduces semantic conflicts. Structural conflicts are resolved at query-time by translating from the semantic integrated view to structural queries. The system provides both logical and physical access transparency by mapping user queries on high-level concepts to schema elements in the underlying data sources. The architecture automatically integrates relational databases, and its application of standardization to the integration problem is unique. The architecture may be deployed in a centralized or distributed fashion, and preserves full database autonomy while allowing transparent access to all databases participating in a global federation without the user's knowledge of the underlying data sources, their location, and their structures. Thus, the contribution is a system which provides system transparency to users, while preserving autonomy for all systems. A distributed deployment allows integration using a web browser, and would have a major impact on how the Web is used and delivered. The integration software, Unity, is the bridge between concept and implementation. Unity is a complete software package for the construction and modification of standard dictionaries, parsing of database schema and metadata to construct X-Specs, combining X-Specs into an integrated view, and for transparent querying. Integration results obtained using Unity illustrate the usefulness of the approach.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".