UPDATE TRANSLATION IN INSTANCE MAPPED HETEROGENEOUS PEER DATABASES
Bibliographic record
Abstract
In data sharing systems, peers are acquainted through pair-wise data sharing settings/mappings for sharing and exchanging data. Besides query processing, supporting update exchange for interchanging data between peers is one of the challenging problems in data sharing systems. In update exchange, an update action posed to a peer is applied to the peer's local database instance and then the update is propagated to the related peers. Previous work on update exchange have considered update propagation considering schema-level mappings between peers, which are conceptually similar to the view maintenance problem. However, there are data sharing systems, where peers are acquainted by instance-level mappings. In such a system, peers use different schemas and data vocabularies to represent semantically same real world entities. The instance-level mappings express how data in one peer relate to data in another peer. One of the problems in exchanging updates in instance-mapped data sharing systems is to translate updates correctly between heterogeneous peers. The translation should be such that insertions, deletions, and modifications of the tuples made by an update in a peer and by the translated version of the update in an acquainted peer are related through the mappings between them. In this paper, we investigate such a mechanism for translating update actions between heterogeneous peer data sources. Before discussing the translation mechanism, the paper first formalize the notion of update translation and derive conditions under which the translation mechanism will produce correct translations of updates.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.030 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.005 | 0.011 |
| Open science | 0.003 | 0.005 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".