MétaCan
Menu
Retour à la cohorte
Enregistrement W2608676843 · doi:10.11575/prism/30313

UNITY - A DATABASE INTEGRATION TOOL

2000· article· en· W2608676843 sur OpenAlexaboutno aff
Ramon Lawrence, Ken Barker

Notice bibliographique

RevuePRISM (University of Calgary) · 2000
Typearticle
Langueen
DomaineComputer Science
ThématiqueAdvanced Database Systems and Queries
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésDatabaseComputer science

Résumé

récupéré en direct d'OpenAlex

The World-Wide Web (WWW) provides users with the ability to access a vast number of data sources distributed across the planet. Internet protocols such as TCP/IP and HTTP have provided the mechanisms for exchanging the data. However, a fundamental problem with distributed data access is the determination of semantically equivalent data. Ideally, users should be able to extract data from multiple Internet sites and have it automatically combined and presented to them in a usable form. No system has been able to accomplish these goals due to limitations in expressing and capturing data semantics. This paper details the construction, function, and deployment of Unity, a database integration software package which allows database semantics to be captured so that they may be automatically integrated. Unity is the tool that we use to implement our integration architecture detailed in previous work. Our integration architecture focuses on capturing the semantics of data stored in databases with the goal of integrating data sources within a company, across a network, and even on the World-Wide Web. Our approach to capturing data semantics revolves around the definition of a standardized dictionary which provides terms for referencing and categorizing data. These standardized terms are then stored in semantic specifications called X-Specs which store metadata and semantic descriptions of the data. Using these semantic specifications, it becomes possible to integrate diverse data sources even though they were not originally designed to work together. The centralized version of the architecture is presented which allows for the independent integration of data source information (represented using X-Specs) into a unified view of the data. The architecture preserves full autonomy of the underlying databases which are transparently accessed by the user from a central portal. Distributing the architecture would by-pass the central portal and allow integration of web data sources to be performed by a user's browser. Such a system which achieves automatic integration of data sources would have a major impact on how the Web is used and delivered. Unity is the bridge between concept and implementation. Unity is a complete software package which allows for the construction and modification of standardized dictionaries, parsing of database schema and metadata to construct X-Specs, and contains an implementation of the integration algorithm to combine X-Specs into an integrated view. Further, Unity provides a mechanism for building queries on the integrated view and algorithms for mapping semantic queries on the integrated view to structural (SQL) queries on the underlying data sources. Notes: Join released technical report. Released as TR-00-17 for the University of Manitoba, and 2000-664-16 for the University of Calgary.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,004
score de la tête « metaresearch » (Gemma)0,010
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: aucune
GenreSignal candidat: Logiciel · Signal consensuel: aucune
Score de désaccord entre enseignants0,012
Score d'incertitude au seuil0,042

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0040,010
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0010,002
Bibliométrie0,0030,002
Études des sciences et des technologies0,0010,001
Communication savante0,0040,006
Science ouverte0,0040,007
Intégrité de la recherche0,0010,002
Charge utile insuffisante (le modèle a refusé de juger)0,0120,008

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,011
Tête enseignante GPT0,201
Écart entre enseignants0,189 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreLogiciel

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations6
Publié2000
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revuePRISM (University of Calgary)Même sujetAdvanced Database Systems and QueriesTravaux en français237 207