Leveraging best practices in data governance: An organization-wide data inventory and mapping project to support a five year data strategy
Notice bibliographique
Résumé
IntroductionThe professional regulation sector is moving toward risk-informed approaches that require high quality data. A key component of a corporate 2017 Data Strategy is the implementation of a data inventory and mapping project to catalogue, centralize, document and govern data assets that support regulatory decisions, programs and operations.
 Objectives and ApproachIn a data rich organization, the goals of the data inventory are to: enhance authoritative data that support programs; identify data duplications/gaps; identify data sources, owners and users; and, apply consistent data management and standards organizationally. Routinely used data assets outside the large enterprise workflow system (excel/word files; databases; paper collections) were catalogued. Using data governance principles and a facilitated questionnaire, departmental data stewards were interviewed about their generated data. Questions included data purpose/sources/types/formats/owners, retention rates, analytical products, gaps and visions for a desired data state. A data mapping methodology highlighted data set and variable connections within and across departments.
 ResultsTo date, over 40 staff members in 10 departments were identified as data content experts. In addition to data in the corporate enterprise system, over 80 unique datasets were identified. In 1 large department, over 2,000 data elements across 26 datasets were inventoried. Data mapping analysis revealed thematic data domains, including member demographics, outcomes, certifications, tracking and financial data, collected and held in multiple formats ((Microsoft Access, Excel, Word), SPSS, PDF, e-mails and paper documents). While 72% of the data elements were formatted numerically, approximately 8% were free text. Significant data redundancies across staff members and departments were revealed, as well as unstandardized variable naming conventions. Gaps analysis highlighted need for standardized, electronic data, where not available and data management training.
 Conclusion/ImplicationsCustomized data mapping reports to data users will facilitate the development of local, standardized departmental data hubs that will centrally link to a centralized data repository to facilitate seamless organization-wide analytics, improvements in current data management practices and greater data collaboration with the ultimate goal of supporting risk-informed approaches.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,019 | 0,042 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,004 | 0,036 |
| Science ouverte | 0,023 | 0,016 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».