Vérification de la qualité de la base de données périnatales Niday pour 2008 : rapport sur un projet d’assurance de la qualité
Bibliographic record
Abstract
Introduction Le présent projet d’assurance qualité vise à déterminer la fiabilité, l’intégralité et l’exhaustivité des données saisies dans la base de données périnatales Niday. Méthodologie La qualité des données a été mesurée en comparant les données réextraites des dossiers des patients aux données entrées à l’origine dans la base de données périnatales Niday. Un échantillon représentatif des hôpitaux de l’Ontario a été sélectionné et un échantillon aléatoire de 100 dossiers mère-enfant appariés a été vérifié pour chaque site. Un sous-ensemble de 33 variables (représentant 96 champs de données) de la base Niday a été choisi pour la réextraction. Résultats Parmi les champs de données pour lesquels le coefficient Kappa de Cohen ou le coefficient de corrélation intraclasse (CCI) a été calculé, 44 % présentaient une concordance excellente ou presque parfaite (au-delà de ce qui pourrait être imputé au hasard). Cependant, environ 17 % d’entre eux ont affiché une concordance inférieure à 95 % et un coefficient Kappa ou un CCI de moins de 60 %, signe d’une concordance presque nulle, médiocre ou modérée (au-delà de ce qui pourrait être imputé au hasard). Analyse L’article présente des recommandations pour améliorer la qualité de ces champs de données.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.059 | 0.095 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".