Notice bibliographique
Résumé
BACKGROUND: In de-identified data, exact birth dates are suppressed to maintain the confidentiality of research participants. When the partial birth date includes only the month and year, researchers who need exact dates must impute a day of the month. In some deidentified datasets, the day of the week is also provided, but this variable is uncommonly incorporated into the imputation of partial dates. OBJECTIVE: To examine the extent to which misclassification is reduced by incorporating the day of the week into partial birth date imputation. METHODS: We simulated a population of 594,677 people using the distribution of birthdays in England and Wales in 2024. We imputed birth dates using four methods: (1) the first day of the month, (2) the 15th of the month, (3) randomly selecting a day of the month and (4) randomly selecting a day of the month conditional on the day of the week. We quantified misclassification as the median number of days between the imputed and true birth date and as the cumulative percentage of the population whose imputed birth date fell within a given number of weeks of their true birth date. RESULTS: Incorporating the day of the week reduced misclassification, with a median of 7 days between the imputed and exact birth dates compared to 8-15 for the other methods. For nearly a quarter of the population, their imputed birth date was their true birth date, compared to 3% in other methods. However, using the 15th day of the month was the best method to ensure that no misclassification was greater than 3 weeks. CONCLUSION: Incorporating day of the week into random birth date selection reduced misclassification. This method is easily accomplished in standard statistical software.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,037 | 0,024 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,001 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».