Time-related Bias in Administrative Health Database Studies of Disease Incidence
Notice bibliographique
Résumé
To the Editor: Administrative health databases are becoming a common source of data to measure disease occurrence. However, database-specific sources of bias (including data inaccuracy, missing data, or misclassification)1–3 may affect the estimates of disease occurrence. Therefore, researchers frequently use case-identifying algorithms, which are validated through medical chart review, surveys, or linkage with other sources of data. Despite such validation, however, the methodological approach to data analysis itself may also introduce bias for diseases where multiple physician contacts or prescriptions over relatively extended time periods are needed to fulfill the criteria for case diagnosis. The factors leading to bias arise from temporal aspects of administrative databases, such as the time and number of physician visits required to fulfill the criteria, the timing of disease diagnosis, and the time span of the database observation period. We used the provincial administrative health databases of Québec, Canada (1996–2009) to illustrate and quantify the impact of time-related bias on estimates of the incidence of inflammatory bowel disease, including Crohn disease and ulcerative colitis, in this Canadian province of 7 million people. Cases were identified using 3 case definitions requiring increasing numbers of physician visits with an inflammatory bowel disease diagnosis (from 2 to 4 to 6 visits) (see also eAppendix, https://links.lww.com/EDE/A832). The Table shows that improper timing of disease diagnosis as the first of the series of visits rather than the case-defining one, caused biases of up to 53% in inflammatory bowel disease incidence. The magnitude of the bias varied with the time span of the observation period and with the number of visits required to fulfill the case definition criteria. The span of the observation period caused biases of up to 21% in incidence when disease diagnosis was timed at the first inflammatory bowel disease contact, but had no impact when diagnosis was timed at the case-defining contact, when the number of visits required was reached.TABLE: Estimating Inflammatory Bowel Disease (IBD) Incidence (cases/100,000 person-years) for the Year 2001 in Québec: Effect of Timing of Diagnosis and Duration of Database Observation Period on Time-related Bias for a Case Definition Involving an Increasing Number of IBD ContactsBias from timing of disease diagnosis was reduced to approximately 3%–5% when a time period to meet the number of visits was specified. (see eTable 1, https://links.lww.com/EDE/A832 illustrating bias when the criteria were met in specified 2-year period.) Differentiating between incident and prevalent cases may be difficult when using administrative databases, because information before the start of the study period is not available. Nevertheless, bias in annual incidence estimates can in turn cause bias in prevalence estimates for the same period. Indeed, Büsch et al4 showed that inflammatory bowel disease prevalence varied with the span of observation period and the number of events required to meet the criteria. The use of a disease-free period before the first disease contact helps avoid an overestimation of incidence rates. We found a small decrease in inflammatory bowel disease incidence when a 2-year disease-free period was used, and the magnitude of bias from timing of disease diagnosis and span of observation period changed accordingly (see eTables 2 and 3, https://links.lww.com/EDE/A832 illustrating bias from time-related factors using a mandatory 2-year disease free period before first inflammatory bowel disease contact). In conclusion, time-related biases in estimating incidence rates can be minimized if the case diagnosis is considered when all criteria are met and if case definitions involve a specified time span. It is important to avoid these biases in studies of disease incidence because they will inherently introduce immortal time bias in subsequent studies of disease prognosis.5 The time required to fulfill the criteria needs to be considered as unexposed.6 As administrative health databases are more often used to estimate disease occurrence for the assessment of the burden of disease and for projecting healthcare expenditures, proper account for the described time-related issues can reduce bias. Maria Vutcovici Alain Bitton McGill University Health Centre Division of Gastroenterology Montréal, Canada [email protected] Maida Sewitch McGill University Faculty of Medicine Montréal, Canada Paul Brassard Valérie Patenaude Samy Suissa Lady Davis Institute for Medical Research Centre for Clinical Epidemiology Jewish General Hospital Montréal, Canada
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,012 | 0,032 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,004 | 0,000 |
| Bibliométrie | 0,001 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,001 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,001 | 0,006 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».