Methodological challenges when carrying out research on CKD and AKI using routine electronic health records
Bibliographic record
Abstract
Research regarding chronic kidney disease (CKD) and acute kidney injury (AKI) using routinely collected data presents particular challenges. The availability, consistency, and quality of renal data in electronic health records has changed over time with developments in policy, practice incentives, clinical knowledge, and associated guideline changes. Epidemiologic research may be affected by patchy data resulting in an unrepresentative sample, selection bias, misclassification, and confounding by factors associated with testing for and recognition of reduced kidney function. We systematically explore the issues that may arise in study design and interpretation when using routine data sources for CKD and AKI research. First, we discuss how access to health care and management of patients with CKD may have an impact on defining the target population for epidemiologic study. We then consider how testing and recognition of CKD and AKI may lead to biases and how to potentially mitigate against these. Illustrative examples from our own research within the UK are used to clarify key points. Any research using routine renal data has to consider the local clinical context to achieve meaningful interpretation of the study findings. Research regarding chronic kidney disease (CKD) and acute kidney injury (AKI) using routinely collected data presents particular challenges. The availability, consistency, and quality of renal data in electronic health records has changed over time with developments in policy, practice incentives, clinical knowledge, and associated guideline changes. Epidemiologic research may be affected by patchy data resulting in an unrepresentative sample, selection bias, misclassification, and confounding by factors associated with testing for and recognition of reduced kidney function. We systematically explore the issues that may arise in study design and interpretation when using routine data sources for CKD and AKI research. First, we discuss how access to health care and management of patients with CKD may have an impact on defining the target population for epidemiologic study. We then consider how testing and recognition of CKD and AKI may lead to biases and how to potentially mitigate against these. Illustrative examples from our own research within the UK are used to clarify key points. Any research using routine renal data has to consider the local clinical context to achieve meaningful interpretation of the study findings. Acute kidney injury (AKI) and chronic kidney disease (CKD) each have a substantial global health care burden.1Mehta R.L. Cerdá J. Burdmann E.A. et al.International Society of Nephrology's 0by25 initiative for acute kidney injury (zero preventable deaths by 2025): a human rights case for nephrology.Lancet. 2011; 385: 2616-2643Abstract Full Text Full Text PDF Scopus (605) Google Scholar The elderly with multiple morbidities are at most risk, and the burden of both diseases is set to rise in line with the demographic transition to older populations.2Department of Economic and Social Affairs Population Division. World Population Ageing 2013. United Nations, 2013. Available at: http://www.un.org/en/development/desa/population/publications/pdf/ageing/WorldPopulationAgeing2013.pdf. Accessed February 3, 2016.Google Scholar There is very limited understanding on how best to manage patients with multiple health conditions.3National Institute for Health and Care Excellence. Acute kidney injury: prevention, detection and management 2013. Available at: https://www.nice.org.uk/guidance/cg169. Accessed February 3, 2016.Google Scholar Historically, epidemiological research on this group was challenging. Large cohorts with many years of follow-up and high rates of retention were needed, resulting in logistical complexities and substantial financial costs. This has changed with the availability of computerized health care records. Data from electronic health records (EHRs) have facilitated large epidemiological studies to answer important questions that would otherwise not have been possible.4McDonald H.I. Thomas S.L. Millett E.R. Nitsch D. CKD and the risk of acute, community-acquired infections among older people with diabetes mellitus: a retrospective cohort study using electronic health records.Am J Kidney Dis. 2015; 66: 60-68Abstract Full Text Full Text PDF Scopus (49) Google Scholar Specific characteristics of kidney disease can present challenges when using routine data for research. Kidney disease is often asymptomatic, and diagnosis relies on blood and urine tests. There are wide disparities in the availability of records of renal function in EHRs.5de Lusignan S. Nitsch D. Belsey J. et al.Disparities in testing for renal function in UK primary care: cross-sectional study.Fam Pract. 2011; 28: 638-646Crossref PubMed Scopus (16) Google Scholar The patchy nature of renal data means that epidemiologic research may be affected by biases. There is increasing use of EHR data in research and as part of performance management in health care. Recently, a reporting guideline for observational studies of routinely collected health care data was published, but no specific guidance exists how to apply this guideline in the context of renal research.6Benchimol E.I. Smeeth L. Guttmann A. et al.for the RECORD Working CommitteeThe REporting of studies Conducted using Observational Routinely-collected health Data (RECORD) Statement.PLoS Med. 2015; 12: e1001885Crossref Scopus (1953) Google Scholar The aim of this mini review is to systematically explore the issues that may arise in study design and interpretation of kidney disease in routine data. We focus on primary care data, as most patients with CKD are seen in the community setting. We will not discuss ethical issues. An excellent in-depth discussion of AKI research using secondary care data has been provided by others.7James M.T. Pannu N. Methodological considerations for observational studies of acute kidney injury using existing data sources.J Nephrol. 2009; 22: 295-305PubMed Google Scholar In general, when reviewing EHR studies, it is useful to consider what the perfect study to address the question would be. Who should the perfect study focus on, and which data items would be needed to control for confounding? Contrasting this perfect scenario with the reality of the databases then clarifies the main limitations and potential biases. How patients access health care impacts on who is recorded within routine data. For renal patients in particular, this may reflect the cause and severity of their renal disease. For example, patients with early CKD are more likely to be identified and monitored in the primary care setting if they have known risk factors such as diabetes mellitus, while those receiving dialysis and under specialist care may only be recorded in secondary care datasets. Some health systems incentivize routine health checks (e.g., in occupational settings), which may include kidney function markers. If appropriate ethics approvals are in place, patient identifiers can be used to collect further information from other routine health data, such as hospitalization or dialysis registry data, to investigate long-term outcomes according to baseline function (examples given in Table 1).Table 1UK example: Potential sources of anonymized information about kidney diseasePrimary care computerized health recordsDatabases are provided from specific primary care software providers; contain information on clinical diagnoses, prescriptions, medical procedures, and laboratory tests; and are traditionally coded with the Read clinical coding system and include feedback from secondary care. Examples include the Clinical Practice Research Datalink (CPRD) and The Health Improvement Network (THIN).Hospital recordsHospital Episode Statistics (HES) provide information on diagnoses (coded with the International Classification of Diseases, Tenth Revision [ICD-10]) and medical procedures related to all National Health Service–funded hospital admissions in England. These data are entered by coders who look at notes and discharge letters. They do not contain information on laboratory results or inpatient prescriptions. Similar data to HES are collected in Wales, Northern Ireland, and Scotland. HES data can be linked to primary care records to provide more complete patient information.Laboratory recordsIt is possible to extract data directly from laboratories in the UK, as is being done for acute kidney injury detection by the National Think Kidney’s program. These data lack detailed patient information such as underlying diagnoses or prescriptions.Pharmacy dispensing recordsSuch data may provide valuable information on whether a patient collected a prescribed drug from the pharmacy. However, there is little information on comorbidity and no information on laboratory test results. Data are usually linked to data containing information on renal patients.Disease registries and auditAs part of National Audit used for monitoring the quality of care and commissioning, a range of disease registries have been set up to capture key features of routine clinical care for specific disease entities. This includes the UK Renal Registry, the National CKD Audit, and other registries that may collect some renal data (e.g., Diabetes Registry). Open table in a new tab In some settings, primary care practitioners are the gatekeepers for access to specialized care. Therefore, in theory, these data should be complete for every patient even without linkage. However, this assumes good information flow between primary and secondary care, and accurate recording of secondary/tertiary care episodes. Whether this is indeed the case is often unknown: for example, there are situations where patients may be able to access specialist care directly without first seeing a primary care practitioner (e.g., patients on renal replacement therapy). In health insurance–financed settings, the most complete data on individual patient care may be found in medical claims data held by the insurance company. However, laboratory results will often not be available within these data. Initiation of renal replacement therapy is often used as a proxy for kidney failure (estimated glomerular filtration rate [eGFR] <15 ml/min/1.73 m2) or end-stage renal disease. However, it is important to understand the characteristics of people who do not have access to care or are not recorded in selected EHRs. For example, in some health-care settings, patients have to pay for renal replacement therapy, for example, in South Africa <50% of patients with end-stage renal disease were accepted onto dialysis because of factors related to poverty.8Moosa M.R. Kidd M. The dangers of rationing dialysis treatment: the dilemma facing a developing country.Kidney Int. 2006; 70: 1107-1114Abstract Full Text Full Text PDF PubMed Scopus (105) Google Scholar, 9White S.L. Chadban S.J. Jan S. Chapman J.R. Cass A. How can we achieve global equity in provision of renal replacement therapy?.Bull World Health Organ. 2008; 86: 229-237Crossref PubMed Scopus (193) Google Scholar Following the example of the US registry, many registries define dialysis continuing beyond 3 months as “chronic dialysis.” The UK renal registry initially only collected data on patients who started and stayed on dialysis for at least 3 months; consequently, patients with end-stage renal disease who initiated dialysis but died before 3 months elapsed were missed. Patients may not want to start renal replacement therapy and instead undergo conservative management, but most renal registries do not collect this information. In Canada, among those aged 85 years or older, the proportion of untreated kidney failure was estimated to be up to 10-fold higher than those treated.10Hemmelgarn B.R. James M.T. Manns B.J. et al.for the Alberta Kidney Disease NetworkRates of treated and untreated kidney failure in older vs younger adults.JAMA. 2012; 307: 2507-2515Crossref PubMed Scopus (121) Google Scholar In order to avoid bias, only patients expected to be tested as part of routine care for baseline kidney function and who would have outcomes of interest recorded if these happened should form part of the target population for study. Having identified this target population of interest, the next step is to check what information can be readily extracted and whether the research question of interest can be addressed. Testing is often triggered by acute illness and therefore will not reflect “baseline” renal function. This differs from large epidemiological studies that include patients who volunteer to and are not However, in the older population in most routine are are usually not M. et al.for the Network of with a reduced using laboratory Nephrol. PubMed Scopus Google Scholar Testing may be by the clinical for example, guidance testing for CKD in those at most risk with and Institute for Health and Care Excellence. kidney disease in and management Available at: Accessed February 3, 2016.Google Scholar In the UK, in to guideline there was an in recording of over time in a cohort of people with diabetes in primary care the characteristics of patients who renal function tested in the may be very from patients routinely tested from example: of recording among patients with diabetes aged years for each financial Patients were for the study if they were within Clinical Practice Research Datalink quality of aged years or and with H.I. The of among with Diabetes and Kidney of and Large There is a to be between recorded test results and a coded diagnosis of coded diagnosis relies on the about CKD as as to the There is wide in with on only of patients for CKD being coded as et of coding for kidney a J Kidney Dis. 2011; Full Text Full Text PDF PubMed Scopus Google Scholar patient using CKD results in the potential to a proportion of patients to a over D. et of databases for disease cohorts is with a Full Text Full Text PDF PubMed Scopus Google Scholar, Nitsch S. The National Kidney Disease Audit Available at: Accessed February 3, 2016.Google Scholar, S. Lusignan S. et of routinely collected practice data to patients with chronic kidney disease a review of medical PubMed Scopus Google Scholar, M.R. J. et of The Health Improvement Network for epidemiologic studies of chronic kidney 2011; PubMed Scopus Google Scholar Therefore, most epidemiologic research in relies on results. for laboratory of have changed over In the UK, laboratory in of was with the of from laboratories have started to to a using which to be in Therefore, from in may for in on the time in a N. A. Lusignan S. the of in renal detection of in laboratory Health 2015; 22: Scopus Google Scholar If results are care to be on which of in Renal Disease Kidney Disease was used for kidney disease is not and there is substantial because results are more likely to be than Lusignan S. Nitsch D. Belsey J. et al.Disparities in testing for renal function in UK primary care: cross-sectional study.Fam Pract. 2011; 28: 638-646Crossref PubMed Scopus (16) Google Scholar, Nitsch S. The National Kidney Disease Audit Available at: Accessed February 3, 2016.Google Scholar, S. Lusignan S. et of routinely collected practice data to patients with chronic kidney disease a review of medical PubMed Scopus Google Scholar For example, on test may be recorded as in the medical notes but not coded in the and therefore the information is to In routine care, more and such as the are usually limited to known to be at high risk such as those with Nitsch S. The National Kidney Disease Audit Available at: Accessed February 3, 2016.Google Scholar studies using routinely collected data have used time of or often of H.I. Thomas S.L. Millett E.R. Nitsch D. CKD and the risk of acute, community-acquired infections among older people with diabetes mellitus: a retrospective cohort study using electronic health records.Am J Kidney Dis. 2015; 66: 60-68Abstract Full Text Full Text PDF Scopus (49) Google Scholar, B.R. Manns B.J. A. et al.for the Alberta Kidney Disease between kidney and PubMed Scopus Google Scholar studies using laboratory data will not be able to the of when the was (e.g., urine and system for AKI was first in et al.for the the Acute renal therapy and information the International of the Acute PubMed Google Scholar aim was to present AKI as a detection and and of the for AKI among in primary and secondary care is but has been limited it has been to it is likely that coding for AKI acute renal in primary care the in secondary care: coding a of AKI by and is likely to S. A. et acute kidney injury a cross-sectional with the clinical 2015; Scopus Google Scholar of coding of renal function relies on the being able to between AKI and The for clinical recognition can be by AKI to for Clinical and for for Acute Kidney on with time 2013. Available at: Accessed February 3, 2016.Google Scholar However, AKI may CKD as in a hospital of those identified CKD than S. A. et acute kidney injury a cross-sectional with the clinical 2015; Scopus Google Scholar the AKI was for use with inpatient laboratory data, it should be used with in community data. with a hospital in the community are likely to be by This the of a in renal function as In it is not possible to apply AKI to a without of a and Patients with available baseline renal function are to be of the they are likely to be people with chronic that have testing or those who more with health care. are likely to have baseline results Similar to what was for renal guideline and performance may in to how information on other key is recorded over For example, a to the management of patients with practitioners to use diagnosis and use et in rates of recorded in primary care time of of the and the quality outcomes 2015; Full Text Full Text PDF PubMed Scopus Google Scholar If such issues are not may in capture of key items may not be recorded at all as they are the of care. For example, in the UK most are by primary care key of interest to renal as and are and prescribed in secondary care, which means that primary care data on these will be example is drug which be in where these are available over the when of diagnoses of it is important to consider to with a Disease may lead some patients to with a new rates for the time may be entered without from new diagnoses early patient in which medical is et The between time and rates in the Practice Research PubMed Scopus Google Scholar found that this rate new patient to baseline within months for most acute and within a for most chronic This means that of disease to have a start time of at least months and a is In this the most observational study will be used as a to discuss biases in renal with as a baseline and AKI as an of potential biases that may arise in an example study of the between baseline and time time that not have systematically to of because specific do not in the study population or that of follow-up time by is not systematically identified and may be or data data coding patients with of as CKD and follow-up at first This assumes that patients do not have AKI a with AKI who do not chronic dialysis may not be recorded in renal registry, or hospitalization or and therefore are not recorded in the AKI rates in the very elderly consider that study with more of CKD may before AKI on discharge is used as an if may be more of AKI if patient has underlying known kidney is by coded AKI and the primary care blood to the acute and then patient to for or define CKD at the with people who have a time of follow-up between the first and study population to those who can health care and use laboratory at and with before risk studies of AKI coding to explore if coding is on study are likely to be an of the underlying in before hospital control to selection in study setting if of control is to are not from the population that case are are not of the population who have to selection in the may contain some some may not have data from those in are identified in a hospital and with control from primary care, some of to the hospital that was used to AKI recorded in registry, but the registry has and is more likely to capture with example and example and for control from a population of hospital sources of that may be are given but are not Potential confounding factors are not in this table but be acute kidney chronic kidney estimated glomerular filtration Open table in a new tab Potential that may be are given but are not Potential confounding factors are not in this table but be addressed. acute kidney chronic kidney estimated glomerular filtration Testing for and in the population is often by clinical as the these characteristics and testing over it is possible to the study design and for these confounding If all test results are for health is needed to avoid but of health may be in EHRs. of the epidemiologic study to a cohort with higher risk who are expected to be routinely tested that the for the baseline renal function test is from test how is key to and potential sources of For example, may be prescribed more or in patients with or of a drug may be a of underlying kidney function. for example, be prescribed more in people with on However, if is recorded than prescriptions, then using recorded data will be to confounding by In epidemiological studies of the of CKD on the of CKD is often on a of at baseline when the study is in However, CKD on a of in clinical records will to CKD because a may be by AKI or because the test may have been by Lusignan S. et has a than the to glomerular filtration rate on the of chronic kidney Pract. 2011; PubMed Scopus Google Scholar Therefore, the clinical of CKD that renal function for at least 3 months for CKD However, if results are to define CKD in epidemiological studies, the time between the first and results be with care to avoid time Patients have to be in order to and have the test 3 and therefore CKD should not be at the time of the first study people who both a test and a follow-up with people who not the follow-up and found that this may an for those who follow-up of at least The of time in epidemiologic Nephrol. 2008; PubMed Scopus Google Scholar the should be between people who have CKD as by the with at least the of follow-up in the The used by James et M.T. M. et al.for the Alberta Kidney Disease and risk of hospitalization and with J Kidney Dis. 2009; Full Text Full Text PDF PubMed Scopus Google Scholar may be more for studies with follow-up time for CKD to be an In this CKD is at given time using the by the most This of the as CKD will in of CKD the will be at the next test The best to avoid selection is to the study population to those for it is that if the it will be recorded in the health and that control in a study are from the population as the case If the study is then the start of dialysis to be to avoid For example, it is important to clarify whether for those on dialysis the start is the of of a dialysis or the start of dialysis is that the may a of time over which example, and and recording of CKD and AKI et in patients with acute renal to Nephrol. 2006; PubMed Scopus Google Scholar how those records may be of patients with AKI recorded at hospital may over time because AKI is an in et of acute kidney injury not dialysis in from to retrospective of hospital J Pract. 70: Scopus Google Scholar of the nature of these issues may in of the cohort by may be for AKI in This can be by studies or by to routine data from potential for large cohorts to be with detailed records to a wide range of kidney research. However, challenges in of availability and of of renal function. what testing and recording of data is when research using routinely collected data. of observational studies using should to the of Conducted Observational Health Data (RECORD) E.I. Smeeth L. Guttmann A. et al.for the RECORD Working CommitteeThe REporting of studies Conducted using Observational Routinely-collected health Data (RECORD) Statement.PLoS Med. 2015; 12: e1001885Crossref Scopus (1953) Google Scholar
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".