Optimizing Administrative Datasets to Examine Acute Kidney Injury in the Era of Big Data: Workgroup Statement from the 15 <sup>th</sup> ADQI Consensus Conference
Bibliographic record
Abstract
PURPOSE OF REVIEW: The purpose of this review is to report how administrative data have been used to study AKI, identify current limitations, and suggest how these data sources might be enhanced to address knowledge gaps in the field. OBJECTIVES: 1) To review the existing evidence-base on how AKI is coded across administrative datasets, 2) To identify limitations, gaps in knowledge, and major barriers to scientific progress in AKI related to coding in administrative data, 3) To discuss how administrative data for AKI might be enhanced to enable "communication" and "translation" within and across administrative jurisdictions, and 4) To suggest how administrative databases might be configured to inform 'registry-based' pragmatic studies. SOURCE OF INFORMATION: Literature review of English language articles through PubMed search for relevant AKI literature focusing on the validation of AKI in administrative data or used administrative data to describe the epidemiology of AKI. SETTING: Acute Dialysis Quality Initiative (ADQI) Consensus Conference September 6-7(th), 2015, Banff, Canada. PATIENTS: Hospitalized patients with AKI. KEY MESSAGES: The coding structure for AKI in many administrative datasets limits understanding of true disease burden (especially less severe AKI), its temporal trends, and clinical phenotyping. Important opportunities exist to improve the quality and coding of AKI data to better address critical knowledge gaps in AKI and improve care. METHODS: A modified Delphi consensus building process consisting of review of the literature and summary statements were developed through a series of alternating breakout and plenary sessions. RESULTS: Administrative codes for AKI are limited by poor sensitivity, lack of standardization to classify severity, and poor contextual phenotyping. These limitations are further hampered by reduced awareness of AKI among providers and the subjective nature of reporting. While an idealized definition of AKI may be difficult to implement, improving standardization of reporting by using laboratory-based definitions and providing complementary information on the context in which AKI occurs are possible. Administrative databases may also help enhance the conduct of and inform clinical or registry-based pragmatic studies. LIMITATIONS: Data sources largely restricted to North American and Europe. IMPLICATIONS: Administrative data are rapidly growing and evolving, and represent an unprecedented opportunity to address knowledge gaps in AKI. Progress will require continued efforts to improve awareness of the impact of AKI on public health, engage key stakeholders, and develop tangible strategies to reconfigure infrastructure to improve the reporting and phenotyping of AKI. WHY IS THIS REVIEW IMPORTANT?: Rapid growth in the size and availability of administrative data has enhanced the clinical study of acute kidney injury (AKI). However, significant limitations exist in coding that hinder our ability to better understand its epidemiology and address knowledge gaps. The following consensus-based review discusses how administrative data have been used to study AKI, identify current limitations, and suggest how these data sources might be enhanced to improve the future study of this disease. WHAT ARE THE KEY MESSAGES?: The current coding structure of administrative data is hindered by a lack of sensitivity, standardization to properly classify severity, and limited clinical phenotyping. These limitations combined with reduced awareness of AKI and the subjective nature of reporting limit understanding of disease burden across settings and time periods. As administrative data become more sophisticated and complex, important opportunities to employ more objective criteria to diagnose and stage AKI as well as improve contextual phenotyping exist that can help address knowledge gaps and improve care.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.032 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".