A Practical, Global Perspective on Using Administrative Data to Conduct Intensive Care Unit Research
Bibliographic record
Abstract
Various data sources can be used to conduct research on critical illness and intensive care unit (ICU) use. Most published studies derive from randomized controlled trials, large-scale clinical databases, or retrospective chart reviews. However, few investigators have access to such data sources or possess the resources to create them. Hospital administrative data, also called health claims data, constitute an important alternative data source that can be used to address a broad range of research questions, including many that would be difficult to study in interventional studies. Such data often contain information that allows identification of ICU care, specific types of critical illness, and ICU-related procedures. The strengths of using administrative databases are that many are population-based, cover broad geographic regions, and are large enough to provide high statistical power and precise effect estimates. Linking hospital data to other databases regarding chronic care facilities, home care services, or rehabilitation services, for example, can expand the scope of research questions that can be answered. However, the limitations of administrative data must be recognized. They are not collected for research purposes; thus, data elements may vary in accuracy, and key clinical variables such as ICU-specific physiologic and laboratory data are usually lacking. Specific efforts should be made to validate the data elements used, as has been done in several world regions. As with any other research question, it is imperative that the analysis plan be carefully defined in advance and that appropriate attention be paid to potential sources of bias and confounding.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.458 | 0.436 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.005 | 0.004 |
| Bibliometrics | 0.013 | 0.013 |
| Science and technology studies | 0.006 | 0.035 |
| Scholarly communication | 0.028 | 0.041 |
| Open science | 0.009 | 0.016 |
| Research integrity | 0.021 | 0.042 |
| Insufficient payload (model declined to judge) | 0.009 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".