MétaCan
Menu
Back to cohort

Comparison of MIMIC-III and MIMIC-IV for big data analytics of health informatics

2023· article· en· W4391093163 on OpenAlexafffund
Dillon Chrimes, Chanhee Kim

Bibliographic record

Venuenot available
Typearticle
Languageen
FieldComputer Science
TopicMachine Learning in Healthcare
Canadian institutionsUniversity of Victoria
FundersUniversity of Victoria
KeywordsComputer scienceBig dataHealth informaticsUsabilityInformaticsInformation retrievalData miningData scienceHealth care

Abstract

fetched live from OpenAlex

The use of a health big data via Medical Information Mart for Intensive Care (MIMIC) data sets has significant advancements in health informatics and clinical research of electronic health records (EHRs) in hospital systems. MIMIC-III and MIMIC-IV data sets are the two latest publicly available iterations of data from electronic records that possess distinct features, variables, and structures in two separate large relational databases. This study aimed to provide big data analytics comparisons of MIMIC-III and MIMIC-IV for experiential learning of health informatics from EHRs.Both data sets were quantitatively and qualitatively evaluated based on dataset properties (data, list of tables, timeline, patient encounters, software, system), data mining and visual categories (i.e., data size, data structure and schema, features/variables, data quality and completeness, clinical focus, access and use, dashboards, and visualization types), data usability heuristics and utilization (i.e., usability, complexity, granularity, applicability, and capacity). Results showed significant difference of 100 patients across 26 tables for MIMIC-III compared to 2,520 patients across 31 tables for MIMIC-IV data sets, respectively. There were 1716 diagnoses (ICD-9) and 503 procedures (ICD-9) for MIMIC-III. There were 262 diseases categories with >1000 disease instances and >2000 treatments for MIMIC-IV with no ICD codes. Moreover, these data sets contained large (big data) charting events in MIMIC-III of 758,356 rows, and MIMIC-IV eICU with charting events of 1,477,163 for nursing, and respiratory of 176,089 rows, respectively. The results suggest that MIMIC-III provided detailed information for retrospective clinical studies and operations in critical care with high data granularity in terms of re-admission (calculated fields from its admission table), length of stay, prescriptions, caregivers, and diagnosis and procedure (ICD-9). However, it lacked clinical capacity because of no diagnosis or charting event offset times or APACHE IV (Acute Physiology and Chronic Health Evaluation) scores that were in the MIMIC-IV eICU dataset. Hence, MIMIC-IV showed higher data granularity and capacity. MIMIC-IV eICU dataset introduces enhanced data attributes, more sophisticated patient trajectory tracking at the ICU unit level, and improved detailed information from electronic records. Moreover, MIMIC-IV has high usability, complexity, and applicability to critical care but lacking hospital operational data of re-admissions, caregivers, and ICD codes that MIMIC-III contained. However, MIMIC-IV dataset contained complex data schemas of treatment strings, nurse care plans, lab results, medications, microbiology, as well as infusion drug and respiratory charting. The big data analytics of MIMIC-III and MIMIC-IV needs to be further investigated for AI applications to demonstrate its usefulness for dynamic decision-making in health care.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.975
Threshold uncertainty score0.369

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0010.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.275
GPT teacher head0.437
Teacher spread0.162 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2023
Admission routes2
Has abstractyes

Explore more

Same topicMachine Learning in HealthcareFrench-language works237,207