MétaCan
Menu
Back to cohort
Record W4414740698 · doi:10.1093/clinchem/hvaf086.413

B-015 Neonatal Bilirubin Method Comparison on Three General Chemistry Platforms Using Pooled Neonatal Samples

2025· article· en· W4414740698 on OpenAlexaffabout
Yury Butorin, Isolde Seiden Long, Dustin Proctor, Heather A. Paul, Joshua E. Raizman, Fangze Cai, Lily Olayinka, Allison Vennera

Bibliographic record

VenueClinical Chemistry · 2025
Typearticle
Languageen
FieldMedicine
TopicNeonatal Health and Biochemistry
Canadian institutionsCalgary Laboratory ServicesUniversity of AlbertaAlberta Hospital EdmontonRed Deer Polytechnic
Fundersnot available
KeywordsSiemensBilirubinDirect bilirubinSignificant differenceStatistical analysisReference valuesMean difference

Abstract

fetched live from OpenAlex

Abstract Background Neonatal total bilirubin (NBIL) results are used to guide the medical need for phototherapy or exchange transfusion in babies. When sample analysis occurs across multiple vendors and instrument platforms, clinical result interpretation challenges can occur. This study evaluated the performance of NBIL methods across multiple general chemistry analyzers - Roche Cobas, Siemens Atellica, and Ortho Vitros - using pooled neonatal plasma samples. NBIL was measured as total bilirubin on Roche Cobas and Siemens Atellica, whereas on Ortho Vitros, it was measured as the sum of conjugated and unconjugated bilirubin fractions (Bc and Bu) using dry slide chemistry. Methods Residual neonatal plasma samples were collected at selected hospital laboratories in Alberta, Canada, and pooled. This was a process requiring meticulous effort due to the small volume typically available from each neonate. Pooled samples were divided into five sets and analyzed in duplicates at separate clinical laboratories using the NBIL method on the respective platform (two laboratories using Roche Cobas (8000 vs Pro), two laboratories using Ortho Vitros XT3400 BuBc slides (two slide generations) and one laboratory using Siemens Atellica. Results from each platform were compared to assess method performance and bias. Graphical and statistical analysis were completed, including Passing Bablok regression analysis. Total allowable difference used for the data analysis was ±20% or 6.84 umol/L. Results Method comparisons reveal up to 29% bias for neonatal bilirubin measurements between Roche Cobas and Siemens Atellica platforms, with an average bias of +27%. Ortho Vitros XT3400 showed an average bias of +11% compared to Roche Cobas. Among the three platforms studied, Siemens Atellica demonstrated the highest bias, while Roche Cobas consistently ran lower compared to the other platforms, with Vitros XT3400 positioned in between. Passing Bablok regression revealed the following relationships:Ortho Vitros (X) vs. Atellica (Y): y=1.12x-0.34Roche Cobas (X) vs. Atellica (Y): y=1.27x-0.26Ortho Vitros (X) vs. Roche Cobas (Y): y=0.86x+2.98 The average bias between two generation of slides installed on two Vitros analyzers was 0.3%, while the average difference between Cobas Pro and Cobas 8000 was 2.9%. Conclusions: Among the three vendor platform families studied, Siemens Atellica demonstrated the highest bias, while Roche Cobas was consistently lower compared to the other two platforms. These discrepancies underscore the critical need for mitigation strategies and assay performance assessment by the vendor to significantly reduce inter-platform bias, to better support patient care, minimize unnecessary or excessive patient management, and ensure consistent and reliable neonatal care. Given the significant variability observed between different vendor platforms for NBIL it is crucial for manufacturers to minimize within-platform inconsistencies. While inter-platform differences can pose challenges for clinical interpretation, reducing variability within a single manufacturer*s systems can help improve result reliability and standardization.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.568
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.001
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0010.000
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.098
GPT teacher head0.455
Teacher spread0.357 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes2
Has abstractyes

Explore more

Same venueClinical ChemistrySame topicNeonatal Health and BiochemistryFrench-language works237,207