In Reply
Bibliographic record
Abstract
In Reply.—Current US guidelines that endorse universal alanine aminotransferase (ALT) thresholds to define “normal” values1 are problematic and could result in unnecessary clinical interventions. In our 5-year national Veterans Health Administration analysis, we reported clinically significant differences in ALT measurements in analyzers by 5 manufacturers. We analyzed more than 22 000 samples using external ALT proficiency testing with standardized College of American Pathologists samples which, as we point out in our manuscript,2 are not commutable with human samples.3 Our findings closely align with a Canadian national study (performed with pooled human blood samples) and an Indiana statewide study demonstrating that ALT measurements vary significantly across North American analyzers.4,5 Taken in the context of this prior work, our observed lack of harmonization in ALT assays affirms that universal ALT thresholds should be avoided as a trigger for clinical action until analyzers and ALT proficiency testing materials are standardized.Proficiency testing material allows comparisons of analytical performance across systems and is used in the daily calibration and equipment maintenance procedures to enable measurement of patient samples to guide treatment decisions. We note that if proficiency testing material were entirely noncommutable, this would not be possible. Nevertheless, our aim was not to quantify the absolute differences between manufacturer systems, which would require the use of split samples and parallel testing. Rather, our goal was to illustrate the problems inherent in uncritically applying universal ALT thresholds given the lack of a sufficiently commutable “standardized” testing material.While a European certified reference material for ALT is commercially available, the use of a single “international reference” material is impractical from an operational viewpoint. As occurred with the World Health Organization thromboplastin for the international normalized ratio, the standard may eventually be depleted. Reliance on a global certified reference standard produced by a single source is inadvisable.The test platforms for the 5 analyzers we studied have differing reference ranges within the same patient population. This key finding affirms the need to examine the reference range when interpreting ALT data for each analyzer, as they are not standardized. We fully endorse the authors' concluding statement, “By appropriately establishing reference ranges, medical decisions will not be hindered by misinterpretations…”.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.004 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".