What’s in a Name? Comparative Analysis of Laboratory Test Naming Guidelines as Applied to Common Confusing Test Names
Bibliographic record
Abstract
Abstract Laboratory test names frequently do not enable easy understandability or promote correct test utilization, which leads to difficulty for providers in finding the correct test and results in unnecessary cost and medical errors. As a further complication, laboratory test names are largely unstandardized and are not named based on a consistent set of conventions. To address these issues, the TRUU-Lab (Test Renaming for Understanding & Utilization) initiative aims to generate a consensus laboratory test naming guideline for better human understandability of laboratory test names. These studies address the first and second aims of the TRUU-Lab initiative: 1) to identify root causes and challenges in understanding and using laboratory test names, and 2) to share resources related to potential solutions. We initially conducted survey studies to capture the most commonly problematic laboratory test names, then performed analysis of these names to identify aspects of these names that led to confusion among providers. 274 survey responses yielded ~100 unique laboratory tests that respondents felt were confusing, and highlighted substantial diversity both in the names of these tests between institutions and in respondent opinion on the best alternative names, with the top 10 most commonly-cited tests having at least 3 unique names, and the top 2 tests (Vitamin D and anti-factor Xa) having at least 10 unique names. Post-survey analysis identified eight common characteristics associated with poor understandability of a test name, including ambiguity, abbreviations, homophones, multiple indications for a single test, non-descriptive proprietary names, synonyms, truncation due to software limitations, and €œpanels where test components are obfuscated. A subset of the survey-identified confusing test names were used to evaluate existing laboratory test naming guidelines for their ability to produce understandable test names. Five guidelines, including LOINC, ONC TigerTeam, Pan-Canadian iEHR Viewer Name, Standards for Pathology Informatics (Australia), and ARUP Laboratories internal style guides, were evaluated, and produced highly variable names given the same test name prompt. Further, existing guidelines also varied in their ability to avoid pitfalls previously identified as associated with poor understandability. Together, these studies highlight the aspects of existing laboratory test names that lead to confusion among ordering providers, and identify the inability of existing laboratory test naming practices to adequately address these issues. Efforts are ongoing within TRUU-Lab to use these results to inform novel laboratory test naming guidelines that promote universal human understandability. Work is also ongoing to apply these novel guidelines to generate new candidate test names, and conduct survey analysis to evaluate the effects of new test naming guidelines on understandability and correct test utilization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".