What’s in a Name? Comparative Analysis of Laboratory Test Naming Guidelines as Applied to Common Confusing Test Names
Bibliographic record
Abstract
Abstract Introduction/Objective Laboratory test names frequently do not enable easy understandability or promote correct test utilization, which leads to difficulty for providers in finding the correct test and results in unnecessary cost and medical errors. Laboratory test names are also largely unstandardized and are not named by a consistent set of conventions. To address these issues, the TRUU-Lab (Test Renaming for Understanding & Utilization) initiative aims to generate a consensus test naming guideline for better human understandability of laboratory test names. These studies address the first aim of the TRUU-Lab initiative: to identify root causes and challenges in understanding and using laboratory test names. Methods We conducted survey studies to capture the most problematic laboratory test names, then performed analysis of these names to identify aspects of these names that led to confusion among providers. A subset of these test names were used to evaluate five existing laboratory test naming guidelines (LOINC, ONC TigerTeam, Pan- Canadian iEHR Viewer Name, Standards for Pathology Informatics (Australia), and ARUP Laboratories internal style guides) for their ability to produce understandable test names. Results 274 survey responses yielded ~100 unique laboratory tests cited as confusing, and highlighted substantial diversity both in the names of these tests between institutions and in respondent opinion on the best alternative names. The top 10 most commonly-cited tests yielded ≥ 3 unique names, and the top 2 tests (Vitamin D and anti- factor Xa) yielded ≥ 10 unique names. Post-survey analysis identified eight characteristics associated with poor understandability of a test name, including ambiguity, abbreviations, homophones, multiple indications for a single test, proprietary names, synonyms, truncation, and “panels” where components are obfuscated. Existing guidelines produced highly variable names given the same prompt, and varied in their ability to avoid pitfalls associated with poor understandability. Conclusion These studies highlight aspects of existing laboratory test names that lead to confusion among ordering providers, and identify the inability of existing laboratory test naming practices to adequately address these issues. Efforts are ongoing within TRUU-Lab to use these results to inform novel laboratory test naming guidelines to promote universal human understandability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.097 | 0.458 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.007 | 0.007 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.005 | 0.008 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".