Data Mining Approaches to Reference Interval Studies
Bibliographic record
Abstract
Both laboratories and in vitro diagnostic manufacturers alike struggle with performing the adequate studies required for producing high quality reference intervals (RIs). The most common approach for RI determination is a priori (direct) sampling of healthy individuals for each RI partition based on age, sex, and preanalytical factors such as diurnal, postural, and postprandial variations. With this approach, healthy reference individuals are selected using specific, well-defined criteria to resemble the patient population being evaluated. According to Clinical Laboratory Standards Institute (CLSI) EP28-A3c, best practice for establishing RIs is to obtain measurements in >120 healthy individuals per each RI partition. However, this is a time-consuming, expensive, and impractical endeavor for a single laboratory. Most RIs in clinical practice are adopted from either manufacturer package inserts, literature, other laboratories, or historical patient data. There are many caveats and pitfalls to consider when adopting RIs from literature sources: (a) methods as well as population types used to produce RIs maybe unknown or may use methods that are no longer in use today; (b) methods of measurement may not be harmonized; and (c) broad population RIs may not always apply to individuals or subgroups of individuals. To this latter point, growing awareness of marginalized and underrepresented populations in medicine has further highlighted the need for population-specific RIs. When inaccurate or inappropriate RIs are applied to laboratory results, it can lead to either under- or overdiagnosis of disease with detrimental downstream consequences. Informatics approaches are increasingly being used for epidemiological investigations, diagnostic modeling, as well as deriving population-specific RIs. Statistical approaches to RI determinations using stored laboratory data were described at least as early as the 1960s by Robert Hoffmann. Such indirect sampling approaches using stored laboratory data have traditionally been used when it is difficult to collect samples from a healthy reference population such as pediatric and geriatric populations. Indirect sampling approaches have historically been limited in clinical use due to the challenges of selecting and obtaining data from presumably healthy individuals. A posteriori methods coupled with indirect sampling can be viewed as a hybrid approach, where results stored in a laboratory information system (LIS) are combined with clinical data that is increasingly available and easily mined from the electronic medical record (EMR) to select a reference population with well-defined health characteristics. With recent successes in analytical method harmonization for several common laboratory analytes and increased integration of EMR data across multiple centers, there is now an opportunity to establish well-annotated sources of RIs that are not only stratified for specific patient populations but also specific to instrument or manufacturer platforms. Furthermore, it is now possible to intersect laboratory data with other clinical information in the EMR as well as external datasets such as geographical and meteorological information to establish precision RIs that take into account both biological and environmental factors (as in the example of impact of sunshine hours on vitamin D and parathyroid hormone RIs). As the field of data science matures, there is a new path on the horizon for high quality RI studies. In this Q&A, 5 experts with diverse experience in using large datasets for RI studies will offer their perspectives on the vast opportunities and challenges that exist in this growing, but still largely unexplored area of laboratory medicine. James Boyd: I am a classicist in terms of the methods I have applied. Much of my work was conducted earlier in my career and is contained in various publications and a book on the topic. I have adhered to the classical definition of the RI that refers to a pair of numbers (the reference limits) that bound the central 95% of a collection of values obtained from a specified group of individuals (the reference participants). The specified groups are selected carefully from a larger population according to preestablished criteria and usually are selected to be “healthy.” RIs are usually defined at a population or at a subpopulation level (e.g., age, gender, body mass index), but can also be evolved for individuals. They most often apply to just a single analyte, but can be extended to multivariate reference regions, covering multiple analytes simultaneously. As outlined in the introductory paragraphs for this Q&A, the classical approach of carefully recruiting reference participants is a “time-consuming, expensive and impractical endeavor” for most laboratories. In the modern era of rich databases, healthcare information systems, and powerful computational facilities, better approaches are needed and are currently being evolved. Julia Drees: In my laboratory, we use data mining alongside more traditional approaches to establish reference ranges for suitable analytes. We use data mining to identify healthy reference populations using both a posteriori methods and more traditional a priori methods. In each case, we start by making an exclusion list containing the diagnoses, medications, and other laboratory results that would indicate that an individual would not belong in a healthy reference population for the analyte in question. We usually do not exclude individuals with unrelated disease conditions even if that makes the population less healthy overall because we want to avoid a super-healthy cohort with an unnaturally narrow range of results. For the a posteriori method, we generate a report with all test results for the analyte in a relevant time period minus those from individuals with any of our exclusions noted in their EMR. For thyroid stimulating hormone, for example, we excluded individuals with diagnosis codes related to thyroid disease or pregnancy, prescriptions for thyroid medications, and any abnormal thyroid laboratory results other than thyroid stimulating hormone. Once we have a list of results from a suitable reference population, we can use the traditional nonparametric method recommended by CLSI EP28-A3c for establishing the intervals. Daniel Holmes: In our laboratory, we have never implemented RIs obtained entirely through data mining. Generally, we have used these techniques to confirm existing RI information or to address information-gaps for age categories that would be logistically challenging for us to investigate. Speaking specifically to indirect RI determination, naturally, we have looked at the popular Hoffmann and Bhattacharya methods but both have a manual component and are therefore subject to human error. Both methods can also lead to “finding-what-you-want-to-find” depending on how the data set is preprepared and the analysis executed. For this reason, modern mixture-model decomposition methods seem to be a better idea since the problem of Gaussian (or other distribution) mixture deconvolution has been extensively addressed in the physical sciences—for example, in spectroscopy and chromatography. Personally, I have used a number of R packages for this purpose: mixdist, mclust, and mixtools, all of which employ expectation maximization and maximum likelihood. I have also looked at noncommercial and commercial software such as Farhad Arzideh’s Reference Limit Estimator and the peak deconvolution tools in OriginLab. We do not have sufficient access to the diverse electronic health records in our jurisdiction to meaningfully use entirely informatic approaches to filtering patient populations to presumably healthy cohorts, although there are numerous published examples of successful strategies. It is worth briefly mentioning that for age-dependent RIs, we have used the raw data from other studies combined with our own to inform our analyses in age categories for which we have a paucity of data. Age-dependent RI fitting has a large body of literature but we looked at or used Altman’s method, Royston’s method, the Lambda Mu Sigma method, and quantile regression. John Ioannidis: I have worked with various statistical tools, ranging from plain descriptive, nonparametric methods all the way to deep learning methods. Even in well-characterized datasets, e.g., those collected with explicit research orientation rather than for routine care and administrative documentation in electronic health records or claims, it becomes very quickly obvious that the definitions used for “normal,” the eligibility criteria, exclusion of outliers, and stratification plans (preconceived or variously data driven) often make a substantial difference. More complex methods, if anything, have more analytical degrees of freedom. More analytical choices may lead to even greater uncertainty and variability in the generated RIs. Arjun Manrai: Reference interval estimation is a crucial task for ensuring proper diagnosis and treatment across nearly all areas of medicine. Modern machine learning and statistical approaches to RI estimation, including those that we are developing, draw heavily on a rich literature that dates back decades. One helpful way to partition the growing collection of approaches is whether a method operates by “direct” vs “indirect” calculation. With direct sampling approaches, the laboratory or investigator selects healthy individuals and directly computes an interval that captures “normal” variation (e.g., 2.5th to 97.5th percentiles). However, the definition of “healthy” is elusive and subjective or may be difficult to ascertain. By contrast, indirect approaches use statistical modeling to calculate a reference range using data from both “healthy” and “sick” individuals such as those from the general hospital population, likely containing data from individuals with overt or subclinical disease. Research is being pursued in the development of both new direct and indirect approaches including incorporating new machine learning approaches. This work is enabled by public datasets such as the National Health and Nutrition Examination Survey, large-scale biobanking efforts (e.g., UK Biobank), and stored laboratory data across institutions. James Boyd: The chief advantage is data availability. There are few limits on the number of data values that can be collected and analyzed with truly “big data” that allows for high precision in reference limit endpoints. Data collected from specific subpopulations (e.g., ICU patients) may be more useful for the evaluation of those populations than data collected from “healthy” individuals. Such needs can easily be served using stored laboratory data. Julia Drees: The number of results available in the a posteriori method, even after excluding results from unsuitable individuals, is generally far more than can be gathered and assayed prospectively in the traditional a priori method. The high volume of results is especially helpful if the analyte requires age- and sex-specific RI partitions since each partition requires 120 results if following CLSI EP28-A3c. Fewer resources are required since no volunteers, tech time, instrument time, or reagents are needed. Additionally, results in the LIS are from samples that were collected and stored in real-world settings and therefore more likely to contain the variation seen in practice and not include any phlebotomy or storage artifacts from a single a priori RI study. Daniel Holmes: There is an abundance of data available at no cost that can often be filtered to remove unsuitable participants based on normal results from non-target analytes. At some institutions it is also possible to cross-reference against the EMR to exclude participants taking certain medications or having relevant medical conditions. John Ioannidis: Most laboratories until now had been using mostly literature-based reference standards along with some limited testing that they would do on their own. Anecdotally, they had not even used the CLSI-recommended 120 samples, but far fewer. Therefore, their RI practices had resembled more of a verification procedure (likely to pick only major anomalies) rather than a full development and validation of RIs. The availability of large collections of stored samples and of large datasets has the advantage of transforming the previous verify-only approach to a more genuine develop-and-validate approach. Often the available collections and datasets are large enough that they allow also full exploration of relevant stratifications. They also come at no or limited cost per data item and, in theory at least, some of them can be shared across many laboratories, allowing also better harmonization and comparisons between laboratories. The advantages become more prominent, when the same resources can be tapped from and compared across more teams and more laboratories. Arjun Manrai: The first major advantage of using stored laboratory data, such as the properly-consented records from a hospital’s testing laboratory, is the number of patient records that may be available for study. Hundreds of thousands or millions of patient laboratory values, across dozens or hundreds of laboratory analytes, may be available from a single institution. A second advantage of using stored laboratory data is representativeness, where the data can be used to investigate how well existing normal ranges capture the variation in the population over time, across clinical and demographic groups, and across regions served by different institutions. To the extent that data can be analyzed across institutions, both advantages will improve. James Boyd: Large datasets do not guarantee high quality RIs. Owing to diagnostic and clinical data in most databases, it is often difficult to the population being or how it is of the population in which the RIs will be applied. values that bound the central 95% of the population can be if the data are not carefully or contain they may not be useful for the Statistical methods for exclusion of reference values are and may not apply generally to It may be difficult to account for in the analyte measurement methods measurement or across large Julia Drees: all analytes are to the a posteriori or indirect data mining approaches. are the best since they are on healthy individuals. other such as some are only on with disease not be of a healthy reference Additionally, methods be used to RIs when to a new the laboratory can that results from the new with results from the historical very data mining is heavily on a EMR if diagnosis codes are not or if a patient is new to the hospital samples or stored results may be from individuals have exclusion To the latter we exclude samples and results from individuals have been of our health for less than Daniel Holmes: the results obtained from data mining methods are not by the existing body of literature, that the “healthy” subpopulation by the statistical is not of the healthy As it specifically to mixture-model decomposition methods (e.g., the of the population may be and may not address this Additionally, if the do not the of the the results may have a subjective on of the For example, there may be no in the normal of the Hoffmann method or a maximum method may to an are John Ioannidis: There are still major and many of them are A problem is the quality of the data. not quality and may even major data mining methods are often to major quality in the data, and they may be even more in that data rather than genuine in any other data it is to how the data have been the preanalytical the of the the analyzed samples, and the sampling and However, these are often or methods that to exclude and individuals may when they work with datasets that include many individuals and limited information to identify them Arjun Manrai: Data mining approaches for RI estimation often data from administrative records (e.g., in is and both the and into these data. A in data mining approaches for RI studies is therefore the general by in the data and machine learning whether a statistical method or machine learning is a the in an administrative (e.g., or some of all of James Boyd: that any data has human or as a quality by the data likely will be the most useful for of RIs that most RIs. To avoid make that hospital that data RIs are being used and that medical the for and of these Julia Drees: CLSI and the of Clinical and Laboratory currently direct methods for selecting a reference population, that the health are not well-defined when using indirect methods. our use of data mining clinical data from the EMR with stored results to well-characterized healthy reference populations. In this of the a posteriori method, no are needed to healthy and populations and we can use the in CLSI just as we would with more traditional methods. As become more and data mining becomes more of individuals, I CLSI and will their Daniel Holmes: There is a from and the on RIs and published in Clinical and Laboratory in This is a for of these approaches on their own data. I would also a by and in Clinical that a on the of indirect RI determination from through to the CLSI only limited on this topic. John Ioannidis: There are many most of them still at a or rather than I it would be to to we have in the from data with data For example, in the there is the that in and is an example of There are also some proper for data, but most of them on the and data that can be rather on the need to the and of the data being to these analytical methods. Arjun Manrai: from the CLSI (e.g., CLSI are often in establishing or RIs. contain for data approaches to RI estimation, including the statistical between establishing and a in terms of or specifically for using data and data mining approaches, the area is As the to establish new I that data and data work to the analytical and clinical into large-scale datasets as well as new modeling approaches. James Boyd: RIs are statistical to a of the healthy population values are the central 95% of the their clinical be on any other than that a sampling of individuals in the healthy population not a of values in of the of the 95% reference I do not of other to their clinical but not very powerful statistical have been described to RIs. of RIs in different populations are Julia Drees: We work with a clinical this and our RIs with published intervals to make on To our intervals with the a posteriori method, we a number of samples using a more traditional a priori method. However, data mining can also a in this The is that we start with recent samples that had a test that is not related to the analyte of but the same we for with normal We can apply the same exclusion criteria of diagnosis medications, and laboratory results related to the analyte in and also exclude individuals that had the test in by a in the In this method, we samples on the report and prospectively them for the analyte of We have used this a priori method for RI and to establish intervals for an analyte that is unsuitable for the a posteriori method. Daniel Holmes: statistical are used depending on whether is RIs or from a large data set filtered for healthy or whether is using mixture decomposition (e.g., In terms of of of the RI on the filtered data, it is possible to from machine learning such as Most we are to the to if I these the and and are the and medical there any in the by age, gender, do the RIs to own and with John Ioannidis: This is the most difficult validation needs studies to that the RIs meaningfully between have clinical and those A needs studies that whether data approach is (or they both have in the most is whether this testing and diagnostic make a in in the This not only in the to identify values, but also in the to have and and work has in this but availability of datasets allow some of these At a first these are likely to be using existing data with sufficient to clinical studies are even but is a Arjun Manrai: Clinical is to first the challenges described is usually to the RIs of any new method against the published RIs for A second of validation from a of approaches, both direct and to the same data, and the of the intervals across the approaches. A of validation reference ranges with clinical A of validation from the of is often an (e.g., into “normal” vs values, the as a proper may analyses to be both more and less to The most powerful of clinical likely these approaches. James Boyd: that the and of RIs for laboratory test is a of laboratory use of RIs is bound to into the data approaches to RI will but will need to be carefully for In the classical is an clinical it has including the of results as to in an analyte the and between the reference population used to it and the population in which it is applied. Once several measurements of the same test have been in an reference values defined for that test are more powerful of in health than RIs. described methods to reference values seem to to be the most approach for on RIs. Julia Drees: I will to become more and in data which in will make data mining more successful at healthy reference populations. I a posteriori methods and data mining in general will become more used laboratories the analytical tools and to access and data in the EMR Daniel Holmes: I am that the Clinical will Data as a for practice and take advantage of the opportunities that exist for quality and medical in routine data. we do we may marginalized in our own by those from other are to do This of will not be through more use on our of but with Data John Ioannidis: It is to data and methods for their analysis are to and to become even more I am less on whether all of this will clinical but some are Arjun Manrai: One of the major challenges for the field is when and how to RIs across groups, where groups may be defined by many including clinical (e.g., individuals with or demographic (e.g., I am that we will increasingly better of and clinical to I machine learning approaches will a in this care and that new data mining approaches will our of it to be “healthy” by collections of many laboratory healthcare and in they have to the of this and have the following (a) to the and of data, or analysis and of (b) or the for (c) of the published and to be for all of the ensuring that related to the or of any of the are and all the of Clinical
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | no category Domain: not available · Genre: Methods About the Canadian research system: no · About a Canadian topic: no | Not applicable | low |
| gpt | no category Domain: not available · Genre: Methods About the Canadian research system: no · About a Canadian topic: no | Other design | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.067 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.007 | 0.009 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.004 | 0.002 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".