Bibliographic record
Abstract
One of the most frequent questions asked by the new clinical chemist is, “How do I set up the quality control for a particular analyzer?” Two new documents described below will readily help both the neophyte and practicing chemist. The practice of quality control in the clinical laboratory has been evolving in fits and starts. Unfortunately, today's laboratory quality control practices are obscured by undeserved complexity which robs the quality control practitioner (primarily the bench medical technologist) of tangible quality control intuition and in Deming's words, also robs her of pride of workmanship (1). In US clinical laboratories, quality control practices are hugely divergent in addition to being costly (2). In a recent review, Kazmierczak stated, “the costs associated with performing quality control testing and the costs associated with evaluating, reviewing and maintaining quality control records are not trivial” (3). He cited a developing nation hospital study that found the costs of quality to be 22% of total direct laboratory expenses, with 89% of the costs for maintaining quality associated with calibration and analysis of quality control material necessary to confirm the accuracy and reliability of test results (4). Simplicity is required in today's quality control practices and in this month's issue of Clinical Chemistry, Yago and Alcover, Spanish clinical chemists, provide highly instructive nomograms that permit the selection of simple quality control rules to yield appropriately low frequencies of defective results in analytic runs of 100 (5). These surprisingly orderly, ready-to-use nomograms provide a welcome contrast to the complex rule sets employed by many practicing clinical chemists and elaborated with or without the aid of quality control selection software. The nomograms of Yago and Alcover yield the number of defective analyses per analytical run of 100 samples. These nomograms visually demonstrate the importance of selecting analyzers with high σ assays (exceeding 4) to keep the number of undetected unacceptable patient results reported under 2 or 3 per 100. The authors have presented the utility of the 12s, 12.5s, 13s, 13.5s, and 14s control rules to detect systematic error. With assay sigmas above 4, the 13s control rule using 3 quality control materials will yield only 1 defect per 100. Assays with sigmas >5 monitored using 3 concentrations of QC material, the 13.5s and 14s control rules yield between 0 and 1 defective samples per 100 samples. The use of more sensitive rules (i.e., 12s, 12.5s) would not be of significant additional benefit and would only increase the probability of false rejection. While not explicitly stated by Yago and Alcover, analytical methods with sigmas of 2 or 3 require selection of a combination of fairly nonspecific but sensitive rules, e.g., the 13s, 22s, and the R4s. The newly remade CLSI document C24, Statistical Quality Control for Quantitative Measurement Procedures, serves as an excellent complement to the Yago and Alcover nomograms (6). C24 emphasizes that instrument performance will not be improved by the use of more sensitive rules. These rules will tend to detect smaller errors resulting in increased QC rejections, increased troubleshooting, repeated measurements, and delayed reporting. Furthermore, some of these more sensitive rules, the 41s and 10x rules have been found not to contribute significantly to error detection but may present higher probabilities of false rejection in highly precise analyzers (7). Over a clinical chemist's long career, she must strive to introduce instruments and assays with lower and lower intrinsic analytic imprecision (higher σ). In this manner, quality control becomes simpler and less mysterious. The calculation of σ is very important. The usual calculation, including that used by Yago and Alcover, invokes the subtraction of bias. The subtraction of the bias may render sigmas so low that the chemist/technologist will attempt to use multiple control rules to manage assay quality, which will invariably increase the probability of false rejection. It is our belief that there is little that a laboratory system can do with bias once interinstrument variation is minimized by using similar analytic systems and occasionally employing a conversion factor to make a test on one analyzer resemble the same test on a different analyzer. When we reached out to one of the coauthors of C24, Dr. Greg Miller wrote, “bias is difficult to assess since proficiency testing data are subject to unpredictable noncommutability influences and laboratories do not have resources to do patient sample comparisons to reference methods (if they could locate a laboratory that offers reference methods). In general, we rely on the manufacturers to minimize bias when they provide calibration traceability to the best available reference system as part of the manufacturing process. Consequently, in most cases bias can be ignored when estimating σ” (G. Miller, personal communication, March 31, 2016). This topic is well developed in C24. The calculation of σ is also well developed in C24. The document provides criteria for selecting the best estimate of imprecision for calculating σ. In the past, the laboratory community has made serious errors in the calculation of σ. It would be highly desirable if the manufacturer provided for every analyte on their specific analyzers, the 15th, 50th, and 85th percentile CV. The user could confidently calculate the sigmas of poorly performing, average performing, and well-performing analyzers. If the sigmas for these various analyzers uniformly exceeded 4 or 5, we would be very happy. A few years ago, one of us was provided with all of the 3-month quality control summaries for a specific hematology system and was allowed to summarize and publish the findings (8). When we attempted to extend this work into critical care chemistry, our efforts were stymied. The laboratory community must laud any manufacturer who openly allows the presentation of such data. We are now witnessing the proliferation of blood gas electrolyte systems that employ disposable reagent cartridges and self-contained calibrating fluids that sometimes serve as internal quality control (9). We caution users not to use the imprecision of these ersatz fluids to calculate σ. These fluids lack commutability with patient samples and there may be little relationship between the very low imprecision of these fluid measurements and the imprecision the user encounters when they analyze patient whole blood (10). To accurately assess imprecision of a new analyzer, we advise the user to run an external quality-control sample 3 times daily over the lifetime of the cartridge. We have found that while the σ of a single system might be adequate, the effective σ will degrade when multiple analyzers are used to report sequential patient samples (11). Again, before acquiring these cartridge-based analyzers, the user should run patient replicates across their multiple analyzers over a suitable period to assure acceptable between-analyzer variation. We caution the users of Yago and Alcover's nomograms to be aware that their choice of total allowable error (TEa) limits will affect the calculation of σ and thus detection of systematic error. As an example, we calculated the σ value for a sodium ion selective electrode in one of our laboratories using the CLIA TEa of 4.0 mmol/L. The resulting mean σ at two concentrations of QC material was 4.99. Using a σ equivalent to 5 and 2 concentrations of QC a 13S rule would be expected to result in a maximum of <1 unacceptable patient result per 100 samples. If we were to use a small TEa of 3.0 mmol/L the resulting mean σ would be approximately 3.8. Using the same 13S rule we would expect to find between 4 and 5 unacceptable patient results per 100 samples. The selection of overly permissive TEa could allow for clinically significant variation to go undetected. Similarly, the selection of overly restrictive TEa limits would result in significant rework, reanalysis and correction of patient results that would have little to no impact on patient care. Yago and Alcover's nomograms for the detection of systematic error are just the beginning of their simplification of laboratory quality control. In a subsequent manuscript, they will provide analogous nomograms for random error. Their nomograms plus the elegant CLSI C24 document should help with our comprehension and implementation of more straightforward and intuitive quality control systems (12).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.047 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.004 | 0.007 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".