The Forgotten Mountain Plot—An Illustration of Bias across Lipase Methods
Bibliographic record
Abstract
When clinical laboratories evaluate methods (example X and Y) for agreement or compatibility, two types of analysis are performed. First, a regression analysis (Fig. 1A and B), and second a Bland–Altman (difference) plot (Fig. 1C and D). The regression analysis provides information on the extent of the method correlation and if the methods under comparison are in agreement by relating the line of identity. Method comparison studies for lipase assays comparing Sentinel (A) and Roche (B) methods against the Sekisui method using regression analysis, where the dashed black line indicates line of identity, the red line corresponds to the fitted line based on Passing–Bablok regression, and the blue shaded region is the ±14.2% desirable allowable total error on the line of identity. The middle panels are the %difference plot for Sentinel (C) and Roche (D) against Sekisui as the reference method, with the red solid line corresponding to the %mean bias. The folded cumulative distribution plot or simply the mountain plot (E) for the %difference between Sentinel and Sekisui or (F) %difference between Roche and Sekisui are in the bottom panel. The 5th percentile is marked by the horizontal dashed line in gray to indicate outliers or data points with the largest %difference, outside of 90% of the data. The mountain plots have smoothing by penalized cubic regression splines (gray solid line). Smoothing functions are not necessary for utilizing mountain plots. To quantitatively assess the extent of bias, a difference plot is generated. However, in the case of proportional bias, the data may not be normally distributed (as in our case below) making the difference plot less useful. While this can be alleviated in some cases by performing mathematical transformations (e.g., log transformation), such transformations on skewed data may not always force the data to conform to normality (as in our case). In addition, the difference between the methods can be normalized to the reference method (Y − X)/X, but again the overall mean %bias is often indicated on non-gaussian distribution of data. While this latter approach provides some insight into bias evaluation, it should be noted that normalization to the reference method and log transformations are infrequently done (1). A more suitable alternative to the difference plot that is useful in such cases is the so-called folded empirical cumulative distribution plot, or simply the “mountain plot” that Krouwer et al. published more than 3 decades ago (2). This approach is underutilized and infrequently taught to trainees (our collective impression), but is informative when the data collected are not normally distributed. In brief, the mountain plot requires first determining the difference between the methods (as one would for generating a difference plot), then ranking the difference (rank 1 = lowest), computing the percentile for each difference (rank × 100/(N + 1), where N = number of differences) and then folding the plot. The latter is simply done by subtraction of the percentiles > 50 from 100 (3). Here, we illustrate the value of the mountain plot during our evaluation of a new lipase assay on the Abbott Alinity c instrument from Sentinel (based on the methylresorufin method). This method was compared against a working lipase assay from Sekisui (also on the Alinity c) and another methylresorufin method on the Roche cobas c502 analyzer. Importantly, the data were not normally distributed (evidenced from histograms and by the Shapiro–Wilk test). While the methods are highly correlated (Fig. 1A and B), there is clearly a proportional bias that is more evident when comparing the Sentinel method to that of the Sekisui method. The bias between the Alinity methods can also be appreciated from the difference plot (Fig. 1C) which would suggest an average bias of −22%, but is skewed by lack of normality, as often the case when there is proportional bias. One can, of course, mark the median bias in the difference plot, as it is a more robust statistic for skewed distribution, but this is rarely done. In general and in such cases with proportional biases, a nonparametric assessment by the mountain plot immediately provides a more reliable indication of the middle ground on relative difference or overall reflection of the bias across lipase methods (Fig. 1E and F). Mountain plots are also good at showing outliers, whereas the traditional parametric methods would ideally require the exclusion of outlier(s). In the case of lipase, the Sentinel method has a higher median bias of −25% (or −12 U/L if absolute difference were plotted) relative to the Sekisui method, which is unacceptable, given the desirable (±14.2%) and optimal (±7.1%) allowable total error for lipase (4). The mountain plots also reveal at a given percentile (e.g., 5th percentile), an insight into some of the largest differences or outliers that can be observed (Fig. 1E and F), which may be more easily visualized by drawing a horizontal line at the required percentile on the graph. While the Roche method is in better agreement with Sekisui, at higher values, discrepancies are similarly evident, typically near and above the linearity limit (Fig. 1D and F). It should be noted that lipase assays in general continue to lack standardization and gaps in distribution of lipase results have been well described (4, 5). These gaps can be attributable to worsening linearity as lipase activity approaches the manufacturers’ linearity limit (4). Here we show that inter-method biases in lipase are appropriately evaluated, independent of how data is distributed through mountain plots. While complementary to the difference plot, mountain plots make it easier to find the central 95% of the data and enable a direct comparison of the distribution between methods (3). Author Contributions:The corresponding author takes full responsibility that all authors on this publication have met the following required criteria of eligibility for authorship: (a) significant contributions to the conception and design, acquisition of data, or analysis and interpretation of data; (b) drafting or revising the article for intellectual content; (c) final approval of the published article; and (d) agreement to be accountable for all aspects of the article thus ensuring that questions related to the accuracy or integrity of any part of the article are appropriately investigated and resolved. Nobody who qualifies for authorship has been omitted from the list. Felix Leung (Conceptualization-Equal, Methodology-Equal, Project administration-Equal, Resources-Equal, Writing—review & editing-Equal), Samantha Logan (Data curation-Equal, Investigation-Equal, Validation-Equal, Writing—review & editing-Equal), Anselmo Fabros (Investigation-Equal, Methodology-Equal, Project administration-Equal, Validation-Equal), Rajeevan Selvaratnam (Conceptualization-Lead, Data curation-Lead, Formal analysis-Lead, Investigation-Lead, Methodology-Lead, Supervision-Lead, Visualization-Lead, Writing—original draft-Lead, Writing—review & editing-Lead) Authors’ Disclosures or Potential Conflicts of Interest:No authors declared any potential conflicts of interest.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.030 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".