Osmolality Gaps: Diagnostic Accuracy and Long-Term Variability
Bibliographic record
Abstract
BACKGROUND: The osmolal gap (OG) is a screening test for the detection of toxic volatiles such as methanol and ethylene glycol. We used mean values of patient data to assess the diagnostic accuracy and long-term stability of OG measurements. METHODS: In a prospective study period in 2003, all requests for volatiles had OGs calculated and quality-control samples were analyzed for OG. ROC curves were constructed to determine whether OG could predict the presence of toxic volatiles in serum. This was also done in a retrospective study for data from 1996 to 2004. Our laboratory database was searched for all emergency room patients for the period of 1996 to 2004 who had tests ordered that allowed us to calculate OGs. RESULTS: For the prospective study period in 2003, the ROC areas indicated that we could accurately predict the presence of toxic volatiles but at markedly different decision cutpoints depending on the formula used. These cutpoints ranged from +10 to +33 mosmol/kg. In the retrospective study, the mean OGs in the patient population for each of the 3 formulas increased by 12 mosmol/kg from 1996 to 2004. For this reason, the diagnostic accuracy was poor when all data were analyzed together. CONCLUSIONS: Under properly controlled conditions, the OG has high sensitivity and specificity for detection of poisoning with some volatiles. Over the long term, however, use of the reference interval of -10 to +10 mosmol/kg yields poor diagnostic accuracy because mean OGs are not constant over time. Bedside calculation is not advisable.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".