In Reply
Bibliographic record
Abstract
We published our findings to alert the diabetes-research community that multiple lots of SureStep®Flexx® strips (LifeScan/Johnson & Johnson) demonstrated increased glucose sensitivity in anemic patients in the intensive care unit. Given that these nonoptimally performing strip lots were used during the NICE-SUGAR (Normoglycemia in Intensive Care Evaluation and Survival Using Glucose Algorithm Regulation) trial, their use may have compromised the study's primary conclusion, that intensive insulin therapy was harmful to intensive care unit patients. We did not hypothesize but merely quoted the 2010 study of Pidcoke et al., which demonstrated that treating artifactual hyperglycemia can produce iatrogenic hypoglycemia (1). In their prospective observational study, the authors compared the hypoglycemia rates in a burn intensive care unit before and after they corrected artifactually increased SureStep®Flexx® glucose measurements with a calculation that incorporated the hematocrit and the original whole-blood glucose measurement. During the 4 months of providing lowered but corrected glucose measurements, the authors achieved a 78% decrease in the prevalence of hypoglycemia in critically ill anemic patients treated with insulin and tight glucose control (P < 0.001)! As for the LifeScan argument about trueness, we deemed that arterial blood gas glucose measurements were truth; Pidcoke et al. used plasma glucose measurements produced by a highly accurate central laboratory analyzer. Finally, the allowable error limits for in-hospital testing of whole-blood glucose will be tightening. It is almost fallacious to invoke the ±20% limits; these limits were devised for patient self-monitoring, not for intensive insulin treatment. At the recent US Food and Drug Administration meeting in Gaithersburg, one of us suggested narrowing these 95% limits for in-hospital whole-blood glucose testing to ±12.5% (2).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.051 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.005 | 0.003 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.027 | 0.034 |
| Insufficient payload (model declined to judge) | 0.021 | 0.012 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".