Comparison of Accuracy Between <sup>13</sup>C- and <sup>14</sup>C-Urea Breath Testing: Is an Indeterminate-Results Category Still Needed?
Bibliographic record
Abstract
Helicobacter pylori infection is the leading cause of peptic ulcer disease. The purpose of this study was, first, to assess the difference in the distribution of negative versus positive results between the older 14C-urea breath test and the newer 13C-urea breath test and, second, to determine whether use of an indeterminate-results category is still meaningful and what type of results should trigger repeated testing. Methods: A retrospective survey was performed of all consecutive patients referred to our service for urea breath testing. We analyzed 562 patients who had undergone testing with 14C-urea and 454 patients who had undergone testing with 13C-urea. Results: In comparison with the wide distribution of negative 14C results, negative 13C results were distributed farther from the cutoff and were grouped more tightly around the mean negative value. Distribution analysis of the negative results for 13C testing, compared with those for 14C testing, revealed a statistically significant difference between the two. Within the 13C group, only 1 patient could have been classified as having indeterminate results using the same indeterminate zone as was used for the 14C group. This is significantly less frequent than what was found for the 14C group. Discussion: Borderline-negative results do occur with 13C-urea breath testing, although less frequently than with 14C-urea breath testing, and we will be carefully monitoring differences falling between 3.0 and 3.5 %Δ. 13C-urea breath testing is safe and simple for the patient and, in most cases, provides clearer positive or negative results for the clinician.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.043 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".