Evaluation of pheochromocytoma diagnostic accuracy using plasma free metanephrines versus 24-hour urinary catecholamines
Bibliographic record
Abstract
Since the first biochemical characterization of catecholamine-secreting tumors in the 1950s, the optimal laboratory approach for diagnosing pheochromocytoma has been debated. This retrospective diagnostic accuracy research compared the performance of Plasma Free Metanephrines (PFM) against 24-hour Urinary Catecholamines (UC) in 247 consecutive patients evaluated for suspected pheochromocytoma at the Toronto Institute of Clinical Sciences between January 2015 and December 2019. All patients underwent both biochemical tests as part of a standardized diagnostic protocol, and final diagnosis was confirmed by histopathology in operated cases or by clinical follow-up of at least 24 months in non-operated cases. Of 247 patients, 34 (13.8%) received a confirmed diagnosis of pheochromocytoma or paraganglioma. PFM demonstrated a sensitivity of 97.1% (33/34) compared to 85.3% (29/34) for UC (p=0.031). Specificity was 91.4% for PFM versus 87.9% for UC (p=0.284). The positive predictive value of PFM was 64.7% compared to 54.7% for UC, while the negative predictive value was 99.4% for PFM versus 97.2% for UC. ROC curve analysis yielded an area under the curve of 0.968 (95% CI 0.941-0.995) for PFM and 0.892 (95% CI 0.837-0.947) for UC, with a statistically significant difference between curves (p=0.004). Among the 18 patients with false-positive PFM results, the most common confounding factors were concurrent tricyclic antidepressant use (n=5), obstructive sleep apnea (n=4), and acute physiological stress from recent hospitalization (n=4). The single false-negative PFM case involved a dopamine-secreting paraganglioma. Time-to-diagnosis was shorter with PFM-first protocols (median 14 days vs. 28 days for UC-first workup, p<0.001). These findings confirm that plasma free metanephrines offer superior diagnostic sensitivity for pheochromocytoma and support their role as the preferred first-line screening test, though awareness of false-positive triggers remains important for avoiding unnecessary imaging and surgery.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.074 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".