Diagnostic accuracy and safety of fine‐needle aspiration biopsy of the parapharyngeal space
Bibliographic record
Abstract
Fine-needle aspiration (FNA) biopsy of the parapharyngeal space (PPS) is a diagnostic challenge and sampling is often done without image guidance, often trans-orally. Primary PPS tumors are rare, and there is a broad differential diagnosis. The accuracy of PPS FNA, in particular without CT-guidance and using liquid-based cytology (LBC), has not been well studied. Pathology records from our institution (a 1,100 bed Canadian academic tertiary care centre) were searched to identify all patients who underwent PPS FNA from September 1991 to August 2009. The FNA diagnosis was compared to the gold standard of subsequent histopathology or long-term clinical follow-up. Of 36 FNAs, 3 employed image guidance. Eleven (31%) FNAs were nondiagnostic. In the 25 diagnostic FNAs, there was sensitivity 89%, specificity 94%, PPV 89%, NPV 94%, and accuracy 92% for the diagnosis of positive or negative for malignancy. A correct specific diagnosis was made in 9/25 (36%) cases. The nondiagnostic rate was significantly higher (P < 0.025) in FNAs prepared as conventional smears (9/17 = 53%) versus LBC (2/19 = 11%). A specific diagnosis was made significantly more often (P < 0.05) with LBC (8/19 = 43%) versus conventional smear (1/17 = 5.9%). One minor complication from FNA occurred. In conclusion, PPS FNA is safe and accurate for the diagnosis of malignancy. The rate of reporting a specific diagnosis is low. Nondiagnostic FNAs are common. There are more specific diagnoses and fewer nondiagnostic tests with LBC than with conventional smears. Improved specimen quality with LBC is likely a factor.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".