Are Bethesda III Thyroid Nodules More Aggressive than Bethesda IV Thyroid Nodules When Found to Be Malignant?
Bibliographic record
Abstract
The Bethesda classification system for thyroid fine needle aspirate (FNA) is used to predict the risk of malignancy and to guide the management of thyroid nodules. We postulated that thyroid malignancies characterized as Bethesda III on FNA have more aggressive features than those classified as Bethesda IV. A retrospective chart review was performed to identify those who underwent thyroid surgery at a single tertiary hospital setting between 2015 and 2020. Associations between Bethesda category, molecular genetic test results, and histopathologic findings were examined. Out of 628 surgeries that were performed, 199 (54.2%) Bethesda III nodules and 216 (82.8%) Bethesda IV nodules were malignant. Of those that were malignant, 37 (18.6%) and 22 (10.2%) Bethesda III and Bethesda IV nodules showed aggressive features, respectively (p value = 0.014). There was a proportionally increased number of aggressive features in extra-thyroidal extension, lymph nodes metastasis, and all aggressive subtypes of papillary thyroid cancer in the Bethesda III category. Although Bethesda IV nodules are much more likely to be malignant (p value = 0.002), our study suggests that Bethesda III nodules that are resected are more likely to have aggressive features than Bethesda IV nodules, with a statistically significant increase in the solid variant of papillary thyroid cancer and lymph node metastasis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".