MétaCan
Menu
Back to cohort
Record W3174967186 · doi:10.1210/clinem/dgab435

Predicting Malignancy in Pediatric Thyroid Nodules: Early Experience With Machine Learning for Clinical Decision Support

2021· article· en· W3174967186 on OpenAlexafffund
Lebohang Radebe, Daniëlle C M van der Kaay, Jonathan D. Wasserman, Anna Goldenberg

Bibliographic record

VenueThe Journal of Clinical Endocrinology & Metabolism · 2021
Typearticle
Languageen
FieldMedicine
TopicThyroid Cancer Diagnosis and Treatment
Canadian institutionsVector InstituteCanadian Institute for Advanced ResearchHospital for Sick ChildrenUniversity of Toronto
FundersGenome CanadaCanadian Institutes of Health ResearchGarron Family Cancer CentreRare Disease FoundationNatural Sciences and Engineering Research Council of CanadaCanadian Institute for Advanced Research
KeywordsThyroid nodulesMalignancyThyroidMedicineComputer scienceMedical physicsRadiologyPathologyInternal medicine

Abstract

fetched live from OpenAlex

OBJECTIVE: To develop a machine learning tool to integrate clinical data for the prediction of non-benign thyroid cytology and histology. CONTEXT: Papillary thyroid carcinoma is the most common endocrine malignancy. Since most nodules are benign, the challenge for the clinician is to identify those most likely to harbor malignancy while limiting exposure to surgical risks among those with benign nodules. METHODS: Random forests (augmented to select features based on our clinical measure of interest), in conjunction with interpretable rule sets, were used on demographic, ultrasound, and biopsy data of thyroid nodules from children younger than 18 years at a tertiary pediatric hospital. Accuracy, false-positive rate (FPR), false-negative rate (FNR), and area under the receiver operator curve (AUROC) are reported. RESULTS: Our models predict nonbenign cytology and malignant histology better than historical outcomes. Specifically, we expect a 68.04% improvement in the FPR, 11.90% increase in accuracy, and 24.85% increase in AUROC for biopsy predictions in 67 patients (28 with benign and 39 with nonbenign histology). We expect a 23.22% decrease in FPR, 32.19% increase in accuracy, and 3.84% decrease in AUROC for surgery prediction in 53 patients (42 with benign and 11 with nonbenign histology). This improvement comes at the expense of the FNR, for which we expect 10.27% with malignancy would be discouraged from performing biopsy, and 11.67% from surgery. Given the small number of patients, these improvements are estimates and are not tested on an independent test set. CONCLUSION: This work presents a first attempt at developing an interpretable machine learning based clinical tool to aid clinicians. Future work will involve sourcing more data and developing probabilistic estimates for predictions.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.014
metaresearch head score (Gemma)0.035
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.014
Threshold uncertainty score0.073

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0140.035
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0000.001
Scholarly communication0.0010.002
Open science0.0010.001
Research integrity0.0010.003
Insufficient payload (model declined to judge)0.0010.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.057
GPT teacher head0.402
Teacher spread0.345 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations16
Published2021
Admission routes2
Has abstractyes

Explore more

Same venueThe Journal of Clinical Endocrinology & MetabolismSame topicThyroid Cancer Diagnosis and TreatmentFrench-language works237,207