Identifying Important Features for Clinical Diagnosis of Thyroid Disorder
Bibliographic record
Abstract
Abnormal production of thyroid hormones in our body causes thyroid disorders such as hypothyroidism, hyper-thyroidism, Hashimoto's disease, Graves' disease, and thyroid nodules. Undiagnosed thyroid disorders can affect the quality of life of an individual both physically and mentally. Thyroid disorders are common but sometimes become difficult to diagnose since the symptoms can be easily associated with other health conditions. Clinicians identify thyroid disorders by measuring the levels of thyroid hormones in our blood stream. This work aims to help clinicians by carefully investigating if thyroid diagnosis improves when all important features (a complete thyroid panel) is measured as opposed to a select few. Much of previous work has focused on the performance of classifiers, supervised and unsupervised, for the prediction of this disorder. Departing from this tradition, we focus on the concept of feature importance and its clinical implications. We identify the top-4 important features that predict the presence of thyroid disorder and show that these can be measured by clinicians cost-effectively. We also identify the pitfalls of current clinical practice of not checking the entire thyroid panel, prevalent in many countries with universal health care. Finally, we show that our results are quite robust and are unlikely to change with the choice of classifier or due to the inherent nature of a dataset in hand like imbalance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.017 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".