A Novel Computational Method for Biomedical Binary Data Analysis: Development of a Thyroid Disease Index Using a Brute‐Force Search with <scp>MLR</scp> Analysis
Bibliographic record
Abstract
The thyroid disease index ( TDI ), which estimates thyroid disease progress based on hormone concentration measurements and hormone pattern changes, was developed. In this study, we measured concentrations of hormone profiles in the androgen and estrogen metabolic pathways from 23 patients with thyroid disease, as well as 20 unaffected people. We illustrated that the hormones 2‐hydroxyestrone (2‐ OH‐E1 ), 2‐hydroxyestradiol (2‐ OH‐E2 ), 2‐methoxyestrone (2‐ MeO‐E1 ), 2‐methoxyestradiol (2‐ MeO‐E2 ), and 2‐methoxyestradiol‐3‐methylether (2‐ MeO‐E2 ‐3‐methylether) are related to the development of thyroid disease through t ‐tests. Though the concentration levels of these hormones generally increase as the disease progresses, big fluctuations cause the determining of a disease's progress by measuring hormone levels to be difficult. The differing patterns between the correlation matrices of the disease and control groups possibly indicates changes in hormone releasing patterns during the thyroid disease's progress. Because of a lack of progressive experimental data on thyroid disease, binary data for the two categories (the thyroid disease patients and the control group) was utilized. Binary logistic regression was used to analyze five risk factors associated with thyroid disease, and the highest overall accuracy was 97.7% with three risk factors. Logistic regression models, however, are unable to describe disease progress. Hence, the TDI was developed to estimate thyroid disease progress. An arbitrary ranking of disease progress was generated for the TDI equation. The ranking contained a total number of 29 030 400 entries with six stages from the control group and eight stages from the disease group. Multiple linear regression ( MLR ) analysis was performed with a brute‐force search. The best result among the MLR runs presented strong correlation ( r 2 values of 0.840 and q 2 values of 0.663) between the selected hormones and the values of the disease progress in the training set. Overall accuracy of our novel method was 90.7%, which is worse than the 97.7% of logistic regression models. Brute‐force search with MLR analysis might classify different types of thyroid disease progress such as thyroid mass (0.8055), goiter (0.8806), thyroid mass which was a thyroid cancer before operation (0.8951 and 0.9112), and cancer (1.001–2.144). The results show that the TDI is a good indicator of thyroid disease progress and that brute‐force search with MLR analysis is useful for biomedical binary data analysis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".