Genetic Deletions in Sputum as Diagnostic Markers for Early Detection of Stage I Non–Small Cell Lung Cancer
Bibliographic record
Abstract
PURPOSE: Analysis of molecular genetic markers in biological fluids has been proposed as a powerful tool for cancer diagnosis. We have characterized in detail the genetic signatures in primary non-small cell lung cancer, which provided potential diagnostic biomarkers for lung cancer. The aim of this study was to determine whether the genetic changes can be used as markers in sputum specimen for the early detection of lung cancer. EXPERIMENTAL DESIGN: Genetic aberrations in the genes HYAL2, FHIT, and SFTPC were evaluated in paired tumors and sputum samples from 38 patients with stage I non-small cell lung cancer and in sputum samples from 36 cancer-free smokers and 28 healthy nonsmokers by using fluorescence in situ hybridization. RESULTS: HYAL2 and FHIT were deleted in 84% and 79% tumors and in 45% and 40% paired sputum, respectively. SFTPC was deleted exclusively in tumor tissues (71%). There was concordance of HYAL2 or FHIT deletions in matched sputum and tumor tissues from lung cancer patients (r = 0.82, P = 0.04; r = 0.84, P = 0.03), suggesting that the genetic changes in sputum might indicate the presence of the same genetic aberrations in lung tumors. Furthermore, abnormal cells were found in 76% sputum by detecting combined HYAL2 and FHIT deletions whereas in 47% sputum by cytology, of the cancer cases, implying that detecting the combination of HYAL2 and FHIT deletions had higher sensitivity than that of sputum cytology for lung cancer diagnosis. In addition, HYAL2 and FHIT deletions in sputum were associated with smoking history of cancer patients and smokers (both P < 0.05). CONCLUSIONS: Tobacco-related HYAL2 and FHIT deletions in sputum may constitute diagnostic markers for early-stage lung cancer.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".