Bibliographic record
Abstract
Cancer is a complex disease affecting many genes and metabolites. While large amounts of data have been generated using genomic sequencing and mass spectrometry, it can be hard to find relevant features in large datasets and to determine their functions in cancer development. This thesis describes methods for finding annotated metabolites in liquid chromatography-mass spectrometry (LC-MS) datasets and new mechanisms in acute myeloid leukemia (AML). To find relevant features in LC-MS data, three algorithms were assessed for ranking annotated peaks. The new "binned" method had comparable or better accuracy than ranking the peaks by intensity or randomly. Also, the filtered dataset with the top ranked peaks significantly improved the robustness and accuracy of a cancerous vs. non-cancerous sample classifier. Hence, the new ranking method may help find new metabolites in cancer. One metabolite, D-2-hydroxyglutarate (D2HG), is produced by mutated isocitrate dehydrogenase 1 (IDH1) in AML. D2HG inhibits epigenetic processes but it's unknown how it promotes myeloproliferative disease and AML. By comparing IDH1 mutated vs. unmutated samples, this study shows that the DNA damage and repair (DDR) gene ataxia telangiectasia mutated (ATM) was downregulated in IDH1 mutated AML and that the ATM downregulation and genomic instability can contribute to cancer development. Finally, to establish whether defects in DDR can be targeted for treatment, this research uncovered AML patients with low expression of O-6-methylguanine-DNA methyltransferase (MGMT), an enzyme associated with resistance to the alkylating agent temozolomide (TMZ) in solid tumours. The MGMT downregulation was closely related to epigenetic dysregulation. Furthermore, low MGMT was associated with TMZ sensitivity in hematopoietic cell lines, indicating TMZ as a potential targeted treatment for AML with epigenetic mutations. In summary, this thesis highlights the relevant causes, mechanisms, and treatments for cancer. After accumulating large amounts of data, the described data-driven approaches should be valuable for treating this difficult disease. 癌症是影響許多基因和代謝物的複雜性疾病,雖然應用基因定序和質譜測定已經產生大量的數據,但是從這些數據中找到與癌症發展有密切相關的標的物與其作用機制則是相當困難。本論文敘述從液相層析與質譜 (liquid chromatography-mass spectrometry, LC-MS) 的數據庫中尋找出與癌症相關的代謝物的方法,以及新發現的急性髓性白血病 (acute myeloid leukemia, AML) 的作用機制。 本研究測試三種計算法來尋查LC-MS數據庫中的數據特徵,並評估它們的識別精度。本論文所呈現的新“分箱”(binned) 方法,與“強度”(intensity) 或“隨機”(random) 的方法相比較,具有相同,甚至更好的準確度。此外,過濾後的數據也明顯的改善癌症與非癌症樣本分類器的穩定性和準確性。因此,這種新的排序方法有助於在癌症研究中找到新的代謝物。 其中有一種代謝物是由AML中突變的異檸檬酸脫氫酶1 (isocitrate dehydrogenase 1, IDH1) 所產生的D-2-羥基戊二酸 (D-2-hydroxyglutarate, D2HG)。 D2HG會抑制表觀遺傳的過程,但是D2HG促進骨髓增生性疾病和AML的作用機制仍然未明,所以研究比較IDH1發生突變與沒有發生突變的AML患者,發現DNA損傷和修復 (DNA damage and repair, DDR) 基因ataxia telangiectasia mutated (ATM) 在IDH1突變的AML患者出現下調的現象,顯示癌症的發展與基因組的不穩定性有關。 為了確定DDR的缺陷是否可以作為治療AML的標的,本研究發現部分的AML患者的O-6-甲基鳥嘌呤-DNA甲基轉移酶 (O-6-methylguanine-DNA methyltransferase, MGMT) 減少,這是一種與實體瘤對於烷化劑 (alkylating agent) 替莫唑胺 (temozolomide, TMZ) 產生耐藥性相關的酶。 MGMT的減少與表觀遺傳的失調有關,並影響造血細胞對於TMZ的敏感性。因此,TMZ可能成為表觀遺傳突變的AML患者的標靶治療藥物。 總結本論文的重點介紹了癌症的病因,作用機制和治療方法。當累積大量的數據後,本研究敘述的數據驅動的新方法對於治療這種難治性疾病應有重要的價值。
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.026 | 0.059 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.006 |
| Bibliometrics | 0.005 | 0.004 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.005 | 0.005 |
| Research integrity | 0.003 | 0.006 |
| Insufficient payload (model declined to judge) | 0.010 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".