Molecular diversity in ductal carcinoma <i>in situ</i> (DCIS) and early invasive breast cancer
Bibliographic record
Abstract
Ductal carcinoma in situ (DCIS) is a non-invasive form of breast cancer where cells restricted to the ducts exhibit an atypical phenotype. Some DCIS lesions are believed to rapidly transit to invasive ductal carcinomas (IDCs), while others remain unchanged. Existing classification systems for DCIS fail to identify those lesions that transit to IDC. We studied gene expression patterns of 31 pure DCIS, 36 pure invasive cancers and 42 cases of mixed diagnosis (invasive cancer with an in situ component) using Agilent Whole Human Genome Oligo Microarrays 44k. Six normal breast tissue samples were also included as controls. qRT-PCR was used for validation. All DCIS and invasive samples could be classified into the "intrinsic" molecular subtypes defined for invasive breast cancer. Hierarchical clustering establishes that samples group by intrinsic subtype, and not by diagnosis. We observed heterogeneity in the transcriptomes among DCIS of high histological grade and identified a distinct subgroup containing seven of the 31 DCIS samples with gene expression characteristics more similar to advanced tumours. A set of genes independent of grade, ER-status and HER2-status was identified by logistic regression that univariately classified a sample as belonging to this distinct DCIS subgroup. qRT-PCR of single markers clearly separated this DCIS subgroup from the other DCIS, and contains samples from several histopathological and intrinsic molecular subtypes. The genes that differentiate between these two types of DCIS suggest several processes related to the re-organisation of the microenvironment. This raises interesting possibilities for identification of DCIS lesions both with and without invasive characteristics, which potentially could be used in clinical assessment of a woman's risk of progression, and lead to improved management that would avoid the current over- and under-treatment of patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".