MétaCan
Menu
Back to cohort
Record W6929819926 · doi:10.5061/dryad.cjsxksn3p

Data from: Identification of intraductal carcinoma of the prostate on tissue specimens using Raman micro-spectroscopy: A diagnostic accuracy case-control study with multicohort validation

2020· dataset· en· W6929819926 on OpenAlexaffabout

Bibliographic record

VenuePolyPublie (École Polytechnique de Montréal) · 2020
Typedataset
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicBacterial Genetics and Biotechnology
Canadian institutionsUniversity Health NetworkCentre Hospitalier de l’Université de MontréalPolytechnique MontréalCentre hospitalier universitaire de QuébecUniversité de Montréal
Fundersnot available
KeywordsProstateProstate cancerCancerCarcinomaHistopathologyRaman spectroscopy

Abstract

fetched live from OpenAlex

Background Prostate cancer (PC) is the most frequently diagnosed cancer in North American men. Pathologists are in critical need of accurate biomarkers to characterize PC, particularly to confirm the presence of intraductal carcinoma of the prostate (IDC-P), an aggressive histopathological variant for which therapeutic options are now available. Our aim was to identify IDC-P with Raman micro-spectroscopy and machine learning technology following a protocol suitable for routine clinical histopathology laboratories. Methods and findings We used Raman micro-spectroscopy to differentiate IDC-P from PC, as well as PC and IDC-P from benign tissue on formalin-fixed paraffin-embedded first-line radical prostatectomy specimens (embedded in tissue microarrays, TMAs) from 483 patients treated in three Canadian institutions between 1993 and 2013. The main measures were the presence or absence of IDC-P and of PC, regardless of the clinical outcomes. Most of the 483 patients were pT2 stage (44–69%), and pT3a (22–49%) was more frequent than pT3b (9–12%). After approval of the construction of the TMAs by local ethics review board, the diagnostic accuracy study was approved by the Centre hospitalier de l’Université de Montréal (CHUM) ethics review board. Briefly, two consecutive sections of each TMA block were cut. The first section was transferred onto a glass slide to perform immunohistochemistry with H&E counterstaining for cell identification. The second section was placed on an aluminum slide, dewaxed, and then used to acquire an average of 7 Raman spectra per specimen (between 4 and 24 Raman spectra, 4 acquisitions / TMA core). Raman spectra of each cell type were then analyzed to retrieve tissue-specific molecular information and to generate classification models using machine learning technology. Models were trained and cross-validated using data from one institution. Accuracy, sensitivity and specificity were respectively of 87 ± 5%, 86 ± 6% and 89 ± 8% to differentiate PC from benign tissue, and of 95 ± 2%, 96 ± 4% and 94 ± 2% respectively to differentiate IDC-P from PC. The trained models were then tested on data from two independent institutions, reaching accuracies, sensitivities and specificities of 84 and 86%, 84 and 87%, and 81 and 82%, respectively to diagnose PC, and of 85 and 91%, 85 and 88%, and 86 and 93% respectively for the identification of IDC-P. IDC-P could further be differentiated from high-grade prostatic intraepithelial neoplasia (HGPIN), a pre-malignant intraductal proliferation which can be mistaken as IDC-P, with accuracies, sensitivities and specificities >95% in both training and testing cohorts. As we used stringent criteria to diagnose IDC-P, the main limitation of our study is the exclusion of borderline, difficult to classify lesions from our datasets. Conclusions In this study, we developed classification models for the analysis of Raman micro-spectroscopy data to differentiate IDC-P, PC and benign tissue, including HGPIN. Raman micro-spectroscopy could be a next-generation histopathological technique used to reinforce the identification of high-risk PC patients and lead to more precise diagnosis of IDC-P.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.009
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Dataset · Consensus signal: none
Teacher disagreement score0.157
Threshold uncertainty score0.312

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.009
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.002
Science and technology studies0.0020.001
Scholarly communication0.0010.000
Open science0.0010.001
Research integrity0.0010.000
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.015
GPT teacher head0.262
Teacher spread0.247 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2020
Admission routes2
Has abstractyes

Explore more

Same venuePolyPublie (École Polytechnique de Montréal)Same topicBacterial Genetics and BiotechnologyFrench-language works237,207