MétaCan
Menu
Back to cohort
Record W3013749053 · doi:10.1101/2020.03.24.005975

Towards CNN Representations for Small Mass Spectrometry Data Classification: From Transfer Learning to Cumulative Learning

2020· preprint· en· W3013749053 on OpenAlexaff
Khawla Seddiki, Philippe Saudemont, Fŕed́eric Precioso, Nina Ogrinc, Maxence Wisztorski, Michel Salzet, Isabelle Fournier, Arnaud Droit

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2020
Typepreprint
Languageen
FieldChemistry
TopicMass Spectrometry Techniques and Applications
Canadian institutionsUniversité Laval
Fundersnot available
KeywordsComputer scienceArtificial intelligencePreprocessorTransfer of learningConvolutional neural networkMachine learningRepresentation (politics)Data pre-processingPattern recognition (psychology)Raw dataFeature learningExternal Data RepresentationData mining

Abstract

fetched live from OpenAlex

Abstract Rapid and accurate clinical diagnosis of pathological conditions remains highly challenging. A very important component of diagnosis tool development is the design of effective classification models with Mass spectrometry (MS) data. Some popular Machine Learning (ML) approaches have been investigated for this purpose but these ML models require time-consuming preprocessing steps such as baseline correction, denoising, and spectrum alignment to remove non-sample-related data artifacts. They also depend on the tedious extraction of handcrafted features, making them unsuitable for rapid analysis. Convolutional Neural Networks (CNNs) have been found to perform well under such circumstances since they can learn efficient representations from raw data without the need for costly preprocessing. However, their effectiveness drastically decreases when the number of available training samples is small, which is a common situation in medical applications. Transfer learning strategies extend an accurate representation model learnt usually on a large dataset containing many categories, to a smaller dataset with far fewer categories. In this study, we first investigate transfer learning on a 1D-CNN we have designed to classify MS data, then we develop a new representation learning method when transfer learning is not powerful enough, as in cases of low-resolution or data heterogeneity. What we propose is to train the same model through several classification tasks over various small datasets in order to accumulate generic knowledge of what MS data are, in the resulting representation. By using rat brain data as the initial training dataset, a representation learning approach can have a classification accuracy exceeding 98% for canine sarcoma cancer cells, human ovarian cancer serums, and pathogenic microorganism biotypes in 1D clinical datasets. We show for the first time the use of cumulative representation learning using datasets generated in different biological contexts, on different organisms, in different mass ranges, with different MS ionization sources, and acquired by different instruments at different resolutions. Our approach thus proposes a promising strategy for improving MS data classification accuracy when only small numbers of samples are available as a prospective cohort. The principles demonstrated in this work could even be beneficial to other domains (astronomy, archaeology…) where training samples are scarce.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.004
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.005
Threshold uncertainty score0.011

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.004
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0000.001
Scholarly communication0.0010.002
Open science0.0010.001
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.069
GPT teacher head0.298
Teacher spread0.229 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations5
Published2020
Admission routes1
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)Same topicMass Spectrometry Techniques and ApplicationsFrench-language works237,207