MétaCan
Menu
Retour à la cohorte
Enregistrement W3024527028 · doi:10.1149/ma2020-01302259mtgabs

Unsupervised Idealization of Nano-Electronic Sensors Recordings with Concept Drifts: An Information Theory Approach for Non-Stationary Single-Molecule Data Analysis

2020· article· en· W3024527028 sur OpenAlexaff
Mohamed Ouqamra, Delphine Bouilly

Notice bibliographique

RevueECS Meeting Abstracts · 2020
Typearticle
Langueen
DomaineChemistry
ThématiqueAnalytical Chemistry and Chromatography
Établissements canadiensUniversité de Montréal
Organismes subventionnairesnon disponible
Mots-clésBiological systemMoleculeChemistryNanotechnologyComputer scienceMaterials science

Résumé

récupéré en direct d'OpenAlex

Single-molecule nanocircuits based on field-effect transistors (smFETs) have known a rapid development and promising results for the functional detection of biomolecular structures and dynamics at the single-molecule scale [1]. In fact, thanks to the size compatibility between the target analyte and the transducer, most often a carbon nanotube, this label-free and amplification-free single-molecule sensing technique allows real-time monitoring of the rapid transitions between different biochemical conformational or interaction states, such as hybridization or folding in nucleic acids. The stability of smFET signals also enables long acquisition periods of such single-molecule interactions with high throughput, allowing to reveal individual reaction pathways usually hidden in average measurements of ensemble methods, and to record time trajectories containing rare or short-lived biomolecular events. Detecting and modeling the kinetics and thermodynamics of biochemical events from smFET recordings requires robust data analysis tools that can idealize these signals into discrete state trajectories, corresponding to successive biochemical states and transitions between them. However, most of the available single-molecule data analysis techniques have been developed for fluorescence and force-based single-molecule experiments, which differ significantly from FET-based experiments, and are thus difficult to adapt to smFET signals. Analysis of smFET time series requires to handle the following set of challenging signal specificities: 1) the stochastic nature of the biomolecular system, 2) the possible non-stationarities in the molecular dynamics of the reaction system, such as changes between transient and steady-state conformations, 3) the multi-source composition of the sensor response, aggregating all contributions from the biochemical system with those from the environment and sensor components in a single input, 4) the mixed noises (AWGN, flicker, and impulse) characteristic of FET devices, 5) the slow baseline drift observed in long acquisitions, and 6) the sizable amount of data generated by such recordings. Here, we propose a new approach for smFET data idealization based on information theory and machine learning for signal processing. We present computational methods designed to achieve automatic detection of molecular events without prior knowledge on the data generating process or signal pre-filtering, which are especially tailored for large, drifting and non-stationary time series that are typical of smFET experiments. First, we address the problem of compensating slow baseline drift, due to sensor degradation or variations in environmental parameters. Such drift can introduce systematic errors in the conversion of signal into discrete states. We developed a 3-step adaptive blind source separation algorithm to decompose multicomponent signals into their embedded layers, thus allowing to separate the contribution of the drift from those of signal and noise [2]. To do so, our algorithm is composed of the following steps: 1) iterative multiscale signal compression based on a minimum description length objective function, 2) unsupervised dynamic drift node positioning, and 3) adaptive piecewise cubic Hermite interpolation. Second, we address the task of trace idealization as a piecewise regression fit of the recorded signal, in order to extract the different biochemical states and rates and thus learn kinetic parameters. To this end, we propose a 4-step hybrid algorithm called cc-EM (compressed clustered Expectation-Maximization): 1) A model-free and unsupervised compression stage is applied to transform the raw signal into a piecewise trajectory. The compression is based on a minimum description length objective function that encodes the entropy of the stochastic sources to achieve minimum redundancies for maximum relevance [3], [4]. 2) A clustering step, based on k-medoids with a swapping cost function, is then applied on compression patterns to gather, without supervision, similar sub-states into common parent states. In traces with concept drift, such soft-clustering enables to reach the lower bound of entropy, corresponding to the optimal compression ratio and learning rate, which drastically decreases the false positive event detection rate and the risk of overfitting. 3) An expectation-maximization (EM) refinement is added to infer possible missed states and to correct the location of transitions to obtain a more accurate idealized trace. 4) Finally, a model selection algorithm screens the compressed-clustered states space domain to select the best fitting model. The proposed algorithms were tested on simulated smFET signals covering a wide range of parameters in noise, baseline and concept drifts. We show that the proposed approach enables to dissociate the discrete source signals emitted by the hidden molecular states from the continuous parasitic signal corresponding to the baseline drift, without any supervision nor prior knowledge on the sensor features or the underlying kinetics of the sensed phenomenon. We also demonstrate that our method is able to recover hidden molecular kinetic parameters, even under large noises and concept drifts. We report improved performances than model-based idealization in precision, recall, computational time, and than Bayesian non parametric approaches in terms of robustness to non-stationary and noisy signals. [1] C. Gu, C. Jia, and X. Guo, “Single-Molecule Electrical Detection with Real-Time Label-Free Capability and Ultrasensitivity,” Small Methods, vol. 1, no. 5, p. 1700071, 2017. [2] M. OUQAMRA and D. BOUILLY, “Unsupervised Drift Compensation Based on Information Theory for Single-Molecule Sensors,” , IEEE SigPort, 2019. [Online]. Available: http://sigport.org/4858. [3] Peter Grünwald, “Introducing the Minimum Description Length Principle,” in Advances in Minimum Description Length: Theory and Applications, chapter 1, pp. 3–22.MIT Press, 2005 [4] D. A. Huffman, “A Method for the Construction of Minimum-Redundancy Codes,” Proceedings of the IRE, vol. 40, pp. 1098–1101, Sept 1952.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,002
score de la tête « metaresearch » (Gemma)0,005
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Simulation ou modélisation · Signal consensuel: Simulation ou modélisation
GenreSignal candidat: Empirique · Signal consensuel: aucune
Score de désaccord entre enseignants0,002
Score d'incertitude au seuil0,010

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0020,005
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0010,001
Études des sciences et des technologies0,0000,002
Communication savante0,0010,002
Science ouverte0,0020,001
Intégrité de la recherche0,0010,002
Charge utile insuffisante (le modèle a refusé de juger)0,0010,001

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,018
Tête enseignante GPT0,235
Écart entre enseignants0,217 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSimulation ou modélisation
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2020
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueECS Meeting AbstractsMême sujetAnalytical Chemistry and ChromatographyTravaux en français237 207