Unsupervised Idealization of Nano-Electronic Sensors Recordings with Concept Drifts: An Information Theory Approach for Non-Stationary Single-Molecule Data Analysis
Bibliographic record
Abstract
Single-molecule nanocircuits based on field-effect transistors (smFETs) have known a rapid development and promising results for the functional detection of biomolecular structures and dynamics at the single-molecule scale [1]. In fact, thanks to the size compatibility between the target analyte and the transducer, most often a carbon nanotube, this label-free and amplification-free single-molecule sensing technique allows real-time monitoring of the rapid transitions between different biochemical conformational or interaction states, such as hybridization or folding in nucleic acids. The stability of smFET signals also enables long acquisition periods of such single-molecule interactions with high throughput, allowing to reveal individual reaction pathways usually hidden in average measurements of ensemble methods, and to record time trajectories containing rare or short-lived biomolecular events. Detecting and modeling the kinetics and thermodynamics of biochemical events from smFET recordings requires robust data analysis tools that can idealize these signals into discrete state trajectories, corresponding to successive biochemical states and transitions between them. However, most of the available single-molecule data analysis techniques have been developed for fluorescence and force-based single-molecule experiments, which differ significantly from FET-based experiments, and are thus difficult to adapt to smFET signals. Analysis of smFET time series requires to handle the following set of challenging signal specificities: 1) the stochastic nature of the biomolecular system, 2) the possible non-stationarities in the molecular dynamics of the reaction system, such as changes between transient and steady-state conformations, 3) the multi-source composition of the sensor response, aggregating all contributions from the biochemical system with those from the environment and sensor components in a single input, 4) the mixed noises (AWGN, flicker, and impulse) characteristic of FET devices, 5) the slow baseline drift observed in long acquisitions, and 6) the sizable amount of data generated by such recordings. Here, we propose a new approach for smFET data idealization based on information theory and machine learning for signal processing. We present computational methods designed to achieve automatic detection of molecular events without prior knowledge on the data generating process or signal pre-filtering, which are especially tailored for large, drifting and non-stationary time series that are typical of smFET experiments. First, we address the problem of compensating slow baseline drift, due to sensor degradation or variations in environmental parameters. Such drift can introduce systematic errors in the conversion of signal into discrete states. We developed a 3-step adaptive blind source separation algorithm to decompose multicomponent signals into their embedded layers, thus allowing to separate the contribution of the drift from those of signal and noise [2]. To do so, our algorithm is composed of the following steps: 1) iterative multiscale signal compression based on a minimum description length objective function, 2) unsupervised dynamic drift node positioning, and 3) adaptive piecewise cubic Hermite interpolation. Second, we address the task of trace idealization as a piecewise regression fit of the recorded signal, in order to extract the different biochemical states and rates and thus learn kinetic parameters. To this end, we propose a 4-step hybrid algorithm called cc-EM (compressed clustered Expectation-Maximization): 1) A model-free and unsupervised compression stage is applied to transform the raw signal into a piecewise trajectory. The compression is based on a minimum description length objective function that encodes the entropy of the stochastic sources to achieve minimum redundancies for maximum relevance [3], [4]. 2) A clustering step, based on k-medoids with a swapping cost function, is then applied on compression patterns to gather, without supervision, similar sub-states into common parent states. In traces with concept drift, such soft-clustering enables to reach the lower bound of entropy, corresponding to the optimal compression ratio and learning rate, which drastically decreases the false positive event detection rate and the risk of overfitting. 3) An expectation-maximization (EM) refinement is added to infer possible missed states and to correct the location of transitions to obtain a more accurate idealized trace. 4) Finally, a model selection algorithm screens the compressed-clustered states space domain to select the best fitting model. The proposed algorithms were tested on simulated smFET signals covering a wide range of parameters in noise, baseline and concept drifts. We show that the proposed approach enables to dissociate the discrete source signals emitted by the hidden molecular states from the continuous parasitic signal corresponding to the baseline drift, without any supervision nor prior knowledge on the sensor features or the underlying kinetics of the sensed phenomenon. We also demonstrate that our method is able to recover hidden molecular kinetic parameters, even under large noises and concept drifts. We report improved performances than model-based idealization in precision, recall, computational time, and than Bayesian non parametric approaches in terms of robustness to non-stationary and noisy signals. [1] C. Gu, C. Jia, and X. Guo, “Single-Molecule Electrical Detection with Real-Time Label-Free Capability and Ultrasensitivity,” Small Methods, vol. 1, no. 5, p. 1700071, 2017. [2] M. OUQAMRA and D. BOUILLY, “Unsupervised Drift Compensation Based on Information Theory for Single-Molecule Sensors,” , IEEE SigPort, 2019. [Online]. Available: http://sigport.org/4858. [3] Peter Grünwald, “Introducing the Minimum Description Length Principle,” in Advances in Minimum Description Length: Theory and Applications, chapter 1, pp. 3–22.MIT Press, 2005 [4] D. A. Huffman, “A Method for the Construction of Minimum-Redundancy Codes,” Proceedings of the IRE, vol. 40, pp. 1098–1101, Sept 1952.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".