AutoSmarTrace: Automated Chain Tracing and Flexibility Analysis of Biological Filaments
Bibliographic record
Abstract
Abstract Single-molecule imaging is widely used to determine statistical distributions of molecular properties. One such characteristic is the bending flexibility of biological filaments, which can be parameterized via the persistence length. Quantitative extraction of persistence length from images of individual filaments requires both the ability to trace the backbone of the chains in the images and sufficient chain statistics to accurately assess the persistence length. Chain tracing can be a tedious task, performed manually or using algorithms that require user input and/or supervision. Such interventions have the potential to introduce user-dependent bias into the chain selection and tracing. Here, we introduce a fully automated algorithm for chain tracing and determination of persistence lengths. Dubbed “AutoSmarTrace”, the algorithm is built off a neural network, trained via machine learning to identify filaments within images recorded using atomic force microscopy (AFM). We validate the performance of AutoSmarTrace on simulated images with widely varying levels of noise, demonstrating its ability to return persistence lengths in agreement with the ground truth. Persistence lengths returned from analysis of experimental images of collagen and DNA agree with previous values obtained from these images with different chain-tracing approaches. While trained on AFM-like images, the algorithm also shows promise to identify chains in other single-molecule imaging approaches, such as rotary shadowing electron microscopy and fluorescence imaging. Statement of Significance Machine learning presents powerful capabilities to the analysis of large data sets. Here, we apply this approach to the determination of bending flexibility – described through persistence length – from single-molecule images of biological filaments. We present AutoSmarTrace, a tool for automated tracing and analysis of chain flexibility. Built on a neural network trained via machine learning, we show that AutoSmarTrace can determine persistence lengths from AFM images of a variety of biological macromolecules including collagen and DNA. While trained on AFM-like images, the algorithm works well to identify filaments in other types of images. This technique can free researchers from tedious tracing of chains in images, removing user bias and standardizing determination of chain mechanical parameters from single-molecule conformational images.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".