A semi-supervised deep learning solution to cell registration in video data from calcium imaging studies
Bibliographic record
Abstract
Calcium imaging is a technique that detects the transient changes of calcium ions during an action potential.It enables researchers to record thousands of neurons simultaneously.However, researchers require a method of tracking individual neurons across different recordings for analyses; this problem is known as cell registration.Current approaches rely on image alignment and statistical correlations to perform cell registration.However, these approaches are often slow and require tuning when used on novel datasets.Our end-to-end approach begins with the generation of semi-synthetic data from a few real calcium recordings examples.This allows the option of creating a dataset of any size to train a deep learning model that performs cell registration.We tested three network architectures that are specialized for image patch comparisons to perform cell registration: Siamese, 2-channel, and center-surround networks.These networks were trained once using only semi-synthetic data and were tested on real held-out datasets.The three models were evaluated on a hand-annotated Chronic Social Defeat Stress dataset and were compared to an existing cell registration approach, CellReg.Our best i model, center-surround, achieved an average accuracy of 80.17% while maintaining a precision score of 0.8982.Unlike previous methods, our end-to-end method introduces a new approach of generating synthetic neuronal data that mimics real world data.Next, we show that the cell registration problem can be structured as a binary classification problem to be solved by a deep learning model.Therefore, with our generated synthetic data, we train deep learning models that reliably map neurons across recordings in a scalable manner while requiring minimal parameter tuning.Our contribution includes a versatile approach to cell registration that introduces a novel way of generating semi-synthetic data which is used to train a deep learning model to reliably match cells from raw imaging data.I would first like to thank both Dr. Benjamin Fung and Dr. Tak Pan Wong for allowing me to pursue this research project.Thank you both for being extraordinary, knowledgeable, and patient mentors across my entire Master's journey.Both of you have taught me valuable skills that I will carry in my future academic, and even non-academic, endeavours.I would also like to thank Amanda and Mia-Lin Gromko for their help in obtaining our datasets.Extra thanks to Mia-Lin Gromko
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".