Towards asteroid discovery with deep learning in large datasets
Bibliographic record
Abstract
As astronomical surveys evolve to capture ever-larger volumes of data, innovative computational tools are increasingly critical for extracting meaningful signals from petabyte-scale datasets. Deep learning – machine learning algorithms that involve artificial neural networks – offers one such tool. Here, we present convolutional neural network-based approaches to enhance the discovery and recovery of asteroids in survey data. Our research utilizes two very different datasets: two decades of archival crowded field data from the Microlensing Observations in Astrophysics (MOA) survey and the ongoing Classical and Large A Solar System (CLASSY) trans-Neptunian objects survey.Though designed to detect microlensing events in the Galactic Bulge and Magellanic Clouds, the MOA survey has incidentally observed several thousand asteroids in two decades of high-cadence imaging data. However, the extremely dense star fields posed a significant challenge to effectively identifying moving sources. To address this, we developed a novel approach that leverages the sky motion of asteroids in consecutive exposures to reveal its ‘tracklet’ – the linear motion path that highlights the asteroids’ movement against the static stellar background (Figure 1). These tracklets formed the basis of our labelled datasets of known asteroids, which we used to train several custom-designed convolutional neural networks (CNNs). We then ensembled the predictions from the best performing models to maximize accuracy and generalization, achieving a recall of 97.67%. In addition, we trained the YOLOv4 object detector to precisely localize asteroid tracklets, achieving a mean Average Precision (mAP) of 90.97%. We are now deploying these trained models across the full MOA data archive to identify both known and previously undetected asteroids – transforming the archival data into a powerful tool for asteroid discovery. In parallel, we applied these deep learning techniques to the CLASSY survey, a Canada France Hawaii Telescope (CFHT) Large Program focused on finding distant TNOs. We labelled over ~75,000 composite images from nightly MegaCam observations, creating a training dataset that spans a variety of asteroid populations, including near-Earth objects, main belt asteroids, centaurs, as well as both real and simulated fast-moving TNOs. Our custom CNNs successfully detected tracklets across these diverse sources, and we once again combined the models to enhance predictive performance and minimize false negatives, achieving a recall of 98.15%. The labelling process highlighted the exceptional depth and clarity of the CLASSY observations as well as the effectiveness of the tracklet approach to identify a diverse range of solar system objects. We are now focusing our efforts on recovering centaurs – which are difficult to isolate because of the vast region they inhabit – from the observations. While our work with CLASSY offers a framework for applying deep learning to future surveys like the Legacy Survey of Space and Time (LSST), the MOA archive uniquely demonstrates the untapped potential of archival microlensing datasets. Our results demonstrate the effectiveness of building targeted training datasets and applying model ensembling to maximize discovery. Together, these strategies offer a practical blueprint for integrating artificial intelligence into the data pipelines of future surveys, ensuring that the scientific potential of next-generation observatories is fully realized.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.010 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.004 | 0.004 |
| Research integrity | 0.003 | 0.005 |
| Insufficient payload (model declined to judge) | 0.002 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".