THE POPULATION OF GALAXY–GALAXY STRONG LENSES IN FORTHCOMING OPTICAL IMAGING SURVEYS
Bibliographic record
Abstract
Ongoing and future imaging surveys represent significant improvements in depth, area, and seeing compared to current data sets. These improvements offer the opportunity to discover up to three orders of magnitude more galaxy–galaxy strong lenses than are currently known. In this work we forecast the number of lenses that will be discoverable in forthcoming surveys and simulate their properties. We generate a population of statistically realistic strong lenses and simulate observations of this population for the Dark Energy Survey (DES), the Large Synoptic Survey Telescope (LSST), and Euclid surveys. We verify our model against the galaxy-scale lens search of the Canada–France–Hawaii Telescope Legacy Survey, predicting 250 discoverable lenses compared to 220 found by Gavazzi et al. The predicted Einstein radius distribution is also remarkably similar to that found by Sonnenfeld et al. For future surveys we find that, assuming Poisson limited lens galaxy subtraction, searches of the DES, LSST, and Euclid data sets should discover 2400, 120000, and 170000 galaxy–galaxy strong lenses, respectively. Finders using blue-minus-red ( ) difference imaging for lens subtraction can discover 1300 and 62000 lenses in DES and LSST. The uncertainties on the model are dominated by the high-redshift source population, which typically gives fractional errors on the discoverable lens number at the level of tens of percent. We find that doubling the signal-to-noise ratio required for a lens to be detectable approximately halves the number of detectable lenses in each survey, indicating the importance of understanding the selection function and the sensitivity of future lens finders in interpreting strong lens statistics. We make our population forecasting and simulated observation codes publicly available so that the selection function of strong lens finders can easily be calibrated.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.010 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.003 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".