MétaCan
Menu
Back to cohort
Record W4414868065 · doi:10.2196/68376

Real-World Performance of COVID-19 Antigen Tests: Predictive Modeling and Laboratory-Based Validation

2025· article· en· W4414868065 on OpenAlexvenueno aff
Miguel Bosch, Dawlyn García, Lindsey Rudtner, Nol Salcedo, Raul Colmenares, Sina Hoche, Jose Arocha, Adriana Moreno, Irene Bosch

Bibliographic record

VenueJMIRx Med · 2025
Typearticle
Languageen
FieldMedicine
TopicSARS-CoV-2 detection and testing
Canadian institutionsnot available
Fundersnot available
KeywordsObserver (physics)Probability density functionFunction (biology)PopulationModel validationPattern recognition (psychology)

Abstract

fetched live from OpenAlex

Background: Rapid and safe deployment of lateral-flow antigen tests, coupled with uncompromised quality assurance, is critical for outbreak control and pandemic preparedness, yet real-world performance assessment still lacks laboratory and quantitative approaches that remain uncommon in current regulatory science. The approach proposed here can help standardize and accelerate early phase appraisal of antigen tests in preparation for clinical validation. Objective: The aim of this study is to present a quantitative, laboratory-anchored framework that links image-based test line intensities and the population distribution of naked-eye limits of detection (LoD) to a probabilistic prediction of positive percent agreement (PPA) as a function of viral-load-related variables (eg, quantitative real-time polymerase chain reaction [qRT-PCR] cycle thresholds [Cts]). Using dilution-series calibrations and a Bayesian model, the predicted PPA-vs-Ct curve closely tracks the observed PPA in a real-world self-testing cohort. Methods: The proposed methodology combines: (1) a quantitative evaluation of the test signal response to concentrations of target protein and inactive virus or active virus, (2) a statistical characterization of the LoD using the observer's visual acuity of the test band, and (3) a calibration of a gold-standard method (eg, qRT-PCR cycles) against virus concentration. We elaborate these quantitative methods and unfold a Bayesian-based predictive model to describe the real-world performance of the antigen test, quantified by the probability of positive agreement as a function of viral-load variables like qRT-PCR Cts. Results: We applied the methodology by characterizing each brand of COVID-19 antigen test and estimating its real-world probability of agreement with qRT-PCR. We aligned protein and inactivated-virus standard curves at matched signal intensities and fit a linear calibration linking protein to viral concentrations. Using logistic regression, we modeled the PPA as a continuous function of qRT-PCR Ct, then integrated this curve over a predefined reference Ct distribution to obtain the expected sensitivity. This standardization enables consistent performance comparisons across sites. Conclusions: Modeling performance under real-world conditions requires coupling laboratory evaluation with the population's ability to perceive the test's visual signal. We represent observer capability as a probability density function of the LoD over the signal-intensity domain. Rather than reporting bin-based sensitivity, we summarize performance with the PPA as a continuous function of qRT-PCR Ct. Our framework produces PPA-Ct curves by composing (1) normalized signal-to-concentration models from the laboratory, (2) the observer LoD distribution, and (3) a Ct-to-viral-load calibration. The resulting inferences are inherently context-bound-disease-, assay-, and setup-specific. External validity depends on the particular antigen lateral-flow test, the user population (visual acuity and interpretation), and cross-laboratory qRT-PCR calibration. Comprehensive clinical studies under intended-use conditions are still required before making generalized claims.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.498
Threshold uncertainty score0.368

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.037
GPT teacher head0.346
Teacher spread0.309 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJMIRx MedSame topicSARS-CoV-2 detection and testingFrench-language works237,207