Phylogeny-based selection of representative variants with Navargator: Proof of principle using humoral cross-reactivity data from two immunization studies
Bibliographic record
Abstract
Abstract Background The accessibility to the immune system of bacterial surface proteins makes them attractive targets for subunit vaccines. However, this same property also means they tend to exhibit high sequence variability. Achieving broad cross-protection usually necessitates that antigens from multiple isolates are included, but the choice of sequence variant is a non-trivial problem. Visual inspection of phylogenetic trees is the norm, but this is subjective and can be greatly influenced by the choice of viewing software. This has real-world implications, as groups have shown that the selection of non-optimal antigens likely led to lower cross-protection in the commercially available vaccines against Neisseria meningitidis serogroup B. Aim / Methods To address this problem, we have developed Navargator, bioinformatics software that takes a phylogenetic tree as input and identifies the variants that are the most similar to the greatest number of other sequences. The underlying premise is that cross-reactivity will be correlated with phylogenetic distances extracted from the tree; this was validated by several rodent immunization studies with the proteins transferrin-binding protein B and factor H binding protein from N. meningitidis and N. gonorrhoeae , measuring antibody-based cross-reactivity between an antigen panel using a custom high-throughput ELISA. Results Navargator has been made freely available both as an online tool and as source code for local installation. We implemented several different clustering methods, with exact algorithms for smaller datasets, and heuristics suitable for large trees of thousands of sequences. Our immunization studies have shown that this approach is sound, and that cross-reactivity is predicted well by phylogenetic distances in a sigmoidal manner. Conclusions The complexity of vaccine development rises sharply with each additional antigen included, so using the minimal number required is an important consideration. Navargator attempts to facilitate this in a systematic and generalizable manner. The user can run the analysis by selecting their desired number of representatives, or they can provide any form of cross-reactivity data and have the program identify a minimum reactivity threshold via correlation with the phylogenetic tree. The program will then identify the smallest number of representatives required to satisfy this threshold.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".