MétaCan
Menu
Back to cohort
Record W3034412992 · doi:10.1093/infdis/jiaa108

Human Immunodeficiency Virus Phylogenetics in the United States—and Elsewhere

2020· letter· en· W3034412992 on OpenAlexaff
Hope R. Lapointe, P. Richard Harrigan

Bibliographic record

VenueThe Journal of Infectious Diseases · 2020
Typeletter
Languageen
FieldImmunology and Microbiology
TopicHIV Research and Treatment
Canadian institutionsAIDS Vancouver
Fundersnot available
KeywordsHuman immunodeficiency virus (HIV)VirologyPhylogeneticsBiologyMedicineGeneticsGene

Abstract

fetched live from OpenAlex

(See the Major Article by Dawson et al on pages 1997–2006.) Phylogenetic tracking of pathogens, sometimes called “molecular epidemiology,” is becoming a critical way of understanding and tackling infectious diseases. If we are to have any success at all in delaying the spread of the 2019 novel coronavirus (2019-nCoV) outbreak, it will be in part thanks to the early and open dissemination of virus genetic information. It is notable that the rapid release of the genome sequence by Wu et al [1] was essential in the development of diagnostic tests. Additional sequences generated worldwide continue to be made available in the public domain, with linked place of origin and patient characteristics. The immediate sharing of methods, software, and even reagents (!) by the UK “Artic Network” soon after the outbreak began allowed researchers to generate and analyze whole-genome sequences for local patient samples. Phylogenetic visualizations for 2019-nCoV were quickly made available on Nextstrain (https://nextstrain.org/ncov) and continue to be updated as the outbreak evolves [2]. This global effort is done in accordance with the World Health Organization (WHO)’s code of conduct on urgent data sharing in public health emergencies [3]. The accomplishments of these groups and others are persuasive evidence of what immediate, global, and public sharing of data and methods can accomplish in a short time [4]. Human immunodeficiency virus (HIV) lies at the other end of the spectrum in terms of data access and sharing. Public sharing of individual sequences or small genotype-phenotype databases of HIV sequences is obviously acceptable and even required to advance scientific progress, but there are a host of issues associated with wholesale analyses of HIV sequences from everyone in a given area. Indeed, the use of phylogenetics in the field of HIV faces many legal, ethical, virological, and social considerations that render an open data-sharing approach inappropriate. It is unfortunate that misuse and/or misinterpretations of HIV phylogenies can result in severe ramifications. This is particularly true in countries where HIV infection is deeply stigmatized, HIV risk behaviors are criminalized, and vulnerable populations are oppressed. A review of the practical considerations has been long overdue. In this issue of the Journal of Infectious Diseases, Dawson et al [5] do an excellent job of covering the ramifications for phylogenetic testing of HIV in the United States. Rather than taking a narrow view, the National Institutes of Health group included experts in virology, molecular epidemiology, public health, bioethics, community engagement, social work, community-based HIV research, and law. This interdisciplinary team presented clear recommendations in 5 areas, from study design to communication of results. An open question is how to apply these recommendations outside the United States. Economic, social, and political climates vary internationally, and legal protections for HIV-infected individuals differ greatly around the world. At a minimum, most jurisdictions likely do not have the United States’ “certificates of confidentiality”, which prohibits researchers from releasing research data to outside parties and are critical in forming the authors’ recommendations. In addition, the United States had a respect for the rule of this type of law. Although research ethics should be broadly similar in many places, community trust and legal protections vary considerably from country to country. Stigma and legal and physical risks of disadvantaged groups infected with or at risk of HIV demand that phylogenetic-related studies should sometimes simply not be undertaken. Researchers should not perform these without considering the legal and ethical risks in their own setting. Another consideration applies to many countries with universal healthcare. Here, these phylogenetic studies may be done in tandem with patients’ ongoing medical care, and therefore they include much or all the longitudinal data from HIV-infected patients. This contrasts in jurisdictions (like the United States) where phylogenetic studies are usually performed on a single diagnostic sample. The greater abundance of sequence data generated longitudinally, per individual, may influence the considerations, and the direct linkage with routine medical care increases the risks associated with driving patients away. There also remain unanswered questions regarding how to share data and methods from phylogenetic studies. There is no WHO mandate for urgent HIV data sharing (unlike with severe acute respiratory syndrome coronavirus 2 [SARS-CoV-2]). Many journals require the underlying data for publication, which is not feasible in large phylogenetic studies. Still, readers of scientific papers need some way to investigate the reproducibility of the results. Even if the underlying data cannot be shared, there is a clear need to publish the software that has been used to analyze the data, and it may be possible to include phylogenetic distances, rather than sequences themselves. There will also likely be technical challenges with regards to how sequence data is analyzed and shared. It is possible that the use of high-throughput “next-generation” sequence data (whereby detailed nucleotide abundances can be assessed) rather than “consensus” Sanger sequencing of HIV can add additional research insight and matching ethical complications. These technical challenges may require the panel described above to reconvene at some point. Regardless, these recommendations provide excellent and comprehensive guidelines for the United States, and a broad set of recommendations is available to provide context internationally [6]. Still, individual countries would be wise to establish similar multidisciplinary panels appropriate for their own legal and cultural context. Researchers have a duty not to proceed with studies that will endanger their participants, and only guidance from such multidisciplinary groups can lower the risk of this happening accidentally. Potential conflicts of interest. All authors: No reported conflicts of interest. All authors have submitted the ICMJE Form for Disclosure of Potential Conflicts of Interest.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.002
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.034
Threshold uncertainty score0.068

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.002
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.001
Science and technology studies0.0020.001
Scholarly communication0.0010.001
Open science0.0000.001
Research integrity0.0040.003
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.014
GPT teacher head0.263
Teacher spread0.248 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations6
Published2020
Admission routes1
Has abstractno

Explore more

Same venueThe Journal of Infectious DiseasesSame topicHIV Research and TreatmentFrench-language works237,207