Whole genome sequencing for epidemiological studies of tuberculosis: a systematic review of reporting practices and factors associated with reporting quality of STROME-ID
Bibliographic record
Abstract
Background: Whole-genome sequencing (WGS) has the potential to improve the understanding of tuberculosis (TB) epidemiology.However, standardized reporting is necessary to enhance the reproducibility and interpretation of WGS results in genomic epidemiology studies of TB to better inform public health decision-making.In 2014, guidelines called STROME-ID were published to provide recommendations for reporting in genomic epidemiology studies.Reporting practices before and after its publication were compared, and the correlation between STROME-ID reporting quality and study-level characteristics were also explored. Methods: This study is registered on PROSPERO (CRD42017064395). MEDLINE, EmbaseClassic and Embase were searched on May 3, 2017 (updated April 23, 2019).976 titles and abstracts were screened, with 114 full-texts eligible for inclusion.The proportion of STROME-ID criteria reported was tabulated for each article, and differences in means were compared before and after STROME-ID's publication date using a t-test.A 6-month lag period after STROME-ID was included to account for articles in-press; sensitivity analyses were also performed.Quasi-Poisson and tobit regression were used to assess whether h-index (HI), journal impact factor (IF), sample size (SS), and geographic region of the senior author's primary affiliation were correlated with the count and proportion of STROME-ID criteria met.Results: The proportion of applicable criteria met in included articles ranged from 16.3-75.0%(mean 49.9%, ± 11.88%), with no difference between mean proportions of criteria comparing before and after guideline publication.HI was not included in the adjusted regression analysis.Only SS was significantly associated with a greater proportion of STROME-ID criteria met.Conclusion: Reporting quality in genomic epidemiology studies of tuberculosis is variable, despite publication of STROME-ID guidelines.Future studies should investigate factors affecting adherence to these guidelines to improve the value and utility of evidence.Journal endorsement may be needed to support this.Résumé Context: Le séquençage du génome entier (SGE) possède le potentiel d'améliorer la compréhension de l'épidémiologie de la tuberculose (TB).Cependant, des rapports standardisés sont nécessaires pour améliorer la reproductibilité et l'interprétation des résultats du SGE dans les études épidémiologie génomiques de la tuberculose.Cela peut mieux éclairer la prise de décision en matière de santé publique.En 2014, des lignes directrices appelées STROME-ID ont été publiées pour fournir des recommandations sur la notification dans les études d'épidémiologie génomique.Les pratiques de déclaration avant et après sa publication ont été comparées, et la corrélation entre la qualité de la déclaration STROME-ID et les caractéristiques au niveau de l'étude a également été explorée.Methodes: Cette étude est enregistrée sur PROSPERO (CRD42017064395).Les bases de données MEDLINE, Embase Classic, et Embase ont été cherchées le 3 mai, 2017 (mise à jour le 23 avril, 2019).976 titres et résumés ont été évalués, dont 114 textes-complètes étaient inclus.La proportion de critères STROME-ID rapporter s'est tabulée pour chaque publication, et les différences entre les moyennes étaient comparées avant et après la publication de STROME-ID, à l'aide d'un test t.Une période de 6 mois après la publication de STROME-ID était inclus pour inclure les articles sous presse.En plus, des analyses de sensibilité ont été réalisés.Les régressions quasi-poisson et tobit ont été utilisés pour déterminer si l'indice-h (IH), le facteur d'impact du journal (FI), la taille de l'échantillon (TE), et la région géographique de l'affiliation de l'auteur principal ont été corrélés avec le nombre et proportion de critères STROME-ID effectués.Resultats: La proportion des critères applicables effectués dans les articles était 16,3-75,0% (taux moyenne 49,9% ± 11,88%), sans différence entre les taux moyennes de la proportion des critères comparées avant et après la publication des lignes directrices STROME-ID.IH n'était pas inclus dans l'analyse de régression ajustée.Seulement TE était associée significativement avec une plus grande proportion des critères STROME-ID effectués.Conclusion: La qualité de rapport des études épidémiologiques génomiques est variée, malgré la publication des lignes directrices STROME-ID.Les études dans le futur devraient investiguer les facteurs responsables pour maintenir l'adhérence des lignes directrices STROME-ID afin d'augmenter la qualité et l'utilité des données.L'appui des journaux pourrait être nécessaire pour augmenter l'adhérence des lignes directrices STROME-ID. Preface and contribution of authorsThis thesis contains 7 chapters.Chapter 1 provides a rationale for the research and outlines the main objectives of the thesis.Chapter 2 is a literature review summarizing the epidemiology of tuberculosis (TB), whole-genome sequencing (WGS) and its epidemiological applications, and the current reporting issues in TB epidemiology studies using WGS.Results are elaborated upon in Chapter 3. Chapter 4 explains the study methodology.The results of the thesis are presented in the form of a manuscript in Chapter 5, which will be submitted to The Lancet Microbe.Chapter 6 reports additional findings to those in the manuscript, which their interpretation is discussed in Chapter 7. The master reference list is provided at the end of the thesis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.117 | 0.329 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.007 | 0.009 |
| Bibliometrics | 0.018 | 0.023 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.004 | 0.005 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".