MétaCan
Menu
Retour à la cohorte
Enregistrement W3044508150 · doi:10.1105/tpc.20.00358

Araport Lives: An Updated Framework for Arabidopsis Bioinformatics

2020· letter· en· W3044508150 sur OpenAlexaff
Asher Pasha, Sabarinath Subramaniam, Alan Cleary, Xingguo Chen, Tanya Berardini, Andrew Farmer, Christopher D. Town, Nicholas J. Provart

Notice bibliographique

RevueThe Plant Cell · 2020
Typeletter
Langueen
DomaineAgricultural and Biological Sciences
ThématiquePlant nutrient uptake and metabolism
Établissements canadiensUniversity of Toronto
Organismes subventionnairesAgricultural Research ServiceBiotechnology and Biological Sciences Research CouncilDirectorate for Biological SciencesU.S. Department of AgricultureNational Science Foundation
Mots-clésArabidopsisArabidopsis thalianaFoundation (evidence)Resource (disambiguation)Computational biologyComputer scienceBiologyData scienceBioinformaticsGeographyGeneGenetics

Résumé

récupéré en direct d'OpenAlex

Conceived as a replacement for the anticipated retirement of The Arabidopsis Information Resource (TAIR), the Araport project was funded by the U.S. National Science Foundation (NSF) in 2013 to develop a new, extensible framework for Arabidopsis (Arabidopsis thaliana) bioinformatics that would facilitate data integration through the federation of distributed informatics resources. Recommended as a 5-year award, funds were initially provided for 2 years with a subsequent 6-month extension to allow for the development of a plan for the continuation of funding that was acceptable to NSF. When funds were exhausted in 2016 with the continuation award still awaiting a decision, the Araport site continued in a maintenance mode with minimal input from legacy personnel at each institution and with no data updates. The renewal request was finally declined in December 2018, leaving Araport with an uncertain future. In light of the critical importance of these services to the scientific community, a group of interested researchers (see Appendix) met in March 2019 to discuss options and propose a solution. A working group evolved from those in attendance at that meeting and has since met monthly to solidify and coordinate the execution of these nascent plans. The results have been encouraging and are described here to inform and inspire the larger plant science community. Given the complete absence of external funding, it was agreed that, rather than try to perpetuate the entire Araport ecosystem, efforts should be directed toward maintaining the most attractive and most used features, namely ThaleMine and JBrowse, by transferring them to new ownership for perpetuation. Thus it was agreed that an updated version of ThaleMine would be established at the Bio-Analytic Resource (BAR) for Plant Biology at the University of Toronto under the leadership of Nicholas Provart and an updated version of JBrowse would be established as part of TAIR under the Phoenix Bioinformatics umbrella overseen by Tanya Berardini and Eva Huala. The JBrowse functionality provided by Araport has been successfully moved to TAIR. Araport had been running version 1.11.6 of the software with a set of tracks that included community submissions. Some of the tracks at the legacy location were no longer functioning after the underlying software (ADAMA) connecting them to outside resources lost support. TAIR installed the latest JBrowse version (1.16.6), replicated the tracks that were functional at Araport, and restored access to the nonfunctional tracks. In addition, two sets of newly integrated community-submitted tracks are now visible in this genome browser. One is a set of 41 tracks representing a multipronged gene expression experiment to track the response to various abiotic stresses from Lee and Bailey-Serres (2019). The other is a set of 4 tracks based on Cap Analysis of Gene Expression (CAGE) experiments to determine promoter bidirectionality performed by Thieffry et al. (2019). The CAGE data are visualized using the Stranded View plugin (Hofmeister and Schmitz, 2018), which allows separation of the display of expression values into plus and minus strands in a single track. New community tracks continue to be added, and existing track information is updated as new data become available. ThaleMine at the BAR was completely rebuilt using the latest InterMine software. The legacy ThaleMine version had not been updated since 2016 and was using the InterMine version 1.8.5, which was not forward-compatible with the latest InterMine version (4.2.0). At the time of writing, the most recent versions of publicly available data have been loaded, as listed in Table 1. Data Sources for a New Instance of ThaleMine Data Sources for a New Instance of ThaleMine As with any instance of InterMine, the BAR's version of ThaleMine at https://bar.utoronto.ca/thalemine/ continues to support application programming interface functionalities, in addition to the extensive web-based query options. It is also compatible with the InterMine BlueGenes interface. As part of the unsuccessful renewal proposal, some aspects of the Araport Comparative Genomics functionalities were planned to be addressed with an instance of the Genome Context Viewer (GCV) software developed at the National Center for Genome Resources (NCGR) by Andrew Farmer and Alan Cleary (Cleary and Farmer, 2018). This viewer was originally developed as part of an NSF-funded initiative for federating disparate legume-focused information resources. It provides services to enable the dynamic comparison of multiple genomes on the basis of their shared functional elements (e.g., genes) and provides an intuitive and powerful user interface for exploring similarities and differences among a set of genomic segments with respect to element content and arrangement. A version of the GCV has been installed and is now running from NCGR (https://gcv-arabidopsis.ncgr.org) as the third component of the “second-generation” Araport (see figure). This version of the GCV provides integration of the Arabidopsis Columbia reference genome (TAIR10/Araport11) with genomes from several other data sources, including two sets of newly assembled Arabidopsis genomes of various accessions (colloquially often called ecotypes) from Jiao and Schneeberger (2020) and from the 1001 Genomes project from Detlef Weigel and colleagues (Felix Bemm, Christian Kubica, and Detlef Weigel, personal communication), as well as a number of Brassicaceae genomes from Phytozome and the Brassicaceae Map Alignment Project initiative. The viewer provides convenient links to related resources for genes and genomic regions, thereby facilitating traversal into the other components of the reconfigured Araport project as well as other relevant tools. The gene family classifications utilized by the current instance are based on PANTHER 14.1 (Mi et al., 2013), and links are provided to the trees developed for these families by the PhyloGenes project (phylogenes.org). Screenshot of New Arabidopsis GCV Showing a Region with Two Clusters of Germin-Like Proteins (PTHR31238 Gene Family, Denoted Here as Purple). The central cluster shows extensive copy number variation among annotations from 14 Arabidopsis genomes and the closely related Arabidopsis lyrata genome (labeled araly.scaffold_7), as highlighted by the asterisks along the bottom. Other apparent copy number variations and presence/absence events can easily be observed. To establish continuity between the original Araport and these new functionalities, http://araport.org/ is now hosted at BAR and visitors are then presented with links to the new and maintained versions of ThaleMine, JBrowse, and the GCV. With these new sites operational, the original Araport site hosted at the Texas Advanced Computing Center has been shut down because of security issues related to the legacy versions of the packages used by the original site. We expect that the new Araport in its various component parts will continue to be widely used not just by Arabidopsis researchers but by the wider plant community. In summary, a grassroots effort by committed community members has built upon the resources developed by the Araport project to provide continuity of Araport's most used and useful features. It is gratifying to see that the vision of the 2012 white paper (International Arabidopsis Informatics Consortium, 2012) suggesting a future for Arabidopsis informatics as a community effort accomplished by a federation of independent community members has, in a modest way, come to pass. March 2020 saw 10,376 views of the ThaleMine landing page, showing a wide uptake by the community. That said, this rescue effort is not really a sustainable solution. Data curation and database maintenance are of vital importance and, notwithstanding TAIR's successful subscription model, is something that is worthy of support by national funding agencies for the continued success of plant research in the United States and worldwide. We are especially grateful to the scientists at the Texas Advanced Computing Center, especially Erik Ferlanti, John Fonner, and Matt Vaughn, for continuing to host and maintain Araport long after its “use-by” date. We thank Vivek Krishnakumar for providing insights and advice on the inner workings of Araport during the transition period. Araport was supported by grants from the National Science Foundation (grant DBI-1262414) and the Biotechnology and Biological Sciences Research Council (grant BB/L027151/1). Development of the GCV was supported by USDA-ARS project funding for the Legume Information System and the National Science Foundation (grant IOS-1444806 to A.F.). The J. Craig Venter Institute workshop that launched the Araport recovery effort was supported by the U.S. National Science Foundation (MCB Award 1062348, made to the U.S. members of the International Arabidopsis Informatics Consortium Steering Committee). The BAR is supported through a grant to N.P. from the National Sciences and Engineering Research Council of Canada and from Genome Canada/Ontario Genomics. TAIR is managed by the nonprofit Phoenix Bioinformatics Corporation and is supported through institutional, lab, and personal subscriptions. List of participants, “Future of Araport” meeting held at the J. Craig Venter Institute (JCVI) in Rockville, Maryland, March 25 and 26, 2019. Tanya Berardini, Phoenix Bioinformatics Agnes Chan, JCVI Yongwook Choi, JCVI Andrew Farmer, NCGR Erik Ferlanti, Texas Advanced Computing Center Eva Huala, Phoenix Bioinformatics Vivek Krishnakumar, formerly JCVI Sean May, University of Nottingham Asher Pasha, University of Toronto Nicholas Provart, University of Toronto David Somers, Ohio State University Chris Town, JCVI Eve Wurtele, Iowa State University Remotely via BlueGenes: Sam Hokin, NCGR Eric Lyons, University of Arizona Todd Michael, JCVI

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,017
score de la tête « metaresearch » (Gemma)0,048
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Méthodes · Signal consensuel: Méthodes
Score de désaccord entre enseignants0,020
Score d'incertitude au seuil0,092

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0170,048
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0020,002
Études des sciences et des technologies0,0020,002
Communication savante0,0050,014
Science ouverte0,0050,005
Intégrité de la recherche0,0070,015
Charge utile insuffisante (le modèle a refusé de juger)0,0200,036

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,038
Tête enseignante GPT0,215
Écart entre enseignants0,177 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations49
Publié2020
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueThe Plant CellMême sujetPlant nutrient uptake and metabolismTravaux en français237 207