MétaCan
Menu
Back to cohort

EukProt: a database of genome-scale predicted proteins across the diversity of eukaryotes

2022· dataset· en· W6901900520 on OpenAlexaff

Bibliographic record

VenueFigshare · 2022
Typedataset
Languageen
Field
Topic
Canadian institutionsUniversity of Alberta
Fundersnot available
KeywordsPhylogenomicsIdentifierPhylogenetic treeSequence databaseSet (abstract data type)Web serverGenomicsTree (set theory)Biological databaseGenome

Abstract

fetched live from OpenAlex

<strong>Version 3</strong> (22 November, 2021)<br> <br> See https://doi.org/10.24072/pcjournal.173 for a detailed description of the database. See http://evocellbio.com/eukprot/ for a BLAST database, interactive plots of BUSCO scores and ‘The Comparative Set’ (TCS): A selected subset of EukProt for comparative genomics investigations. Protein sequence FASTA files of the TCS are available at https://doi.org/10.6084/m9.figshare.21586065. See https://github.com/beaplab/EukProt for utility scripts, annotations, and all the files necessary to build the tree in Figures 1 and 3 (from the DOI above).<br> <br> Scroll to the end of this page for changes since version 2.<br> <br> Are we missing anything? Please let us know! <br> <strong>EukProt is a database of published and publicly available predicted protein sets selected to represent the breadth of eukaryotic diversity, currently including 993 species from all major supergroups as well as orphan taxa. The goal of the database is to provide a single, convenient resource for gene-based research across the spectrum of eukaryotic life, such as phylogenomics and gene family evolution. Each species is placed within the UniEuk taxonomic framework in order to facilitate downstream analyses, and each data set is associated with a unique, persistent identifier to facilitate comparison and replication among analyses. The database is regularly updated, and all versions will be permanently stored and made available via FigShare. The current version has a number of updates, notably ‘The Comparative Set’ (TCS), a reduced taxonomic set with high estimated completeness while maintaining a substantial phylogenetic breadth, which comprises 196 predicted proteomes. A BLAST web server and graphical displays of data set completeness are available</strong> <strong>at http://evocellbio.com/eukprot/. We invite the community to provide suggestions for new data sets and new annotation features to be included in subsequent versions, with the goal of building a collaborative resource that will promote research to understand eukaryotic diversity and diversification.</strong><br> <br> This release contains 5 files:<br> <br> <strong>EukProt_proteins.v03.2021_11_22.tgz</strong>: 993 protein data sets, for species with either a genome (375) or single-cell genome (56), a transcriptome (498), a single-cell transcriptome (47), or an EST assembly (17).<br> <br> <strong>EukProt_genome_annotations.v03.2021_11_22.tgz</strong>: gene annotations, in GFF format, as produced by EukMetaSanity (https://github.com/cjneely10/EukMetaSanity) for 40 genomes lacking publicly available protein annotations. The proteins predicted from these annotations are included in the proteins file.<br> <br> <strong>EukProt_included_data_sets.v03.2021_11_22.txt</strong> and <strong>EukProt_not_included_data_sets.v03.2021_11_22.txt</strong>: tables of information on data sets either included (993 data sets) or not included (163) in the database. Tab-delimited; multiple entries in the same cell are comma-delimited; missing data is represented with the “N/A” value. With the following columns:<br> <br> EukProt_ID: the unique identifier associated with the data set. This will not change among versions. If a new data set becomes available for the species, it will be assigned a new unique identifier.<br> <br> Name_to_Use: the name of the species for protein/genome annotation/assembled transcriptome files.<br> <br> Strain: the strain(s) of the species sequenced.<br> <br> Previous_Names: any previous names that this species was known by.<br> <br> Replaces_EukProt_ID/Replaced_by_EukProt_ID: if the data set changes with respect to an earlier version, the EukProt ID of the data set that it replaces (in the included table) or that it is replaced by (in the not_included table).<br> <br> Genus_UniEuk, Epithet_UniEuk, Supergroup_UniEuk, Taxogroup1_UniEuk, Taxogroup2_UniEuk: taxonomic identifiers at different levels of the UniEuk taxonomy (Berney et al. 2017, DOI: 10.1111/jeu.12414, based on Adl et al. 2019, DOI: 10.1111/jeu.12691).<br> <br> Taxonomy_UniEuk: the full lineage of the species in the UniEuk taxonomy (semicolon-delimited).<br> <br> Merged_Strains: whether multiple strains of the same species were merged to create the data set.<br> <br> Data_Source_URL: the URL(s) from which the data were downloaded.<br> <br> Data_Source_Name: the name of the data set (as assigned by the data source).<br> <br> Paper_DOI: the DOI(s) of the paper(s) that published the data set.<br> <br> Actions_Prior_to_Use: the action(s) that were taken to process the publicly available files in order to produce the data set in this database. Actions taken (see our manuscript for more details):<br> ‘assemble mRNA’: Trinity v. 2.8.4, http://trinityrnaseq.github.io/<br> ‘CD-HIT’: v. 4.6, http://weizhongli-lab.org/cd-hit/<br> ‘extractfeat’, ‘seqret’, ‘transeq’, ‘trimseq’: from EMBOSS package v. 6.6.0.0, http://emboss.sourceforge.net/<br> ‘translate mRNA’: Transdecoder v. 5.3.0, http://transdecoder.github.io/<br> ‘gffread’: v.0.12.3 https://github.com/gpertea/gffread<br> ‘predict genes’: EukMetaSanity https://github.com/cjneely10/EukMetaSanity (cloned on 21 September, 2021)<br> All parameter values were default, unless otherwise specified.<br> <br> Data_Source_Type: the type of the source data (possible types: EST, transcriptome, single-cell transcriptome, genome, single-cell genome).<br> <br> Notes: additional information on the data set (including why it is replaced by/is replacing another data set, or why it was not included).<br> <br> Columns_Modified_Since_Previous_Version: column(s) in this file modified for the data set since the previous release. Not listed: modifications to the Notes column or to new columns added in this version.<br> <br> Alternative_Strain_Names: non-exhaustive list of alternative names for the sequenced strain for this data set.<br> <br> 18S_Sequence_GenBank_ID: GenBank identifier for the strain sequenced in the data set. When multiple strains were sequenced, identifiers are separated with a comma, in the same order as the Strain column. Ranges of identifiers for the same strain are separated by a hyphen. ‘N/A’ indicates either that there is no GenBank sequence for the strain or that all available sequences are not full-length (&lt; 1,500 bp).<br> <br> 18S_Sequence: 18S for the strain derived from publicly available sequences associated with the data set, in the case where a GenBank sequence is not available.<br> <br> 18S_Sequence_Source: the source for the sequence in the 18S_Sequence column, if any.<br> <br> 18S_Sequence_Other_Strain_GenBank_ID: GenBank identifier for 18S sequence(s) from other strains of the same species as the data set.<br> <br> 18S_Sequence_Other_Strain_Name: strain name(s) for the sequences in the 18S_Sequence_Other_Strain_GenBank_ID column.<br> <br> 18S_and_Taxonomy_Notes: additional information on the values in the 18S_Sequence columns.<br> <br> <strong>Changes since version 2</strong><br> <br> There are 324 new data sets included. 57 of these replace data sets from version 2.<br> <br> 40 newly published data sets were added to the list that are not included in the database (annotated in the Notes column with the reasons they were not included).<br> <br> Instead of unannotated genomes (for published genomes lacking protein predictions), we now include predicted proteins and gene annotations (in GFF3 format).<br> <br> All sequences within each file are now assigned a standardized, unique identifier based on the data set’s EukProt_ID and on the type of data (protein or transcriptome). Illegal characters are removed from sequences.<br> <br> In the UniEuk_Taxonomy field, single quotes are now used instead of double quotes, to be consistent with other UniEuk databases (EukMap, EukRibo).<br> <br> Changes to metadata of individual data sets (in the included and not_included tables) with respect to the previous version are now listed in the Columns_Modified_Since_Previous_Version column.<br> <br> The Taxogroup_UniEuk column has been split into the Taxogroup1_UniEuk and Taxogroup2_UniEuk columns. This resulted in the Supergroup_UniEuk column changing for Opisthokonta.<br> <br> In addition, the following new columns have been added (see our manuscript for details): Alternative_Strain_Names, 18S_Sequence_GenBank_ID, 18S_Sequence, 18S_Sequence_Source, 18S_Sequence_Other_Strain_GenBank_ID, 18S_Sequence_Other_Strain_Name, 18S_and_Taxonomy_Notes.<br> <strong>EukProt_assembled_transcriptomes.v03.2021_11_22.tgz</strong>: assembled transcriptome contigs, for 126 species with publicly available mRNA sequence reads but no publicly available assembly. The proteins predicted from these assemblies are included in the proteins file. <br> Sequence names in the proteins and transcriptomes files have standardized, unique identifiers with the following format:<br> <br> &gt;[EukProt ID]_[Name_to_Use]_[Type abbreviation][Counter] [Previous header contents]<br> <br> Type abbreviations are P (protein) and T (transcriptome).<br> <br> All characters not in the following list are removed from nucleic acid sequences:<br> ACGTNUKSYMWRBDHV<br> All characters not in the the following list are removed from protein sequences:<br> ABCDEFGHIKLMNPQRSTUVWYZX*<br> <br> Lists of legal characters are from: https://blast.ncbi.nlm.nih.gov/Blast.cgi?CMD=Web&amp;PAGE_TYPE=BlastDocs&amp;DOC_TYPE=BlastHelp<br> <br>

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.002
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Open science, Insufficient payload (model declined to judge)
Consensus categoriesInsufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.747
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.002
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.001
Science and technology studies0.0010.000
Scholarly communication0.0000.000
Open science0.0040.016
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.7490.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.054
GPT teacher head0.286
Teacher spread0.232 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueFigshareFrench-language works237,207