PrionOme: A database of prions and other sequences relevant to prion phenomena
Bibliographic record
Abstract
Abstract Prions are units of propagation of an altered state of a protein or proteins. Prions can propagate from cell to cell, and from organism to organism, through cooption of other protein copies. Prions contain no necessary nucleic acids, and are important both as both pathogenic agents, and as a potential force in epigenetic phenomena. The original prions were derived from a misfolded form of the mammalian Prion Protein PrP. Infection by these prions causes neurodegenerative diseases. Other prions cause non-Mendelian inheritance in budding yeast, and sometimes act as diseases of yeast. We have compiled a database of >2000 prion-related sequences, called the PrionOme. The database comprises seven PrionOme classification categories: prionogenic sequences (i.e., sequences that can make prions), ‘prionoids’ (i.e., phenomena that have some prion characteristics), orthologs, paralogs, pseudogenes, prion interactors, and prion-like molecules. Database entries list: supporting information for PrionOme classifications, prion-determinant areas (where relevant), and disordered and compositionally-biased regions. Also included are original references for the PrionOme classifications, transcripts and genomic coordinates, and structural data (including comparative models). We provide database usage examples for both vertebrate and fungal prion contexts. As development of this resource is on-going, we will be very happy to receive and act on any constructive comments from peer scientists in the areas of prion biology and protein misfolding, either by email or using the feedback form provided on the PrionOme website. We hope that this database will be a valuable experimental aid and reference resource. It is freely available at: http://libaio.biol.mcgill.ca/prion.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".