Adapting Genome-wide microRNA Discovery and Target Prediction to Specific Species
Bibliographic record
Abstract
microRNAs (miRNAs) are small non-coding ribonucleic acids that post-transcriptionally regulate gene expression through the targeting of messenger RNA (mRNAs).miRNAs have been implicated in numerous biological processes in animals and plants, so discovering miRNAs within unannotated genomes and determining which mRNA they may target are important challenges.MiRNA and miRNA targets can be identified experimentally through costly and time consuming wet-lab verification techniques or computationally using a variety of techniques.In the case of miRNA discovery, several models exist, however, they have been tested and trained data with a class imbalance not indicative of the real-life problem distribution.Several ML miRNA target predictors exist; however, they primarily focus on the Animal Kingdom.Several rule-based miRNA target predictors have been developed in plant species, but they often fail to discover new miRNA targets with non-canonical miRNA-mRNA binding.In this thesis, we focus on two areas of miRNA research: miRNA discovery and miRNA target prediction.Specifically, we develop machine learning techniques to discover miRNA in the Soy Cyst Nematode genome -a destructive pathogen of soybean and a species for which very few miRNAs are currently known.Considering that, in some species, it has been shown that miRNAs are differentially expressed in response to pathogen stress, we also predict gene targets for each putative miRNA within both SCN and soybean.Additionally, we develop a plant-specific miRNA targeting model and webserver rigorously tested across four plant species.Finally, we explore the applications of domain adaptation on miRNA targeting.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".