14. Mapping, Curation, and Evolutionary Conservation of Macaque miRNA
Bibliographic record
Abstract
MicroRNAs (miRNAs) are small regulatory RNA molecules that switch off gene expression. Their main function is to degrade or stop the translation of target messenger RNAs through binding to their 3’ untranslated regions. miRNAs are excellent disease biomarkers due to their cell-type specificity, abundance, and stability. However, the sequences and locations of miRNAs within the human genome are a source of confusion in miRNA diagnostics. Here, I am defining the genomic locations and examining the specificity of miRNA expression in the Rhesus macaque tissues, using evolutionary conservation to guide our understanding of miRNA biology. First, I mapped the human miRNA precursor sequences in the macaque genome through the UCSC Genome Browser. Next, I expect to assess the validity of a miRNA by aligning macaque small RNA sequences against their corresponding precursor sequences. Lastly, I will calculate miRNA tissue specificity using existing data generated from 65 tissues obtained during a macaque necropsy. Through this approach, I expect to generate miRNA expression profiles using matching human miRNA expression profiles. These profiles were preprocessed through data normalization, outlier removal, and filtering of low expressed miRNAs. Feature selection and tissue specificity measures will be used to identify tissue-specific miRNA and an atlas of miRNA expression will be generated. miRNA conservation between humans and macaques will be assessed and macaque segments that did not align with the human genome will be investigated separately as they may be new miRNA.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.005 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".