Genome assembly for the Sierra Nevada Parnassian ( <i>Parnassius behrii</i> ) and a brief review of butterfly genome sizes
Bibliographic record
Abstract
The Sierra Nevada Parnassian (Parnassius behrii W.H. Edwards, 1870) (Lepidoptera: Papilionidae) is a high-elevation specialist butterfly endemic to the Sierra Nevada, California. We present a genome assembly for P. behrii, representing the first major genomic resource for the species and greater Parnassius phoebus species complex. The assembly consists of two haplotypes, 1.59 Gb and 1.46 Gb in length, with contig N50 values of 10.93 Mb and 11.84 Mb, scaffold N50 values of 52.56 Mb and 51.90 Mb, scaffold L50 values of 13 and 14, and BUSCO completeness scores of 98.7% and 94.4%, respectively. Both haplotypes are highly contiguous, with 31 chromosome-length scaffolds, including putative Z and W sex chromosomes. We annotated the genome with National Center for Biotechnology Information's (NCBI's) EGAPx pipeline, integrating database and novel transcript alignment with Hidden Markov Model-based gene predictions, yielding 17,191 genes with a BUSCO score of 98.1%. RepeatMasker identified that 26.68% (424.97 Mb) of the genome consists of repetitive elements. We also assembled a mitochondrial genome for P. behrii (15,391 bp) containing 2 rRNAs, 22 unique transfer RNAs, and 13 protein-coding genes. We finally reviewed 514 high-quality butterfly genomes available from the National Center for Biotechnology Information (NCBI). Parnassius species were observed to have the largest genomes, with P. behrii being the largest. This assembly provides a foundational resource for whole-genome research on P. behrii and the broader P. phoebus complex, enabling analyses of evolutionary differentiation, local adaptation, inbreeding, gene flow, taxonomy and conservation practices.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".