Additional file 14 of Re-examination of two diatom reference genomes using long-read sequencing
Bibliographic record
Abstract
Additional file 14: Supplementary Figure 1. Workflow of sample preparation, MinION sequencing, de novo genome assembly and downstream analyses, including methods to compare the de novo long-read reference assemblies and gene prediction. Supplementary Figure 2. Bivariate scatterplots showing the relationship of MinION read lengths and average Phred read quality scores for Phaeodactylum tricornutum (A & B) and Thalassiosira pseudonana (C & D). The unfiltered datasets (A & C) include all generated MinION data and the filtered datasets (B & D) include a subset of reads filtered by length and quality. Supplementary Figure 3. Pulsed-field gel electrophoresis of P. tricornutum (CCMP632) DNA using settings to optimize resolution of large (A) and small (B) fragments. The red arrowheads (<) indicate potential chromosome-sized fragments. Ladders for sizing fragments include Saccharomyces cerevisiae (Sc) and Hansenula wingei (Hw). Lanes labelled 1–4 and 8–13 are not relevant to this study. Supplementary Figure 4. Relationship between 16,491 predicted protein models from the Flye Thalassiosira pseudonana assembly to reference protein set and the matches to known proteins of the new predicted protein models (2010). Supplementary Figure 5. PloidyNGS plot of the frequency of the two most abundant alleles in the Phaeodactylum tricornutum genome indicates that it is a diploid organism. Supplementary Figure 6. Stacked histograms showing the different categories of P. tricornutum haplotypes supported by Bionano hybrid scaffolding. Characterizations are based on a total of 66 haplotypes expected based on a diploid genome with 33 chromosomes as estimated for the reference genome. Supplementary Figure 7. ‘Hybrid-hybrid’ Bionano-Canu super-scaffolds that may represent mis-assemblies owing to segmental duplications (A & B) or LTR-RT insertions (C) located at the ends of Canu contigs. Mauve [81] schematics of the syntenic regions between a Bionano-Canu super-scaffold and the reference chromosomes that it contains demonstrate the hybrid nature of each super-scaffold. Inserts provide a more detailed illustration of the high sequence identity between the hybrid-hybrid scaffolds and the reference chromosomes at the areas of the genome containing segmental duplications or LTR-RT insertions. Supplementary Figure 8. Bionano-Canu super-scaffolds that were identified as ‘hybrid-hybrid’ scaffolds most probably owing to errors of the Bionano mapping process. For each example (A-F), a schematic of the super-scaffold is annotated with its respective Canu contigs shown in purple blocks. Gap regions inserted by Bionano are indicated by solid black lines. Syntenic regions between each ‘hybrid-hybrid’ super-scaffold and the two reference chromosomes it contains were evaluated by Mauve [81] and shown as colored blocks above the super-scaffold schematic. Blastn results are reported for each of the Canu contigs resolved to the ‘hybrid-hybrid’ super-scaffold against the appropriate reference genome chromosome. Supplementary Figure 9. Multiple sequence alignment showing a region of the P. tricornutum Canu assembly that is represented by two contigs (haplotype 1 = tig94 & haplotype 2 = tig92) while the reference genome is only represented by a single scaffold (chr3). The blue and red boxes indicate SNPs between the reference scaffold and Canu haplotigs. While the reference scaffold and haplotype 1 match at the first three SNPs, the reference scaffold disagrees with haplotype 1 at the following few SNPs, matching haplotype 2, instead. That pattern is strongly suggestive that the reference is an amalgamation of the two haplotypes resolved by the Canu assembly. Asterisks represent mapped Illumina short-reads with green boxes representing areas where individual ~ 120 bp Illumina reads supported the SNPs captured for each Canu haplotype. Supplementary Figure 10. IGV schematic showing the location of four SNPs between the reference genome and two Canu haplotigs. The SNPs are indicated by the four bi-colored columns, which correspond to the number of alternative bases detected at each site. The four SNPs are consistent across the mapped reads with the blue boxes representing haplotype 1 and the red boxes representing haplotype 2. Roughly equivalent numbers of reads were found to support each haplotype, which is consistent with P. tricornutum as a diploid genome. Supplementary Figure 11. Full-length CoDi long-terminal repeat retrotransposon content resolved for T. pseudonana. The number of previously reported & overlooked loci are reported for the reference genome (A) as well as the Flye assembly (B), which also included novel LTR insertions. The number of LTR-RTs detected for each CoDi group in the reference genome is compared to the number of LTR-RTs detected for each CoDi group in the Flye de novo assembly (C). LTR-RTs are characterized as either “previously reported loci” (i.e., loci homologous to those previously reported by Maumus et al. [15] in the reference genome), “overlooked loci” (those homologous to those present in the reference genome but not reported) or “novel loci” (i.e., loci detected in our long-read assembly but without a homologous insertion in the reference genome).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.020 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.004 | 0.006 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.802 | 0.192 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".