Additional file 8: Table S2. of Reversing chromatin accessibility differences that distinguish homologous mitotic metaphase chromosomes
Bibliographic record
Abstract
CNV and gene expression data in context of regions with DA and equivalent accessibility. Single copy probe locations of regions with DA (in bold) and equivalent accessibility (i.e. without DA) are from indicated GRCh37 genomic coordinates. Overlapping CNVs in the population from Healthy Sample (HS) and Ontario Population Genomics Platform (OPGP) are separated by a “//.” No recurrent CNVs (diploid copy number state) were observed in regions overlapped by single copy probes, and are indicated by a “0”. A 179 kb gain in region of single copy probe within C9orf66 was seen in only 1 out of ~400 HS individuals. This observation is a rare CNV and the region in which it resides, does not exhibit DA. SC probe locations do not overlap locations of CNVs reported in cell line GM06326 (see methods). GM10958 cell line CNV data have not been analyzed by genomic microarray analysis. This cell line has been characterized as phenotypically normal (Coriell Cell Repository). Compared to other tissues, mRNA abundance among human lymphocytes cells (EBV transformed, n = 54 individuals) showed low or no expression (most values below 0 on log10 scale, Genotype-Tissue Expression – GTEx database) for each of the single copy probes tested. C9orf66 expression was not catalogued in GTEx. The EMBL expression atlas (Illumina body map, http://www.ebi.ac.uk/gxa/experiments/E-MTAB-513) confirmed that it is not expressed in leukocytes. Marks of transcriptionally active chromatin (i.e. H3K36me3, H4K20me1 - reported values represent integrated intensities) were not significantly present (p = 0.91) in either DA (in bold, μ = 170.86) or equivalent accessible genomic regions (μ = 161.15). Integrated intensity values were processed from ENCODE ChIP-seq data on lymphoblastoid cell line GM12878 using the Broad histone signal intensity ‘StdSig’ setting within the UCSC Genome Table browser.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.023 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.885 | 0.228 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".