Data files for manuscript "Elucidating the clinical and molecular spectrum of SMARCC2-associated NDD in a cohort of 65 affected individuals"
Bibliographic record
Abstract
# 2023-06-28 # Data files for manuscript "Elucidating the clinical and molecular spectrum of SMARCC2-associated NDD in a cohort of 65 affected individuals" # Summary This ZIP-file contains the supplementary files of our SMARCC2 study "Elucidating the clinical and molecular spectrum of SMARCC2-associated NDD in a cohort of 65 affected individuals". Suppl. File S2 contains comprehensive clinical data Suppl. File S3 contains comprehensive genetic data Suppl. File S4 contains files of the SMARCC2 N-terminal homology model # Folder structure ./ (parent directory containing this README file and all subfolders) ./Files/ (contains Excel Suppl. File S2 and Suppl.File S3, and ZIP Suppl.File S4) # Files and checksums Algorithm Hash Path --------- ---- ---- MD5 003A879AA75CF8B5C4E3E05F76C4EB4C SMARCC2-Supplementary\Files\FileS2_cases_clinical-table.xlsx MD5 73B4F9E83419A8404345101CAA4D2205 SMARCC2-Supplementary\Files\FileS3_variants-and-domains.xlsx MD5 B0A12F36B4801EB6C4D21BA02B4270BB SMARCC2-Supplementary\Files\FileS4_SMARCC2 N-terminal homology model.zip
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.029 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.003 | 0.006 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.794 | 0.405 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".