Mining for single nucleotide variants (SNVs) at the kallikrein locus with predicted functional consequences
Bibliographic record
Abstract
Kallikreins (KLKs) are a group of 15 serine proteases encoded by the KLK locus on chromosome 19. Certain single nucleotide variants (SNVs) within the KLK locus have been linked to human disease. Next-generation sequencing of large human cohorts enables reexamination of genomic variation at the KLK locus. We aimed to identify all KLK-related SNVs and examine their impact on gene regulation and function. To this end, we mined KLK SNVs across Ensembl and Exome Variant Server, with exome-sequencing data from 6503 individuals. PolyPhen-2-based prediction of damaging SNVs and population frequencies of these SNVs were examined. Damaging SNVs were plotted on protein sequence and structure. We identified 4866 SNVs, the largest number of KLK-related SNVs reported. Fourteen percent of noncoding SNVs overlapped with transcription factor binding sites. We identified 602 missense coding SNVs, among which 148 were predicted to be damaging. Nine missense SNVs were common (>1% frequency) and displayed significantly different frequencies between European-American and African-American populations. SNVs predicted to be damaging appeared to alter tertiary structure of KLK1 and KLK6. Similarly, these missense SNVs may affect KLK function, resulting in disease phenotypes. Our study represents a mine of information for those studying KLK-related SNVs and their associations with diseases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".