Genetic association studies for complex traits: relevance for the sports medicine practitioner
Bibliographic record
Abstract
In this issue of BJSM , September et al 1 report DNA variants within the COL5A1 gene among patients with Achilles tendinopathy (the cases) and among controls with no tendinopathy, matched for age and country of origin. Their findings suggest that a common DNA variation in the COL5A1 gene may be a risk factor for Achilles tendinopathy. Replication of their results among larger cohorts will be necessary to validate this finding. It can be a challenge for the busy sports medicine practitioner to distil the clinical relevance of association studies such as this one. Although genetic testing for mendelian (single-gene) disorders is widely available in many countries, genetic testing for single-nucleotide polymorphisms (SNPs) is generally unavailable outside research laboratories. With certain notable exceptions (eg, the link between apoE variants, cardiovascular and neurological disease),2-4 statistical associations between common SNPs and complex diseases have not been borne out by further study. Numerous pitfalls occur in both the design and interpretation of these studies, which has resulted in a poor track record for independent replication. Thus, a brief summary of these pitfalls may be useful to the BJSM ’s readership. By way of background, there are just over 14 million SNPs known to exist in the human genome.5 Most of these exist in two possible forms, reflecting variation in the sequence that arose as a new mutation many generations ago. At some loci, any of three or even all four DNA bases (adenine, guanine, cytosine and thymine; A,G, C and T) may occur at measurable frequency in a population. Once a variant’s frequency reaches 1% of all alleles in the population, such a variant is no longer considered a “mutation,” but rather a “polymorphism.” Recent advances in the rapidity and cost-effectiveness of DNA sequencing technologies have enabled the assessment …
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".