Impact of gene expression data pre-processing on expression quantitative trait locus mapping
Bibliographic record
Abstract
We evaluate the impact of three pre-processing methods for Affymetrix microarray data on expression quantitative trait locus (eQTL) mapping, using 14 CEPH Utah families (GAW Problem 1 data). Different sets of expression traits were chosen according to different selection criteria: expression level, variance, and heritability. For each gene, three expression phenotypes were obtained by different pre-processing methods. Each quantitative phenotype was then submitted to a whole-genome scan, using multipoint variance component LODs. Pre-processing methods were compared with respect to their linkage outcomes (number of linkage signals with LODs greater than 3, consistencies in the location of the trait-specific linkage signals, and type of cis/trans-regulating loci). Overall, we found little agreement between linkage results from the different pre-processing methods: most of the linkage signals were specific to one pre-processing method. However, agreement rates varied according to the criteria used to select the traits. For instance, these rates were higher in the set of the most heritable traits. On the other hand, the pre-processing method had little impact on the relative proportion of detected cis and trans-regulating loci. Interestingly, although the number of detected cis-regulating loci was relatively small, pre-processing methods agreed much better in this set of linkage signals than in the trans-regulating loci. Several potential factors explaining the discordance observed between the methods are discussed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.046 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".