Methods for high throughput validation of amplified fragment pools of BAC DNA for constructing high resolution CGH arrays
Bibliographic record
Abstract
BACKGROUND: The recent development of array based comparative genomic hybridization (CGH) technology provides improved resolution for detection of genomic DNA copy number alterations. In array CGH, generating spotting solution is a multi-step process where bacterial artificial chromosome (BAC) clones are converted to replenishable PCR amplified fragments pools (AFP) for use as spotting solution in a microarray format on glass substrate. With completion of the human and mouse genome sequencing, large BAC clone sets providing complete genome coverage are available for construction of whole genome BAC arrays. Currently, Southern hybridization, fluorescent in-situ hybridization (FISH), and BAC end sequencing methods are commonly used to identify the initial BAC clone but not the end product used for spotting arrays. The AFP sequencing technique described in this study is a novel method designed to verify the identity of array spotting solution in a high throughput manner. RESULTS: We show here that Southern hybridization, FISH, and AFP sequencing can be used to verify the identity of final spotting solutions using less than 10% of the AFP product. Single pass AFP sequencing identified over half of the 960 AFPs analyzed. Moreover, using two vector primers approximately 90% of the AFP spotting solutions can be identified. CONCLUSIONS: In this feasibility study we demonstrate that current methods for identifying initial BAC clones can be adapted to verify the identity of AFP spotting solutions used in printing arrays. Of these methods, AFP sequencing proves to be the most efficient for large scale identification of spotting solution in a high throughput manner.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.009 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.007 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".