High Throughput Crystallography at SGC Toronto: an Overview
Bibliographic record
Abstract
The completion of the human genome allows the analysis, for the first time, of biological systems in the context of entire gene families. For enzymes, this approach permits the exploration of complex substrate specificity networks that often exhibit considerable overlap within and between protein families. The case for a family-based approach to protein studies is compelling, given the prospect of exploiting these specificities for various purposes, such as the development of therapeutic reagents. The Structural Genomics Consortium (SGC) was created to determine the structures of proteins with relevance to human health and place the structures into the public domain without restriction on use. The SGC operates out of the Universities of Toronto and Oxford, and Karolinska Institutet, each working on nonoverlapping protein target lists. The SGC focus on human protein families requires a repertoire of crystallography methods that differ from those adopted by structural genomics projects that are focused on filling out protein fold space. The key differences are heavier reliance on in house x-ray sources for diffraction data collection and predominant use of molecular replacement for phase determination. As projects such as the US Protein Structure Initiative and others fill the PDB with representatives of most major fold families, the SGC approach will become an increasingly useful model for many structural biology laboratories in the future. Technical details of the flow of samples and data within the high throughput (HTP) environment at SGC Toronto are presented, and provide a useful paradigm for the organization of collaborative or shared x-ray instrumentation facilities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.002 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.009 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.005 | 0.002 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.021 | 0.021 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".