MétaCan
Menu
Back to cohort
Record W790473025

Enforcing secondary and tertiary structure for crystallographic phasing. Developing ARCIMBOLDO and BORGES

2015· dissertation· en· W790473025 on OpenAlexaboutno aff
Massimo Sammito

Bibliographic record

VenueDipòsit Digital de la Universitat de Barcelona (Universitat de Barcelona) · 2015
Typedissertation
Languageen
FieldMaterials Science
TopicEnzyme Structure and Function
Canadian institutionsnot available
FundersUniversität Konstanz
KeywordsPhaserCrystallographyMaterials scienceChemistryPhysicsOptics
DOInot available

Abstract

fetched live from OpenAlex

ARCIMBOLDO is an ab initio phasing method for macromolecular crystallographic X-ray diffraction data, which combines location of model fragments such as polyalanine α- helices with the program PHASER and density modification and main chain autotracing with the program SHELXE.\n\n\t\t\t\t The method has been named after the Italian painter Giuseppe Arcimboldo (1526-1593), who used to compose portraits out of common objects such as fruits and vegetables. Following the analogy, ARCIMBOLDO composes an unknown structure by assembling small secondary structure elements, which are conserved across families of unrelated tertiary structure. Exploiting this method requires a multi-solution approach due to the difficulty to recognize correct solutions at early stages.\n\n\t\t\t\t Moreover, phasing a structure starting from partial information provided by such a small percentage of the total model (around 10% of the main chain atoms) is challenging and requires evaluation of alternative hypotheses under statistical constraints to avoid combinatorial explosion.\n\n\t\t\t\t ARCIMBOLDO methods have proven successful in many cases of previously unknown structures[3] and also on a pool of test structures[4]. The program can accept any Sohnke space group and all the most frequent ones are represented in the pool of structures solved so far. In both studies data were collected in the most common protein space groups.\n\n\t\t\t\t Data quality is crucial for phasing methods, and particularly sensitive for ARCIMBOLDO, where low resolution (worse than 2.1 Å) and lack of completeness (less than 98%) drastically decrease the chance of success.\n\n\t\t\t\t Location of secondary structure elements is not indicated as phasing method for large structures or complexes (over 400 residues) unless very long helices are present and high resolution data are available. Such cases would require the placement of many fragments in order to assemble 10% of the main chain, which can lead to an unmanageable number of solutions. To approach correctly this different scenario we have implemented dedicated methods in ARCIMBOLDO_BORGES[7] and ARCIMBOLDO_SHREDDER[8]. These programs exploit libraries of folds or large search models and are described later in the text.\n\n\t\t\t\t The current implementation[4], coded in Python, is deployed as a standalone binary, freely available under registration from http://chango.ibmb.csic.es/download. The binary is compatible with common Linux distributions and latest versions of the Mac OSX operating system. Users can find online manuals, tutorials and documentation in our website. As of 30th April 2015, it has been downloaded 664 times and distributed to 121 research groups; furthermore, it has been installed in many European synchrotron facilities such as the Alba Synchrotron in Spain, the Diamond Light Source in United Kingdom and SOLEIL Synchrotron in France. The software is also available through SBGrid Consortium (https://sbgrid.org), a network of institutions across 19 countries, which provides a distributed grid network of computers to run structural biology software. We have recently started a collaboration with the San Diego Supercomputer Center (http://www.sdsc.edu) in California (USA), to develop optimized and dedicated versions of the programs for their platform with the aim of addressing difficult phasing cases. Due to this recent spread in the crystallographic community ARCIMBOLDO has been presented in many international conferences such as the International Union of Crystallography Meeting in Madrid (ES) 2011 and in Montreal (CA) 2014; the European Crystallographic Meeting in Bergen (NO) 2012, Warwick (UK) 2013; and many schools and workshops such as the International School of Crystallography in Erice (IT) 2012 and Macromolecular Crystallography School in Madrid (ES) 2014.\n\n\n\t\t\t\t This thesis is organised in the standard scientific format comprising five main parts: \n\n\t\t\t\t 1. INTRODUCTION: introducing the theoretical topics directly or indirectly related to the contents of the thesis and also discussing the state of the art of current scientific production related to the objective proposed.\n\n\t\t\t\t 2. OBJECTIVES: listing all general goals and particular aims of the doctoral project conducted.\n\n\t\t\t\t 3. MATERIALS AND METHODS: detailing the hardware and software environment, including third party software and algorithms employed in the project.\n\n\t\t\t\t 4. RESULTS AND DISCUSSION: presenting all the produced algorithms, software, experiments and tests that correspond to the prefixed objectives.\n\n\t\t\t\t 5. CONCLUSION: summarising the whole project and listing its achievements by the end of the doctoral studies.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.010
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.010
Threshold uncertainty score0.033

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.010
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.002
Science and technology studies0.0010.001
Scholarly communication0.0030.002
Open science0.0020.003
Research integrity0.0020.004
Insufficient payload (model declined to judge)0.0100.007

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.004
GPT teacher head0.224
Teacher spread0.220 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2015
Admission routes1
Has abstractyes

Explore more

Same venueDipòsit Digital de la Universitat de Barcelona (Universitat de Barcelona)Same topicEnzyme Structure and FunctionFrench-language works237,207