SHELXT: Integrating space group determination and structure solution
Bibliographic record
Abstract
SHELXT is intended for robust routine solution of small molecule crystal structures. It makes the simple but powerful assumption that the structure consists of resolved atoms, but unlike classical direct methods it is not required that the atoms are 'equal'. This enables it to succeed with poor or incomplete data but makes it unsuitable for structures that are twinned, modulated or contain severe (e.g. 'whole molecule') disorder. SHELXT is a dual-space program that starts with a Patterson minimum superposition and iteratively applies the random omit procedure (also used in SHELXD) with data expanded to space group P1, but does not use phase probability relations or solvent flipping. In the SHELX system it will probably obsolete SHELXS but not SHELXD, which is better for large equal-atom and twinned structures. SHELXT reads any legal SHELX format .ins and (HKLF3 or 4) .hkl files. It extracts the Laue group and tries to find space groups in this Laue group and origin shifts to fit the phases from the best P1 solution, and makes an approximate assignment of element types using the elements specified on the SFAC instruction (and maybe a couple more). This is followed by an isotropic refinement and an attempt to assign the absolute structure if the space group is non-centrosymmetric. It is hoped to release SHELXT as part of the SHELX system (http://shelx.uni-ac.gwdg.de/SHELX/index.php) in time for the 2014 Montreal IUCr Meeting.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.009 |
| Meta-epidemiology (narrow) | 0.004 | 0.004 |
| Meta-epidemiology (broad) | 0.012 | 0.003 |
| Bibliometrics | 0.006 | 0.009 |
| Science and technology studies | 0.004 | 0.001 |
| Scholarly communication | 0.006 | 0.007 |
| Open science | 0.014 | 0.009 |
| Research integrity | 0.003 | 0.014 |
| Insufficient payload (model declined to judge) | 0.099 | 0.092 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".