Analytical Differentiable Finite-Resolution Density Map Calculation in CCTBX/Phenix
Bibliographic record
Abstract
Beyond validation, publication, and other analyses, the final stage of structure determination using cryo-EM typically involves atomic model refinement against experimental data. This refinement is most naturally performed in real space - bypassing Fourier space - because all objects at this stage, including models and maps, exist in real space. However, many tools currently used in cryo-EM structure determination originate from and remain anchored to crystallography, which primarily operates in reciprocal (Fourier) space. While Phenix tools specifically designed for cryo-EM during the resolution revolution were tailored to operate in real space - such as phenix.real_space_refine for atomic model refinement against maps, which accounts for 95% of structures deposited in the PDB using cryo-EM - they are still suboptimal in at least two aspects. First, the refinement target for coordinate refinement in phenix.real_space_refine is perhaps the simplest and fastest to compute, as it focuses on fitting atoms to the nearest density peaks without considering the overall shape of the density map. The advantage of this approach is that it enables refinement of very large molecules on relatively modest computing resources (e.g., laptops). However, the drawback is the need for excessive geometric restraints to compensate for the simplicity of this refinement target, which does not account for the shape of the map. The second limitation is that ADP (B-factor) and occupancy refinements still require a bypass through Fourier space, as they rely on spatial integration of density peaks for fitting. Here, we will focus our discussion on implementing resolution-truncated density map calculations in CCTBX and Phenix using accurate, differentiable analytic approximations of density maps. This implementation will enable more accurate real-space refinement in Phenix for all atomic model parameters (coordinates, ADPs, occupancies, etc.), eliminating the need for Fourier space entirely in this process. Additionally, making this approach available in the freely accessible CCTBX framework will provide the broader community with a uniform method for computing finite-resolution density maps and their derivatives with respect to atomic model parameters. This is particularly valuable for machine learning-based model building and refinement approaches, where algorithms for computing differentiable finite-resolution density maps are essential (e.g., qFit, ROCKET).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.040 | 0.007 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".