Bibliographic record
Abstract
The last few years have seen spectacular growth in the widespread adoption of cryoelectron microscopy (cryo-EM).With the advent of direct electron detectors, the number of depositions in the EM Data Bank (EMDB; https://emdb-empiar.org/) has increased steadily at a rate of ~30% every year since 2014, crossing the 10 000 mark in early 2020.What may not be widely appreciated is that this extraordinary rate of growth in cryo-EM is not new.The number of depositions in the EM data repository increased on average by ~40% each year between 2004 and 2014, quickly expanding beyond the 52 entries listed in 2004.Besides the sheer increase in numbers of structures being determined using cryo-EM methods, an important aspect to this growth is the application of cryo-EM to determine structures of heterogeneous and dynamic assemblies.Protein complexes of this kind will most likely continue to remain largely intractable to crystallographic approaches that require the generation of ordered two-or three-dimensional crystals.The rapid progress in the application of cryo-EM methods to study SARS-CoV-2 proteins provides a convincing demonstration of the power of cryo-EM in the arsenal of structural biology.Within a month or so of the availability of the gene sequence, Wrapp et al. (2020) and Walls et al. (2020) were able to obtain the first structures of the soluble part of the trimeric SARS-CoV-2 spike protein.Nearly a dozen groups worldwide have extended these studies in short order to decipher many conformational variants of the spike protein, including structures of native spikes displayed on intact viral membranes.These are major achievements that are also a testament to the speed, quality and biomedical relevance of present day cryo-EM methods.The structures provide important insights into the complexity of spike architecture, especially in the conformational spread of the domain that binds the human ACE2 receptor.These studies, in combination with emerging structural information on the binding sites of various antibodies on the S-protein are fundamentally relevant to the design of effective therapeutics and vaccines against COVID-19.Before January 2020, of the 84 coronavirus cryo-EM structures deposited in the repository, 82 were of trimeric spike proteins.In contrast, of the 34 entries related to SARS-CoV-2 that have been deposited in 2020, only ~50% are of spike proteins, with the rest coming from entries such as the ORF3a membrane protein, ACE2 protein in complex with the membrane protein BOAT1 and RNA polymerase complexes in different conformational states [Fig.1(a)].What is noteworthy about these advances is that protein complexes of this kind are often flexible and conformationally heterogeneous, making them challenging to study by X-ray crystallography.The advent of 'single particle' methods such as cryo-EM and XFELs offer the exciting prospect to transcend the description of proteins in terms of static structures, and instead derive conformational landscapes that may provide a much better way to understand their biological function (Ourmazd, 2019).We are now publishing an increasing number of important papers in this area in IUCrJ, and are particularly keen to welcome papers in IUCrJ and other IUCr journals that increase understanding of SARS-CoV-2.The realization of the potential impact of cryo-EM led to the nucleation of several national and regional facilities in different countries to nurture the growth of the cryo-EM field (Subramaniam, 2019).While these facilities continue to play a key role, we are now seeing growth in the cryo-EM field that is fueled by the increased availability of highend cryo-EM instrumentation within individual institutions.For this trend to continue, it is critical that the cost of cryo-EM instrumentation necessary for high-resolution structure determination decreases significantly.Funders of biomedical research need to
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.016 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.006 | 0.005 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.015 | 0.028 |
| Insufficient payload (model declined to judge) | 0.009 | 0.010 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".