Bibliographic record
Abstract
Compartmentalization is essential for all complex forms of life. In eukaryotic cells, membrane-bound organelles, as well as a multitude of protein- and nucleic acid-rich subcellular structures, maintain boundaries and serve as enrichment zones to promote and regulate protein function. Consistent with the critical importance of these boundaries, alterations in the machinery that mediate protein transport between these compartments has been implicated in a number of diverse diseases. Understanding the composition of each cellular “compartment” (be it a classical organelle or a large protein complex) remains a challenging task. For soluble protein complexes, approaches such as affinity purification other biochemical fractionation coupled to mass spectrometry provides important insight, but this is not the case for detergent-insoluble components. Classically, both microscopy and organellar purifications have been employed for identifying the composition of these structures, but these approaches have limitations, notably in resolution for standard high-throughput fluorescence microscopy and in the difficulty in purifying some of the structures (e.g. p-bodies) for approaches based on biochemical isolations. Prompted by the recent implementation in vivo biotinylation approaches such as BioID, we report here the systematic mapping of the composition of various subcellular structures, using as baits proteins (or protein fragments) which are well-characterized markers for a specified location. We defined how relationships between “prey” proteins detected through this approach can help understanding the protein organization inside a cell. We will discuss our low-resolution map of a human cell containing major organelles and non-membrane bound structures, but also a higher resolution map of RNA-containing cellular structures, including the p-bodies and the stress granules that regulate mRNA stability. This will be presented alongside new computational tools that will help the scientific community to make use of our dataset. Support or Funding Information We acknowledge the support of the Canadian Institutes of Health Research and the Natural Sciences and Engineering Research Council of Canada. A draft map of a human cell using proximity biotinylation A draft map of a human cell using proximity biotinylation
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".