Author response: Endogenous tagging using split mNeonGreen in human iPSCs for live imaging studies
Bibliographic record
Abstract
The human body contains around 20,000 different proteins that perform a myriad of essential roles. To understand how these proteins work in healthy individuals and during disease, we need to know their precise locations inside cells and how these locations may change in different situations. Genetic tools known as fluorescent proteins are often used as tags to study the location of specific proteins of interest within cells. When exposed to light, the fluorescent proteins emit specific colours of light that can be observed using microscopes. In a fluorescent protein system known as split mNeonGreen, researchers insert the DNA encoding two fragments of a fluorescent protein (one large, one small) separately into cells. The large fragment can be found throughout the cell, while the small fragment is attached to specific host proteins. When the two fragments meet, they assemble into the full mNeonGreen protein and can fluoresce. Researchers can use split mNeonGreen and other similar systems to generate large libraries of cells where the small fragment of a fluorescent protein is attached to thousands of different host proteins. However, so far these libraries are restricted to a handful of different types of cells. To address this challenge, Husser et al. inserted the DNA encoding the large fragment of mNeonGreen into human cells known as induced pluripotent stem cells, which are able to give rise to any other type of human cell. This then enabled the team to quickly and efficiently generate a library of stem cells that express the small fragment of mNeonGreen attached to different host proteins. Further experiments studied the locations of host proteins in the stem cells just before they divided into two cells. This suggested that there are differences between how induced pluripotent stem cells and other types of cells divide. In the future, the cells and the method developed by Husser et al. may be used by other researchers to create atlases showing where human proteins are located in many other types of cells.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.007 | 0.006 |
| Insufficient payload (model declined to judge) | 0.112 | 0.056 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".