A systematic screen for genes expressed in definitive endoderm by Serial Analysis of Gene Expression (SAGE)
Bibliographic record
Abstract
BACKGROUND: The embryonic definitive endoderm (DE) gives rise to organs of the gastrointestinal and respiratory tract including the liver, pancreas and epithelia of the lung and colon. Understanding how DE progenitor cells generate these tissues is critical to understanding the cause of visceral organ disorders and cancers, and will ultimately lead to novel therapies including tissue and organ regeneration. However, investigation into the molecular mechanisms of DE differentiation has been hindered by the lack of early DE-specific markers. RESULTS: We describe the identification of novel as well as known genes that are expressed in DE using Serial Analysis of Gene Expression (SAGE). We generated and analyzed three longSAGE libraries from early DE of murine embryos: early whole definitive endoderm (0-6 somite stage), foregut (8-12 somite stage), and hindgut (8-12 somite stage). A list of candidate genes enriched for expression in endoderm was compiled through comparisons within these three endoderm libraries and against 133 mouse longSAGE libraries generated by the Mouse Atlas of Gene Expression Project encompassing multiple embryonic tissues and stages. Using whole mount in situ hybridization, we confirmed that 22/32 (69%) genes showed previously uncharacterized expression in the DE. Importantly, two genes identified, Pyy and 5730521E12Rik, showed exclusive DE expression at early stages of endoderm patterning. CONCLUSION: The high efficiency of this endoderm screen indicates that our approach can be successfully used to analyze and validate the vast amount of data obtained by the Mouse Atlas of Gene Expression Project. Importantly, these novel early endoderm-expressing genes will be valuable for further investigation into the molecular mechanisms that regulate endoderm development.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".