ProtoCell: a computational protocell model implemented using <i>String</i>
Bibliographic record
Abstract
A computer language called String (Islam et al., 2022) was originally designed for the purpose of manual authoring and computational evolution of the active entities (called ‘ribozymes’) within a new computational protocell model (called ProtoCell). String’s source code (or ‘phenotype’) has an assembly-like format, while its encoding (or ‘genotype’) has the form of an RNA sequence. In this paper, we present, in brief, the complete ProtoCell model, comprising three essential subsystems, all utilizing ribozymes written in String. The first or, Genomic subsystem is made of a single loop of RNA, which has the encodings of all the ribozymes of the model. The Membrane subsystem is made of a self-assembling lipid, with embedded trans-membrane transporters, realized as ribozymes. The last sub-system is the Metabolism, the factory of the cell, where all the necessary building blocks, and energy molecule, are synthesized in reactions catalyzed by ribozymes. We combined all three subsystems to achieve a stable ProtoCell simulation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.017 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".