Competition and coevolution drive the evolution and the diversification of CRISPR immunity
Bibliographic record
Abstract
Data for the paper : "Competition and coevolution drive the evolution and the diversification of CRISPR immunity" Does NOT contain sequencing data (they are available on NCBI, BioProject PRJNA843584) Bacteria_density.csv Density of bacteria (cfu/ml) in all replicates at all times. B: Treatment A ; Pw: Treatment B ; PR: Treatment C Phages_density.csv Density of phages (pfu/ml) in all replicates at all times. B: Treatment A ; Pw: Treatment B ; PR: Treatment C NC_007019.1.fasta Genome of bacteriophage 2972 downloaded from https://www-ncbi-nlm-nih-gov.inee.bib.cnrs.fr/nuccore/66391759 NC_007019.1_T0.fasta Genome of bacteriophage 2972, updated to include the mutations observed in the mix of phages at T0 in treatment B. Mutations.xlsx This file shows the mutations in the genome of the 16 starting resistant bacteria compared to the sensitive ancestor DGCC 7710. All mutations were confirmed via Sanger sequencing of PCR products amplified from the chromosome of the BIMs. The sheet ‘BIM mutation details’ shows for each resistant bacteria details on the mutations detected, as well as the annotation of the protein produced by the mutated gene or the name of the gene when possible. As some mutations are observed in several bacteria, the sheet ‘Simplified combined mutations’ synthetically shows the presence/absence of mutations to the locus tag/gene level precision in each bacterial genotype. The effect of the mutation is shown with a color code, described below the table. Processed data: The following files are here as starting point to analyze data without having to process of raw data. Bacteria_data.csv Sequencing of bacteria populations: This file contains the frequency of each host genotype through time in all replicates. For the Genotypes, ‘start-end’ corresponds to the susceptible DGCC 7710. Any ‘PAM\_XXX’ between ‘start’ and ‘end’ indicates the presence of spacer XXX in this genotype. This pattern can be present several times for multi-resistant genotypes. Spacers are named according to the middle position of the corresponding protospacer in the phage. Phage_data.csv Sequencing of phage populations: This file contains the frequency of each phage mutation through time in all replicates. It contains the type, the position on the genome, the reference allele, the mutated allele and the frequency of each mutations. The column with time 0 do not show replicate number next to the treatment as the sequencing was done for the phage mix used at the beginning of all replicates for each treatment. Matching_data.csv Dynamics of phage mutations that escape CRISPR immunity: This file contains the host spacer frequency and the corresponding phage mutation frequency through time in all replicates. Spacers are named according to the middle position of the corresponding protospacer in the phage. Many lines contain frequencies of 0 as they were not filtered for computation purposes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.192 | 0.079 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".