<scp>RNA</scp> editing in acute myeloid leukaemia with normal karyotype
Bibliographic record
Abstract
RNA editing is a post-transcriptional modification, the most common one being the deamination of adenosine into inosine (Galeano et al, 2012). Over the last few years, high throughput sequencing methods of DNA and RNA have enabled the prediction of RNA editing sites. In acute leukaemia, only a few studies have investigated the presence of editing events (Beghini et al, 2000; Ma et al, 2011). Furthermore, an edited form of protein tyrosine phosphatase (PTPN6) has been reported in acute myeloid leukaemia (AML) (Beghini et al, 2000). The present study aimed to identify, by bioinformatics analysis from RNA sequencing data, recurrent putative RNA editing sites (A to G conversion) in cases of AML with normal karyotype, a subgroup of leukaemia composed of tumours with relative genetic stability. The cohort comprised 40 patients classified as AML-M0 (n = 1), AML-M1 (n = 13), AML-M2 (n = 10), AML-M4 (n = 12), AML-M5 (n = 3) and AML-M6 (n = 1) with a median age of 67 years. Among them, 14 patients carried the fms-related tyrosine kinase 3 mutation (FLT3), 14 patients the nucleophosmin (NPM1) mutation; the remainder had no FLT3 or NPM1 mutations. Total RNA was extracted from fresh and thawed samples obtained after informed consent in accordance with the Declaration of Helsinki and stored at the Hémopathies Inserm Midi-Pyrénées (HIMIP) collection. Two human bone marrow CD33-positive cell samples from healthy donors (StemCell technologies, Vancouver, BC, Canada) were used as normal counterpart. RNA samples were sequenced using 91 paired-end strand-specific RNA sequencing, which generated around 80M of clean reads per sample. The alignment of reads was performed with Spliced Transcripts Alignment to a Reference (STAR; Dobin et al, 2013) against the hg19 reference genome. Only unique alignments were retained (mapping quality = 255). Mapping quality was then set to 40, duplicates were removed (samtools rmdup: http://samtools.sourceforge.net/samtools.shtml) and read group information added (PicardTools: http://picard.sourceforge.net/). Next, we proceeded to the variant calling pre-process step by realigning and recalibrating the alignments with Genome Analysis Toolkit (GATK; https://www.broadinstitute.org/gatk/). Variants for the 42 samples were detected using GATK UnifiedGenotyper tool with standard options and merged with vcf-merge (http://vcftools.sourceforge.net/index.html). The variants present in healthy samples were discarded with vcftools. We applied the following filters: minimum depth: 10 and minimum variant quality: 30. A Perl Script enabled us to extract the strand orientation of the variant in order to identify A to G editing sites. Sites with too many mismatches in the first ten bases of each read were discarded to avoid artificial mismatches derived from random-hexamer priming. Site variations occurring in less than 3 patients were also discarded. We then selected variants with A->G sites and applied a Minimum Allele Frequency filter >0·1. Finally, variant annotation was performed with Variant Effect Predictor (McLaren et al, 2010), and each variation was compared to the Single Nucleotide Polymorphism database (dbSNP; http://www.ncbi.nlm.nih.gov/snp) and reported into the final annotation file. For the bioinformatics analyses, we focused our attention on A to G conversion leading to non-synonymous coding changes on reference sequence. Results were matched with that of normal bone marrow CD33+ cells used as controls. We selected variations with minimal depth of 20 reads in at least one patient. Among the non-synonymous coding alterations we distinguished two groups: those already listed in the dbSNP database and those not yet described. Of note, the dbSNP database reports all of the variations described in the literature, including SNPs and mutations, but also variations found on RNA. Among the non-synonymous coding variations recorded in dbSNP, we excluded those corresponding to SNP and mutations acquired from genomic DNA sequences and obtained a list of 12 putative editing sites (Table 1). Nine of these variations were already reported in the RNA editing database RADAR (Rigorously Annotated Database of A-to-I RNA editing; http://rnaedit.com) (Ramaswami & Li, 2014). Interestingly, 3 positions had not been previously described and corresponded to putative RNA editing sites distributed on three distinct mRNA sequences: SEC24 family member B (SEC24B), zinc finger protein 12 (ZNF12) and SMC5-SMC6 complex localization factor 2 (SLF2, also termed FAM178A) (Table 1). In order to confirm that these three variations, observed on patients’ RNA, were real RNA editing sites and not polymorphisms or mutation sites, we verified the absence of these variations on the corresponding genomic DNA by polymerase chain reaction (Table 1, Fig 1), including three already described RNA editing sites, Cyclin I (CCNI), Insulin-like growth factor binding protein 7 (IGFBP7) and Cyclin-dependent kinase 13 (CDK13), as positive controls. As shown in Fig 1, this experiment confirms that they are real editing sites. We did not observe any association of RNA editing sites with NPM1 or FLT3 mutations. No edited form of PTPN6 was found but this can be explained by the fact that our series is restricted to cases with normal karyotype. Modification of proteins consecutive to the RNA editing process could play some role in leukaemogenesis. For example, the A to G conversion in CCNI mRNA, leading to an R/G substitution in the cyclin box of the protein, could have consequences on the regulation of cell cycle progression. The newly discovered RNA editing site on ZNF12 mRNA drives a Y to C amino acid change in one of the zinc finger C2H2 type domain and could change binding specificities of this transcription factor. Indeed, among the different genes with RNA editing sites, IGFBP7 is an interesting target. Taking into account that unedited IGFBP7 induces apoptosis of AML cells and synergizes with chemotherapy in suppression of leukaemia cell survival (Godfried Sie et al, 2012; Verhagen et al, 2014), it is tempting to suggest that the edited form significantly enhances this phenotype. Interestingly, some of the RNA editing sites found in our cohort have also been found in solid cancers, such as hepatocellular carcinoma [Coatomer protein complex, subunit alpha (COPA), U3 small nucleolar ribonucleoprotein, homolog C (UTP14C) and CCNI] (Chan et al, 2014; Hu et al, 2014). These findings require further experiments to test the impact of A-to-I conversions on the function of the corresponding proteins. This work was supported by grants from the ARC, Institut Universitaire de France, ANR (Labex Toucan). According to the French law, HIMIP collection has been declared to the Ministry of Higher Education and Research (DC 2008-307 collections 1) and obtained a transfer agreement (AC 2008-129) after approbation by ethical committees (Comité de Protection des Personnes Sud-Ouest et Outremer II and APHP ethical committee). Clinical and biological annotations have been declared to the CNIL (Comité National Informatique et Libertés). Cathy Quelen performed the experiments and wrote the paper. Yaelle Eloit performed some experiments, Céline Noirot performed bioinformatical analysis, Marina Bousquet designed the study and participated to the redaction of the paper and Pierre Brousset designed the study and wrote the paper. The authors declare no conflict of interest. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".