A transcriptomic dataset used to derive biomarkers of chemically induced histone deacetylase inhibition (HDACi) in human TK6 cells
Bibliographic record
Abstract
Transcriptomic biomarkers facilitate mode of action analysis of toxicants by detecting specific patterns of gene expression perturbations. We identified an 81-gene transcriptomic biomarker of histone deacetylase inhibitors (HDACi) using whole transcriptome data sets of TK6 human lymphoblastoid cells generated by Templated Oligo-Sequencing (TempO-Seq) after 4 h of exposure to 20 reference compounds (10 HDACi and 10 non-HDACi) [1]. The biomarker, named TGx-HDACi, was derived using the nearest shrunken centroid (NSC) method and can distinguish HDACi from non-HDACi compounds based on the expression pattern across the 81 genes. The classification capability of TGx-HDACi was evaluated by NSC probability analysis of 11 external validation compounds (4 HDACi and 7 non-HDACi) with a probability cut-off of 90%. Thus far, TGx-HDACi has demonstrated 100% accuracy in classifying the reference and validation compounds as HDACi or non-HDACi. Of the 81 TGx-HDACi genes, 19 genes are part of the S1500+ gene panel containing 2753 genes, developed for toxicological assessments [2]. Herein, we assessed the classification performance of the biomarker with this reduced gene set to determine if TGx-HDACi can be applied to analyze S1500+ gene expression profiles. The 20 reference compounds and 11 validation compounds were correctly classified as HDACi or non-HDACi by the NSC probability analysis, principal component analysis, and hierarchical clustering based on the expression of the 19 genes, demonstrating 100% accuracy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".