MétaCan
Menu
Back to cohort
Record W4411236837 · doi:10.1016/j.htct.2025.103847

Multi-cohort gene expression model enhances prognostic stratification in diffuse large B-cell lymphoma

2025· letter· en· W4411236837 on OpenAlexaff
Valbert Oliveira Costa Filho, Felipe Pantoja Mesquita, Erick Figueiredo Saldanha, Pedro Robson Costa Passos, Mariana Macambira Noronha, Sílvia Helena Barem Rabenhorst

Bibliographic record

VenueHematology Transfusion and Cell Therapy · 2025
Typeletter
Languageen
FieldMedicine
TopicLymphoma Diagnosis and Treatment
Canadian institutionsPrincess Margaret Cancer Centre
Fundersnot available
KeywordsMedicineStratification (seeds)LymphomaRisk stratificationDiffuse large B-cell lymphomaCohortPrognostic modelOncologyExpression (computer science)Internal medicineComputational biologyCancer researchOverall survivalBiologyComputer science

Abstract

fetched live from OpenAlex

Multi-cohort gene expression model enhances prognostic stratification in diffuse large B-cell lymphomaDear Editor, Diffuse large B-cell lymphoma (DLBCL), the most common type of lymphoma, in most cases is marked by significant heterogeneity and aggressive clinical behavior.While standard chemotherapy often achieves initial responses, these are short-lived, and resistance and relapse are frequent challenges. 1 Traditionally, risk stratification has relied on clinical tools, including the International Prognostic Index (IPI) and its variation.2 However, molecular stratification is promising to predict outcomes with greater accuracy, though gene-based approaches are still preliminary.3 Progress in this field is hindered by limited sample sizes and the substantial intra-and inter-regional variability of DLBCL.4,5 Consequently, large-scale studies are essential to refine risk stratification and optimize patient outcomes.This study aimed to establish a prognostic gene expression signature for patients with DLBCL based on tumor transcriptome patterns.To achieve this, we analyzed transcriptome and survival data from 11 diverse cohorts worldwide.Given the variability in RNA sequencing or microarray platforms across the 11 datasets, we focused on the genes common to all datasets, resulting in a panel of 11,425 genes.Detailed information regarding the datasets can be found in Supplementary Table 1.Due to platform-specific differences in scale, the gene expression values were transformed into z-scores.Datasets with fewer than 100 patients were combined into a cohort referred to as the Merged Cohort.In total, six cohorts were used in this study: the National Cancer Institute Cohort (GSE10846), University of York Cohort (GSE181063), University of York II Cohort (GSE32918), Univer-sit€ atsmedizin Berlin Cohort (GSE4475), University of Leeds Cohort (GSE69053), and the Merged Cohort (GSE69053, E_TABM_346, GSE11318, GSE21846, GSE23501, GSE57611, and TCGA-DLBC).For each cohort, a univariate Cox regression was performed employing all genes in the panel, identifying those with a p-value <0.05 as prognostic.Genes were defined as core prognostic genes (CPGs) if they consistently predicted either favorable prognosis in at least 5 out of 6 cohorts or unfavorable prognosis in at least 5 out of 6 cohorts, with no conflicting outcomes.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.005
metaresearch head score (Gemma)0.013
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Other · Consensus signal: none
Teacher disagreement score0.005
Threshold uncertainty score0.024

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0050.013
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0000.001
Science and technology studies0.0010.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.017
GPT teacher head0.262
Teacher spread0.245 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueHematology Transfusion and Cell TherapySame topicLymphoma Diagnosis and TreatmentFrench-language works237,207