Abstract WP412: Use of Interactive Cluster Heatmaps and Text Analytics in the Escape-na1 Phase 3 Clinical Trial for Pharmacovigilance of Adverse Events and Site Performance
Bibliographic record
Abstract
Introduction: Stroke trial subjects sustain numerous adverse events (AE’s) which may be related to the disease under study or to the therapeutic intervention. In studies involving pharmacovigilance of multiple AE’s, ongoing oversight of AEs by investigators and trial sponsors is critical to ensure the safety and well being of trial subjects. Static cluster heatmaps are used extensively in the field of gene expression studies and text analytics are used widely to monitor text on social media sites such as Twitter. Here we present novel dynamic tools using static cluster heatmaps and text analytics to aid with the visualization of AEs and of clinical site performance in AE reporting in the ESCAPE-NA-1 trial (Clinicaltrials.gov NCT 02930018). Methodology: In an interactive cluster heatmap display, the frequency of AEs is color coded in cells, with the rows corresponding to either sites or patients and columns correspond to unique AEs. Clustering is visualized on a heatmap by addition of row and column dendrograms. Quantification of AEs in the clinical trial is performed through text analytics and the results are visualized using a word cloud. Difference of AE across sites is visualized using a comparison cloud, a cloud that compares frequencies across sites. A cluster heatmap and word cloud are made interactive by formatting into HTML. Results: The main results obtained from interactive AE cluster heatmaps are patient ,site and AE clusters (Figure 1). The results obtained from text analytics are word clouds that convey overall AEs frequency and AEs frequency across sites (Figure 1). Due to active status of the trial, we will share the methodology of constructing interactive AE cluster heatmaps and word clouds and broad visual results without providing specifics of sites or AEs. Conclusion: Interactive cluster heatmaps and word clouds constructed via text analytics are a novel way of visualizing complex multidimensional AE data from clinical trials.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.022 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.044 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".