Mining Twitter data to #educate the public about #sepsis
Bibliographic record
Abstract
IntroductionSepsis is not well known. Only 58% of Americans know the word sepsis, less than 1% can identify its symptoms and one-third wrongly believe the disease is contagious (1). What if social media educated users about sepsis? There are at least 500 million tweets worldwide per day on Twitter.
 Objectives and ApproachEarly detection of sepsis with early treatment is associated with a decrease in mortality (2). The current study aims to use Twitter to share sepsis patients’ experiences. The approach consisted of using text data mining techniques by randomly extracting tweets (N =150) with the hashtag #sepsissurvivor (3) using R software (4). The study retrieved and quantified sepsis patients’ tweets into word frequency distributions using documentation summarization and word cloud techniques (5) for visual representation of Twitter data. Sepsis patients used images symptoms cards (6) to raise awareness. The study used the R package "tesseract" to extract text from images (6).
 ResultsPatients sharing their experiences frequently used the word “sepsis.” Cardiorespiratory compromise (septic shock—the highest mortality risk) was illustrated in the words "my heart stops” or elevated "heart" rate or “low blood pressure.” Several studies have reported increased mortality associated with delays in antibiotic administration (7). Many sepsis survivors had antibiotics exposure, both in a timely manner or delayed in use. Sepsis patients experienced long stays in the hospital. Tweets mentioned "infection" 39 times (8), which supports patients’ diagnoses in addition to high rates of "fever." The clustering technique using word association indicated infection was highly correlated with sepsis (9). Sepsis survivors shared the "pain" they went through.
 Conclusion/ImplicationsTwitter presents an opportunity for patients to disseminate information about sepsis raising awareness about important symptoms. The information tweeted explores the impact of this diagnosis, and the need for early treatment. The current study demonstrates the opportunity to raise awareness through the learned experiences of patients in a novel medium.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.028 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.001 | 0.004 |
| Open science | 0.008 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".