Monitoring Freshman College Experience Through Content Analysis of Tweets: Observational Study
Bibliographic record
Abstract
Background: Freshman experiences can greatly influence students’ success. Traditional methods of monitoring the freshman experience, such as conducting surveys, can be resource intensive and time consuming. Social media, such as Twitter, enable users to share their daily experiences. Thus, it may be possible to use Twitter to monitor students’ postsecondary experience. Objective: Our objectives were to (1) describe the proportion of content posted on Twitter by college students relating to academic studies, personal health, and social life throughout the semester; and (2) examine whether the proportion of content differed by demographics and during nonexam versus exam periods. Methods: Between October 5 and December 11, 2015, we collected tweets from 170 freshmen attending the University of California Los Angeles, California, USA, aged 18 to 20 years. We categorized the tweets into topics related to academic, personal health, and social life using keyword searches. Mann-Whitney U and Kruskal-Wallis H tests examined whether the content posted differed by sex, ethnicity, and major. The Friedman test determined whether the total number of tweets and percentage of tweets related to academic studies, personal health, and social life differed between nonexam (weeks 1-8) and final exam (weeks 9 and 10) periods. Results: Participants posted 24,421 tweets during the fall semester. Academic-related tweets (n=3433, 14.06%) were the most prevalent during the entire semester, compared with tweets related to personal health (n=2483, 10.17%) and social life (n=1646, 6.74%). The proportion of academic-related tweets increased during final-exam compared with nonexam periods (mean rank 68.9, mean 18%, standard error (SE) 0.1% vs mean rank 80.7, mean 21%, SE 0.2%; Z=–2.1, P=.04). Meanwhile, the proportion of tweets related to social life decreased during final exams compared with nonexam periods (mean rank 70.2, mean 5.4%, SE 0.01% vs mean rank 81.8, mean 7.4%, SE 0.01%; Z=–4.8, P.05). However, during the final-exam periods, the percentage of academic tweets was significantly lower among African Americans than whites (χ24=15.1, P=.004). The percentages of tweets related to academic studies, personal health, and social life were not significantly different between areas of study during nonexam and exam periods (P>.05). Conclusions: The results suggest that the number of tweets related to academic studies and social life fluctuates to reflect real-time events. Student’s ethnicity influenced the proportion of academic-related tweets posted. The findings from this study provide valuable information on the types of information that could be extracted from social media data. This information can be valuable for school administrators and researchers to improve students’ university experience.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".