Bibliographic record
Abstract
Publicly available social data has been adoptedwidely to explore language of crowds and leverage themin real world problem predictions. In microblogs, usersextensively share information about their moods, topics ofinterests, and social events which provide ideal data resourcefor many applications. We also study footprints of socialproblems in Twitter data. Hidden topics identified fromTwitter content are utilized to predict crime trend. Since ourproblem has a sequential order, extracting meaningful patternsinvolves temporal analysis. Prediction model requiresto address information evolution, in which data are morerelated when they are close in time rather than further apart. The study has been presented into two steps: firstly, a temporaltopic detection model is introduced to infer predictivehidden topics. The model builds a dynamic vocabulary todetect emerged topics. Topics are compared over time to havediversity and novelty in each time consideration. Secondly, apredictive model is proposed which utilizes identified temporaltopics to predict crime trend in prospective timeframe. The model does not suffer from lack of available learningexamples. Learning examples are annotated with knowledgeinferred from the trend. The experiments have revealed, temporal topic detection outperforms static topic modelingwhen dealing with sequential data. Topics are more diversewhen are inferred in different time slices. In general, theresults indicate temporal topics have a strong correlationwith crime index changes. Predictability is high in somespecific crime types and could be variant depending on theincidents. The study provides insight into the correlation oflanguage and real world problems and impacts of social datain providing predictive indicators.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".