A Probabilistic Reasoning Framework to Detect Fake News on Social Media
Bibliographic record
Abstract
The rapid growth of social media has made the Internet a critical platform for spreading misinformation, which shapes public opinion and harms society. Despite significant research in fake news detection, most probabilistic efforts rely heavily on Naive Bayes, with limited exploration of other probabilistic models. This paper introduces a Bayesian network (BN) modelling-based framework for fake news detection, offering a probabilistic estimate of the likelihood of news being false. Unlike binary classification, this approach reflects human decision-making by evaluating three key questions: “who”, “what”, and “when”. Each module corresponds to a specific feature set, enabling nuanced reasoning about the news's credibility. The framework is flexible, allowing adjustments through expert input, knowledge bases, or real-world data. We validate the approach using a semi-synthetic dataset containing features from news content, user behaviour, and social context. The results highlight the framework's capacity to leverage expert knowledge, providing a more reliable and adaptive solution compared to traditional classifiers. The BN-based method demonstrates enhanced robustness, positioning it as a promising tool for tackling misinformation in an evolving digital landscape.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.021 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.005 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".