Real-time credible online health information inquiring: a novel search engine misinformation notifier extension (SEMiNExt) during COVID-19-like disease outbreak
Bibliographic record
Abstract
Abstract Public health-related misinformation spread rapidly in online networks, particularly, in social media during any disease outbreak. Misinformation of coronavirus disease 2019 (COVID-19) drug protocol or presentation of its treatment from untrusted sources have shown dramatic consequences on public health. Authorities are utilizing several surveillance tools to detect, and slow down the rapid misinformation spread online, still millions of misinformation are found online. However, there is no currently available tool for receiving real-time misinformation notification during online health or COVID-19 related inquiries. Our proposed novel combinational approach, where we have integrated machine learning techniques with novel search engine misinformation notifier extension (SEMiNExt), helps to understand which news or information is from unreliable sources in real-time. The extension filters the search results and shows notification beforehand; it is a new and unexplored approach to prevent the spread of misinformation. To validate the user query, SEMiNExt transfers the data to a machine learning algorithm or classifier which predicts the authenticity of the search inquiry and sends a binary decision as either true or false. The results show that the supervised learning algorithm works best when 80% of the data set have been used for training purpose. Also, 10-fold cross-validation demonstrate a maximum accuracy and F1-score of 84.3% and 84.1% respectively for the Decision Tree classifier while the K-nearest-neighbor (KNN) algorithm shows the least performance. The SEMiNExt approach has introduced the possibility to improve online health communication system by showing misinformation notifications in real-time which enables safer web-based searching while inquiring on health-related issues.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".