MétaCan
Menu
Back to cohort
Record W4415707048 · doi:10.1109/access.2025.3627187

Boosting Arabic Fake Reviews Detection by Integrating Textual and Metadata Features: A Transformer-Based Model

2025· article· en· W4415707048 on OpenAlexafffund
Ibrahim Amin, Ismail Fakhr, Mohamed Waleed Fakhr, Rasha Kashef

Bibliographic record

VenueIEEE Access · 2025
Typearticle
Languageen
FieldComputer Science
TopicSpam and Phishing Detection
Canadian institutionsToronto Metropolitan University
FundersShanghai University of Engineering and ScienceToronto Metropolitan University
KeywordsMetadataArabicBoosting (machine learning)Language modelVariety (cybernetics)Modern Standard ArabicUser-generated content

Abstract

fetched live from OpenAlex

Fake reviews present a significant threat to e-businesses and content providers, lowering consumer trust and damaging brand reputation. As such, the detection and prevention of fake reviews is essential for maintaining the integrity and success of e-businesses. On the other hand, the Arabic language presents unique challenges due to its complex linguistic structure and the wide variety of dialects spoken across different regions. However, the availability of Arabic datasets for fake review detection remains limited, where the available ones either suffer from small sample sizes or are translated from English to Modern Standard Arabic, failing to capture the natural, colloquial language typically used in reviews. Moreover, most Arabic fake reviews research has focused mainly on the textual content of the reviews and has not considered the metadata. Therefore, there has been no comprehensive research investigating the benefits and effects of integrating metadata features with textual content for classifying Arabic fake reviews. To this end, this paper is two-fold. Firstly, a balanced Egyptian Arabic dataset has been created, translated from the YelpZip English dataset using transformers, which includes the metadata. Secondly, a comprehensive and comparative study is conducted to investigate the effects of augmenting the textual content with the metadata features. Baseline experiments with textual content only and fine-tuned pre-trained Arabic BERT models achieved an F1-score of around 69%. Combining pre-trained Arabic language model embeddings with handcrafted metadata features significantly boosts performance, with the best-performing system achieving an F1-score of 77% without the user and product IDs as features and as high as 87% with the inclusion of user and product IDs.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.003
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.007
Threshold uncertainty score0.014

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.003
Meta-epidemiology (narrow)0.0020.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.001
Science and technology studies0.0000.000
Scholarly communication0.0010.002
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0020.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.030
GPT teacher head0.309
Teacher spread0.279 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2025
Admission routes2
Has abstractyes

Explore more

Same venueIEEE AccessSame topicSpam and Phishing DetectionFrench-language works237,207