Illuminating the dark web market of fraudulent identity documents and personal information: An international and Australian perspective
Bibliographic record
Abstract
From the beginnings of Silk Road in 2011, anonymous online marketplaces have continued to grow despite the best efforts of law enforcement. While these ever-present marketplaces remain flooded with illicit drugs and related paraphernalia, the sale and distribution of fraudulent identity documents remains a persistent problem, with these items consistently appearing for sale on both the open and dark web. While fraudulent Australian documents are some of the most popular products for sale, there is still much that is unknown about the Australian criminal market and its place within anonymous online marketplaces. Given the success of previous research in understanding the illicit drug trade through examining these marketplaces, this work examines two markets to gain an understanding of where Australian document fraud sits within this digital ecosystem. Two anonymous online marketplaces were crawled across 2020 and 2021, White House Market (WHM), and Empire Market. This data was extracted and examined to identify trends within both the international online market and the online market specifically for Australian documents, both of which have been relatively underexplored in the online space. To help illuminate the features of the market, the types of documents for sale, supply and demand trends, and trafficking flows along with vendor-related trends (e.g. product diversification and presence across markets) were examined. Each market was examined individually and then, where possible, comparisons were drawn to gain a more holistic understanding of the online fraudulent document market, with a specific focus on Australian products. Results indicate that, while the fraudulent document portion of the market is small, it is diverse, with numerous different identity-related products for sale, the most common being driver's licences from the United States (U.S.) and Australia, with digital documents dominating the whole marketplace. Overall, the most popular U.S. products were those that could be used to facilitate identity fraud, with the most popular Australian products being driver's licences and ID packs, likely linked to the presence of the 100-point identity check system used in Australia. This study demonstrates that anonymous online marketplaces have thus far been under-utilised in the study of the fraudulent document market, and that to properly understand the illicit market for fraudulent documents and personal information both the online and physical sides of the market should be considered. This information, if properly utilised, can improve the current understanding of this persistent criminal environment, building on previous research and assisting policymakers in making informed decisions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.011 | 0.012 |
| Science and technology studies | 0.003 | 0.005 |
| Scholarly communication | 0.012 | 0.016 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".