Bibliographic record
Abstract
The Organized Crime and Corruption Reporting Project (OCCRP) came into possession of a secret dataset of property owners of the Palm Islands, the elite high-end artificial islands on the coast of Dubai.\nWith over 250 neighborhoods on Dubai’s waterfront, a group of journalists around the world has been investigating who these individuals are that can afford the posh and pricey real estate. While most fall into the uber-rich category, some also have corrupt to criminal backgrounds leading to questions such as if the Palm Islands are truly a real-estate paradise, or instead a refuge for the corrupt.\nThe task for each journalist was to dig up any leads of corruption, money laundering and criminal acts to find just exactly who can afford to be - and how they can afford to be - Palm Island property owners. I was given two different lists: a list of approximately 900 individuals and companies affiliated with the United States; and, a list of approximately 700 individuals and companies affiliated with Canada.\nFor each name in my two lists, I spent no more than 10 minutes backgrounding each person using: Google, Pipl and Spokeo, the OFAC sanctions list, LexisNexis clip search and LinkedIn. I would also search each person’s affiliated email address and company. Each week I wrote up my findings in story memos.\nOnce I found prospective money launderers or corrupt individuals, I began reporting out with a more extensive clip search, using the Public Access to Court Electronic Records (PACER) to look up court cases and making calls to potential sources.\nAmong the findings were: a businessman who committed health care fraud and fled to the United Arab Emirates after serving prison time; a CEO of global construction company with U.S. federal - including military - contracts in Afghanistan and Africa who was sued for breach of contract; a bank on the Office of Foreign Assets Control (OFAC) sanctions list whose executives slipped away -- all of whom own properties on the Palm Islands.\nLink to capstone project: https://www.occrp.org (direct link will be available March 2018) and http://www.nicolerothwell.com/reporters-notebook/.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.004 | 0.001 |
| Research integrity | 0.004 | 0.003 |
| Insufficient payload (model declined to judge) | 0.020 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".