Bibliographic record
Abstract
In the year 2012, a generation ago in digital technology, the person who generated the most internet searches in India was not a cricketer or a Bollywood star. Nor was it a politician or a religious figure. None of them were close. The person most Indians were curious about that year—as measured by the total number of Google searches—was Canadian-Indian Karenjit Kaur Vohra, a.k.a. Sunny Leone, a former porn star and Penthouse Pet of the Year. It wasn’t the case only in 2012. As hundreds of millions of Indians continued to discover the internet through 2013, 2014, 2015, 2016, and even 2017, Sunny Leone remained the most-searched-for person in India. People simply couldn’t get enough. (Prime Minister Narendra Modi made it to number two in 2014, the year he was elected, but Leone remained the clear favorite.) Prudish, conservative, family-values India . . . and a porn star? Leone was no longer even performing; she had stopped around 2010 and started her own production company with her husband and manager, Daniel Weber. In 2011, she came to India as a guest on the reality TV show Bigg Boss, a local version of the Big Brother franchise. Leone’s appearance was predictably controversial (by design, of course: it was good for the ratings). Although most Indians hadn’t heard of her, it didn’t take long for word to spread: “A porn star—from America—here in India?” At the time, parliamentarian Anurag Thakur complained to the Ministry of Information and Broadcasting, arguing that Leone’s presence on a nationally telecast program would “have a negative impact on the mindset of children.” Thakur added: “When children see these porn stars on TV and then do a Google search, it shows a vulgar site. It will have a bad impact in the long run.” There were no laws, however, to stop Leone from appearing on TV. While the production of pornography was officially illegal in India, Leone could justifiably argue she was no longer involved in the industry. She was trying to pivot to general entertainment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.007 | 0.009 |
| Open science | 0.000 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.220 | 0.090 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".