The potential use of social media and other internet-related data and communications for child maltreatment surveillance and epidemiological research: Scoping review and recommendations
Bibliographic record
Abstract
Collecting child maltreatment data is a complicated undertaking for many reasons. As a result, there is an interest by child maltreatment researchers to develop methodologies that allow for the triangulation of data sources. To better understand how social media and internet-based technologies could contribute to these approaches, we conducted a scoping review to provide an overview of social media and internet-based methodologies for health research, to report results of evaluation and validation research on these methods, and to highlight studies with potential relevance to child maltreatment research and surveillance. Many approaches were identified in the broad health literature; however, there has been limited application of these approaches to child maltreatment. The most common use was recruiting participants or engaging existing participants using online methods. From the broad health literature, social media and internet-based approaches to surveillance and epidemiologic research appear promising. Many of the approaches are relatively low cost and easy to implement without extensive infrastructure, but there are also a range of limitations for each method. Several methods have a mixed record of validation and sources of error in estimation are not yet understood or predictable. In addition to the problems relevant to other health outcomes, child maltreatment researchers face additional challenges, including the complex ethical issues associated with both internet-based and child maltreatment research. If these issues are adequately addressed, social media and internet-based technologies may be a promising approach to reducing some of the limitations in existing child maltreatment data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".