Finding tools for data documenting racism and the Black experience internationally and guides for searching and using data with an antiracism lens
Bibliographic record
Abstract
In the wake of the murders of George Floyd, Breonna Taylor and Ahmaud Arbery in 2020, among other (on-going) acts of violence against African Americans, a group of IASSIST members felt compelled to gather materials that can help all of us better recognize, acknowledge and combat inherent racial bias. Their work led to the formation of the current IASSIST Anti-Racism Resources Action Group. This Group has three subgroups that have been working to compile a variety of resources that might otherwise be difficult to find, such as: (1) sources of data and datasets on a variety of topics that document racism and the Black experience internationally; (2) tools, articles and rubrics for building anti-racism into the process of working with data across the research lifecycle. This panel consists of members of the subgroup that is working to developing a guide for finding these race and race-related data and resources. We will discuss the challenges and considerations in finding data documenting racism and the Black experience internationally, as well discrimination based on indigenous, national, and cultural/ethnic origins. We will share strategies and examples of how to apply the strategies. We will answer questions about the finding tool and the related list of data sources and guides for working with an anti-racism lens. We may begin to explore how these tools might be adapted and expanded over time to address other types of discrimination, such as by migrant/refugee status, religion, gender identity, or sexuality." The panel welcome participants feedback on strategies and their knowledge and experience that can contribute to making these tools applicable internationally and inclusive.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.136 | 0.270 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.047 | 0.028 |
| Science and technology studies | 0.008 | 0.007 |
| Scholarly communication | 0.018 | 0.027 |
| Open science | 0.005 | 0.015 |
| Research integrity | 0.004 | 0.005 |
| Insufficient payload (model declined to judge) | 0.024 | 0.018 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".