SDSS-IV from 2014 to 2016: A Detailed Demographic Comparison over Three Years
Bibliographic record
Abstract
Abstract The Sloan Digital Sky Survey (SDSS) is one of the largest international astronomy organizations. We present demographic data based on surveys of its members from 2014, 2015 and 2016, during the fourth phase of SDSS (SDSS-IV). We find about half of SDSS-IV collaboration members were based in North America, a quarter in Europe, and the remainder in Asia and Central and South America. Overall, 26%–36% are women (from 2014 to 2016), up to 2% report non-binary genders. 11%–14% report that they are racial or ethnic minorities where they live. The fraction of women drops with seniority, and is also lower among collaboration leadership. Men in SDSS-IV were more likely to report being in a leadership role, and for the role to be funded and formally recognized. SDSS-IV collaboration members are twice as likely to have a parent with a college degree, than the general population, and are ten times more likely to have a parent with a PhD. This trend is slightly enhanced for female collaboration members. Despite this, the fraction of first generation college students is significant (31%). This fraction increased among collaboration members who are racial or ethnic minorities (40%–50%), and decreased among women (15%–25%). SDSS-IV implemented many inclusive policies and established a dedicated committee, the Committee on INclusiveness in SDSS. More than 60% of the collaboration agree that the collaboration is inclusive; however, collaboration leadership more strongly agree with this than the general membership. In this paper, we explain these results in full, including the history of inclusive efforts in SDSS-IV. We conclude with a list of suggested recommendations based on our findings, which can be used to improve equity and inclusion in large astronomical collaborations, which we argue is not only moral, but will also optimize their scientific output.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".