Making Visible the Invisible: Why Disability-Disaggregated Data is Vital to “Leave No-One Behind”
Bibliographic record
Abstract
People with disability make up approximately 15% of the world’s population and are, therefore, a major focus of the ‘leave no-one behind’ agenda. It is well known that people with disabilities face exclusion, particularly in low-income contexts, where 80% of people with disability live. Understanding the detail and causes of exclusion is crucial to achieving inclusion, but this cannot be done without good quality, comprehensive data. Against the background of the Convention for the Rights of Persons with Disabilities in 2006, and the advent of 2015’s 2030 Agenda for Sustainable Development there has never been a better time for the drive towards equality of inclusion for people with disability. Governments have laid out targets across seventeen Sustainable Development Goals (SDGs), with explicit references to people with disability. Good quality comprehensive disability data, however, is essential to measuring progress towards these targets and goals, and ultimately their success. It is commonly assumed that there is a lack of disability data, and development actors tend to attribute lack of data as the reason for failing to proactively plan for the inclusion of people with disabilities within their programming. However, it is an incorrect assumption that there is a lack of disability data. There is now a growing amount of disability data available. Disability, however, is a notoriously complex phenomenon, with definitions of disability varying across contexts, as well as variations in methodologies that are employed to measure it. Therefore, the body of disability data that does exist is not comprehensive, is often of low quality, and is lacking in comparability. The need for comprehensive, high quality disability data is an urgent priority bringing together a number of disability actors, with a concerted response underway. We argue here that enough data does exist and can be easily disaggregated as demonstrated by Leonard Cheshire’s Disability Data Portal and other studies using the Washington Group Question Sets developed by the Washington Group on Disability Statistics. Disaggregated data can improve planning and budgeting for reasonable accommodation to realise the human rights of people with disabilities. We know from existing evidence that disability data has the potential to drive improvements, allowing the monitoring and evaluation so essential to the success of the 2030 agenda of ‘leaving no-one behind’.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.087 | 0.268 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.008 |
| Science and technology studies | 0.006 | 0.025 |
| Scholarly communication | 0.021 | 0.043 |
| Open science | 0.005 | 0.015 |
| Research integrity | 0.007 | 0.012 |
| Insufficient payload (model declined to judge) | 0.014 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".