Making the most of data: Data skills training in English universities
Bibliographic record
Abstract
The collection and analysis of quantitative data is becoming increasingly important across a range of sectors.As business and research interest in data expands, so too does the demand for workers able to analyse and interpret datasets.The potential to utilise data hinges on the supply of skilled individuals.However, research suggests that employers are struggling to find suitable candidates for data roles. 1 In recognition of this challenge, the government asked Universities UK to 'review how data analytics skills are taught across different disciplines and assess whether more work is required to further embed these skills across disciplines.'This report aims to engage with both the immediate shortage of data analysts, and the need for greater data literacy.As organisations become more data driven there is a need for all workers to be able to interpret data and to undertake basic analysis.Taken together, this report and Nesta's report, Skills of the datavores: talent and the data revolution, set out a coherent picture of both the supply and the demand for data analysts and data-literate graduates.In a joint briefing statement in July 2015, Nesta and Universities UK will present findings, implications for policy makers, and recommendations. Findings1 Although the skills shortage is widely reported the skills that entry-level data analysts should have is not clearly set out.This has restricted positive action.In order to move forward these skills should be clearly set out, both by employers and by educators in their description of course content. 2The data skills shortage is not simply characterised by a lack of recruits with the right technical skills, but rather by a lack of recruits with the right combination of skills.The shortage of technical skills widely reported in the media is an over simplification of what is, in reality, a more complex issue.Employers report that there is a shortage of graduates with the right combination of skills.The combination of skills required includes a range of technical skills and domain knowledge, but also the ability to transform data outputs into something valuable to employers.3 Usually, a combination of technical skills is achieved through multi-disciplinary teams, with every team member possessing deep skills in several areas and basic knowledge in others.This shows that data skills needs cannot be boiled down to a simple list of skills that all undergraduates should acquire.Rather, there may be a number of core skills that should be shared by all members of a data team, and individual, specialist skills that may be developed in particular disciplines.The development of data teams emphasises the need for data analysts to possess strong teamwork and communication skills.4 There is no consistent method for identifying the extent of data analysis teaching within undergraduate programmes.A scheme to identify courses with significant data analysis components would provide valuable information to both prospective students and employers.5 Many undergraduate degree programmes teach the basic technical skills needed to understand and analyse data.Data can be gathered and analysed to enhance knowledge and understanding.This is largely reflected in undergraduate degree courses, where data analysis skills are taught across many programmes.This is also recognised by employers, who recruit data analysts from a range of subject areas, most commonly from those science, technology, engineering and mathematics (STEM) and social science courses where data analysis training is most prevalent and advanced.1 McKinsey Global Institute (2011) Big data: The next frontier for innovation, completion and productivity available at: http://www.mckinsey.com
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".