Beyond the Dataset: Understanding Sociotechnical Aspects of the Knowledge Discovery Process Among Modern Data Professionals
Bibliographic record
Abstract
Data professionals are among the most sought-out professionals in today’s industry. Although the skillsets and training can vary among these professionals, there is some consensus that a combination of technical and analytical skills is necessary. In fact, a growing number of dedicated undergraduate, graduate, and certificate programs are now offering such core skills to train modern data professionals. Despite the rapid growth of the data profession, we have few insights into what it is like to be a data professional on-the-job beyond having specific technical and analytical skills. We used the Knowledge Discovery Process (KDP) as a framework to understand the sociotechnical and collaborative challenges that data professionals face. We carried out 20 semi-structured interviews with data professionals across seven different domains. Our results indicate that KDP in practice is highly social, collaborative, and dependent on domain knowledge. To address the sociotechnical gap, the need for a translator within the KDP has emerged. The main contribution of this thesis is in providing empirical insights into the work of data professionals, highlighting the sociotechnical challenges that they face on the job. Also, we propose a new analytic approach to combine thematic analysis and cognitive work analysis (CWA) on the same dataset. Implications of this research will improve the productivity of data professionals and will have implications for designing future tools and training materials for the next generation of data professionals.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.002 |
| Open science | 0.009 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".