Code, camera, action: how software developers document and share program knowledge using YouTube
Bibliographic record
Abstract
Creating documentation is a challenging task in software engineering and most techniques involve the laborious and sometimes tedious job of writing text. This thesis explores an alternative to traditional text-based documentation, the screencast, which captures a developer’s screen while they narrate how a program or software tool works. This thesis presents a study investigating how developers produce and share developer-focused screencasts using the YouTube social platform. First, a set of development screencasts were identified and analyzed to determine how developers have adapted to the medium to meet the demands of development-related documen- tation needs. These videos raised questions regarding the techniques and strategies used for sharing software knowledge. Second, screencast producers were interviewed to understand their motivations for creating screencasts, and to uncover the perceived benefits and challenges in producing code-focused videos. From this study a theory was developed describing the techniques used by devel- opers in screencasts. This thesis also discusses YouTube’s role in the social developer ecosystem, and presents a list of best practices for future screencast creators. This work lays the groundwork for future studies exploring how screencasts can play a role in sharing software development knowledge.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.016 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.003 | 0.005 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".