Founders Online: Early Access: Reflections on Open Access, Crowd Sourcing, and Metadata Standards
Bibliographic record
Abstract
Founders Online, a digital initiative of the National Historical Publications and Records Commission (NHPRC) of the U.S. National Archives, launched in June 2013. Since its debut, the site has attracted over a million visitors interested in learning more about the creation of the United States of America in the words of six of its Founding Fathers. Founders Online contains 177,000 letters or other writings of these men and their contemporaries. Widely used by academics and the general public, the site has demonstrated the value of digital humanities’ emphasis on free access. As a former assistant editor at Documents Compass, a program of the Virginia Foundation of the Humanities, I served as a project manager on the Early Access portion of the project. We worked directly with the staffs of the currently active Founding Fathers documentary editing projects to make preliminary versions of unpublished documents available for early viewing on Founders Online. These Early Access documents will eventually be replaced by fully vetted and annotated versions to be completed later by the documentary editing projects. Relying on a large staff of over thirty people, we transcribed or proofread over 50,000 Early Access documents from 2012 to 2015. My Early Access experience demonstrated the need to give employees constant feedback, to reward them for good work, and to encourage specialization among project staff. My experience also reemphasized the need for unified metadata standards when aggregating different sets of data from multiple projects into a single digital platform. Founders Online, une initiative numérique de la National Historical Publications and Records Commission (NHPRC) des Archives nationales des États-Unis, lancée en juin 2013. Depuis ses débuts, le site a attiré plus d’un million de visiteurs intéressés à en apprendre davantage au sujet de la création des États-Unis d’Amérique d’après six des pères fondateurs. Founders Online contient 177,000 lettres ou autres écrits de ces hommes et de leurs contemporains. Largement utilisé par les universitaires et le public en général, le site a démontré la valeur de l’emphase des humanités numériques sur le libre accès. En tant qu’ancien rédacteur en chef adjoint à Documents Compass, un programme de la Virginia Foundation of the Humanities, j’ai travaillé comme gestionnaire de projet pour la partie d’accès anticipé du projet. Nous avons travaillé directement avec les membres du personnel des projets de montage documentaire de Founding Fathers actifs à l’heure actuelle, pour rendre disponibles en accès anticipé des versions préliminaires de documents non publiés sur Founders Online. Ces documents en accès anticipé seront éventuellement remplacés par des versions entièrement approuvées et annotées qui seront complétées plus tard par les projets de montage documentaire. Comptant sur un personnel nombreux de plus de trente personnes, nous avons transcrit ou relu plus de 50,000 documents d’accès anticipé entre 2012 et 2015. Mon expérience de l’accès anticipé a démontré le besoin de donner aux employés une rétroaction constante, de les récompenser pour leur bon travail, et d’encourager la spécialisation parmi le personnel du projet. Mon expérience a de plus souligné davantage le besoin de normes de métadonnées communes en transposant différents ensembles de données de projets multiples en une plateforme numérique unique. Mots-clés: Founders Online; Libre accès; transcription; métadonnées; externalisation à grande échelle; Histoire numérique
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.202 | 0.216 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.006 |
| Science and technology studies | 0.034 | 0.069 |
| Scholarly communication | 0.045 | 0.077 |
| Open science | 0.008 | 0.036 |
| Research integrity | 0.015 | 0.023 |
| Insufficient payload (model declined to judge) | 0.009 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".