Recovering a Future Artifact: The Winnifred Eaton Archive and a Genealogy of Care
Bibliographic record
Abstract
Recovering a Future ArtifactThe Winnifred Eaton Archive and a Genealogy of Care Sydney Lines and Joey Takeda The Winnifred Eaton Archive (WEA)1 is a digital repository that seeks to digitize, transcribe, and edit the wide-ranging oeuvre of Winnifred Eaton Reeve, an early Canadian author and American literary celebrity; the first Chinese American—and Chinese Canadian—novelist; an early Hollywood screenwriter; and, most infamously, an ethnic impostor. Though born to a Chinese mother (Achuen Amoy) and a British father (Edward Eaton), Eaton was most popularly known throughout the early twentieth century as the best-selling half-Japanese American novelist Onoto Watanna, an identity she successfully marketed for decades until returning to Canada in 1917, where she established herself as Winnifred Reeve, a Canadian literary luminary.2 Led by Mary Chapman at the University of British Columbia, the WEA has so far collected over three hundred texts by Eaton, two hundred of which have been transcribed for a total of over 1.3 million transcribed words; the archive includes dozens of uncataloged articles from Eaton's life in Alberta and many previously unpublished or uncredited screenplays produced during her career as a screenwriter during Hollywood's Golden Age. While version 1.0 of the Winnifred Eaton Archive was officially launched in September 2020, the project's earliest manifestation had appeared nearly twenty years earlier as Jean Lee Cole's Winnifred Eaton Digital Archive (WEDA). Initiated by Cole following the publication of her book The Literary Voices of Winnifred Eaton (2002), the WEDA featured nearly three dozen newly recovered texts alongside supplementary biographical and editorial material and was hosted by the University of Virginia's Electronic Text Center (UVA EText Center). In Traces of the Old, Uses of the New (2015), Amy Earhart [End Page 233] identifies the WEDA as emblematic of the early history of "digital literary recovery projects," noting that it serves as an example of both a "successful recovery collaboration" between institutional centers and individual scholars (74) and the now-all-too-familiar history of "haphazard" digital preservation (78). When the UVA EText Center was decommissioned in 2010, the standalone WEDA project was dismantled, and Eaton's texts (without accompanying images or Cole's critical paratexts) were subsequently ingested into UVA's digital repositories as individual entries. This effectively dissolved the WEDA and, as Earhart notes, left Eaton's "recovered body of work obscured once more" (79). As of 2023 the UVA catalog continues to host Eaton's texts, but the original WEDA is nearly impossible to find, and search results outside of UVA's VIRGO catalog trigger redirects to a partially and imperfectly archived version in the Internet Archive; as Cole reflected during the launch of the WEA, the WEDA became, like so many digital recovery projects, "another ghost on the Wayback Machine" ("Recovering Onoto Watanna" 2:45–2:50). As our work with the WEA has made clear, the disappearance and spectral afterlife of the WEDA serve not just as a parable for the challenges of creating sustainable infrastructure but also as a productive occasion for recognizing the collaborative and generative forms of labor involved in recovering, repairing, and rebuilding a digital archive. In considering the relationship between the WEDA and the WEA, it has become clear that the latter is not an iteration or new edition of the former but instead something of a re-recovery project. This is an important distinction, we think, as it recognizes both continuity and discontinuity between the two; the WEA, while a direct descendant of the WEDA, is its own edition and necessarily engages with the politics of recovering both Eaton's oeuvre and the original WEDA. Indeed, just as the work of recovery (both analog and digital) is, as Brigitte Fielder explains in "Recovery," an ongoing process that entangles practical, theoretical, and ideological concerns "that extend beyond textual location and inclusion" (18), the work of "re-recovery," as we outline below, is more than an exercise in excavating and rescuing digital artifacts. It may be better understood as an act of what Steven J. Jackson terms "repair": "the subtle acts of care by which order and meaning in complex sociotechnical systems are maintained and transformed, human value...
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.020 | 0.020 |
| Scholarly communication | 0.011 | 0.009 |
| Open science | 0.002 | 0.007 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.018 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".