MétaCan
Menu
Retour à la cohorte
Enregistrement W2559433214 · doi:10.5749/movingimage.16.1.0057

Who's Trending in 1910s American Cinema? Exploring ECHO and MHDL at Scale with Arclight

2016· article· en· W2559433214 sur OpenAlexaboutno aff
Derek Long, Eric Hoyt, Kevin Ponto, Tony Tran, Kit Hughes

Notice bibliographique

RevueThe Moving Image The Journal of the Association of Moving Image Archivists · 2016
Typearticle
Langueen
DomaineEconomics, Econometrics and Finance
ThématiqueCinema and Media Studies
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésEcho (communications protocol)Movie theaterScale (ratio)ArtArt historyComputer scienceCartographyGeographyComputer security

Résumé

récupéré en direct d'OpenAlex

Who’s Trending in 1910s American Cinema?Exploring ECHO and MHDL at Scale with Arclight Derek Long (bio), Eric Hoyt (bio), Kevin Ponto (bio), Tony Tran (bio), and Kit Hughes (bio) Over the last decade, digital humanities scholars have challenged a diverse set of academic disciplines to reevaluate traditional canons and histories through the use of computational methods. Whether we choose to call such approaches “distant reading” (as Franco Moretti does), “macroanalysis” (as per Matthew Jockers), or simply “text mining,” the basic intervention of these digital methods remains the same: presenting opportunities to harness the power of scale, shifting our attention away from a small number of canonical texts and testing long-held historical assumptions.1 What might such a “distant reading” approach look like for the history of American cinema in the 1910s? Given that so many of the most revealing original texts of that history—the films themselves—cannot be read at all, how might we perform a distant reading of early cinema texts that we do have in abundance, such as industry trade papers and fan magazines?2 In this article, we propose and test two interrelated methods of distant reading for the history of silent cinema: Scaled Entity Search (SES) and query comparison. We developed these approaches as part of Project Arclight, an initiative to create a web-based application allowing users to analyze millions of pages of digitally scanned film industry trade journals and magazines.3 Our initial inspiration came from a twenty-first-century platform: Twitter. If researchers can use algorithms to track celebrities “trending” on Twitter—becoming suddenly high profile and gaining an explosive increase in mentions on the service—could we imagine the 2 million digitized pages of the Media History Digital Library (MHDL) as a giant social media stream, itself rich with trending stars, directors, and other personnel? Given a large enough list of such personnel—all of the twenty thousand credited, named entities from more than thirty-five thousand filmographic records in the Early Cinema History Online (ECHO) data set, for example—we realized that it was absolutely possible to measure the “top trending” names of the film industry trade press in the 1910s. Even beyond trends, we realized that this work allowed us to analyze relationships—in particular, the relationships between filmworkers and the trade press and between workers’ representations in the press and their prolificacy (measured by number of films worked on). Our subsequent analysis revealed the outsized influence of certain entities and the curious underinfluence of others. The article is divided into three parts. First, we examine ECHO as a data set, attending to its strengths and limitations for historical research. We also describe the method by which it was processed and transformed to generate lists of unique historical entities (e.g., cinematographers’ names) to be searched at scale in the MHDL. In the second section, we present the results of our SES of those entities using Arclight and offer case studies that highlight both the affordances and the drawbacks of SES as a historiographical method to explore the trending of, and relationships between, [End Page 58] entities. Finally, we discuss the revelations of query comparison between ECHO and MHDL, showing how the interplay between these two large-scale corpora allows us to expand the canon of what Charlie Keil and Shelley Stamp have called the “transitional era” of film history.4 By shedding light on forgotten personnel and recontextualizing well-known stars, quantitative and computational approaches such as those offered by SES and the Arclight app offer a significant avenue for scaled research on the history of a period whose study has long been faced with the challenge of access to primary materials. Ultimately, we hope these approaches offer a roadmap for future research using these rich data sets. Our work has been possible thanks to the formal support of academic institutions and funding agencies as well as the informal support of film libraries, archives, and collecting communities. Arclight’s development and our research were enabled by a Digging into Data grant,5 sponsored by the US Institute for Museum and Library Services6 and Canada’s Social Sciences and Humanities Research Council.7 Our research...

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,003
score de la tête « metaresearch » (Gemma)0,001
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: Observationnel
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,093
Score d'incertitude au seuil0,356

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0030,001
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0010,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0010,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,014
Tête enseignante GPT0,209
Écart entre enseignants0,195 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeObservationnel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations5
Publié2016
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueThe Moving Image The Journal of the Association of Moving Image ArchivistsMême sujetCinema and Media StudiesTravaux en français237 207