There Had to Be a Better Way: John Nitti and Julianne Nyhan
Notice bibliographique
Résumé
This oral history conversation was carried out via Skype on 17 October 2013 at 18:00 GMT. Nitti was provided with the core questions in advance of the interview. He recalls that his first encounter with computing came about when a fellow PhD student asked him to visit the campus computing facility of the University of Wisconsin-Madison, where a new concordancing programme had recently been made available via the campus mainframe, the UNIVAC. He found the computing that he encountered there rather primitive: input was in uppercase letters only and via a keypunch machine. Nevertheless, the possibility of using computing in research stuck with him and when his mentor Professor Lloyd Kasten agreed that the Old Spanish Dictionary project should be computerised, Nitti set to work. He won his first significant NEH grant c.1972; up to that point (and, where necessary, continuing for some years after) Kasten cheerfully financed out of his own pocket some of the technology that Nitti adapted to the project. In this interview Nitti gives a fascinating insight into his dissatisfaction with both the state and provision of the computing that he encountered, especially during the 1970s and early 1980s. He describes how he circumvented such problems not only via his innovative use of technology but also through the many collaborations he developed with the commercial and professional sectors. As well as describing how he and Kasten set up the Hispanic Seminary of Medieval Studies he also mentions less formal processes of knowledge dissemination, for example, his so-called lecture ‘roadshow’ in the USA and Canada where he demonstrated the technologies used on the dictionary project to colleagues in other universities. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,003 | 0,010 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,019 | 0,013 |
| Communication savante | 0,010 | 0,011 |
| Science ouverte | 0,001 | 0,006 |
| Intégrité de la recherche | 0,003 | 0,012 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,011 | 0,004 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».