Obecność polskich utworów literackich w zagranicznych bibliotekach cyfrowych
Bibliographic record
Abstract
Obecność polskich utworów literackich w zagranicznych bibliotekach cyfrowych Wprowadzenie Biblioteka cyfrowa to relatywnie nowy cyfrowy twór umożliwiający przechowywanie i udostępnianie wirtualnych treści, niezależnie od miejsca i czasu.Ich funkcjonowanie opiera się w całości na sieci Internet i rozwija równolegle z nim.Nie powinno więc dziwić, iż prototyp wirtualnej biblioteki -"Projekt Gutenberg", powstał już w 1971 roku.Zasady jego działania stały się wzorcem dla licznych naśladowców.Krok naprzód możliwy był dopiero po unowocześnieniu języka programowania HTML i popularyzacji samej technologii komputerowej.Spełnianie podstawowych funkcji bibliotecznych stało się możliwe od 1995 roku.Proces przemian trwał do ok.2002 r., kiedy to rozpoczęła się epoka bibliotek cyfrowych II generacji 1 .Od zarania historii bibliotek cyfrowych, główny trzon ich zasobu stanowiły teksty reprezentujące dziedzictwo kulturowe, w szczególności zaś europejski i amerykański kanon literacki.W niniejszym artykule podjęto próbę zbadania tego zjawiska w odniesieniu do polskiego kręgu kulturowego, czyli obecności polskiego kanonu literackiego w międzynarodowych bibliotekach cyfrowych.Główną ideą było zbadanie ich ilościowego rozkładu w głównych zagranicznych agregatorach treści oraz jakościowa analiza wytypowanych kolekcji.Istotnym problemem napotkanym przez autora na etapie założeń badawczych okazał się fakt niespójności i płynności pojęcia kanonu literackiego.Temat ten podjął Piotr Wilczek 2 który zauważył, że kanon zależy od doświadczeń osoby decydującej o jego kształcie.Inne ujęcie problemu przedstawił Stefan Chwin 3 .
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.007 | 0.012 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.014 | 0.007 |
| Open science | 0.000 | 0.004 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.007 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".