Systematics Compilations on Various Sponge Taxa
Bibliographic record
Abstract
Dataset Size:26 PDF files; a total of 12,864 pages, (17.6 GB in total) Scan details are located in the spreadsheet: 00b_HMR_Sponge_Taxa_Compilation.xlsx Organization: Each PDF file corresponds to one binder of information compiled by Dr. H.M. Reiswig. The entries are arranged alphabetically by taxon, most of which are genera. Most notes are updated, but it appears that Dr. Reiswig added commentary into a given taxon over years of study. There is no specific organization of information within each taxon, nor is the entire series arranged by taxonomic hierarchy. These binders were intended as an aide-memoire in easily located (alphabetical) sections. Where Dr. Reiswig refers to specimens that he examined in the companion dataset, "Notes, Illustrations and Annotations on Sponge Specimens", he uses the same HMR Code Numbers. Content: Over 187 sponge taxa are represented, including one order, 20 families, four subfamilies, and 162 genera. As he examined specimens, he retained his notes comparing specimens, consulting published descriptions and observations from museum visits. Information for each taxon is highly variable in detail and length. Some taxa are accompanied by a single page of basic systematics information, while others have extensive systematics information, drawings of whole mounts and spicules, measurements, and commentary on published information. In some cases, Dr. Reiswig has provided detailed anatomical comparisons between similar sponge species. Within some genera, Dr. Reiswig has provided systematic information to the species level; for some species, he refers to specimens examined in the first dataset (using the same HMR specimen codes).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.013 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.027 | 0.026 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.159 | 0.066 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".