Bibliographic record
Abstract
W e aRe pleased to present the proceedings of the IV Workshop on Information, Data, and Technology (WIDaT 2021).We initially planned this event to be held in Belo Horizonte, Brazil, hosted by the Federal Center for Technological Education of Minas Gerais (CEFET-MG).Still, due to the COVID-19 pandemic, we celebrated it virtually from 20-21 October 2021.In its fourth edition, we aimed at bringing together, through interdisciplinary approaches, researchers and other professionals oriented to data access and use coming from computer science, information science, engineering, mathematics, and related fields.The event also focused on promoting the approximation of research groups that work in the data and information field.We tried to discuss several dispersed initiatives, but of great potential for integration among themselves.Researchers, professors, professionals, and students from different areas of knowledge attended WIDaT 2021.We had the privilege to have presentations from national and international professionals with recognized experience in the event scope.Our attendees were quite international since we registered listeners from Brazil, Argentina, Colombia, Canada, Portugal, and Spain.Twenty-two full articles were accepted out of 38 submissions.For the event to be successful, several efforts were made by the Organizing Committee, mainly due to the difficult moment imposed by the COVID-19 pandemic.For the efforts made, we thank all authors, organizers, reviewers, and collaborators who contributed significantly to improving the quality of the presentations.Special thanks also go to the speakers for sharing their research results and experiences such as the listeners.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.091 | 0.042 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".