Study Describes Research Scientists’ Information Seeking Behaviour, but Methodological Issues Make Usefulness as Evidence Debatable
Bibliographic record
Abstract
A Review of: Hemminger, B.M., Lu, D., Vaughan, K.T.L., & Adams, S. J. (2007). Information seeking behavior of academic scientists. Journal of the American Society for Information Science & Technology, 58(14), 2205-2225. Abstract Objective – To quantify the transition to electronic communication in information-seeking behaviour of academic scientists. Design – Census survey. Setting – University of North Carolina at Chapel Hill, a large public research university. Subjects – Nine hundred two faculty, research staff, and graduate students involved in research in basic or medical science departments. Participants self-selected (26%) from 3523 recruited. The sample reflected the larger population in terms of gender, age, university position, and department. Methods – The authors developed a web-based survey and delivered it via PHP Survey Tool. They developed the questions to parallel similar earlier studies to allow for comparative analysis. The survey included 28 main questions with some questions including further follow-up questions depending on the initial answer. The instrument included three initial questions designed to reveal the participant’s place and role in the university, and further coding classified participants’ department as either basic or medical science. The questions included categorical, continuous, and open-ended types. While most questions focused on the scientists’ information seeking behaviour, the three final open-ended questions asked about their opinions of the library and ideal searching environment. Answers were transferred into a MySQL database, then imported into SAS to generate simple descriptive statistics. Main Results – Participants reported easy access to online resources, and a strong preference for conducting research online, even when access to a physical library is convenient. Infrequent visits to the library predominantly took place to utilize materials not available online, although the third most common answer for visiting was to take advantage of the library building as a quiet reading space (14%). Additional questions revealed both type and specifics of most popular sources for research, preferred journals, current awareness tools, reasons for choice of journal for publication, and use of bibliographic management tools. Conclusion – Scientists prefer online tools for conducting library research, although specific contexts influence the preference, and online articles may be printed out for reading or annotation. The participants are taking advantage of the developing online arena, utilizing databases for research, as well as literature searching, access to journals and conference proceedings, and to keep abreast of current research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.089 | 0.189 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.008 | 0.013 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.004 | 0.007 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".