Government Information Use by First-Year Undergraduate Students: A Citation Analysis
Bibliographic record
Abstract
Objective – The objective of the study was to investigate how first-year undergraduate students in a general education communication course engaged with government information sources in their academic research. The study examined the frequency, types, and access points of cited government information, as well as patterns in secondary citations and topic-based variation, to identify implications for library instruction, discovery systems, and collection strategies. Methods – For the study, the researchers analyzed citations from persuasive papers submitted by 136 students across 14 course sections. A total of 1,704 citations were reviewed, of which 124 were identified as government information sources. A classification scheme was developed to code citations by source type, government level, agency, and access point. Researchers also conducted a secondary citation analysis to identify where students referenced government-produced content through nongovernmental sources and categorized papers by topic to assess variation in government information use. Results – Government sources constituted 7.3% of all citations, with 45.3% of students citing at least one government source. Most cited materials came from U.S. federal agencies, particularly the Department of Health and Human Services and the U.S. Congress. Students predominantly accessed government sources through open Web sources, with minimal use of library databases and materials. The types of government sources most commonly cited were webpages, press releases, and reports. An additional 201 secondary citations referenced government information indirectly. Citation patterns varied by topic, with higher engagement in papers on government, immigration, and environmental issues. Conclusion – The findings suggest that even without explicit instruction or assignment requirements, undergraduate students demonstrated baseline awareness and independent use of government information sources. However, their reliance on open Web access and secondary references highlights gaps in discovery, evaluation, and access. Instructional support could enhance students’ ability to locate and critically engage with more complex and authoritative government documents. Beyond instruction, the findings inform strategies for enhancing discovery, improving visibility, and promoting balanced access to government information.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.042 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.019 | 0.022 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".