Enhancing Public Confidence: The GAO's Peer Review Experience: Even Auditors Need to Be Audited
Bibliographic record
Abstract
In 2004 we at the Government Accountability Office (GAO) arranged to have a multinational team of experienced performance auditors conduct the first-ever peer review of our performance audits of federal government programs (www.gao. gov/peerreviewrpt2005.pdf). We also hired KPMG LLP to conduct a peer review of our financial audit practice their fourth such engagement with us. Both audit teams concluded that during the period reviewed the GAO's quality assurance system was suitably designed and operating effectively to provide reasonable assurance of conforming to applicable professional standards. The international peer review team also said that other national audit offices may want to emulate several of our practices and made suggestions that further enhanced our practice and provided other significant benefits. According to the reviewers, [the] quality assurance system reinforces the GAO's independence, objectivity and reliability. These reviews inform Congress of and give the American people confidence in the quality of our financial and performance audits. UNDER THE MICROSCOPE A team of 16 experienced performance auditors from seven countries reviewed our performance audit practice. The Office of the Auditor General of Canada led the multinational team. Other participants included national audit offices in Australia, Mexico, the Netherlands, Norway, South Africa and Sweden. The KPMG team consisted of experienced financial audit partners and managers with extensive government financial auditing experience. The review teams focused on the elements of our quality assurance system dealing with engagement performance and compliance monitoring. They reviewed our audit policies and process controls, examined a representative sample of our 2004 audit engagement files and reports on government programs, and interviewed senior managers and staff responsible for selected engagements. They also evaluated our internal inspection program, including a representative sample of engagement files that our internal inspectors had examined in 2004 to determine whether their findings were supportable. The performance audit team followed government auditing standards (the Yellow Book) and conducted the review in a manner consistent with the code of ethics and standards issued by the International Organization of Supreme Audit Institutions. The review team's ultimate objective was to determine whether the GAO's system of quality controls provided reasonable assurance that our work is independent, objective and reliable. The financial audit team followed the applicable AICPA peer review standards as well as government auditing standards. GOOD MARKS A key organizational and operational benefit of the peer review was the reviewers' confirmation that some of our key overarching quality control procedures are global better practices (see How to Do It Better,). According to their report, those practices help ensure that we focus our efforts on factors that affect government performance, that we assign engagement resources according to risk, that we develop complete and reliable evidence, that our engagement teams have the guidance and diagnostic tools necessary to perform their work and that our reports are clear, persuasive and fair. The reviewers also said that the GAO could improve its performance audit practice by, for example, enhancing the transparency and efficiency of its quality assurance system and policies. We have implemented some of these recommendations, and we are testing others (see Dividends Earned,). A PLAN FOR EXCELLENCE The GAO's approach to quality assurance was based on applicable professional standards and the agency's core values of accountability, reliability and integrity and it ends with public dissemination of virtually all its products. We had already created a quality assurance framework (see the exhibit) to summarize the policies and procedures we use to ensure compliance with professional standards and our core values. …
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.009 |
| Open science | 0.004 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".