Bibliographic record
Abstract
As clinicians, all practitioners should review their practice. Audit is well mandated, at least for urologists; however, what does the public expect from us? Certainly, we are aware that plenty of other groups, be they government, regulators or others, are watching 1. Increasingly, beyond our control, patients may also rate our performance individually via social media with no recourse concerning ‘fake’ data 2. This is not the best form for patient-reported outcomes. However, there is little else out there to judge if we are at a benchmark, beyond or below it. So how do we control outcome data and present it in a manner that is ethical, confidential and timely? If we asked a patient if it was a good idea to audit results, present them and learn from them, they would probably call this ‘best practice’ 3. The Royal Australasian College of Surgeons deems this practice to be an audit, and it is mandatory for all surgeons. So why, if we want to look at outcomes in a hospital, does it become an expensive, red-tape-filled exercise in low-risk ethics? The answer is probably because of misunderstandings as to what is research vs true audit and ultimately the role of ethics committees. Another contributing factor is a desire by international journals for studies, even retrospective and really simply audit studies, to have a research and ethics approval ‘number’ prior to publication (although this is vague and may just mean ethical principles were followed). Countries such as Canada recognized this conundrum over a decade ago by having algorithm-driven ethics for low-risk endeavours that are approved immediately at no or minimal cost. The universities (and by default hospitals) offer this service as it leads to better practice and surprisingly and refreshingly, more publications. We accept without question that ethical practices must be followed and used to their full extent where appropriate, but the processes need to evolve rather than devolve to a situation discouraging analysis, reflection and ultimately better practice. An alternate view is that independent data are the only way forward; thus, some may argue that registries are the answer. Yes, registries are important in that they are independent – the Prostate Outcomes registries are an example 4, 5 – however, their scope is limited and it often takes considerable time to get results. We need further technology, resources and thoughts regarding how to move the situation so it is almost ‘real-time’. Gaps abound, but it is clinicians and not politicians or bureaucrats that should be leading this process. Zeps et al. 6 make a good point in the commentary in this edition of BJUI that ‘there is still no national or international coordinated data repository that facilitates identification of clinical variation that can be addressed through research’. Whilst this situation remains we must strive to obtain as many data as possible and to learn from them. Data are generally de-identified and very low risk, so is it really an ethical dilemma to attempt to improve clinical practice? Many would argue it is unethical not to be doing this as best practice. None declared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.228 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.004 | 0.009 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".