MétaCan
Menu
Back to cohort

Implementing EBLIP: if it works in Edmonton will it work in Newcastle?

2008· article· en· W2038571553 on OpenAlexaboutno aff
Andrew Booth

Bibliographic record

VenueHealth Information & Libraries Journal · 2008
Typearticle
Languageen
FieldHealth Professions
TopicHealth Policy Implementation Science
Canadian institutionsnot available
Fundersnot available
KeywordsInternal validityExternal validityContext (archaeology)Intervention (counseling)Evidence-based practiceCritical appraisalPsychologyInterpretation (philosophy)Computer scienceApplied psychologyKnowledge managementMedical educationMedicineSocial psychologyAlternative medicine

Abstract

fetched live from OpenAlex

Andrew Booth Recent years have seen a shift in emphasis away from production and interpretation of evidence towards implementation.1 Such a move is evident within the evidence-based practice movement in general and sessions at recent Cochrane Collaboration and Campbell Collaboration colloquia reflect this preoccupation. More specifically within evidence-based library and information practice (EBLIP), with the same evidence being discussed by practitioners from a richness of settings and cultures, attention has increasingly turned to what is to be done with the evidence once it has been appraised and interpreted. Previous features have highlighted difficulties in translating findings from research into practice. Critical appraisal focuses on the internal validity of a study. Put simply, internal validity means the ability of a study to establish that the changes measured in a study are a direct result of the intervention and only the result of that intervention.2 In contrast, external validity refers to the extent to which the results obtained are likely to apply to other similar programmes or approaches.2 Linking investigation to internal validity and implementation to external validity we identify four possible models: Ideally the evidence base will include both rigorous research studies to establish cause and effect and implementation studies to demonstrate applicability in other settings, including our own. Such situations are comparatively rare. Where a research project has previously established internal validity, an implementation study could then explore the impact of an intervention in a specific context. ‘Good practice’ or ‘innovation’ typically describe situations where assumptions are made about the likely effectiveness of an intervention, i.e. internal validity is taken as a ‘given’ and efforts concentrate, instead, on exploring implementation locally. This is the ‘evaluation bypass’ described in a previous feature.3 ‘Evaluation in innovation’, i.e. evaluation alongside innovation3–5 is where internal and external validity are explored simultaneously. This has the advantage of not putting too large a brake on progress but preventing the adoption of unevaluated practice. Previously I have drawn attention to one deficiency that affects both investigation and innovation—that is the ‘hero innovator’.6 Planning and conduct of a new research study, or for that matter introduction of good practice, requires considerable effort and commitment from both the project leader and affected staff. Implementation either makes the assumption that this is matched by corresponding effort and commitment from the team leader and staff at an implementation site, or assumes that the effect of this input on the overall effectiveness of the intervention is negligible. While one might not wish to go as far as I do (in suggesting that the entire effect of an intervention might be attributed to the innovator), clearly either of these alternative assumptions is potentially flawed. Another factor, more frequently referred to in the literature, is the ‘Hawthorne effect’.6 This explains why staff members who are closely monitored and controlled within a rigorous investigation may sustain a level of performance that would perhaps not be present were they not being observed. Once monitoring mechanisms are removed, so the argument goes, their performance returns to more ‘typical’ levels. A key consideration with regard to implementation of anything, but particularly where this relies on human activities, is the concept of ‘fidelity’. We tend to think that an intervention such as taking a tablet is fairly straightforward. The patient takes a tablet, the tablet works to the degree that it is able and the researcher measures the effect. However, is this really as straightforward as it seems? In the context of a clinical trial, a patient is usually provided with information regarding what the tablet is intended to do. What is being measured is therefore the effect of the tablet itself plus this information (and the context in which it is given). If we prescribe the same tablet in clinical practice, the quality of supporting information may be correspondingly poorer, compliance may be suboptimal and the effect may therefore be less pronounced than under experimental conditions. If this is so for a ‘mechanical’ procedure such as taking a tablet, how much more so for a human-mediated process such as information skills training or provision of information services? Recently, our research team has examined the effects of implementation fidelity on the outcomes of human resource management programmes and policies.7 This led me to consider the same issues in connection with the programmes, policies and services delivered by our libraries. For example, why might clinical librarian programmes in local NHS Trusts perform worse (or indeed better) than the published results of rigorously evaluated similar programmes? The conceptual framework that informed our thinking likely applies to EBLIP. Broadly speaking, implementation fidelity refers to the degree to which an intervention or programme is delivered as intended. An obvious analogy is to the fidelity of sound produced through a pair of speakers compared with how digital representation of that sound is encoded on a compact disc (CD). In this analogy, differences could be attributed to the quality of the disc itself, the laser that reads it, the amplifier that augments the sound, the speakers that produce the sound—even the quality of the hearing of the listener. What are the conceptual equivalents of these elements? In our framework, the equivalent of the CD itself is adherence. Adherence to a programme is similar to compliance with a drug. Four factors determine adherence. First is adherence to the content of a programme. Clearly, if a clinical librarian post is introduced specifically to support production of clinical guidelines, then its success is not directly comparable with that of a post that involves attendance at ward rounds. Although the ‘package’ rejoices in the title of ‘clinical librarian’, content of the package is demonstrably different. Second comes coverage—a programme that provides a specialist service to one or two clinical teams is not comparable with a more general service that supports clinicians across a wide range of clinical specialties. This may reflect generic differences in the intensity of support or opportunities to specialize in sources and terminology. However, differences may also exist between specialties, for example ophthalmology and surgery, which may themselves explain differences in success between a published evaluation and a service that, in all other respects, imitates the published model. The third factor relates to frequency. Clearly, a personalized service such as a clinical librarian has more opportunity for interaction with its users where involvement is weekly rather than monthly. At the other end, however, overfamiliarity may result in underutilization of the services and a greater degree of ‘downtime’. Just as clinical researchers calculate an optimal dose for a drug, we cannot assume that either the published frequency or the frequency of the implementation is already optimal. Differences between the two frequencies may explain away differences in effects. Finally comes the important variable of duration. Are we to expect that a 2-year clinical librarian implementation will perform better or worse than a published 1-year evaluation? As with frequency, we may assume that there is an optimal cut-off point between allowing time to demonstrate an effect and not allowing that effect to be dissipated or lost. This is particularly an issue whenever we upgrade an intervention from a time-limited project to a service that is to continue indefinitely. The framework that we have outlined has four other factors in addition to adherence.7 These are: intervention complexity; facilitation strategies; quality of delivery; participant responsiveness. These factors are known as ‘moderators’ because they are not necessarily a function of the package itself (compare the CD) but rather relate to the context in which it is implemented (compare the hi-fi equipment or the receptivity of the listener). First on this list is intervention complexity. On the one hand, delivering multi-faceted interventions (e.g. training, current awareness, literature search services) provides multiple routes by which a clinical librarian post might achieve effectiveness (‘many strings to the bow’, as it were). On the other hand, this may dissipate energies in delivering the interventions or create confusion in what the service is attempting to achieve. A related issue concerns whether multiple interventions are conceived with an underlying rationale (as a package) or simply thrown together opportunistically (as a bundle). In a worst-case scenario, they may act antagonistically, rather than synergistically. For example, training users to conduct their own literature searches might result in a decreasing demand for mediated literature searches. Of course, complexity may not relate simply to the different components of a service, but could equally refer to the complex nature of a single component, for example appraising the results of a literature search before sending them to a user. Next comes facilitation strategies. Within a team environment, as at the Vanderbilt University, it may be possible to accompany the induction of a new member of staff with targeted training. In addition, there may be detailed manuals, procedures or service specifications. Clearly, adherence to such standardization may result in a more uniform output, one more closely related to deliverables from a rigorous research project. However, where a clinical librarian sets up a service from scratch, there is the possibility of greater autonomy and variation from per protocol strictures. An added complication, of course, is that the addition of too many standards and procedures could itself result in a more complex intervention! Quality of delivery is a further area of variation. Those with experience of working with surgeons will recognize debates around volume and outcome (whereby a surgeon must achieve a certain threshold of procedures to achieve proficiency) and skill mix (differences in complexity of individual cases). To these we may add variation in practitioner experience. If a rigorous evaluation examines a clinical librarian with 12 years’ experience, but a post is implemented with a post holder with only 2 years’ practice, we would expect a discernible difference in proficiency, at least during progression up the ‘learning curve’. Finally, participant responsiveness may explain away differences in the uptake or impact of a service once implemented. One clinical team may be very keen to participate in a clinical librarian project, another may engage with the process grudgingly, if at all. Differing levels of engagement were found to result in differential impact between clinical teams in a within-study comparison.8 Surely such a difference in motivation will be even more detectable between an experimental study and a subsequent service implementation? Given the differences identified above, we may wonder how any implementation could ever expect to replicate the success of a rigorous evaluation. The answer, of course, is that, as with equities, our effect size can either go up or down. An implementation may start with more modest content, coverage and complexity than a formal evaluation. Its participants may prove more responsive and an absence of protocols or procedures may give the clinical librarian more scope for innovation or ‘intrapreneurialism’.9 This framework aims to identify possible sources of departure from the published study to explain differential impacts or outcomes. Of course, an implementation study may also require a broader range of evaluation outcomes than those employed in a tightly focused rigorous research study. Closely allied to the above discussion is the aforementioned issue of ‘Evaluation in innovation’, i.e. evaluation alongside innovation. Within health services research there is a specific research design called the n-of-1 study10. This allows the investigator to change specific variables during the conduct of the treatment (i.e. implementation) and to measure the effect of each change. Because each variable is being changed within the same implementation context, the investigator can gain a clearer idea of the impact of each change. Changes with a more marked effect provide a stronger ‘signal’ that is more likely to survive the ‘noise’ of implementation. While such approaches are rare in the information world, we could draw an analogy with our FOLIO e-learning course programme where, because courses follow thick and fast, we can change one variable between courses (e.g. size of buddy group) and instantly introduce questions evaluating the impact of such a change. Hopefully, the above discussion, using a framework tangentially related to EBLIP, will help library and information practitioners to identify why local initiatives may not prove as successful as published evaluations (although the effect may operate in either direction) and encourage managers to implement research-derived services more faithfully. Such differences can exist between Edmonton, Canada (in the Northern hemisphere) and Newcastle, Australia (in the Southern hemisphere). Equally, they can exist between Edmonton (in the South of England) and Newcastle (in the North of England). Indeed, as we have illustrated, differences may even exist within the same organization or library service. Tangible evidence, should we need it, of the wisdom of the proverb that ‘there's many a slip twixt cup and EBLIP’!

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.008
metaresearch head score (Gemma)0.002
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Science and technology studies, Scholarly communication, Insufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.321
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0080.002
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0010.002
Science and technology studies0.0030.000
Scholarly communication0.0000.014
Open science0.0000.000
Research integrity0.0000.002
Insufficient payload (model declined to judge)0.0030.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.414
GPT teacher head0.562
Teacher spread0.147 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations6
Published2008
Admission routes1
Has abstractyes

Explore more

Same venueHealth Information & Libraries JournalSame topicHealth Policy Implementation ScienceFrench-language works237,207