Sampling
Of 33 GPs and 25 appraisers who were identified as eligible, 12 doctors (36%) and 12 appraisers (48%) participated; 16 participants were male. Participating appraisees worked with an average of five other GPs in their practice (range 2–9). Nine doctors had discussed their multisource feedback data in appraisal and all the appraisers had discussed multisource feedback data with one or more appraisees. The age, sex, and summary performance scores for GP participants, and the wider sample of GPs who undertook the GMC multisource feedback process within the two localities, are shown in Table 1.
Table 1 Age, sex, and summary performance scores of 12 GPs contributing to this study and 116 GPs completing GMC multisource feedback processes in the same sites
Data collection
Interviews were completed between November 2009 and February 2010; each interview lasted 30–40 minutes. All participants gave permission for the interview to be recorded and transcribed.
Twelve participants completed a feedback form commenting on this study's findings. Most were satisfied that the summary accurately represented their views and experiences that were expressed at interview. Seven responders provided additional comments on the summary (data not provided) which were jointly considered by the academic lead and two researchers on this study. After checking transcripts, it was agreed that these comments represented diversity within this study's overarching themes. No amendments were made to the summary and the authors' interpretation of the interview data was unchanged.
Two overarching themes were identified:
the utility of multisource feedback in appraisal; and
linking revalidation, multisource feedback, and appraisal systems.
Utility of multisource feedback in appraisal
This theme reflected on the benefits (or otherwise) of using multisource feedback within appraisal. It comprised five sub-themes:
benefits of incorporating multisource feedback in appraisal;
validity of the colleague questionnaire;
validity of the patient questionnaire;
difficulty understanding and accepting ‘low’ benchmark scores;
desire for qualitative feedback.
Benefits of multisource feedback in appraisal
All appraisers were supportive of using some form of multisource feedback to enable GPs to reflect on their practice, and all the doctors viewed multisource feedback as providing a useful opportunity to learn and develop. Some had used GMC multisource feedback data to change their practice:
‘… the bit where I came into the lower quartile … was the record keeping. So one of the things I've done since is that we have voice-activated software where we can dictate into our notes and I actually make much more comprehensive notes … I'm quite up-front and honest and I said, “This is the bit, and this is what I thought about it, and this is what I've done as a result.” To which he [the appraiser] said, “Excellent, that's what appraisals are about; that's what feedback is about”.’ (GP 107)
Although participants were supportive of using multisource feedback to guide learning and development, most expressed some concerns about elements of the GMC multisource feedback methodology, which might limit its use in appraisal. These will be dealt with in the following sections.
Validity of the GMC colleague questionnaire
Participants were concerned about the extent to which colleagues' questionnaire responses accurately reflected a doctor's performance; some colleagues might be unfamiliar with the doctor's clinical practice, making it impossible to provide accurate and informed feedback:
‘I found the hardest part of it [GMC multisource feedback] [was] nominating 20 individuals … many of them would be unable to comment meaningfully on my clinical abilities.’ (GP 113)
Some interviewees expressed concerns that doctors identify colleagues rather than this being done on an independent basis. Interviewees felt that, intentionally or unintentionally, this may influence the feedback given:
‘It's basically very flawed, partly because you're going to select people [colleagues] that are going to give the answers that you want. Maybe not deliberately but because they're the people that you know. So it's not going to give a very true picture, necessarily, of your performance.’ (GP 108)
A few interviewees were concerned also that some colleagues may not complete the questionnaire honestly for fear that their negative (albeit truthful) feedback would have a detrimental impact on their colleague and/or undermine working relations:
‘I was chatting about this … at another surgery, and they said, “Oh yeah. I filled in ‘excellent’ for all my colleagues even though I don't think they are … because I don't want them to feel dispirited because I've got to work with them”.’ (GP 109)
Validity of the GMC patient questionnaire
The extent to which patient responses are a valid reflection of a doctor's performance was raised. Some interviewees commented that unfavourable feedback might be received simply because of the patient group with whom doctors work:
‘If you were in an area where there's an immigrant population with language problems, people who aren't literate, you're going to score very badly in this because the people who are coming in won't be the sort of people who are trying to be nice about you, don't care what they're doing, or they won't understand, and you'll be judged badly, not because your practice is poor but because your practice population is different.’ (GP 108)
One participant believed that patient feedback might be a reaction to the ‘interpersonal style’ of the doctor rather than their skills, so a doctor who is well liked by patients may receive very positive feedback.
Difficulty understanding and accepting ‘low’ benchmark scores
Some doctors and appraisers found it difficult to understand the GMC benchmark data, stating it was unclear if a low benchmark score represented real cause for concern or if, as absolute differences in scores between the upper and lower quartiles were small, those differences had marginal importance:
‘So quite what do those things [absolute scores and the quartile in which the score is located] mean? Does that mean that there was a real issue? Did somebody spot some subtlety … that it would have been useful for me to know about?’ (GP 106)
Some participants reported that they, or a doctor whom they had appraised who ranked in the two quartiles below the sample mean (three or four) in one or more areas of their performance, had been very distressed:
‘Of course, what happens is you get the Summary of Evaluation Results thing and you see “Benchmark Performance Band” and you see a “3” and you think, “They think I'm terrible!” In fact, I stopped reading it. I felt annoyed. I felt peeved. I actually put it away, I just thought, “I can't read this now”.’ (GP 109)
Desire for qualitative feedback
Some interviewees reported that the qualitative feedback was useful to help doctors understand the reasons why they had been scored as they had:
‘The comments are probably as useful or more useful than the scores, really, because they do give you a little bit more flesh on the bones of the scores to look at what's behind them.’ (GP 108)
Narrative comments, particularly those provided by patients, were seen as a rare source of direct encouragement:
‘The free-text feedback was quite helpful. I think when you're working away at the coalface, you actually get very little praise for what you do in, often, quite difficult circumstances. And to actually see in writing that people are grateful for what you've done and grateful for the way you treat them, is actually something that I hadn't actually had in 20 years of being a doctor, so it's actually quite a boost really.’ (GP 110)
Some interviewees reported that comments provided a clear indication of where someone could improve performance:
‘That's the whole point of these issues, that you actually look at something that you thought was not an area of concern and to allow comment and change your practice … I've shared it with the team and we're looking at how we might be more aware of that and what we might change”.’(GP 103)
Linking revalidation, multisource feedback, and appraisal processes
Interviewees reflected on both the link between appraisal and revalidation, and the role of multisource feedback in appraisal and for revalidation. Participants' views were mainly regarding multisource feedback in general, that is, multisource feedback data collection using any appropriate survey instrument. A few participants expressed views that specifically focused on the GMC multisource feedback methods.
The likelihood of a link being established between appraisal and revalidation was widely accepted:
‘I think it is a good way to do revalidation rather than have a test. So I think, I can only speak for our areas where our appraisal has really been quite — we've been in the forefront of appraisal and we have quite a good system in this particular area and people have done quite well. So, it does seem to work in our area. I don't know whether it will work in every area … but I think it's a good thing”.’ (Appraiser 209)
Participants reported that multisource feedback could make a positive contribution to revalidation:
‘I think it is good, politically, that the nation knows that it is being asked what it thinks of doctors, and that it knows that also our colleagues are asked to tell us what they think of us.’ (Appraiser 208)
However, some expressed concerns that surveys may not be able to identify doctors who are performing poorly or those deliberately concealing dangerous performance:
‘I'm just a bit afraid that we're all going to spend half our lives filling in forms about each other and there isn't very robust evidence that they're going to be useful in catching out people with problems. Lots of people will have said this but, you know, Harold Shipman could probably find 20 people to say nice things about him.’ (GP 108)
One participant suggested that if multisource feedback is used for revalidation, it may deter responders from providing critical — albeit honest — feedback. Many interviewees perceived tensions between the formative role of appraisal (including multisource feedback) and the summative function of revalidation. If appraisal data are used as evidence for revalidation, some felt this would inhibit doctors from openly exploring difficulties or professional limitations:
‘In terms of being an appraiser, I'm absolutely certain that, in a sense, the formative value of appraisal is going to go down as a result of revalidation. People are going to be extremely careful what significant events they log and stuff like that.’ (Appraiser 212)
One appraiser was already advising appraisees to exercise caution about what to write in appraisal documentation:
‘I'm advising people not to put anything in [appraisal documentation] that they might find a bit exposing, because the questions are a bit intrusive. Because, at the moment, as far as I can see, if there is a concern about the doctor they can subpoena the appraisal documentation, and I don't want to put anybody in a position where something that was written down … could be used in evidence against them.’ (Appraiser 209)
However, another participant reflected that doctors should be prepared to receive honest, constructive feedback to improve their clinical practice:
‘This is about people looking at their clinical practice, evaluating themselves, imposing change themselves, rather than it being a very judgmental process. Now, yes, somebody's going to look at it and ensure that your practice is safe and you keep up to date, look at what you're doing. I personally think that's only fair and reasonable. We're put in a huge position of trust and respect by patients and I think that's something that we don't deserve. We earn it by demonstrating things like this.’ (GP 107)