|
This is part 5 of my review of "Making good progress?" by Daisy Christodoulou. You can find the review index and my analysis of chapter 1 HERE. Chapter 5: Exam-based assessment In chapter 5, Christodoulou takes a closer look at exam based assessment and makes a fair case for the use of question level analysis to pinpoint student weaknesses. Of course, as she notes, this relies on checking of a good portion of the domain and therefore is not immediately applicable to subjects like history, where extended answers are more common and where pupils might only answer 5 questions rather than 55. Christodoulou also does a good job of exploring some of the issues of question setting and domain sampling which are inherent in all exams. She makes the important point that it is difficult to draw really reliable formative data from summative tests due to the broader focus of many exam questions. This would certainly be a useful lesson for some practitioners to learn when declaring the success or otherwise of their methods based on exam results alone. (In brief defence of exam boards, there are extensive guidelines on making exam texts accessible; whether or not these are followed however, is debateable). Overall, this is a useful summary of the valid use of question level analysis, but once again there is an underlying implication that question level analysis is not happening in schools around the country. Once again it feels that a study of actual classroom practice would have yielded more useful insights into how teachers might move forwards. More interesting is the suggestion that, whilst authentic tasks may provide some summative benefits, they make poor formative assessments. Christodoulou makes the case that formative exams should focus much more on the building blocks of the authentic tasks. Here I would also tend to agree. Formative assessments in history lessons tend to be the timeline activities, dates quizzes, sequencing and inference activities which become the building blocks for final summative assessments. However, the obsession with linear progression models has encouraged the use of inappropriate tests during teaching. However, I am less convinced that pupils’ historical writing would improve if they only focused on comprehension questions for example. If memory is the product of thought, then pupils need to engage more critically with what they read (and that is before we get onto the important issues surrounding motivation).
Where I do think this falls down a little is the suggestion that the more complex a task the less use it is formatively. Christodoulou gives many example of English exams (and I have a whole other rant about English language) butt these do not really reflect the kind of complex task related to history. In fact, some historical misconceptions might only begin to appear when applied to a complex task. It is difficult to assess a pupil’s understanding of the significance of the Renaissance until they begin to place it into wider context and develop their criteria for assessing it for example. Whilst I agree that many shorter, more specific formative tasks might aid in getting pupils to write this final piece, the final essay would still have a lot to reveal I think. In the final two sections of the chapter, Christodoulou explains why grades fail to provide useful formative information. This goes back to earlier worries about linear progression models and the incomparable nature of different exams. Here I find myself in complete agreement with Christodoulou on the limits of these grades and their pernicious effects on the curriculum. She also goes on to note the significant tensions between teachers and senior managers and begins to explore this power dynamic for the first time in any depth. Of course, her interpretation of the senior manager’s concern “is the test valid?” does not reflect my own experiences of the same concerns “does the number go up?” but that’s context for you. She also muses on the potential benefits and limitations of modular exams, concluding that they are worse than final summative ones. Again, I think this sets up modular and final exams as polar opposites, when a mixed methods approach might be a useful compromise. A system in which modular exams are conducted half termly and a final summative exam covers the whole domain, might allow for better triangulation of evidence. Some practical suggestions here or some relevant research might have been nice. This is part 4 of my review of "Making good progress?" by Daisy Christodoulou. You can find the review index and my analysis of chapter 1 HERE. Chapter 4: Descriptor-based assessment Chapter 4 begins with an overview of descriptor based assessment. This is by far and away the most common form of assessment used in schools today, in both its formative and summative uses. Christodoulou notes how the summative descriptors for Key Stage 3 rapidly expanded to become APP criteria, designed for formative use. She also notes how many schools’ post-levels solutions are also based on the generation of generic, linear descriptors. This is a topic I find quite interesting and have written on the subject HERE, HERE and HERE. Christodoulou goes on to explore the uses and limitations of descriptors as a tool for formative assessment. She argues that descriptors do not allow teachers to analyse the performance of students, or distinguish between “fleeting performance and genuine long-term learning” (p.85). Whilst I do take her point here, certainly in terms of KS3 level descriptors or GCSE bands, there seems to be a conflation of descriptors and generic descriptors going on here which is not fully acknowledged. I would argue that it is possible to create useful performance descriptors for individual assessments (an essay for example), and to tailor these for the purpose of assessing both summatively and formatively in that particular task. Indeed, Burnham and Brown have written on this theme in Teaching History 157. It is also possible, I believe, to use descriptors of gold-standard performance to tailor a descriptor-based mark scheme for formative purposes. In a sense, it is about whether the descriptors we use have been properly adapted for the purpose we want to use them for, as well as whether or not they were valid descriptors in the first instance. In the case of APP grids, I would say they answer to both of these questions is in the negative. Christodoulou then uses an example of how summative descriptors cannot help teachers to analyse the performance of their students. Her example of the need for vocabulary to make inferences is a good one and shows the limits of her example mark scheme in diagnosing the formative needs of the child. However, this also assumes that the teacher is only using the summative mark scheme to formatively assess the child in question. I am sure that some teachers only use the mark schemes to form their views on pupils’ progress, however I know that our teacher training course at Leeds Trinity tells trainees that this is a very poor method of assessing formatively. So once again, Christodoulou seems confuse the existence of a particular form of assessment with evidence of how teachers practice assessment on a day to day basis. What is missing here, indeed what this book is crying out for, is a detailed study of what teachers do in the classroom. Again, I suspect that the use of descriptors is more connected with a practical concern (driven by policy) to get the greatest possible advantage in public examinations. At this point I became quite interested in what happens in Ark schools, where Christodoulou serves as head of assessment. Trawling through the websites of all Ark secondaries, I managed to find 2 schools adopting Christodoulou’s suggestions and a large number still using KS3 levels. I did not get around to investigating Key Stage 4, but the use of summative, generic descriptors certainly seemed to still be in place in 4 schools, and APP grids in another. Most did not publish any information about assessment. This may not mean very much in the long run, but it is an interesting study in power dynamics and the ways in which schools respond to changing external demands. In the next part of the chapter, Christodoulou suggests that it might be more valid to test spelling and vocab separately from the final tasks. She glosses over the inherent sampling issue of this however and does not go on to discuss the phenomenon where students can accurately recall information but struggle to apply in in context – the very aspect which APP grids were meant to solve. I am in no way defending APP as an approach here, I tend to agree with Christodoulou that it encouraged lazy assessment based on unhelpful criteria. However, I also think this misuse was driven by that same misuse of power and accountability which I have discussed before, rather than genuine beliefs among teachers that the APP grid was anything more than hoop jumping. The remainder of the chapter deals with the issue of generic feedback, which I feel was covered suitably in Chapters 1 and 2. Here it feels that we are labouring the point by using extreme examples of practice. In the history example for instance, Christodoulou gives an example of using generic descriptors to offer feedback on an essay on the Battle of Hastings. However, I would contend that most teachers would check the knowledge of the essay as well, thereby overcoming much of the damage of the generic mark scheme. Even this might be overcome by following Burnham and Brown’s suggestions referred to above. In the further example given of a history exam on Stalin, Christodoulou suggests that such a question might quickly highlight misunderstandings. But this question actually opens up a continuum of answers for which multiple choice is probably not appropriate, as a case might be made for options A and C and a series of other potential options are missing, narrowing the scope of the potential answer. Would I use this as a short formative task, yes. Would I want it to be the only assessment of this, no! The final section of the chapter goes back to the issue of the validity of descriptors for summative assessment and raises a number of useful and relevant points on their limitations. I found this uncontroversial and fairly measured. However, in subjects like history it is very difficult to depart from descriptors without affecting the validity of that assessment for determining the historical understanding of the student. All written subject will need at least some qualitative aspect, so I am interested to see how Christodoulou deals with this in her next chapter.
OK, I'll admit that was blatant click bait (or a social experiment if you prefer). Over the last few days I have been publishing my thoughts on Daisy Christodoulo's book "Making Good Progress?" In the spirit of AfL, I have actually done more of an analysis than a summative review. However, noticed that engagements with these analyses were far lower than for other posts, despite the apparent popularity of the book.
So the test was this: which link would get the most clicks and from whom (results soon)? - the "book analysis" link with no grade? - the one giving the book a D grade? - the one giving the book 5 stars? I imagine that the love/hate links will prove to be most popular. I wonder if this links back to our natural desire for summative feedback (indeed this is one area which is a constant battle with kids when giving formative feedback). You will note that I have kept the headings deliberately ambiguous: 5 stars out of how many? And D as in GCSE or as in BTEC Distinction? This was very deliberate. Anyway, if you'd like my actual thoughts please follow the link HERE This is part 3 of my review of "Making good progress?" by Daisy Christodoulou. You can find the review index and my analysis of chapter 1 HERE.
Chapter 3: Making valid inferences In chapter 3, Christodoulou addresses the idea of the different purposes of assessment. This is a good introduction to the notion of how purpose affects what data we collect. The overview of validity and reliability is clear and cogent and serves as a useful guide to the teaching novice. In fact, I am seriously considering using this chapter with my PGCE students when we discuss assessment. The section on the trade-offs between reliability and validity in subjects like history is particularly helpful. Certainly it might help students question the merits of allowing students to take assessments home to finish for example. The second part of the chapter deals with the issue of formative assessment. Of particular note, and very important, is the point that formative assessments cannot legitimately use summative grading criteria when the conditions are so different. This is a very direct challenge to all those schools whose post-levels solution has been to bring in GCSE grading for every piece of work from KS3-4. This echoes much work done over the last 2 decades in history education, notably that of Lee, Shemilt, Counsell, Brown, Burnham, and many others. So far, I have found this the most useful chapter, possibly because the underlying idea of a clash of ideologies seems to be less evident here. This is part 2 of my review of "Making good progress?" by Daisy Christodoulou. You can find the review index and my analysis of chapter 1 HERE.
Chapter 2: Curriculum aims and teaching methods Chapter 2 opens with a focus on how the “generic skills” approach to teaching came to dominate. Christodoulou offers some good examples of the emptiness of approaches such as the RSA’s opening minds (although I wonder how many still teach this curriculum). She also uses Ofsted subject reports to show how generic approaches have been promoted. In many senses, this comes back to my own point about chapter 1, namely that I think the power of Ofsted, and their abuse of this power, has much to answer for in terms of promoting poor approaches to teaching. It is good to see this being acknowledged more fully here. However, it is worth noting that most of the examples Christodoulou cites come from several years ago now, especially the work from DfE level. I think there has already been a significant shift in the educational landscape and the “generic skills” voices have certainly lost much of their prior power and influence. A good deal of the chapter makes the case for deliberate practice as a means of targeted improvement. The example is given of a baseball player who might only get to run a particular play once in a game, but might ask for the pitch to be given over and over in practice. This is argument does make logical sense, although I wonder if it runs counter to some of the research set out in Brown et al.’s “Make it Stick”. In the book and using the same baseball analogy, the researchers suggest that deliberate practice is important, but that ultimately the moves need to be practiced in a game-like situation for the learning to fully embed. This does not discount Christodoulou’s point, but I wonder if we place so much emphasis on deliberate practice, that we might in turn encourage teachers to ignore the final outcomes altogether. Pianists may play scales and football players practise drills, but at the end of the day they all play regular concerts or games too. Christodoulou’s also makes the assertion, later in the chapter, that learning does not need to involve significant effort, also seems to run counter to Willingham’s research suggesting that memory is the product of thought: the more thought happening the better the memory. I think her point here may be about cognitive overload, but I would be interested to know more about this debate. I am also interested in Chrostodoulou’s complete rejection of “authentic tasks”. Whilst the cognitive science (and to be frank common sense) supports the notion that students cannot be asked to think through problems for which they have limited knowledge, there is a point in every person’s academic life when they must bridge the gap between knowledge acquisition and knowledge generation. I would hold that it is also important to inculcate pupils in the methods of a discipline as well as its core content. This is especially true of history where the “core content” is completely limitless. Although Christodoulou does not address this point directly, I think there is a disconnect between the view shared by many in the “traditionalist” school, that subjects should not be “dumbed down,” and at the same time, holding the same children off from advancing in their knowledge of their subject as an academic discipline. Christodoulou’s observations on self- and peer- assessment were quite interesting. I was partly expecting a full rejection of these in favour of teacher led assessment. However, Christodoulou seems to suggest that peer- and self- assessment are vital components of teaching and learning. At this point I have found myself in general agreement with most of the rest of the chapter. Christodoulou makes a convincing case for deliberate practice. I am sure such practices are embedded already in many classrooms, even if they are not seen more widely during Ofsted inspections. I do wonder however if there is an implication that direct instruction and deliberate practice were once done more effectively (before the advent of “generic skills” for example) and if so, if there is any reliable evidence to support this notion. Certainly, in terms of history, Cannadine’s investigation into 100 years of the history curriculum would suggest not. Chapter 1: Why didn’t assessment for learning transform our schools?
This chapter gives a very useful summary of the key difference between assessment for learning and assessment of learning. Christodoulou suggests that the perversion of AfL is partially to do with Ofsted and the DfE confusing the two types of assessment. However, she goes on to assert that the main reason why AfL has failed is that many teachers subscribe to a generic skills model of education and therefore believe that assessment for learning and of learning are essentially the same. She argues that teachers need to see improvement as the process of deliberate practice. Therefore, to get better at writing an essay about the First World War, they might engage in answering short questions from a textbook, mastering the chronology etc. I completely agree with Christodoulou's point here, and I certainly think that some aspects of ITT have contributed unhealthily to this over the years. The growing use and perversion of Key Stage 3 levels also contributed to this problem (as Wiliam notes in the preface). However, I think that her implication that the majority of teachers are believers in the "skill based" model of education misses the mark somewhat. Indeed, I could not see any concrete research supporting the notion that the majority of teachers have bought into the "skills" model. From my experience working in schools, most history teachers I know engage in deliberate practice when getting their pupils to improve at history (and I am yet to meet a maths teacher who doesn’t believe in deliberate practice). Where this breaks down, and where I do tend to agree more with Christodoulou’s assertion, is at GCSE. Here the temptation is to practice exam questions or “exam skills” over and over to the detriment of other aspects of deliberate practice. Unlike Christodoulou, I would argue that many teachers have been put in a situation where they have accepted (or feel they have to accept) the generic “skills model”, even when it runs contrary to how they might prefer to teach. This, in my opinion has been driven much more significantly by Ofsted than Christodoulou suggests. Many senior leaders have responded to Ofsted pressures by reducing the professional freedoms of their staff and pushing “generic skills” models of teaching. I can list scores of whole school initiatives which have shoved generic skills and assessment of learning to the fore in schools – notably those which have replaced KS3 levels with generic criteria from the GCSE mark schemes. In many ways I would argue that this is connected to the promotion of effective “leaders” over those whose educational pedagogies might have been more sound. These directives get the support of a small core of people who also subscribe to the “generic skills” model, and so a hegemony is created. Many teachers who are uneasy with this shift either do not have the confidence to challenge such directives from above, or lack the rigorous training and professional knowledge to offer a reasoned challenge (an issue of how teachers are trained on which I have written before). I also think exam boards have played a major role in a way that Christodoulou does not acknowledge here either. The simple fact is that deliberate practice and generic skills have often been synonymous at GCSE. Mark schemes in history have hitherto demanded that students provide 2 points on one side, 2 points on another, and a conclusion, for example. The increasing shift towards limited examinations training has meant that knowledge in such exams is rarely taken fully into account where exam structures are followed. Therefore the pragmatic classroom practitioner teaches “exam skills” in full knowledge of the fact that these are not synonymous with teaching history (or English, or science, or whatever). Again, there are some who miss this distinction, but it is notable that many history teachers see GCSE as an odd deviation from proper history teaching at Key Stages 3 and 5. In essence, I agree completely with Christodoulou’s concerns and her analysis of the problem, however I think that she paints the reason for this problem as one of educational aims in stark blank and white. In reality I would suggest this is a multifaceted, three-dimensional sculpture including significant aspects of power and control, mixed with aspects of pragmatism, ideology, idealised leadership, ignorance, and wrong-headedness. Of course, at this juncture, I accept that she may well cover these other aspects in the following chapters. |
Key FilesArchives
November 2025
Categories |
RSS Feed