← Back to Blog
Research

Rosenshine, Hattie and Wiliam: What Broke Through

Teaching has an unusual relationship with its own evidence base. Thousands of education research papers are published every year. Almost none of them make any discernible impact on what happens in classrooms. The profession is simultaneously over-researched and under-informed — awash in studies that never leave the journals they were published in, while a tiny handful of publications become the common currency of staffroom conversation, CPD sessions, and school improvement plans.

The interesting question is not what the research says. There are plenty of summaries for that. The interesting question is why certain pieces of research — and not others — broke through. Why did Black and Wiliam's paper on formative assessment, published in a practitioner magazine in 1998, reshape how an entire country thought about feedback? Why did Hattie's meta-synthesis, dense with statistical methodology most teachers never engage with, become the single most cited text in school leadership? Why did a short paper by Barak Rosenshine, published in an American teachers' union magazine in 2012, become the dominant framework for lesson planning in British schools six years after it was written — and a decade after the author had retired?

These are not accidents. Each of these publications succeeded because it arrived at the right moment, in the right form, and was carried into schools by the right people. Understanding why they broke through matters, because it tells us something about the conditions under which research changes practice — and something about what gets distorted in the process.

Black and Wiliam: the paper that launched a thousand traffic lights

Paul Black and Dylan Wiliam published "Inside the Black Box: Raising Standards through Classroom Assessment" in 1998 in Phi Delta Kappan — not a peer-reviewed academic journal, but a practitioner-facing magazine for American educators. The choice of publication was deliberate. Black and Wiliam were synthesising a substantial body of research on formative assessment, but they were writing for teachers and school leaders, not for other researchers. The paper was short, accessible, and built around a compelling metaphor: the classroom as a black box, where inputs (curriculum, resources, policy) and outputs (attainment, grades, exam results) were visible to everyone, but what happened inside — the interaction between teaching and learning — was opaque.

The argument was precise. Formative assessment — the process of gathering evidence about pupil learning during instruction and using it to adapt what happens next — had effect sizes that placed it among the most powerful interventions available. But most classrooms were not doing it. Teachers were assessing constantly, but they were assessing summatively — testing what had been learned rather than using assessment to shape what was being learned. The feedback pupils received was often evaluative (grades, marks, ticks) rather than instructive (what to do differently). The assessment-learning loop was broken.

What made "Inside the Black Box" break through was not the novelty of the research — much of it had been published years earlier — but the timing and the framing. In England, the paper arrived into a policy context that was hungry for it. The National Strategies era was in full swing. Ofsted was looking for evidence of effective teaching. School leaders needed something concrete to point to when explaining what good practice looked like. Black and Wiliam gave them a framework — and, crucially, a vocabulary. "Assessment for Learning" became the phrase that launched a decade of CPD.

The tragedy is what happened next. The research was clear, nuanced, and demanding. The implementation was often none of these things. Assessment for Learning became traffic lights. It became "two stars and a wish." It became learning objectives on the board in child-friendly language and self-assessment smiley faces on worksheets and peer marking with green pens. None of these techniques appeared in Black and Wiliam's paper. They were the accretions of a professional development industry that took a sophisticated research finding and reduced it to a set of procedures that could be observed in a twenty-minute lesson observation and ticked on a proforma.

Wiliam himself has been candid about this. His subsequent work — "Embedded Formative Assessment" (2011), his long-running argument that feedback should produce a cognitive rather than emotional response, his insistence that formative assessment is not a set of techniques but a fundamental reorientation of the relationship between teaching and learning — can be read, in part, as a sustained attempt to rescue his own research from what the profession had done with it. The irony of "Inside the Black Box" is that it succeeded too well. The phrase "Assessment for Learning" became so ubiquitous that it ceased to mean anything specific. Every school in the country claimed to be doing it. Very few were doing what Black and Wiliam had actually described.

There is a lesson in this that extends well beyond formative assessment. When research breaks through into practice, it is inevitably simplified. The question is whether the simplification preserves the core insight or destroys it. In the case of AfL, the core insight — that assessment should be used to adapt instruction, not just to measure outcomes — was frequently lost beneath the procedural overlay. Teachers were doing AfL activities without doing AfL. The distinction between the activity and the purpose is the distinction that gets lost in translation, every time.

Hattie: the number that changed everything

John Hattie's "Visible Learning" was published in 2009 and it did something that no previous education research publication had done: it gave school leaders a number. Effect size 0.73 for feedback. Effect size 0.69 for metacognition. Effect size 0.40 for cooperative learning. Effect size -0.34 for repeating a year. The numbers were drawn from over 800 meta-analyses covering millions of pupils, and they were presented as a single ranked list — the most ambitious attempt ever made to quantify what works in education.

This section includes an interactive chart — open the page in a browser to explore it.

The impact was immediate and enormous. School leaders who had never previously engaged with meta-analysis could now walk into a meeting and cite, with apparent authority, that feedback was the most powerful intervention available and that reducing class size was a waste of money. The effect size became the universal currency of education discourse. CPD providers used it to justify their programmes. Ofsted inspectors used it to frame their expectations. Improvement plans across the country were suddenly peppered with references to Hattie's barometer, and to d=0.40 as though it were a threshold to clear rather than the reference point it was offered as.

What made "Visible Learning" break through was partly its scale — the sheer ambition of synthesising that much research into a single volume was impossible to ignore — and partly its form. Hattie did something genuinely valuable: he took the overwhelming complexity of educational research and made it usable. A busy headteacher could scan the ranked list in ten minutes and take something concrete into a report by lunchtime. That accessibility was not a weakness — it was the point. It gave a profession that was being asked to justify everything it did a shared vocabulary grounded in the best available evidence, and it did so with a clarity that no prior publication had managed.

The criticisms, when they came, were serious and they were methodological. Simpson (2017) argued that standardised effect sizes are shaped as much by the design of a study as by the power of the intervention it tests — by how tightly the sample is drawn, by what the comparison group was doing, by how narrowly the outcome was measured. On that argument, effect sizes drawn from different meta-analyses could not meaningfully be ranked on a single scale: 0.73 for feedback and 0.73 for direct instruction were not the same quantity. It is a critique worth taking seriously, and it lands on a technique the whole field uses rather than on one author. It is also worth saying that Hattie's own position was always more careful than the way the barometer came to be read. He offered 0.40 as a hinge point for interpretation, not a pass mark, and he has consistently resisted the league-table reading his audience imposed on him. The averaging tells you something real — feedback matters — while concealing what it depends on. Individual feedback studies ranged from strongly negative to strongly positive, and whether feedback helped or hindered turned on its type, its context, and the learner receiving it. A single number was never meant to be the whole story, and Hattie said so before his critics did.

None of this stopped "Visible Learning" from becoming the most influential education publication of the century so far. And the reason, I think, is that it met a need that the profession had not previously been able to articulate: the need for a shared reference point. Before Hattie, conversations about what works in education were largely impressionistic. After Hattie there was a common language — imperfect, often oversimplified in the retelling, but common, where there had been none at all. The effect size table gave people who disagreed about everything else a shared frame within which to disagree. That is not nothing. It is most of what a field needs in order to start making progress.

What I take from Hattie's work is the principle at its core: that teaching practices have differential effects, that some produce more learning than others, and that it is possible to observe, name, and measure what teachers do in classrooms in a way that connects to outcomes. The individual numbers are contested and context-dependent — Hattie has said as much himself — but the barometer settled the prior question, which was whether we should be looking at all. That conviction, that what happens in classrooms is worth examining systematically rather than left to professional intuition, is what "Visible Learning" embedded in the culture of school improvement. It is a substantial contribution, and the field is still drawing on it.

Rosenshine: the slowest revolution in education

Barak Rosenshine published "Principles of Instruction: Research-Based Strategies That All Teachers Should Know" in the Spring 2012 issue of American Educator, the magazine of the American Federation of Teachers. It was twelve pages long, plainly written, and summarised research that Rosenshine had been conducting and synthesising since the 1970s. It contained ten principles — daily review, small steps, questioning, models, guided practice, checking for understanding, high success rate, scaffolding, independent practice, weekly and monthly review — each grounded in the convergence of three research traditions: cognitive science, classroom observation studies, and research on effective tutoring.

The paper did not immediately reshape anything. In the United States, where it was published, it attracted modest attention. In the United Kingdom, where it would eventually become one of the most influential documents in contemporary education, it was virtually unknown for six years.

The Rosenshine phenomenon in Britain is, as far as I can tell, largely the work of a small number of people — most prominently Tom Sherrington, whose blog posts and subsequent book on the principles brought them to a mass audience, and Oliver Caviglioli, whose visual poster condensed the ten principles into a single page that could be printed, laminated, and stuck on a staffroom wall. The poster became ubiquitous. It appeared in every school I visited between 2019 and 2022. It appeared in NQT folders, in department handbooks, in CPD presentation slides, and — with a frequency that Rosenshine himself might have found ironic — in lesson observation feedback proformas.

Why did it break through? Several reasons, I think, operating together. First, the principles were simple without being simplistic. They described teaching practices that experienced teachers recognised immediately from their own classrooms — daily review, modelling, checking for understanding — but gave those practices a name, a research base, and a rationale. The experience of reading Rosenshine for the first time was, for many teachers, the experience of seeing things they already did described with a precision and authority they had not previously encountered. That recognition is powerful. It validates practice at the same time as it sharpens it.

Second, the principles arrived into a British context that was ready for them. The late 2010s saw a significant shift in the dominant discourse around teaching in England — a movement away from the progressive, child-centred pedagogies that had been influential since the National Curriculum era and towards a more explicit, knowledge-rich, teacher-led approach. Writers like Daisy Christodoulou, David Didau, and Sherrington himself had been arguing for years that direct instruction had been unfairly maligned, that teacher-led explanation was not the enemy of deep learning, and that the research evidence overwhelmingly favoured structured, explicit teaching over discovery-based approaches. Rosenshine's principles were the research base that this movement had been waiting for — a concise, accessible, eminently citable summary of decades of evidence showing that the most effective teachers taught explicitly, checked frequently, and scaffolded systematically.

In Scotland, the reception was slightly different. Curriculum for Excellence — with its emphasis on skills, interdisciplinary learning, and learner agency — sat uncomfortably alongside the directive clarity of Rosenshine's principles. There was, and to some extent still is, a tension between the CfE rhetoric of learner-led, experiential education and the Rosenshine evidence for teacher-led, structured instruction. Most Scottish teachers I have spoken to resolve this tension pragmatically: they use Rosenshine's principles for the instructional core of their practice and CfE's language for the documentation that surrounds it. Whether this represents a productive synthesis or a concealed contradiction is a question that Scottish education has not yet confronted honestly.

Third, and perhaps most importantly, the principles were operational. They told teachers what to do. Not in the abstract — not "use evidence-based strategies" — but specifically: begin with review, present in small steps, ask questions, provide models, guide practice. Each principle was a verb. Each could be enacted in a classroom the following morning. The gap between understanding the research and changing your practice was, for once, almost nonexistent.

What got lost

Each of these publications changed something real about how teachers think and how schools operate. But each was also, in the process of being adopted, distorted in predictable ways.

Black and Wiliam's formative assessment became a set of observable techniques detached from the underlying principle. Schools could demonstrate that they were "doing AfL" without any evidence that they were using assessment to adapt instruction — because the observation systems that monitored compliance could see a traffic light activity but could not see whether the teacher changed what they did next as a result of the information it produced.

Hattie's effect sizes became a shopping list. School leaders cherry-picked strategies based on their position in the ranked table without engaging with the conditions under which those effect sizes were achieved. Feedback was adopted as a priority in every improvement plan in the country — but the feedback that was actually given in classrooms remained overwhelmingly generic praise, because the distinction between effective and ineffective feedback requires a granularity that "d=0.73" does not provide.

Rosenshine's principles became a lesson observation checklist. The poster went on the wall. The ten principles went on the proforma. Observers walked into rooms looking for evidence of daily review, small steps, and checking for understanding — and teachers, knowing this, performed accordingly. Whether the principles were embedded in the teacher's practice or enacted for the benefit of the observer was, once again, a question that the observation system could not answer.

The pattern is consistent. Research breaks through. It is simplified for adoption. The simplification is adopted as orthodoxy. The orthodoxy is monitored through observation. And the observation system — because it captures impressions rather than evidence — cannot distinguish between a school that has deeply engaged with the research and a school that has bought the poster.

What the research asks of us

If there is a common thread across Black and Wiliam, Hattie, and Rosenshine, it is this: what matters is not whether a practice is present, but what it looks like when it is enacted. Feedback is not feedback is not feedback. There is task-specific feedback that advances learning and generic praise that does not. There is questioning and there is questioning — closed recall directed at volunteers is not open inquiry with cold calling and three-second wait time and probing follow-ups that build on pupil responses. Both are "questioning." They are not the same thing.

The research demands specificity. The systems we have built to implement it — the CPD sessions, the improvement plans, the observation proformas — have systematically eliminated that specificity in favour of categories broad enough to be checked but too broad to be useful. We have taken research that distinguishes between types of feedback, types of questioning, types of practice, and reduced it to checkboxes that record presence or absence. The information that makes the research actionable — the qualitative detail, the distinction between effective and ineffective enactment, the specificity that connects a classroom observation to a development conversation — is the information we have designed our systems to discard.

This is, ultimately, why I built a tool that captures not just whether a practice was observed but what it looked like — through descriptive qualifiers grounded in the same research that Black and Wiliam, Hattie, and Rosenshine produced. Not because the research was wrong, but because the systems we built around it were not equal to what it demanded. The research asked for specificity. We gave it checkboxes. That gap — between what the evidence says and what our tools can capture — is the gap that still needs closing.


Jamie Scobie writes from extensive experience in Scottish secondary education, including pastoral care, data for improvement, and school self-evaluation. This blog is an independent publication: he writes in a personal capacity as the creator of Learning Lens, writing about classroom observation, teaching evidence, and education policy. He speaks here only for himself and for Learning Lens, not for any employer or other organisation. He holds an MSt from Cambridge (Distinction) and a Masters from Stirling.