← Back to Blog
Observation practice

Classroom Observation: What the Research Says It's For

There is a version of classroom observation that everybody recognises and nobody much likes. A senior leader appears at the back of the room with a clipboard. The teacher, who has known about this visit for a week, delivers a lesson that bears only a passing resemblance to what they'd normally do. The observer ticks boxes, writes up some notes, and feeds back using a formula that begins with praise, moves to a criticism disguised as a development point, and ends with more praise in case the criticism landed too hard. The teacher nods, files the sheet, and nothing changes.

This is observation as performance — a ritual that exists to demonstrate that quality assurance is happening rather than to assure quality. And while most school leaders would agree, if pressed, that this isn't quite working, the model persists because the alternative requires confronting something genuinely uncomfortable: we are not nearly as good at judging teaching as we think we are, and the systems we have built tend to serve accountability rather than improvement.

The reliability problem

The most important thing to understand about classroom observation is how unreliable it is when used as a measurement tool. Robert Coe's influential work at Durham University, drawing on the $52 million Measures of Effective Teaching (MET) study funded by the Gates Foundation, established this with uncomfortable clarity. The MET project (Kane & Staiger, 2012) tested five different observation instruments across 3,000 teachers and found that inter-rater reliability ranged from 0.24 to 0.68 — figures that, translated into practical terms, mean that if one observer judges a lesson 'Outstanding', there is a 51–78% probability that a second observer watching the same lesson would give a different grade. Coe's own summary of this evidence is worth sitting with: highly trained observers using the best available methodologies can distinguish an above-average teacher from a below-average one about 60% of the time. A coin toss would give you 50%.

This is not a counsel of despair. It is, however, a counsel against overconfidence. When Ho and Kane (2013) examined the reliability of observations conducted by school-based personnel — principals, assistant heads, peer observers — the picture was even more sobering. Scores varied considerably from lesson to lesson for any given teacher, and from observer to observer for any given lesson. The implication is not that observation is useless, but that treating it as a precise measuring instrument — assigning grades, making high-stakes judgements on the basis of one or two visits — involves a degree of false confidence that the evidence simply does not support.

The question, then, is what observation is good for, if not for reliably ranking teaching quality. And this is where the research points somewhere genuinely productive.

What observation can do: making practice visible

Graham Nuthall's remarkable ethnographic work, published posthumously as The Hidden Lives of Learners (2007), demonstrated something that ought to be obvious but rarely gets taken seriously enough: most of what happens in a classroom is invisible to the teacher. Nuthall placed microphones on individual pupils and tracked their experiences minute by minute across thousands of hours of classroom life. What he found was that student learning is shaped by three worlds operating simultaneously — the public world of teacher-managed activity, the semi-private world of peer relationships and interactions, and the private world of each student's own thinking. Teachers see and manage the first of these. The other two — where much of the real learning and, crucially, the real mislearning happens — remain largely hidden. Students already know, on average, about 40–50% of what the teacher intends to teach, but that knowledge is unevenly distributed, idiosyncratic, and often wrong in ways that peer conversation reinforces rather than corrects.

This finding reframes what observation is for. Its value is not as a tool for judging whether teaching is good enough, but as a structured way of seeing what is otherwise invisible — to the teacher, to colleagues, to the school as a whole. Observation is one of the very few mechanisms through which teaching practice, which happens largely behind closed doors, can become an object of professional conversation. And professional conversation about practice — what Timperley et al. (2007) call the 'teacher inquiry and knowledge-building cycle' — is the mechanism through which teaching improves.

Helen Timperley's Best Evidence Synthesis (2007), a systematic review of the international evidence on teacher professional learning, identified that the professional development most likely to produce measurable gains in pupil outcomes shared several features: it was sustained over time, it was grounded in evidence from teachers' own practice, it involved cycles of inquiry where teachers examined the impact of their teaching on student learning, and it took place within a collegial environment characterised by what Timperley describes as 'relationships of respect and challenge.' Observation — when designed well — sits at the heart of this cycle. It provides the evidence. The inquiry cycle provides the structure. The collegial culture provides the conditions.

I've written separately about why most CPD fails the teachers it's meant to serve — observation without a development pathway is surveillance with extra steps.

The accountability trap

When observation is primarily about accountability — about grading lessons, identifying underperformance, or generating evidence for capability proceedings — it creates a set of rational but counterproductive incentives. Teachers prepare showcase lessons rather than typical ones. They avoid pedagogical risk because the cost of a lesson that doesn't come off is higher than the benefit of one that does. They experience observation as something done to them rather than with them, and the feedback conversation becomes a negotiation about the verdict rather than an exploration of practice.

Matt O'Leary's research on observation in the further education sector (2013, 2014) documented these dynamics in detail. O'Leary found that graded observation systems produced 'performativity' — a term borrowed from Stephen Ball — where teachers' primary concern became how they would be perceived by the observer rather than engaging with genuine questions about their teaching. The observation became a performance of competence, disconnected from the daily reality of the classroom.

The most striking illustration of this came when Ofsted clarified in 2014 that it would no longer grade individual lessons during inspections. Schools that had built their entire observation culture around the four-point scale suddenly had to ask themselves: if we're not grading lessons, what are we actually doing when we observe? Many discovered that the grade had been doing all the intellectual work, and without it, the observation conversation collapsed into vagueness. A BlueSky Education survey of 204 schools in 2019 found that 59% had moved away from grading — but the challenge of building something genuinely developmental in its place remained largely unmet.

Scotland's position here is instructive. How Good Is Our School? 4 (HGIOS4), Education Scotland's quality framework, doesn't ask schools to grade individual lessons. It asks, through Quality Indicator 2.3 (Learning, teaching and assessment), whether there is a culture of classroom observation and professional dialogue that leads to improvement. The emphasis is on self-evaluation — on schools developing the capacity to understand their own practice rather than having it measured from outside. The GTCS Professional Standards, similarly, frame classroom observation as a component of professional learning rather than performance management. But the frameworks alone don't solve the problem. The question of how to observe — with what focus, what structure, what follow-through — remains one that most schools answer with inherited habits rather than deliberate design.

What developmental observation requires

The EEF's Effective Professional Development guidance report (Sims et al., 2021), based on a meta-analysis of 104 studies, identified four groups of mechanisms through which professional development produces changes in teacher practice: building knowledge, motivating teachers, developing teaching techniques, and embedding practice. The report's most significant finding for the observation question is that professional development programmes which included all four mechanism groups were more likely to produce measurable gains in pupil attainment than those which included only some of them — and that observation and feedback, when properly structured, can serve as the vehicle through which all four operate.

This gives us a framework for thinking about what developmental observation requires.

This section includes an interactive chart — open the page in a browser to explore it.

Focus. An accountability observation asks how good is this lesson? A developmental observation asks what is happening with [something specific] in this classroom? The specificity matters because it is what makes the subsequent conversation productive. Watching for everything means seeing nothing in particular. But watching for how a teacher uses questioning to check understanding, or how pupils respond during independent work, or what happens during the transition between activities — that level of focus produces evidence that both observer and teacher can examine together. The EEF framework would call this the 'building knowledge' mechanism: observation generates specific, contextualised knowledge about practice that no course or reading can replicate.

Reciprocity. The most powerful professional learning from observation often accrues to the observer rather than the observed. Peer observation — teachers watching teachers, with a shared focus and genuine voluntarism — produces what Timperley calls 'relationships of respect and challenge', where the act of seeing another teacher's practice creates both validation (they struggle with the same things I do) and provocation (I hadn't thought of doing it that way). The MET study's own recommendation, often overlooked, was that schools should use multiple observers and average their judgements — a finding that, reframed for development rather than accountability, becomes an argument for collective observation practices rather than top-down ones.

Structured follow-through. This is where most observation systems fail. The research on feedback effects — from Kluger and DeNisi's (1996) landmark meta-analysis through to Hattie and Timperley's (2007) framework — consistently shows that feedback which triggers an ego-protective emotional response (relief, anxiety, defensiveness) does not improve performance, while feedback that directs attention to the task and how to improve it does. Dylan Wiliam, drawing on this evidence, argues that feedback functions formatively only when the information is actually used by the learner to close the gap between current and desired performance. For teachers receiving observation feedback, the implication is the same: if the primary reaction is emotional rather than cognitive — if the teacher leaves the conversation feeling judged rather than thinking about their practice — the feedback has not functioned as intended.

Aggregation. A single observation tells you about one teacher in one lesson on one day. But observations looked at in aggregate tell you something about how your school teaches — the shared strengths, the common development needs, the patterns that no individual observation can reveal on its own. This is where observation connects to school improvement. In HGIOS4 terms, this is the bridge between QI 2.3 (learning, teaching and assessment) and QI 1.1 (self-evaluation for self-improvement): the observation evidence becomes part of the school's self-evaluation cycle, informing improvement priorities that are grounded in direct evidence of classroom practice rather than proxy measures like exam results or inspection gradings.

Observation evidence is only one part of the picture. Schools are already sitting on data they don't know how to use. The question is whether observation evidence makes it into the improvement plan — or gets filed and forgotten.

Why this matters

The case for getting classroom observation right is not about compliance, and it is not about accountability — both of which have their place but neither of which requires the kind of careful design being argued for here. The case is about something more fundamental: that teaching is a complex, largely invisible practice that improves primarily through contact with other people's practice, structured reflection, and the willingness to examine evidence honestly. Observation is the mechanism through which all of this happens — or doesn't.

Schools that treat observation as a compliance exercise — a box ticked, a form filed — are leaving on the table the most direct source of evidence they have about the thing that matters most. Schools that design observation for development — with focus, reciprocity, structured feedback, and aggregation — are building the conditions in which teaching improves collectively rather than individually, and in which improvement is sustained rather than episodic.

The research is clear enough. The challenge is practical. And that practical challenge — how to structure observation so that it produces insight, supports teachers, and informs school improvement — is one that requires not just good intentions but the right tools, the right processes, and the right professional culture.

If you're exploring tools that treat observation as evidence rather than verdict, Learning Lens is available for Scottish schools — you can book a conversation from the home page.


References

Coe, R. (2013). Improving Education: A Triumph of Hope over Experience. Inaugural lecture, Durham University.

Coe, R., Aloisi, C., Higgins, S., & Major, L.E. (2014). What Makes Great Teaching? Review of the Underpinning Research. Sutton Trust.

Education Endowment Foundation (2021). Effective Professional Development: Guidance Report. London: EEF.

Education Scotland (2015). How Good Is Our School? (4th edition). Livingston: Education Scotland.

Hattie, J. & Timperley, H. (2007). 'The Power of Feedback.' Review of Educational Research, 77(1), 81–112.

Ho, A.D. & Kane, T.J. (2013). The Reliability of Classroom Observations by School Personnel. MET Project, Bill & Melinda Gates Foundation.

Kane, T.J. & Staiger, D.O. (2012). Gathering Feedback for Teaching: Combining High-Quality Observations with Student Surveys and Achievement Gains. MET Project, Bill & Melinda Gates Foundation.

Kluger, A.N. & DeNisi, A. (1996). 'The Effects of Feedback Interventions on Performance: A Historical Review, a Meta-Analysis, and a Preliminary Feedback Intervention Theory.' Psychological Bulletin, 119(2), 254–284.

Nuthall, G. (2007). The Hidden Lives of Learners. Wellington: NZCER Press.

O'Leary, M. (2014). Classroom Observation: A Guide to the Effective Observation of Teaching and Learning. London: Routledge.

Sims, S., Fletcher-Wood, H., O'Mara-Eves, A., Cottingham, S., Stansfield, C., Van Herwegen, J. & Anders, J. (2021). What Are the Characteristics of Effective Teacher Professional Development? A Systematic Review and Meta-Analysis. London: EEF.

Timperley, H., Wilson, A., Barrar, H. & Fung, I. (2007). Teacher Professional Learning and Development: Best Evidence Synthesis Iteration. Wellington: Ministry of Education.

Wiliam, D. (2011). Embedded Formative Assessment. Bloomington, IN: Solution Tree Press.


Jamie Scobie writes from extensive experience in Scottish secondary education, including pastoral care, data for improvement, and school self-evaluation. This blog is an independent publication: he writes in a personal capacity as the creator of Learning Lens, writing about classroom observation, teaching evidence, and education policy. He speaks here only for himself and for Learning Lens, not for any employer or other organisation. He holds an MSt from Cambridge (Distinction) and a Masters from Stirling.