← Back to Blog
Observation practice

Why Lesson Observation Fails — and How to Fix It

I remember the first time I was formally observed. I was a student teacher at Moray House, and I had planned what I believed to be a lesson of extraordinary ambition. The kind of lesson you plan when you have more ideas than experience — when the gap between what you imagine a classroom could look like and what it actually looks like on a Tuesday morning with thirty teenagers who did not ask to be part of your professional development has not yet been properly calibrated. I was nervous in the way that student teachers are nervous: not about the pupils, but about the person at the back of the room. The presence of someone watching you teach does something to the act of teaching. It changes it. It contaminates the very thing it is trying to measure. I knew this instinctively before I could articulate it, and I suspect every teacher who has been observed knows it too.

The observation went fine. I got feedback. The feedback was kind. It was also, with hindsight, largely useless. Not because my observer was unkind or incompetent — she was neither — but because the system within which she was operating did not give her the tools to tell me anything I could act on. She told me the lesson had good pace. She told me the pupils were engaged. She told me my questioning could be stronger. I nodded. I filed the paperwork. I moved on to the next placement. I could not have told you, then or now, what "stronger questioning" meant in practice, what it would look like if I achieved it, or how I would know I had.

That was 2012. In the years since, I have been observed more times than I can count — as a probationer, in classroom roles, in faculty work, and in pastoral leadership — and I have observed others in roughly equal measure. I have sat in classrooms with a laptop, an iPad, a notebook, and once, during a particularly optimistic experiment with efficiency, nothing at all. I have given feedback that I believed was incisive and watched it produce no discernible change. I have received feedback that I believed was wrong and been unable to challenge it because the evidence base was a paragraph of impressionistic prose written by someone who had been in my room for twenty minutes. I have, on more than one occasion, never received the feedback at all.

This is a post about what is wrong with how we observe teaching, what the research says about why it is wrong, and what I think it would take to fix it. It is also, inevitably, a post about why I built Learning Lens — because I did not set out to build a product. I set out to solve a problem that had irritated me for the better part of a career.

This section includes an interactive chart — open the page in a browser to explore it.

The psychology of being watched

There is something particular about lesson observation that distinguishes it from other forms of professional scrutiny. A surgeon does not perform differently because a colleague is watching. A lawyer does not present differently because a partner is in the courtroom. Teaching is different. The presence of an observer fundamentally alters the dynamic between teacher and class — not because teachers are fragile, but because teaching is relational. It is built on the thousands of micro-interactions between a teacher and their pupils that accumulate over weeks and months into something approaching trust. An observer walks into that relationship for twenty minutes and attempts to draw conclusions about its quality.

The consequence is that observation produces anxiety out of all proportion to its stated purpose. Ask any teacher about their experience of being observed and the conversation will turn, almost immediately, to how it made them feel. Not what they learned. Not how their practice changed. How they felt. The emotional response overwhelms the professional one. Kelchtermans (2009) writes about the fragility of teachers' self-esteem — how positive self-evaluations fluctuate and how feedback is filtered and interpreted through the teacher's existing sense of who they are as a professional. Research undertaken at Cambridge on how teachers' professional identities shifted during the pandemic showed this with uncomfortable clarity: teachers described feeling "deskilled," questioning whether they were even doing the same job, their self-esteem eroded not by criticism but by the removal of the conditions that allowed them to feel competent. Observation is the moment where that fragility is most exposed. You are not being assessed on a piece of work you submitted. You are being assessed on an act you are performing, in real time, in front of the children you care about, by a person whose judgement you may or may not trust.

I felt this as a probationer more acutely than at any other stage. The probationary year is a peculiar thing — you are a teacher, fully registered, in front of your own classes, and yet you are also, simultaneously, someone who is being evaluated for their fitness to remain a teacher. The observation visits from your supporter, your head of department, sometimes an external quality assurance visit — these carry a weight that is disproportionate to their duration. A thirty-minute visit can colour your confidence for a week. A critical comment, however constructively intended, can sit in the pit of your stomach for a month. And a positive visit can produce a kind of euphoria that is equally unhelpful, because the lesson that went well when someone was watching is not necessarily the lesson that represents your actual practice.

I do not think this anxiety is something that can be trained out of teachers. It is an inevitable consequence of the fact that teaching, as Nias (1989) observed, is an occupation where the person cannot easily be separated from the craft. When someone tells you that your questioning was weak, they are telling you that something you did — something that felt, in the moment, like an expression of who you are as a teacher — was not good enough. The professional and the personal are inseparable in teaching, and observation is the site where that inseparability is felt most keenly.

The variability problem

The deeper problem with observation is not emotional. It is epistemic. The question is not whether observation makes teachers anxious — it does — but whether observation produces reliable evidence about teaching quality. The answer, according to the research, is that it often does not.

Robert Coe's work on lesson observation reliability is, or should be, required reading for anyone who conducts observations in schools. The research consistently shows that two observers watching the same lesson will frequently disagree about what they saw. Not at the margins — not a minor quibble about whether the pace was good or very good — but fundamentally, about whether the teaching was effective. The MET Project, funded by the Gates Foundation, found that an observation of a single lesson by a single observer was an unreliable measure of teaching quality. You needed multiple observations by multiple observers before the signal separated from the noise.

This landed differently for me than it might for someone who reads it as an abstract finding. I had, by the time I encountered Coe's work properly, been both the observer who was confident in their judgement and the teacher who was confident that the observer's judgement was wrong. I had sat in rooms where the teaching seemed strong to me — clear explanations, good pace, pupils working hard — and later heard a colleague describe the same lesson as lacking challenge. I had, conversely, watched lessons that felt disorganised and uncertain and later read observation feedback that described them as creative and pupil-centred. Both of us were acting in good faith. Both of us believed we were seeing clearly. The problem was that we had no shared framework for what we were looking at, no common language for describing it, and no structured way to compare what we had observed against what someone else might have observed in the same room.

This is the variability problem, and it is not a function of incompetent observers. It is a function of observation systems that ask observers to make holistic judgements — is this lesson good? — without giving them the tools to make specific ones. When your observation framework consists of a blank proforma with headings like "Quality of Teaching" and "Pupil Engagement," the observer fills it with whatever their own experience, biases, and priorities produce. Two observers with different teaching backgrounds, different conceptions of what effective teaching looks like, and different sensitivities to different classroom dynamics will write different things. The variation is not error. It is the inevitable product of a system that has mistaken impression for evidence.

The feedback that never came

If the observation itself is unreliable, the feedback conversation is where the damage compounds — or, more commonly, where it simply fails to occur.

I have lost count of the number of times, across multiple schools and multiple roles, where an observation was conducted and the feedback was either delayed to the point of irrelevance, delivered so briefly as to be meaningless, or never delivered at all. This is not, in my experience, an unusual situation. It is the norm. Teachers joke about it, in the way that people joke about things that would distress them if they took them seriously. You were observed on a Thursday. The feedback was supposed to happen that afternoon. Then it was rescheduled to Friday. Then it was after the holiday. Then it was never mentioned again.

What this communicates to the teacher being observed is clear, even if it is not intended: the observation was not about your development. It was about compliance. Someone needed to record that observations had happened. The box was ticked. You can move on.

Even when feedback is delivered promptly, it often fails the test that Dylan Wiliam sets for effective feedback: it must produce a cognitive response rather than an emotional one. "Your questioning was good" produces an emotional response — a brief glow of validation — and no change in practice. "You used six closed questions in the first ten minutes and one open question in the last five — what would happen if you reversed that ratio?" produces a cognitive response: the teacher has to think about what they do, consider an alternative, and decide whether to try it. The difference between these two pieces of feedback is the difference between observation as compliance and observation as professional learning. The first requires a person with a clipboard. The second requires a person with a framework, a shared language for what they observed, and the specificity to describe it in terms the teacher can act on.

The problem is that most observation systems are designed to produce the first kind of feedback, not the second. A proforma that says "Quality of Teaching: Good" gives the observer no incentive, and no structure, to record the specific practices they observed. A proforma that asks the observer to describe what they saw in narrative prose produces paragraphs of varying quality, none of which are comparable across observations, teachers, or time periods. Neither produces data. Both produce impressions.

The other side of the clipboard

Becoming an observer changed my relationship with all of this. When you are the person walking into someone else's classroom, the psychology reverses but the structural problems remain. You are now the one making judgements that will affect someone's confidence, their professional trajectory, and their relationship with you as a colleague. The weight of that responsibility is not trivial, and most observers — certainly most middle leaders conducting observations for the first time — are given almost no training in how to carry it.

I remember the first observation I conducted as a new faculty head. I walked in with a proforma, sat at the back, and tried to watch everything at once. Questioning. Pace. Differentiation. Pupil engagement. Behaviour management. The quality of the resources. The learning intention on the board. The seating plan. The interaction with the support for learning assistant. I was trying to hold all of this in my head simultaneously while also trying to be inconspicuous, which is difficult when you are the only adult in the room who is not supposed to be there.

What I wrote on the proforma afterwards was a paragraph of generalities. The lesson was well-structured. Pupils were on task. Questioning was used effectively. Resources were good. It was, I realise now, almost identical in both tone and usefulness to the feedback I had received as a student teacher a decade earlier. Not because I did not know what good teaching looked like — I did, at least as well as anyone with my experience could — but because the system I was operating within did not require me to be specific and did not give me the tools to be specific even if I wanted to be.

The problem was not that I could not see what was happening in the room. The problem was that I had no structured way to record it. I could see that the teacher was asking questions — but I had no framework for distinguishing between the closed recall questions that dominated the first half and the open, probing questions that appeared briefly in the last five minutes. I could see that pupils were working — but I had no way to distinguish between the cognitive engagement of the pupils at the front who were wrestling with a challenging task and the surface compliance of the pupils at the back who were copying from the textbook without understanding a word of it. I could see that feedback was being given — but I had no way to record that it was almost entirely generic praise rather than the task-specific, process-focused feedback that the research says changes learning.

Everything I could see, I could describe. But description is not evidence. Evidence requires structure — a shared framework that makes one observation comparable to another, one teacher's practice comparable to a colleague's, and one term's data comparable to the next. Without that structure, I was producing anecdotes. Anecdotes that might be insightful, or might be wrong, and that would certainly be forgotten by the time the next Standards and Quality report needed writing.

What I think observation should be

The research base here is clear, and it converges from multiple directions. Coe's work on reliability says that single observations by single observers are insufficient — you need aggregation across observations. Hattie and Timperley's work on feedback says that feedback must be specific, task-focused, and actionable — not generic, person-directed, and vague. Black and Wiliam's work on formative assessment says that evidence must be gathered, interpreted, and acted upon — the assessment cycle is only complete when it changes what happens next. Rosenshine's Principles, Alexander's dialogic teaching, the EEF's guidance on metacognition and feedback — all of this converges on a set of specific, observable, research-grounded teaching practices that can be identified in a classroom by someone who knows what to look for.

The observation system most schools use was not built on any of this. It was built on a compliance model — observations must happen, observations must be recorded, someone must sign the form — that pre-dates most of the research that would tell us how to do it well. The form is the artefact of a system that values the fact of observation over the quality of what is observed.

What would it look like if we started from the evidence instead?

It would start with a shared framework of observable practices — not "Quality of Teaching" as a single holistic judgement, but the specific, nameable things a teacher does that the research says matter. Open questioning. Wait time. Modelling with worked examples. Scaffolding with gradual release. Retrieval practice. Explicit strategy instruction. These are not mysterious. They are well-documented, extensively researched, and observable by anyone who knows what they look like.

It would capture not just whether a practice was present, but what it looked like — the descriptive detail that distinguishes between a teacher who asks open questions but only takes answers from volunteers and a teacher who asks open questions with cold calling, three-second wait time, and probing follow-ups that build on the pupil's response. Both are "questioning." They are not the same thing. A system that records both as "questioning observed" has lost the information that makes the evidence useful.

It would aggregate across observations, across teachers, across departments, and across time. Not to rank teachers — the variability problem tells us that single observations cannot bear that weight — but to see patterns. Is questioning in the English department different from questioning in Maths? Is feedback across the school predominantly generic praise, or is it task-specific? Has the focus on metacognition that was in the improvement plan actually translated into classroom practice, or is it still a poster on the wall?

It would produce feedback that is specific enough to act on and data that is structured enough to inform the improvement plan. Not a paragraph of impressionistic prose that gets filed and forgotten, but a record of what was seen — described in terms that both the observer and the teacher understand because the framework is shared — connected to the development suggestions that the research provides.

And it would respect the teacher's intelligence and professionalism throughout. The purpose is not to grade. The purpose is to see what is happening in classrooms clearly enough to make it better. Teachers deserve to know what was observed in their room, described in specific terms, connected to research, and followed by a concrete suggestion for development. That is the minimum. Most observation systems fail to deliver even that.

Why I built it

Learning Lens exists because years on both sides of the clipboard showed how often both sides fail. From the observed side, observation should respect practice, describe it specifically, and say something usable. From the observer side, capture should be structured enough to produce evidence rather than impressions. From the improvement side, schools need data that shows what is happening across departments, surfaces patterns no single visit can reveal, and connects observation evidence to the priorities that are supposed to drive the work.

None of the tools I had access to provided any of this. The proformas produced paragraphs. The spreadsheets produced numbers stripped of context. The conversations produced good intentions that evaporated by the following Monday. I was writing Standards and Quality reports that cited observation evidence that did not, if I am honest, amount to much more than a collection of impressions dressed in evaluative language.

Building Learning Lens was, at first, simply building the tool I wished I had. A way to capture what I could see in a classroom — quickly, specifically, in real time — using a framework grounded in the research I had spent years reading. A way to describe the quality of what I observed, not through a rating or a grade, but through descriptive qualifiers that told the story: not "questioning was good" but "open questioning, cold calling, three-second wait time, probing follow-ups building on pupil responses." A way to aggregate that evidence across visits, across departments, across a term — so the school's improvement plan could be built from data rather than from assumption.

The qualifier taxonomy, the observation lenses, the analytics that show which practices are present and which are absent, the trend data that shows whether a professional learning intervention has changed what happens in classrooms — all of this exists because of years experiencing, in granular detail, what happens when observation systems produce impressions instead of evidence.

For the student teacher at Moray House who was told their questioning could be stronger — it tells them what stronger questioning looks like, described in terms they can observe in their own practice and in others'. For the probationer whose confidence is shaped by the weight of a thirty-minute visit — it ensures the evidence is specific, shared, and connected to research rather than dependent on the subjective impression of whoever happened to be in the room. For the teacher whose feedback never arrived — it produces a structured record that exists independently of whether the post-observation conversation happens on schedule. For the observer who walks in with a blank proforma and tries to see everything at once — it gives them a focused lens and a framework that channels their attention toward what the research says matters.

For the school leader writing a Standards and Quality report — it replaces the impressionistic paragraph with aggregated, evidence-based data that can be interrogated, compared, and traced back to specific observations in specific classrooms. It closes the loop between what we say we are doing and what the evidence shows is happening.

Observation should not be something that teachers dread and observers endure. It should be the mechanism through which a school understands its own practice clearly enough to improve it. That requires better evidence, better tools, and a system that treats the complexity of teaching with the specificity it deserves.

I built Learning Lens because I could not find it. And because the problem that began with a student teacher being told to improve their questioning — without being told what that meant, what it would look like, or how to know when they had done it — is the same problem that persists in schools today. The research has moved on. The tools have not. Until now.


Jamie Scobie writes from extensive experience in Scottish secondary education, including pastoral care, data for improvement, and school self-evaluation. This blog is an independent publication: he writes in a personal capacity as the creator of Learning Lens, writing about classroom observation, teaching evidence, and education policy. He speaks here only for himself and for Learning Lens, not for any employer or other organisation. He holds an MSt from Cambridge (Distinction) and a Masters from Stirling.