Learning Lens was built first for Scotland — specifically for How Good Is Our School? (HGIOS4), the self-evaluation language familiar from years in Scottish classrooms and improvement work. The taxonomy of 59 classroom practices at the core of the platform was grounded in international research — Rosenshine, Hattie, Black and Wiliam, Alexander — but the way that evidence was aggregated, reported, and turned into evaluative prose was Scottish to its bones. "Almost all teachers" meant something specific. "The majority of" meant something different. The six-point quality scale mapped to HGIOS4's own architecture. Scotland was not just the first market. It was the starting point.
Then an English school asked whether it would work with Ofsted's Education Inspection Framework.
The honest answer, at that point, was: I don't know. But the process of finding out taught me something I hadn't expected — something that has changed how I think about observation evidence, inspection frameworks, and the relationship between the two. The practices I was capturing in classrooms didn't change when the framework changed. Rosenshine's principles of instruction do not reorganise themselves at the border. A teacher using retrieval practice in Glasgow is doing the same thing as a teacher using retrieval practice in Gloucester. What changed was the lens through which that evidence was interpreted — the evaluative language, the quality thresholds, the way the same data was grouped and reported. HGIOS4 and the Ofsted EIF ask many of the same questions. They just ask them differently, weight them differently, and describe the answers in different words.
Building for two frameworks forced me to separate what I now think of as the two layers of any observation system: the evidence layer and the interpretation layer. The evidence layer — what happened in the classroom, which practices were observed, what the patterns look like across departments and over time — is universal. It doesn't belong to any framework. The interpretation layer — how that evidence is described, which framework vocabulary it maps to, what "good" looks like in the language of a particular inspection body — is local. And once you separate the two, a question follows naturally: how many interpretation layers could the same evidence support?
There was another catalyst too. HGIOS4 — the framework I'd built for, the one I knew best — is likely to change. HMIE's separation from Education Scotland signals a new framework on the horizon, and while nobody yet knows what it will look like, the possibility that the Scottish interpretation layer might shift made me think harder about what sits beneath it. If the framework changes, do the practices change? They don't. The research base doesn't reorganise itself because a government agency restructures. Which meant the question was no longer "does my taxonomy work with this framework?" but something more fundamental: what is the taxonomy describing, and why these practices and not some other set?
Why these practices, and why they grow
The honest answer is that the number isn't fixed. The taxonomy started at fifty-three practices in ten categories — the result of a long process of reading, teaching, observing, and asking what could be reliably identified in a live classroom by a non-specialist observer in real time. It now sits at fifty-nine across eleven categories, after adding a Relationships & Regulation cluster that I deliberately left out at the start and have since brought in. The number will keep changing. What matters is the process by which something earns its place.
That last constraint — reliably identifiable in a live classroom, by a non-specialist observer, in real time — matters more than it sounds. There are aspects of teaching quality that are undeniably important but almost impossible to observe in a single visit: the quality of a teacher's subject knowledge, for instance, or the coherence of their long-term curriculum sequencing. Rob Coe's Great Teaching Toolkit — which distils over 100 frameworks into four dimensions and seventeen elements — includes "understanding the content they are teaching" as a core element of great teaching. He is right. But you cannot see that in a fifteen-minute walkthrough. You can infer it, over time, from the quality of explanations and the precision of examples. But the inference is fragile, and building a data system on fragile inferences is building on sand.
So the taxonomy captures what is observable, reliably, by someone watching a lesson. That means it focuses heavily on the instructional core: how new material is presented, how understanding is checked, how questions are used, how practice is structured, how feedback is given. These are the areas where the research base is deepest and the observational signal is clearest — Rosenshine's principles, Hattie's effect sizes for feedback and metacognition, Black and Wiliam on formative assessment, Alexander on dialogic teaching, the EEF guidance reports, the desirable difficulties literature from Bjork and Bjork, Sweller's cognitive load theory, Chi and Wylie's ICAP framework for engagement.
The practices sit at a level of granularity that is specific enough to be meaningful — you can see the difference between a teacher using cold call and one using hands-up — but general enough to apply across subjects, phases, and national contexts. They are not the only practices that matter. They are the ones where the intersection of research evidence, observational reliability, and practical utility is strongest. The taxonomy will keep evolving as the evidence and the practitioner reality push it to.
What I was less prepared for was discovering how closely this taxonomy — built from a UK classroom perspective — maps to the constructs that researchers in very different traditions have independently identified as central to teaching quality.
The research frameworks most people never see
School leaders tend to know their national inspection framework intimately and the international research base barely at all. What they often don't know is that behind those frameworks sits a body of work on teaching quality measurement that has been converging, gradually and from multiple directions, on a remarkably similar set of constructs.
Kirsti Klette, at the University of Oslo, has spent two decades arguing for what she calls a "shared language of teaching" — a common vocabulary for describing and measuring classroom practice that could work across national boundaries. Her work on the PLATO observation protocol and the wider QUINT project has shown that when researchers from different countries use different observation instruments on the same lessons, they tend to agree on the broad dimensions of quality — even when their specific frameworks differ. The question she poses — whether we can develop a shared professional language for teaching — is one I think about constantly, because it is essentially the question my taxonomy is trying to answer at the practical level.
The Classroom Assessment Scoring System (CLASS), developed by Robert Pianta at the University of Virginia, measures teaching quality across three domains: emotional support, classroom organisation, and instructional support. It was built for US classrooms but has been validated internationally and used in contexts from Germany to Chile. The "three basic dimensions" identified by Klieme and colleagues in the German didactics tradition — cognitive activation, classroom management, and supportive climate — converge on essentially the same structure from a completely different intellectual starting point. Coe's Great Teaching Toolkit synthesised over 100 frameworks and arrived at four dimensions: understanding the content, creating a supportive environment, maximising opportunity to learn, and activating hard thinking. The language is different every time. The territory it describes is recognisably the same.
My taxonomy does not replicate any of these frameworks — it was built for a different purpose. CLASS, PLATO, and the GTT are designed primarily for research, training, or professional development. Learning Lens is designed for live capture during observation, which imposes different constraints on granularity, speed, and interaction design. But the overlap in what we all consider important is not coincidental. It reflects a research consensus that has been building for thirty years — the same meta-analyses, the same experimental studies, the same evidence about what makes a measurable difference to student learning. The practices are not a proprietary invention. They are a practical operationalisation of that consensus, shaped by the additional requirement that everything in the taxonomy must be identifiable by a trained observer in real time.
The following table maps how the major frameworks organise this shared territory:
| Domain | Learning Lens (59 practices, 11 categories) | CLASS (Pianta) | Great Teaching Toolkit (Coe) | Three Basic Dimensions (Klieme) | Danielson FFT |
|---|---|---|---|---|---|
| How content is taught | Presentation of New Material; Prior Knowledge & Review | Instructional Learning Formats | Activating Hard Thinking | Cognitive Activation | Domain 3: Instruction |
| How understanding is checked | Questioning & Dialogue; Checking for Understanding | Quality of Feedback; Language Modelling | Activating Hard Thinking | Cognitive Activation | Domain 3: Instruction |
| How practice is structured | Guided & Independent Practice; Metacognition & Self-Regulation | Analysis and Inquiry | Maximising Opportunity to Learn | Classroom Management | Domain 3: Instruction |
| How feedback works | Feedback | Quality of Feedback | Activating Hard Thinking | Cognitive Activation | Domain 3: Instruction |
| How learners are included | Differentiation & Inclusion; Active Learning & Collaboration | Regard for Student Perspectives | Creating a Supportive Environment | Supportive Climate | Domain 2: Environment |
| How the classroom functions | (Captured via Pupil Responses) | Behaviour Management; Productivity | Maximising Opportunity to Learn | Classroom Management | Domain 2: Environment |
The convergence is imperfect — it could not be otherwise, given the different purposes each framework serves. But the pattern is striking: whatever vocabulary you use, whatever tradition you come from, the aspects of teaching that matter most to student learning cluster around the same set of concerns. My practices are one way of describing that cluster. They will not be the last. But they are grounded in the same evidence that underpins every serious framework in this table, and they were built with a constraint that most of those frameworks don't face — the need to work in a practitioner's hand, in a live classroom, without slowing down the act of observation.
That question — how far the evidence layer extends — led me to spend time with the inspection and evaluation frameworks I hadn't previously studied. What follows is not a definitive guide — I haven't taught in any of these systems, and I'd distrust anyone who claimed deep expertise in six inspection frameworks simultaneously. But mapping the landscape has been instructive, and the patterns that emerge are worth sharing.
British Schools Overseas: The framework you almost already know
There are over a thousand British international schools operating outside the UK — in the Gulf, East Africa, Southeast Asia, and beyond. Those that describe themselves as "British" can seek accreditation through the British Schools Overseas (BSO) scheme, which is administered by the UK Department for Education and monitored by Ofsted. The inspection is carried out by approved inspectorates — ISI, Education Development Trust, and others — and the standards are explicitly benchmarked against independent schools in England.
For anyone already working within the Ofsted EIF, BSO is recognisable territory. The standards cover quality of education, pupil personal development, safeguarding, leadership, and the extent to which the school's "British character" is evident in its ethos and curriculum. The evaluative approach — looking at curriculum intent, implementation, and impact — echoes the EIF directly. The accreditation cycle is three years, and the published reports are structured in ways that any English school leader would find familiar.
What struck me about BSO was not how different it was from what I'd already built, but how close. The observation evidence that supports an Ofsted self-evaluation would support a BSO self-evaluation with relatively minor adaptation. The gap is not in what you capture. It's in how you frame it — and that framing is a configuration problem, not an engineering one.
The UAE Unified Inspection Framework: Six standards, one ambition
The UAE operates a unified inspection framework across its three main regulatory bodies — KHDA in Dubai, ADEK in Abu Dhabi, and SPEA in Sharjah. The framework uses six performance standards covering student achievement, personal and social development, teaching and assessment, curriculum, school protection and support, and leadership and management. Schools are rated on a scale from Outstanding to Very Weak.
Two things about this framework caught my attention. First, the emphasis on self-evaluation prior to inspection. Schools are required to submit self-evaluation evidence aligned to the performance standards before inspectors arrive — the inspection begins with what the school says about itself, and the inspector's job is to test those claims. That is structurally identical to how HGIOS4 operates in Scotland, and it means that the use case for structured observation evidence is the same: a school that can point to systematically collected, framework-aligned evidence of teaching quality is in a fundamentally stronger position than one that relies on anecdote and memory.
Second, the teaching and assessment standard maps surprisingly well to the categories I was already capturing. When inspectors evaluate teaching quality in a Dubai school, they're looking at whether teachers check for understanding, whether questioning promotes thinking, whether feedback helps pupils improve, whether differentiation meets the needs of different learners. These are not UAE-specific pedagogical concerns. They are the concerns that every evidence-informed observation taxonomy is built around — because they are the things that the research says matter most.
Australia: Standards without inspection
Australia does something structurally different from Scotland, England, or the UAE. There is no national inspection regime. Instead, the Australian Professional Standards for Teachers — maintained by the Australian Institute for Teaching and School Leadership (AITSL) — define what teachers should know and be able to do across seven standards and four career stages, from Graduate through to Lead. Teacher evaluation happens at the school level, conducted by principals, using frameworks that vary by state and sector.
This means there is no external inspection to prepare for — but there is an accreditation and registration system that requires teachers to demonstrate their practice against the standards. The gap I kept hearing about in the Australian literature was not a lack of standards but a lack of structured evidence. Schools know what the standards say. What they often lack is a systematic way to connect classroom observation evidence to the standard it relates to — to say, with confidence, "here is what this teacher's practice looks like against Standard 3: Plan for and implement effective teaching and learning."
The observation taxonomy I'd built for HGIOS4 and Ofsted could, in principle, be mapped to the AITSL standards. The practice of checking for understanding relates to Standard 5 (Assess, provide feedback and report on student learning). The practice of differentiation relates to Standard 1 (Know students and how they learn). The mapping is not trivial — the standards are broad and the taxonomy is granular — but the research base is shared. Hattie is Australian. The EEF guidance reports are read in Melbourne as readily as in Manchester. The evidence layer translates. It's the interpretation layer that would need building.
New Zealand: A system redesigning how it reports
New Zealand's Education Review Office (ERO) conducts external reviews of schools on a roughly three-yearly cycle, using a School Evaluation Indicators framework that emphasises responsive curriculum, effective teaching, and educationally powerful connections. ERO recently overhauled its reporting format — moving toward clearer, more parent-friendly language and a focus on how well schools support learner progress.
What interested me about the New Zealand system was the explicit emphasis on school self-review as both an internal improvement tool and an input to external evaluation. ERO doesn't just assess schools — it assesses how well schools assess themselves. The quality of a school's own inquiry into its practice is part of the review. That creates a direct incentive for structured observation evidence, because a school that can demonstrate a systematic cycle of looking at teaching, identifying patterns, and acting on them is telling the ERO story that ERO wants to hear.
The cultural context is different — the bicultural dimension of New Zealand education, the centrality of Māori learner outcomes, the emphasis on whanaungatanga and manaakitanga in the evaluation indicators — and any tool entering that market would need to engage with those dimensions respectfully rather than grafting a UK framework onto a fundamentally different set of values. But the underlying architecture — evidence of teaching quality feeding self-evaluation feeding external review — is the same pattern I'd built for in Scotland.
The United States: Frameworks in competition
The US is the most fragmented market I looked at. There is no national inspection framework. Instead, individual states and districts choose from competing evaluation models — the Danielson Framework for Teaching, Marzano's Focused Teacher Evaluation Model, the NIET TAP system, and dozens of state-specific variations. Danielson organises teaching into four domains and 22 components. Marzano uses four domains and 23 competencies. Both are research-grounded, both are rubric-based, and both are used by thousands of schools.
The scale is extraordinary, but the fragmentation is the challenge. A tool that works for a Danielson district in Pennsylvania may need significant reconfiguration for a Marzano district in Florida. The evidence layer would still translate — the practices being observed are the same practices described in both frameworks — but the interpretation layer would need to be not just configurable but district-configurable. That's a different order of complexity from switching between two national frameworks.
What I found telling, though, was a statistic I kept encountering: despite the post-2009 overhaul of teacher evaluation in virtually every state, the vast majority of districts still rate 95% or more of their teachers as effective or better. The frameworks exist. The rubrics are detailed. But the observation systems feeding them are not producing the differentiation in evidence that the frameworks were designed to support. The tools are not doing the job the frameworks need them to do. I recognise that problem. It's the one I set out to solve.
Canada: Ontario as a case study
Ontario operates a legislated Teacher Performance Appraisal system — classroom observations conducted by principals on a five-year cycle, assessed against 16 competencies. It is structured, mandated, and largely compliance-oriented. The competencies are broad — "Teachers know the curriculum," "Teachers know a variety of effective assessment and evaluation practices" — and the rating scale is binary: satisfactory or unsatisfactory.
The strength is universality: every teacher is appraised, every appraisal follows the same process. The limitation is depth. A binary rating against broad competencies does not produce the kind of granular evidence that drives improvement. It answers "is this teacher performing adequately?" but not "what does this teacher's practice look like in detail, and where are the specific opportunities to develop?" The other Canadian provinces have their own systems, each with different structures and expectations. As with the US, the market is large but the entry cost of localisation is high.
What the map tells me
Six inspection systems. Six different evaluation architectures. And underneath all of them — and underneath the research frameworks that inform them — the same set of questions: Are teachers checking whether pupils understand what's been taught? Is questioning promoting thinking or just recall? Are all learners being appropriately challenged? Is feedback helping pupils improve? Are pupils actively engaged in their learning? Does the school know what its teaching looks like, and is it doing something about it?
There is something philosophically interesting happening here. The history of education is a history of disagreement — about what should be taught, how, to whom, and for what purpose. Curriculum wars, pedagogy wars, the phonics debate, the knowledge-skills divide. But at the level of classroom interaction — at the level of what a teacher does with thirty young people in a room — the international research community has, quietly and without much fanfare, arrived at something approaching consensus. Not perfect consensus. Not consensus on everything. But agreement that a relatively small number of instructional practices account for a disproportionately large share of the variation in student learning. Rosenshine knew this in 1986. What has changed is that the evidence has deepened, the replications have multiplied across contexts and cultures, and the convergence across independently developed frameworks has become difficult to attribute to chance.
The frameworks provide the vocabulary. The observation provides the evidence. And the relationship between the two is, in every system I looked at, the same: the evidence layer is universal, grounded in a shared international research base; the interpretation layer is local, shaped by the history, values, and accountability structures of each system.
I don't know yet how many of these frameworks Learning Lens will eventually support. Building for HGIOS4 taught me what structured observation evidence could look like. Building for Ofsted taught me that the same evidence could serve a different framework without compromising either. The architecture now separates the evidence from the interpretation — and that separation, which began as a technical decision, turns out to be the most important design choice in the product. As the platform develops and as we work with schools in new contexts, the taxonomy itself will develop too — not by abandoning its research foundations, but by testing them against the reality of classrooms in systems I haven't yet seen from the inside.
The practices are universal. The frameworks are local. And the schools, wherever they are, are trying to answer the same question: what is actually happening in our classrooms, and how do we know?
References
Alexander, R. (2020). A Dialogic Teaching Companion. Abingdon: Routledge.
Black, P. & Wiliam, D. (1998). 'Assessment and Classroom Learning.' Assessment in Education, 5(1), 7–74.
Chi, M.T.H. & Wylie, R. (2014). 'The ICAP Framework: Linking Cognitive Engagement Activities to Active Learning Outcomes.' Educational Psychologist, 49(4), 219–243.
Coe, R., Rauch, C.J., Kime, S. & Singleton, D. (2020). Great Teaching Toolkit: Evidence Review. Evidence Based Education.
Hattie, J. (2009). Visible Learning: A Synthesis of Over 800 Meta-Analyses Relating to Achievement. Abingdon: Routledge.
Klette, K. (2023). 'Classroom Observation as a Means of Understanding Teaching Quality: Towards a Shared Language of Teaching?' Journal of Curriculum Studies, 55(1), 49–62.
Pianta, R.C., La Paro, K.M. & Hamre, B.K. (2008). Classroom Assessment Scoring System (CLASS). Baltimore: Brookes.
Rosenshine, B. (2012). 'Principles of Instruction: Research-Based Strategies That All Teachers Should Know.' American Educator, 36(1), 12–19, 39.
Jamie Scobie writes from extensive experience in Scottish secondary education, including pastoral care, data for improvement, and school self-evaluation. This blog is an independent publication: he writes in a personal capacity as the creator of Learning Lens, writing about classroom observation, teaching evidence, and education policy. He speaks here only for himself and for Learning Lens, not for any employer or other organisation. He holds an MSt from Cambridge (Distinction) and a Masters from Stirling.