How we use effect sizes
Learning Lens quotes research figures in a few places — on the landing page, and in the research context attached to each of the eleven practice categories. This page sets out what those numbers mean, where they come from, and the rules we hold ourselves to when we print one.
It exists because effect sizes are unusually easy to misuse, and because we got this wrong ourselves before we got it right.
What an effect size is
An effect size expresses the size of a difference in standardised units, so that results measured on different tests, in different subjects, with different age groups can be placed on a common scale. Cohen's d, the measure used throughout John Hattie's Visible Learning synthesis, expresses a difference between two groups in standard deviations.
The Education Endowment Foundation presents the same kind of underlying statistics differently — as months of additional progress — on the grounds that it is easier for schools to reason about. The two are not interchangeable. Wherever we quote a figure, we say which measure it is.
What 0.40 means — and what it does not
Hattie's much-quoted 0.40 is the average effect size across all the influences in his synthesis. He calls it the hinge point, and it exists to give a reference: an influence above it is above average for the things education researchers have chosen to study.
That is all it is.
- It is not a test of statistical significance.
- It is not a pass mark, and nothing below it is thereby ineffective.
- It is not a measure of what matters most in a particular school. Plenty of things a school must do well sit below the average of everything that has been studied, and several high-scoring influences are impractical, expensive, or depend entirely on context.
We do not use 0.40 as a threshold anywhere in Learning Lens, and we do not colour-code practices by their distance from it.
What these numbers are averages of
Every figure on this site is an average across many studies of a broad influence. Two consequences follow, and we try to keep both visible.
The average conceals the range. Feedback is the standard illustration. The pooled figure is high, but the studies behind it run from strongly positive to actively harmful, depending on what kind of feedback was given, to whom, and whether they had any opportunity to act on it. An average tells you where the research points in general. It tells you very little about a particular lesson.
An influence is not a practice. "Classroom discussion" is a category of study, not a technique a teacher performs on a Tuesday afternoon. Attaching a category's average to one specific classroom practice makes the number look far more precise than it is.
Until August 2026 this product did exactly that: category-level averages were attached to individual practices, displayed against a threshold, and used to rank which areas a teacher should develop next. That was wrong, and removing it is what prompted this page.
Where our figures come from
Research context in Learning Lens sits at category level. Each of the eleven practice categories carries a short summary of what the body of research on that broad approach has found, together with the sources behind it. Individual practices carry no effect size of their own.
Where we quote a figure, we state four things:
- The metric — Cohen's d, or months of progress.
- The scope — what the figure is an average of.
- The source and its edition — figures move between editions, so an unqualified "Hattie (2009)" is not good enough.
- Any serious challenge to it, where one exists.
We hold every figure to one of three tiers:
- Tier A — verified against the primary source, with the exact table or page recorded, the metric named, and the comparison condition noted. Only Tier A figures are shown as numbers.
- Tier B — independent reputable sources agree, but the primary source has not yet been checked. May be described in prose; never printed as a numeral.
- Tier C — unverified, contested, or not traceable to a stated source. Not shown at all.
The criticisms we take seriously
That ranked effect sizes imply a precision the data cannot support. Simpson (2017) and others have argued that figures drawn from different meta-analyses — different methodologies, different outcome measures, different populations — cannot meaningfully be placed on a single ranked scale. We think this criticism is largely right. It is why Learning Lens ranks nothing: not practices against each other, not categories, not teachers.
That the headline feedback figure is too high. The pooled barometer figure of around d = 0.73 comes from a synthesis of existing meta-analyses. Wisniewski, Zierer and Hattie (2020) went back to primary studies — 435 of them — and arrived at a considerably more modest d = 0.48, alongside heterogeneity so large that they concluded feedback "cannot be understood as a single consistent form of treatment". Notably, Hattie is a co-author of the paper that revised his own number downwards. We quote the figure our cited source reports, and we note this one beside it.
So why print numbers at all? Because a figure with its provenance and its limits attached is more useful to a teacher than either a naked number or no number. Removing the figures would make this product easier to defend and less honest about where its claims come from.
What we do not do with these numbers
Learning Lens is descriptive, not evaluative. It records what was observed in a classroom. It does not grade lessons, rate teachers, or rank practices.
Concretely:
- No practice is presented as better or more important than another.
- No effect size is attached to an individual teacher's data, and no observation is scored against one.
- Development suggestions are not ordered by effect size — they follow the taxonomy's own order.
- No colour scale, traffic light, or threshold marks any practice as above or below a line.
Presenting the evidence accurately is our responsibility. What it means in a particular classroom, with particular pupils, is a professional judgement that belongs to the school.
Verification
Every figure in the product is being checked against its primary source. For Hattie's influences that means the printed editions — Visible Learning (2009) and Visible Learning for Teachers (2012) — recording the exact table, the number of meta-analyses and studies pooled, and the comparison condition, since figures differ between editions. Until a figure reaches Tier A it rests on agreement between independent secondary sources, and is described rather than asserted.
EEF figures are re-checked annually and whenever the Teaching and Learning Toolkit is revised; the metacognition and self-regulation strand, for example, has moved between Toolkit versions.
Last reviewed: August 2026.
Sources referred to on this page
- Hattie, J. Visible Learning (2009)
- Hattie, J. Visible Learning for Teachers (2012)
- Hattie, J. & Timperley, H. "The Power of Feedback", Review of Educational Research 77(1), 81–112 (2007)
- Wisniewski, B., Zierer, K. & Hattie, J. "The Power of Feedback Revisited: A Meta-Analysis of Educational Feedback Research", Frontiers in Psychology 10, 3087 (2020)
- Simpson, A. "The misdirection of public policy: comparing and combining standardised effect sizes", Journal of Education Policy 32(4), 450–466 (2017)
- Education Endowment Foundation, Teaching and Learning Toolkit