← All articles

Learning science

Bloom's 2 sigma problem, and whether it still holds

In 1984 Bloom reported that one-to-one tutoring moved average students two standard deviations ahead. The number is widely quoted, rarely replicated, and still points somewhere true.

5 min read

One empty chair facing a teacher's chair in a quiet classroom with rows of desks behind

Key takeaways

  • Bloom reported that students tutored one to one with mastery requirements performed about two standard deviations above conventionally taught students.
  • Two sigma would move an average student above roughly 98 per cent of a conventional class, which is an extraordinary claim.
  • The studies behind it were small, short, and conducted by graduate students under conditions unlike a normal classroom.
  • Later work has rarely reproduced the full effect. Modern tutoring meta-analyses find around 0.3 standard deviations.
  • The durable part is the direction, not the magnitude: individual attention with mastery checks is the strongest lever we know of, and the problem Bloom posed - affording it for everyone - is still open.

In 1984 Benjamin Bloom reported that students taught one to one, with mastery requirements, performed about two standard deviations above conventionally taught students - which would move an average student above roughly 98 per cent of a normal class. The finding is quoted constantly in education technology and has rarely been reproduced at that size. What has held up is the direction rather than the magnitude.

What Bloom actually reported

The 1984 article compared three conditions:

  1. Conventional teaching. A class of about thirty, periodic tests, move on regardless.
  2. Mastery learning. The same class size, but with formative tests and corrective work before progressing.
  3. Tutoring. One to one or very small groups, with the same mastery requirements.

Mastery learning produced roughly one standard deviation over conventional. Tutoring produced roughly two.

Bloom's framing was explicitly a challenge rather than a product claim. He called it a problem because tutoring every child is economically impossible, and asked researchers to find group methods that achieve similar results.

Why two sigma is such a large claim

In education research, 0.2 standard deviations is a modest effect, 0.4 is large, and anything above 0.5 invites questions about the design. Two standard deviations is in a different category altogether: it would take a student at the 50th percentile to roughly the 98th.

Effects that large almost never survive contact with ordinary conditions, and this one largely has not.

What the caveats are

The studies were small. A handful of studies with modest samples, conducted as graduate dissertations under Bloom's supervision.

They were short. Weeks, on specific units, not years of schooling.

The conditions were favourable. Enthusiastic tutors, closely supervised, teaching material designed for the experiment.

The outcomes were narrow. Achievement on tests of the specific taught material rather than broad measures.

None of this means the work was dishonest. It means the number describes those studies rather than the world.

What modern evidence shows

The best current estimate for tutoring comes from a 2024 meta-analysis of 89 randomised experiments, which found a pooled effect of 0.288 standard deviations, with the largest effects for trained tutors, earlier grades and at least three sessions a week.

That is about one seventh of Bloom's figure. It is also, by the standards of this literature, a strong result: tutoring remains one of the most reliable interventions anybody has tested.

So the honest summary is: tutoring works, well, and not miraculously.

What survives

One-to-one beats group. Consistently, across methods and decades.

Mastery requirements matter. The idea that a student should demonstrate competence before moving on is separable from tutoring, and it was worth about a standard deviation on its own in Bloom's data. It remains uncommon in practice because a fixed timetable is easier to administer.

The economic problem is unresolved. Bloom's actual question - how to deliver the benefits of individual attention at the cost of group instruction - is exactly as open now as in 1984. Technology is one proposed answer; it has not yet been shown to be the answer.

What mastery learning was, separately

The two-sigma result is usually quoted as being about tutoring, and Bloom's design had two ingredients. The second gets forgotten and is the more transferable one.

Mastery learning means a student demonstrates competence on a unit before moving to the next, rather than progressing on a fixed timetable regardless of what they absorbed. In Bloom's data this was worth roughly one standard deviation on its own, in ordinary classes of thirty.

It remains uncommon for administrative reasons rather than pedagogical ones. A timetable is easier to run than a readiness check, and a class where students progress at different rates is harder to staff, assess and report on.

The practical consequence for anybody teaching one student - a homeschooling parent, a tutor, a piece of software - is that half of Bloom's effect is available without one-to-one being the expensive part. You just have to be willing not to move on.

Why later studies found less

Several reasons, and none of them is that Bloom made it up.

Regression toward the mean in replication. Striking first results are usually followed by smaller ones, across every field that has looked.

The original conditions were unusually good. Enthusiastic tutors, close supervision, material designed for the study, short duration.

Narrow outcome measures. Tests of exactly the material taught, which flatter any intervention.

Different implementations. "Mastery learning" in a large school programme is not the same thing as in a controlled study, and fidelity drops as scale rises.

What the honest version of the claim is

Three sentences, which is all the evidence supports:

  1. One-to-one teaching with mastery requirements is the strongest intervention in the education literature.
  2. Its size in ordinary conditions is closer to 0.3 standard deviations than to 2.0.
  3. Nobody has yet solved the problem Bloom actually posed, which was how to afford it for everyone.

Anybody citing two sigma for a product is skipping the second sentence, and the second sentence is the one a parent deciding how to spend money needs.

How to read the claim when you see it

When an education product cites two sigma, it is making an argument by association: Bloom found tutoring produces two sigma, this product is like tutoring, therefore. Both steps are doing a great deal of unearned work. The honest version is narrower: individual attention with mastery checking is the strongest lever we know of, and a product built around it is building on something real, which is not the same as having produced the effect.

Where TruLearn fits

We cite Bloom, and this article is how we cite him: as the reason one-to-one teaching with mastery requirements is worth trying to automate, not as a promise about a score.

TruLearn teaches one student at a time and requires a concept to be explained back correctly before it treats it as learned, which is Bloom's two conditions. What we do not have is evidence of our own effect, which we state plainly. Our bar for ourselves is the one Bloom's own critics applied: an external, objective measure, held to honestly.

More: what makes tutoring work, and the explain-it-back method.

Frequently asked questions

What is Bloom's 2 sigma problem?
In a 1984 article, Benjamin Bloom reported that students taught one to one with mastery learning performed about two standard deviations better than students taught conventionally. The problem in the name is economic: one-to-one tutoring for every student is unaffordable, so the challenge Bloom set was to find group methods that achieve similar results.
Has the two sigma result been replicated?
Not at that magnitude. The original studies were small and conducted under favourable conditions. Contemporary meta-analyses of tutoring find effects closer to 0.3 standard deviations, which is still large for education but far from two.
Does that mean Bloom was wrong?
It means the number was specific to its conditions rather than a general law. What has held up is the ordering: one-to-one instruction with mastery requirements outperforms conventional group teaching consistently. What has not held up is the size.
What is mastery learning?
An approach where students must demonstrate competence on one unit before moving to the next, rather than progressing on a fixed timetable regardless of what they absorbed. Bloom's tutored condition combined one-to-one teaching with this requirement, and both parts contributed.
Why is this quoted so often in edtech?
Because it is the most compelling available argument for personalised instruction, and because two standard deviations is a dramatic number. It is frequently cited without the caveats, which is a reason to be sceptical of anyone invoking it as a product claim, ourselves included.

Keep reading

Join the beta.

Create an account and we'll activate it as places open. Free during the beta.

Bloom's 2 sigma problem: what it says and what holds up