6 min read

The 2 Sigma Problem

reading listlearning sciencetutoring

In 1984 Benjamin Bloom put a number on something teachers had always suspected, which is that a good tutor working with one student is worth a great deal more than the same teacher working with thirty. The number turned out to be big enough to be uncomfortable, and what he did with it was turn it into a challenge rather than a boast.

What the paper actually says

Bloom (1984) reported that students taught one-to-one with mastery learning performed about two standard deviations better than students taught the same material in an ordinary class. Two standard deviations is a very long way. It means the average tutored student ended up scoring above roughly 98% of the students in the conventional classroom.

He didn't stop at that headline, though, and the finding sits on a ladder rather than standing alone. An ordinary class is the baseline. Take that same class and run it with mastery learning, meaning teaching to a defined standard, testing often, and giving corrective feedback until each student reaches it, and the average student moves up by about one standard deviation, to roughly the 84th percentile of where they started. It's only when mastery learning gets delivered one-to-one that you reach two sigma.

And that gap is what the paper is really about. One-to-one tutoring is simultaneously the best instruction anyone knows how to give and the least affordable thing to give at scale. So Bloom's question was never whether tutoring works, which is obvious enough. It was whether group methods could be found that get anywhere near it.

He called that the 2 sigma problem, and spent the rest of the paper cataloguing what he called alterable variables: the levers a teacher or a system can genuinely change, like feedback and correction, reinforcement, the quality of cues and explanations, time on task, student participation, each ranked by how much it actually moves the outcome. The bet was that some well-chosen combination of those might approach tutoring without requiring one tutor per child.

Why it's lasted

Bloom wasn't standing outside this field commenting on it, he largely built it. This is the same Benjamin Bloom of Bloom's taxonomy and of mastery learning, writing near the end of a long career about the exact thing he'd spent that career studying, and the paper wears that authority quite lightly.

It's also short. A dozen pages in Educational Researcher, readable in a single sitting. Landmark papers are usually heavy going and this one is almost conversational, with an influence completely out of proportion to its length. It named a problem crisply enough that four decades of researchers have organised their work around the name, which is most of what a great paper ever does.

What reading it changes

Before Bloom, the gap between a tutored student and a classroom student was easy enough to file under talent, or effort, or how much the parents cared. After Bloom it becomes a design problem, because on average the tutored student and the classroom student are the same student. The thing that changed was the instruction: somebody was watching this particular learner, noticing what they hadn't grasped yet, and refusing to move on until they had.

That reframes the classroom's weakness quite precisely, and not as a criticism of teachers. A class isn't worse at teaching because the teachers in it are worse. It's worse because one person can't hold thirty separate models of thirty separate students in their head, correct each error as it appears, and pace each learner independently. The things a tutor does, continuous feedback, correction before a gap compounds, attention aimed at one individual, are exactly the things that don't survive being divided by thirty.

The average tutored student and the average classroom student are the same student. Only the instruction is different.

After Bloom, 1984

The honest part

The 2 sigma figure needs holding with some care, and the paper is better for being clear about where it came from. Bloom drew it out of small, short studies, a few weeks each, run by his doctoral students with schoolchildren on tightly defined topics. It was never a general law and it hasn't been reproduced at full strength since.

When VanLehn (2011) reviewed decades of tutoring research, human tutors came out around 0.8 standard deviations ahead, with intelligent tutoring systems not far behind them. That's real and valuable and it's also roughly a third of Bloom's number.

So the honest reading is this: two sigma is the ceiling Bloom observed under near-ideal conditions, not the effect you should expect in the wild. The direction is solid, because individual attention with fast correction genuinely does beat one-size teaching. The magnitude is contested. Both of those are true simultaneously, and a careful reader holds onto both rather than picking whichever one suits them.

Some of the individual levers Bloom pointed at have since been pinned down on their own, too. Frequent low-stakes testing, one of his feedback-corrective mechanisms, is now among the best-evidenced techniques in the entire field (Roediger and Karpicke, 2006; Dunlosky et al., 2013). The 2 sigma problem was never going to be solved by one idea. It was always a search for the right combination of them.

Where we stand

Two sigma is our target, not our claim. Building tutoring-style attention that reaches many people is the whole ambition, and asserting we'd matched Bloom's ceiling would be exactly the sort of thing this paper teaches you to distrust.

Who should read it

Read it if you build or choose learning tools. Every "personalised tutoring at scale" pitch, mine very much included, is a promise to make progress on Bloom's problem, and you can't judge a promise like that without the paper that framed it.

Read it if you teach, particularly if you're drawn to mastery learning. And read it if you're an ambitious self-learner, because it explains better than any list of study hacks why finding something that gives you real feedback beats grinding away on your own.

Almost nobody should skip it. It's short, it's foundational, and it's genuinely a pleasure to read. The only exception is the casual reader who wants the headline and nothing else, and the headline is one sentence: one-to-one tutoring beat the classroom by about two standard deviations, and the lasting challenge is getting group instruction anywhere close. If that's all you needed, you've got it. Everyone else should spend the hour.

For where this sits in the wider evidence, see the bookshelf behind the method.

Sources

Read more