In 1984 the educational psychologist Benjamin Bloom published a finding that has haunted every reformer since. Give a child one-to-one tutoring, he showed, and the average student rises about two standard deviations — from the middle of the class to the top few per cent. Tutoring works. It simply costs more than any system can afford at scale.
For four decades that has been education’s cruel arithmetic: we know what works, and we cannot pay for it. Then, last year, a careful study in Nigeria suggested the machine might finally break the trade-off — and the result was large enough to make me sit up.
It is also, I should say at once, a result to read carefully. The promise is real. So is the catch.
We have spent forty years trying to scale the one thing that works. The machine may finally let us — for those who can reach it.
What follows: Bloom’s old challenge, the first hard evidence that AI can meet it, the catch nobody should ignore, and what actually works in a classroom.
Bloom’s two-sigma challenge
Bloom’s finding was never really about tutoring. It was a gauntlet thrown to the rest of us: find a way to give thirty children at once what one patient tutor gives one. Mastery learning, better materials, one technology after another — each closed part of the gap; none closed all of it.
The two-sigma prize stayed out of reach for a simple reason. The thing that delivered it — an expert giving full, responsive attention to a single learner — was precisely the thing that would not scale. We could describe the cure. We could not afford to prescribe it.
So a generation of reform learned to aim lower, and to call it realism. That is the backdrop against which last year’s evidence should be read.

We could always describe the cure. We just could never afford to prescribe it.
The first real proof
Which is why the Nigerian trial matters. Working with the World Bank, researchers ran a six-week, after-school programme in Edo State: first-year secondary students, twice a week, working through English with a generative-AI tutor — GPT-4, via Microsoft Copilot — guided by carefully designed prompts. A comparable group carried on as normal.
The treated students gained about 0.3 of a standard deviation. In the currency that matters, that is the equivalent of roughly one and a half to two years of ordinary schooling, compressed into six weeks — placing it among the most cost-effective education interventions ever measured. The more sessions a student attended, the more they gained; the effect even showed up in their end-of-year exams, on material the programme never touched.
Mind the gap
Now the careful reading. A six-week pilot is not a system, and a tutor that needs a device, a connection and reliable power is not yet a tutor for everyone. Some 2.6 billion people are still offline. If AI tutoring scales to those already connected and not to those who are not, it will not close the learning gap — it will widen it, faster than anything before.
There is a subtler risk, too. A tool that hands over a fluent answer can build dependence as easily as understanding. The Nigerian gains did not come from a chatbot left alone with children; they came from structured prompts, trained facilitators and a curriculum. The intelligence was on tap. The judgement stayed human.

Augment the teacher, don’t replace them
So I would resist both the hype and the backlash. The lesson of the evidence is narrow and useful: AI tutoring works when it augments a teacher, and disappoints when it is asked to replace one. UNESCO’s guidance lands in the same place — human-centred, age-appropriate, teacher firmly in the loop.
Think of it as a division of labour. The machine supplies patience and personalisation at a scale no staffroom can muster — endless worked examples, infinite re-explanations, no fatigue. The teacher supplies the relationship, the judgement and the care that no model has. Used that way, the technology does not de-skill the classroom; it gives the teacher back the time to do the human part well.
The intelligence can be on tap. The teaching — the relationship, the judgement, the care — stays human.

For anyone deploying AI in learning
Whether you run a school, a university or a corporate academy, the same five rules hold.
Augment, don’t replace
Use AI as the teacher’s force-multiplier, with a human firmly in the loop — not as a substitute for one.
Mind the gap first
Design for access before scale. An intervention that reaches only the connected widens the inequality it claims to fix.
Structure the use
Prompts, facilitation, curriculum. The Nigerian result came from a designed programme, not an unsupervised chatbot.
Teach the judgement
Build the skill of questioning the machine, not merely accepting it. Dependence is the failure mode to design out.
Measure learning, not usage
Track understanding and outcomes, not screen-time. What you measure is what you will optimise for.
Get this right and Bloom’s forty-year-old challenge has, at last, an affordable answer.
The machine can finally give every child a tutor. Whether it gives every child one is a choice — and the choice is ours.
- Bloom, B. S. “The 2 Sigma Problem.” Educational Researcher, 1984.
- World Bank. “From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria,” 2024–25 (Edo State; ~0.3 SD effect).
- UNESCO. Guidance for Generative AI in Education and Research, 2023.
- ITU. Facts and Figures 2023 — ~2.6 billion people offline.
