AI Tutors: The Old Dream of the Private Tutor, Finally Within Reach?
Randomised trials show that well-designed AI tutors can rival active learning — but the gap between research prototype and product remains wide.
Randomised trials show that well-designed AI tutors can rival active learning — but the gap between research prototype and product remains wide.
In 1984, Benjamin Bloom framed a now-famous challenge: a student taught one-to-one with mastery learning outperformed classroom peers by roughly two standard deviations — the '2 sigma problem' [1]. Human tutoring worked, but was economically impossible to generalise. Forty years on, large language models revive that promise: a patient, always-available tutor for every learner. The evidence is starting to arrive, and it is more nuanced than the marketing demos.
In 2025, a randomised trial at Harvard compared an AI tutor built on pedagogical principles with an active-learning class: students with the AI tutor learned more, in less time, and reported being more engaged [2]. Another trial, in UK secondary classrooms, showed that students guided by a LearnLM-based system did at least as well as those tutored by a human, and were slightly better at solving novel problems [3].
AI can also augment the human tutor rather than replace them. In the Tutor CoPilot trial, an AI co-pilot whispering real-time suggestions to tutors raised student mastery by 4 points on average, and by 9 points for students of the lowest-rated tutors — for about 20 dollars per tutor per year [4].
Not every study finds a positive effect. An evaluation of Khanmigo for learning scientific concepts found no statistically significant difference versus a plain Google search, even though students viewed the tool positively [5]. Above all, models' pedagogical skill remains uneven: in a shared evaluation task, the best AI tutors scored only modestly at reliably identifying and remediating student mistakes [6].
A promising route is to specialise models on genuine tutoring sessions. The TeachLM project, trained on thousands of hours of authentic one-to-one tutoring, doubled learner talk-time and increased the number of exchanges, with a questioning style closer to that of a skilled teacher [7]. In other words, a good AI tutor does not talk more: it makes the student talk and think more.
The goal is not a model that answers well, but a model that makes people learn — two distinct skills.
Learnya synthesis
For institutions — especially in Europe and Switzerland, attentive to privacy — the stakes go beyond performance: an AI tutor processes sensitive learning data and must sit within controlled governance. Bloom's dream is becoming technically credible, but it will be decided in the details of design, evidence and trust, not in model power alone.
One point deserves emphasis: most encouraging results come from systems carefully engineered by research teams and tested under controlled conditions. Moving to a product deployed across hundreds of classrooms, with non-specialist teachers, unstable connections and varied curricula, remains a considerable leap. A tutor that shines in the lab can disappoint in the real world if its integration into the course, its robustness and its maintenance have not been thought through.
The quality of the evidence itself calls for nuance. Shared evaluation tasks show models still struggle to finely diagnose an error or choose the right hint [6]. Many studies focus on a single subject, over a short period, and measure immediate gains. We still lack longitudinal follow-ups that would verify retention several months later, or the effect on motivation and autonomy over the long run.
Finally, an AI tutor never exists alone: it fits into a classroom ecology. The best results combine AI and teacher, as shown by the trial where the system mainly supports the least experienced tutors. The right question is therefore not 'does AI do as well as a human?' but 'what division of roles maximises learning?'. That is a choice of pedagogical design and organisation, not just of model.
The field is entering a maturity phase where the questions change in nature. We no longer only ask whether an AI tutor can help a student progress, but under what conditions, for which profiles and at what real pedagogical cost. Upcoming trials will need to compare not AI against nothing, but different tutor designs against each other, and track effects over several months.
Three work streams will dominate: robustness in real conditions, far from the lab; articulation with the teacher's work, who remains the orchestrator of meaning; and control of learning data, the condition of families' trust. The ideal tutor will not be the most talkative, but the one that knows when to stay silent to let the student search — and that accounts for what it does with their data.