Learnya/ Blog
NAO humanoid robot used in research and education
Jiuguang Wang, Wikimedia Commons  · CC BY-SA 3.0

In 1984, Benjamin Bloom framed a now-famous challenge: a student taught one-to-one with mastery learning outperformed classroom peers by roughly two standard deviations — the '2 sigma problem' [1]. Human tutoring worked, but was economically impossible to generalise. Forty years on, large language models revive that promise: a patient, always-available tutor for every learner. The evidence is starting to arrive, and it is more nuanced than the marketing demos.

Encouraging randomised trials

In 2025, a randomised trial at Harvard compared an AI tutor built on pedagogical principles with an active-learning class: students with the AI tutor learned more, in less time, and reported being more engaged [2]. Another trial, in UK secondary classrooms, showed that students guided by a LearnLM-based system did at least as well as those tutored by a human, and were slightly better at solving novel problems [3].

AI can also augment the human tutor rather than replace them. In the Tutor CoPilot trial, an AI co-pilot whispering real-time suggestions to tutors raised student mastery by 4 points on average, and by 9 points for students of the lowest-rated tutors — for about 20 dollars per tutor per year [4].

But caution remains warranted

Not every study finds a positive effect. An evaluation of Khanmigo for learning scientific concepts found no statistically significant difference versus a plain Google search, even though students viewed the tool positively [5]. Above all, models' pedagogical skill remains uneven: in a shared evaluation task, the best AI tutors scored only modestly at reliably identifying and remediating student mistakes [6].

  • Confidently hallucinating a wrong explanation is still possible.
  • Diagnosing the real source of a mistake is harder than giving the right answer.
  • Many studies are short and measure immediate performance, not durable retention.

The missing ingredient: traces of real tutoring

A promising route is to specialise models on genuine tutoring sessions. The TeachLM project, trained on thousands of hours of authentic one-to-one tutoring, doubled learner talk-time and increased the number of exchanges, with a questioning style closer to that of a skilled teacher [7]. In other words, a good AI tutor does not talk more: it makes the student talk and think more.

The goal is not a model that answers well, but a model that makes people learn — two distinct skills.

Learnya synthesis

For institutions — especially in Europe and Switzerland, attentive to privacy — the stakes go beyond performance: an AI tutor processes sensitive learning data and must sit within controlled governance. Bloom's dream is becoming technically credible, but it will be decided in the details of design, evidence and trust, not in model power alone.

From research prototype to field product

One point deserves emphasis: most encouraging results come from systems carefully engineered by research teams and tested under controlled conditions. Moving to a product deployed across hundreds of classrooms, with non-specialist teachers, unstable connections and varied curricula, remains a considerable leap. A tutor that shines in the lab can disappoint in the real world if its integration into the course, its robustness and its maintenance have not been thought through.

The quality of the evidence itself calls for nuance. Shared evaluation tasks show models still struggle to finely diagnose an error or choose the right hint [6]. Many studies focus on a single subject, over a short period, and measure immediate gains. We still lack longitudinal follow-ups that would verify retention several months later, or the effect on motivation and autonomy over the long run.

Finally, an AI tutor never exists alone: it fits into a classroom ecology. The best results combine AI and teacher, as shown by the trial where the system mainly supports the least experienced tutors. The right question is therefore not 'does AI do as well as a human?' but 'what division of roles maximises learning?'. That is a choice of pedagogical design and organisation, not just of model.

What we will understand better by 2027

The field is entering a maturity phase where the questions change in nature. We no longer only ask whether an AI tutor can help a student progress, but under what conditions, for which profiles and at what real pedagogical cost. Upcoming trials will need to compare not AI against nothing, but different tutor designs against each other, and track effects over several months.

Three work streams will dominate: robustness in real conditions, far from the lab; articulation with the teacher's work, who remains the orchestrator of meaning; and control of learning data, the condition of families' trust. The ideal tutor will not be the most talkative, but the one that knows when to stay silent to let the student search — and that accounts for what it does with their data.

Sources

  1. 1. The 2 Sigma Problem: The Search for Methods of Group Instruction as Effective as One-to-One Tutoring , Bloom, B. S. , Educational Researcher , 1984 https://journals.sagepub.com/doi/10.3102/0013189X013006004
  2. 2. AI tutoring outperforms in-class active learning: an RCT in an authentic educational setting , Kestin, G., Miller, K., Klales, A., et al. , Scientific Reports (Nature) , 2025 https://www.nature.com/articles/s41598-025-97652-6
  3. 3. AI tutoring can safely and effectively support students: an exploratory RCT in UK classrooms , LearnLM Team (Google) & Eedi , arXiv , 2025 https://arxiv.org/abs/2512.23633
  4. 4. Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise , Wang, R. E., Ribeiro, A. T., Robinson, C. D., Loeb, S., & Demszky, D. , arXiv (Stanford) , 2024 https://arxiv.org/abs/2410.03017
  5. 5. Leveraging Khanmigo Generative AI-Powered Tool for Personalized Tutoring to Learn Scientific Concepts , Slijepcevic, N., & Yaylali, A. , Journal of Teaching and Learning (ERIC) , 2025 https://eric.ed.gov/?id=EJ1487444
  6. 6. Findings of the BEA 2025 Shared Task on Pedagogical Ability Assessment of AI-powered Tutors , Kochmar, E., Maurya, K. K., Petukhova, K., Srivatsa, K. V. A., Tack, A., & Vasselli, J. , arXiv (BEA 2025 Workshop) , 2025 https://arxiv.org/abs/2507.10579
  7. 7. TeachLM: Post-Training LLMs for Education Using Authentic Learning Data , Perczel, J., Chow, J., & Demszky, D. , arXiv , 2025 https://arxiv.org/abs/2510.05087
← All articles