Learnya/ Blog

Voice, Image, Text: Multimodal AI Enters the Classroom

Tutors that see a diagram, hear a read-aloud and read a notebook: where multimodal AI in education really stands.

Student working at a laptop computer
Courtcourtwest, Wikimedia Commons  · CC BY-SA 4.0

Until 2023, educational AI mostly spoke text. The arrival of multimodal models — able to process text, image, audio and sometimes video together — changes the nature of possible interactions. A learner can photograph a handwritten geometry exercise, read a passage aloud to practise pronunciation, or show a biology diagram and ask for an explanation. Recent surveys of LLMs in education identify this multimodal shift as a major axis for 2025-2027 [1].

What multimodality makes possible

  • Understanding handwritten work or a photographed diagram, not just typed text.
  • Voice tutors: spoken dialogue, reading practice, language and oral skills.
  • Multimodal learning analytics: gesture, prosody, in-class collaboration.
  • Accessibility: image description for visually impaired students, captioning, read-aloud.

These capabilities rest on recent models that handle speech in and out almost natively, paving the way for credible conversational tutors. On the research side, multimodal learning analytics now exploits sensors, voice and traces to finely understand processes previously invisible, such as collaboration or engagement [2].

A trap: solving well is not teaching well

Enthusiasm must be tempered. Multimodal models are often evaluated on their ability to solve a scientific problem, not to teach it. A recent study shows that a model excellent at solving can be a poor tutor: it gives the answer instead of guiding, or explains a figure badly [3]. Pedagogical skill — asking the right question, diagnosing the error, metering the hint — does not correlate with raw performance.

Teachers themselves voice precise expectations and reservations. A survey of K-12 educators shows strong interest in multimodal uses (grading handwritten work, oral feedback) but concern about reliability, verification workload and student data protection [4].

Privacy: multimodality raises the stakes

Voice, faces, handwriting, classroom video: multimodal data is more sensitive than anonymous text. Voice is biometric data; the image of a minor warrants heightened protection. UNESCO stresses that deploying educational AI must be subordinate to data protection and a human-centred vision [5], and the 2025 AI Index highlights the persistent gap between adoption and governance [6]. For European and Swiss institutions, this argues for architectures where multimodal processing stays under control — ideally sovereign — rather than delegated to opaque services.

Multimodality does not only bring AI closer to how humans communicate; it also brings risks closer to how humans are identifiable.

Learnya synthesis

In practice, the value of multimodal AI in education will depend less on technical prowess than on three conditions: pedagogical design that favours guidance over the solution, human verification of sensitive outputs, and strict governance of multimodal personal data.

Concrete use cases, subject by subject

Multimodality makes full sense once you descend to the level of subjects. In languages, a voice tutor lets learners work on pronunciation and fluency, with feedback impossible in writing. In maths and science, the ability to read a handwritten diagram or a figure is a game-changer for grading and explanation. In art and history, image analysis opens rich dialogues around works and documents.

But each modality adds its own fallibility. Speech recognition stumbles on accents and noise; image reading confuses similar symbols or hallucinates a detail. And these errors are all the more deceptive because they come with a confident answer [3]. In education, where learners trust, an unflagged multimodal error can take durable root. Human verification of sensitive outputs therefore remains essential.

Finally, the data question arises with particular sharpness. Processing a minor's voice, face or handwriting falls under sensitive, often biometric categories. Reference frameworks — UNESCO foremost [5] — subordinate deployment to data protection and a clear pedagogical purpose. Architectures that keep multimodal processing close to the institution, under control, offer the best balance between rich uses and trust.

A richness to govern

Multimodality brings AI closer to how we actually teach and learn: through speech, gesture, image, manipulation. That is real progress for accessibility and for grounding learning in the concrete. But the more AI perceives, the more it knows about the learner — and the greater the responsibility of whoever deploys it.

The 2026-2027 horizon will likely see voice tutors and assistants able to read a notebook multiply. The dividing line between responsible and risky uses will run through governance: which multimodal data is captured, where it is processed, how long it is kept, and who accesses it. The richness of modalities is only valuable if matched by equivalent rigour on protecting people.

Sources

  1. 1. Large Language Models for Education: A Survey and Outlook , Wang, S., et al. , arXiv , 2024 https://arxiv.org/abs/2403.18105
  2. 2. Artificial intelligence in multimodal learning analytics: A systematic literature review , Multiple authors , Computers and Education: Artificial Intelligence , 2025 https://www.sciencedirect.com/science/article/pii/S2666920X25000669
  3. 3. Is Your Multimodal Large Language Model a Good Science Tutor? , Multiple authors , arXiv , 2025 https://arxiv.org/abs/2505.06418
  4. 4. Beyond Text: Probing K-12 Educators' Perspectives and Ideas for Learning Opportunities Leveraging Multimodal Large Language Models , Multiple authors , arXiv , 2025 https://arxiv.org/abs/2507.20720
  5. 5. Guidance for generative AI in education and research , UNESCO , UNESCO , 2023 https://www.unesco.org/en/articles/guidance-generative-ai-education-and-research
  6. 6. The 2025 AI Index Report — Education , Stanford Institute for Human-Centered AI (HAI) , Stanford HAI , 2025 https://hai.stanford.edu/ai-index/2025-ai-index-report/education
← All articles