Learnya/ Blog
Student writing an exam paper in a classroom
Alison wood, Wikimedia Commons  · CC BY 3.0

Grading is one of teaching's most time-consuming tasks. Applying LLMs to it is tempting: short-answer scoring, essay evaluation, even an 'LLM-as-judge' arbitrating the quality of a piece of work. The time savings are real. But 2025 research urges a careful distinction between what AI does well and what it does badly.

Where it works: short answers

On well-scoped short answers, the best models reach high correlations with human graders — up to very strong values in some studies [1]. Constrained-format scoring (definitions, calculations, factual items) is the most favourable ground. But as soon as the answer becomes open-ended, technical or long, variability rises sharply [1].

Where it gets harder: essays and judgement

For essays, the picture is darker. A study across five LLMs and university essays finds low human-machine agreement and poor internal reliability, with models inflating coherence scores in particular [2]. Human + AI composite scoring improves reliability over AI alone, which argues for hybrid models rather than full automation [3].

The blind spot: fairness

The most worrying risk is fairness. Work shows automated scores diverge from human scores especially for second-language learners [6]. Grading deployed without a bias audit can therefore systematically penalise certain groups while wearing the appearance of algorithmic objectivity.

  • Good for: short answers, rapid formative feedback, pre-screening.
  • Handle with care: essays, summative grading, high-stakes decisions.
  • Always audit: linguistic, cultural and format biases.

An automated grade is only reliable if you know its errors. Grading AI should augment the teacher's judgement, not confiscate it.

Learnya synthesis

The emerging consensus is clear: AI excels at generating rapid formative feedback and reducing workload, but high-stakes grading must stay under human control. It is also a regulatory requirement: in the European Union, evaluating learning outcomes falls under the high-risk uses governed by the AI Act. Automate, yes; abdicate judgement, no.

Formative first, summative with caution

The most useful distinction for deploying AI grading contrasts formative with summative assessment. In formative use — to help the student progress — rapid feedback, even imperfect, has value: it arrives at the right moment and can be discussed. In summative use — when the grade counts and shapes the future — the slightest error or bias becomes unacceptable. That is precisely where research recommends keeping humans in charge [5].

One must also understand why models err. An LLM does not 'understand' a paper the way a grader does: it estimates a probability from textual regularities. Hence documented artefacts — over-scoring coherence, sensitivity to criterion order, length or style [4]. These biases are not random; they can favour one kind of writing over another, raising a fairness problem when populations differ.

In practice, responsible assisted grading calibrates the model on papers already scored by humans, explicitly measures agreement and fairness across groups, and reserves final validation of high-stakes decisions for the teacher. Human + AI composition, more reliable than AI alone [3], is no timid compromise: it is the architecture that captures the time savings without abandoning pedagogical and legal responsibility.

Rethinking assessment, not just automating it

The arrival of AI should not only speed up existing grading; it invites us to rethink what we assess. If a model writes a competent essay in seconds, the question is no longer only how to grade the essay, but which competencies we really want to certify — reasoning, arguing, creating, collaborating — and how to observe them reliably.

Rising formats — oral defences, process assessment, portfolios — are both more AI-resistant and more aligned with deep learning. AI finds its full place there as support for the grader: preparing, structuring, flagging, proposing a first draft of feedback. The direction is clear: automate what can be automated without risk, reserve human judgement for decisions that shape learners' futures, and relentlessly measure the fairness of the system.

A further consideration is validity over time. Because models and their behaviour change with each update, an assessment pipeline validated once may drift silently. Institutions that treat AI grading as a living system — periodically re-checking agreement with human raters, re-testing for bias across student groups, and versioning their prompts and rubrics — are far better protected than those who validate once and forget. Assessment is a chain of trust, and every link, human or machine, must be maintained rather than assumed.

Sources

  1. 1. LLM-as-a-Grader: Practical Insights from Large Language Model for Short-Answer and Report Evaluation , Byun, G., Rajwal, S., & Choi, J. D. , arXiv , 2025 https://arxiv.org/abs/2511.10819
  2. 2. Assessing the Reliability and Validity of Large Language Models for Automated Assessment of Student Essays in Higher Education , Gaggioli, A., Casaburi, F., Ercolani, G., Collova, F., Torre, F., & Davide, F. , arXiv , 2025 https://arxiv.org/abs/2508.02442
  3. 3. Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory , Song, D., Lee, W.-C., & Jiao, H. , arXiv , 2025 https://arxiv.org/abs/2507.19980
  4. 4. Evaluating Scoring Bias in LLM-as-a-Judge , Li, Q., Dou, S., Shao, K., Chen, C., & Hu, H. , arXiv , 2025 https://arxiv.org/abs/2506.22316
  5. 5. Evaluating large language models as graders of medical short answer questions: a comparative analysis with expert human graders , Bolgova, O., Ganguly, P., Ikram, M., & Mavrych, V. , Medical Education Online , 2025 https://pmc.ncbi.nlm.nih.gov/articles/PMC12377152/
  6. 6. Evaluating LLM-Based Automated Essay Scoring: Accuracy, Fairness, and Validity , Huang, Y., & Wilson, J. , ACL Anthology (AIME-Con) , 2025 https://aclanthology.org/2025.aimecon-wip.9/
  7. 7. LLM-based Automated Grading with Human-in-the-Loop (GradeHITL) , Chu, H., Li, Y., Yang, T., Copur-Gencturk, Y., & Tang, J. , arXiv / IEEE TALE 2025 , 2025 https://arxiv.org/abs/2504.05239
← All articles