Skip to main content

GenAI cannot accurately mark essays in Higher Education

24 August 2026

Female academic sitting at a desk marking papers

Generative Artificial Intelligence (GenAI) is inadequate at accurately grading essays and cannot replace humans when marking student work in Higher Education, finds new research.

Published in Assessment & Evaluation in Higher Education, new research from Cardiff University and the University of Melbourne has investigated whether GenAI could mimic human marking when evaluating written student work.

Dr William Kay, Cardiff University's School of Biosciences, said: “The rapid rise of Generative Artificial Intelligence – known as GenAI – has promoted interest in if it can support the evaluation of student work in Higher Education. We wanted to understand if large language models (LLMs) could mimic human assessment of extended written assignments well enough to be able to guide students in judging the quality of their work.”

To test the capabilities of LLMs in evaluating student work, the researchers used two versions of a popular GenAI platform, ChatGPT, to mark 50 undergraduate bioscience essays. They asked ChatGPT to mark the essays against seven assessment criteria and applied four different prompting conditions.

LLM-assigned marks were compared to human marks, evaluating both average scores and mark variability.

William Kay
Our findings indicate that marks awarded by LLMs varied considerably and were inadequate predictors of the human mark awarded to essays.
Dr William Kay Senior Lecturer in Statistics (Teaching & Scholarship)

“We found significant discrepancies between GenAI-marked essays and those marked by humans. Overall essay marks were relatively similar when evaluated by humans and GenAI, but when assessing student performance based on individual marking criteria and on an essay-by-essay basis, the differences between LLM- and human-assigned marks were substantial,” said Dr Kay.

In all but one case, LLMs typically returned higher average marks than humans.

The largest difference in the average mark awarded by any LLM compared to a human-awarded mark was 16.1 marks, and 40 marks at the individual essay level.

“We also observed that LLMs reduced the marks for high-scoring essays and inflated them for low-scoring work, resulting in systematic compression of marks towards the middle.”

William Kay
The findings of this study highlight that, at present, GenAI is unable to reliably assign a mark to a subjective piece of written work comparable to that of humans – even with extensive training of the LLM.
Dr William Kay Senior Lecturer in Statistics (Teaching & Scholarship)

“At present, the LLMs tested are not suitable to act as alternatives to human tutors for individual students to provide predicted grades on their work.

“While there is interest across the sector in whether the pattern-recognition capabilities of LLMs could facilitate objective grading of students’ work, making marking more efficient and relieving pressure on staff, the findings of this research indicate that at present this is not advisable. Aside from the ethical issues of submitting student work to GenAI tools without express consent, LLMs cannot, and should not, be relied upon for assigning grades to extended written work by students.

“As LLMs become more sophisticated, it is possible that in the future their ability to mimic human judgement may enhance. But, as we find in this study, aligning marks between humans and GenAI may be hard to achieve,” said Dr Kay.

The research, Can GenAI be trained to mimic human markers of extended written assignments in higher education?, was published in Assessment & Evaluation in Higher Education.