Study: AI Cannot Reliably Grade Student Essays Like Human Teachers
New research from Cardiff University and the University of Melbourne finds that generative artificial intelligence (GenAI) cannot yet reliably evaluate student essays with the same consistency as human educators. ChatGPT's grading showed significant variability.

A new study from Cardiff University and the University of Melbourne has found that generative artificial intelligence (GenAI) is currently unable to reliably evaluate student essays with the same consistency as human teachers. Researchers tested ChatGPT's ability to grade 50 undergraduate essays in biosciences using seven criteria.
The study employed four different prompting methods and compared the AI's scores against the average scores and score fluctuations from human graders. While the overall scores from AI and humans were relatively close, significant discrepancies emerged when examining individual criteria and specific essays.
In most cases, the average scores provided by the AI were higher than those given by human evaluators. The largest difference in average scores was 16.1 points, with individual essay score differences reaching up to 40 points. The AI also tended to compress the score range, lowering scores for high-achieving essays and raising scores for lower-achieving ones.
The findings suggest that AI is not currently suitable for replacing teachers in grading assignments, particularly for subjective written work. While the evaluation capabilities of AI may improve with further development, achieving consistency with human judgment is likely to be challenging.