Doctoral students are turning towards AI to compete in the academic landscape and complete their PhD degrees (Krumsvik, 2024; Nature, 2025). And why not? Completing a PhD by research can be very stressful and difficult; particularly, the first part.

In their probationary year, doctoral students must conjure a research question that no one has answered before. Full-time students have about 12 months to justify why their question is worthy of asking and document how they will answer it.  That document needs to be around 100 pages of rigorous academic thought, including references, and may even include a number of proposed research papers they will publish in the remaining years.

After a year of thinking, research, and editing with one or more academic supervisors their work is sent away to academic experts who will review it and write back with pages of critique and questions.  Students have a few months to respond and reply to every question and suggestions, rewrite portions of your document and present it all before a live audience. Oh, and even though they have likely completed at least four years of undergraduate work with high marks awarded against rubrics and assessment criteria, they are now in the dark and at the mercy of those in power because there are no rubrics or universal learning outcomes to meet when you’re doing a PhD by research. It’s far more… human and subjective.

This human touch is not necessarily a bad thing because research and publication is full of rejection, critique, correction, re-writing, and resubmission. It would be counter-productive to award a PhD to someone without apprenticing them into the real world of research. Good supervisors and graduate schools know that PhD programs are largely formational and gaining feedback literacy is a core part of doctoral formation and socialisation. But could AI play a part in this doctoral formation and assessment? Can it replace assessors, supervisors, reviewers and all the eliminate all the delays, disappointments, and differing opinions that come with involving humans in the candidate review process?

Using AI to generate formative feedback in doctoral education.

Today (3 Aug 2025) I am celebrating my first academic journal publication as lead author!🥳

I’m very grateful to my co-authors Dr. Peter Grainger and Dr. Wayne Graham for their contribution and guidance.

Title: Using AI to generate formative feedback in doctoral education.
Journal: Assessment & Evaluation in Higher Education
7,000 words
Link: https://doi.org/10.1080/02602938.2025.2536558

The paper in plain language:

We explored the validity of using a rubric designed with probationary doctoral confirmation in mind, along with ChatGPT, to compare generative AI feedback to human feedback. Firstly, we took Grainger’s Formative Assessment Criteria-Based Tool (F.A.C.T.) rubric (Grainger et al., 2024) and loaded it to a custom private ChatGPT bot. After some testing to see if the bot could reference the rubric correctly, we uploaded a confirmation thesis and asked it to assess it against the criteria set out in the F.A.C.T. rubric.

The generative AI bot returned flattering results and 100% high marks against the 23,000-word thesis. A series of questions were asked including “This doesn’t seem critical enough. Can you re-assess it?”. The bot returned differing results across the criteria, assessing the work as “Unsatisfactory” in some areas. As prompts were altered, so were the attributed marks and outcomes from the AI bot.

Next, we compared the feedback to a human reviewer. They didn’t use the same F.A.C.T. rubric that the chatbot did but returned with warranted feedback including the need for the candidate to build a robust theoretical foundation, refine their research question and strengthen their methodology.

Observations:
The AI identified some similar areas needing improvement as human reviewers, particularly regarding contribution to knowledge. However, this may have been because the rubric used similar phrasing and LLMs have a habit of extruding and modifying text to solicit meaning (Tensen et al., 2025). However, significant differences were apparent. The human reviewer provided more contextually nuanced feedback, with specific disciplinary insights that the AI did not. Additionally, human reviewers offered more actionable and specific suggestions for improvement. AI’s suggestions, while thematically similar, remained more general and less tailored to disciplinary practices. Also, as mentioned earlier, AI changed the marks it awarded depending on the prompt which is not helpful for students seeking clarification for action and guidance. Much of these findings align with other observations in the field (Bearman & Luckin, 2020; Boud & Bearman, 2024; Carless et al., 2024)

We concluded, that at best, generative AI LLMs like ChatGPT might serve as a kind of sparring partner for doctoral students (as suggested in the literature). AI-driven feedback may one day be as supplemental or complimentary to the formation and socialisation process but by no means could it replace human reviewers. And why should it? Research is largely a field created by, driven by, and managed by passionate, clever, volunteers who love to pursue truth, express thought, and make the world a better place.

Continued research and case studies are warranted in this emerging area.

References:

Bearman, M., & Luckin, R. (2020). Preparing University Assessment for a World with AI: Tasks for Human Intelligence. In M. Bearman, P. Dawson, R. Ajjawi, J. Tai, & D. Boud (Eds.), Re-imagining University Assessment in a Digital World (Vol. 7, pp. 49–63). Springer International Publishing. https://doi.org/10.1007/978-3-030-41956-1_5

Boud, D., & Bearman, M. (2024). The assessment challenge of social and collaborative learning in higher education. Educational Philosophy and Theory, 56(5), 459–468. https://doi.org/10.1080/00131857.2022.2114346

Carless, D., Jung, J., & Li, Y. (2024). Feedback as socialization in doctoral education: Towards the enactment of authentic feedback. Studies in Higher Education, 49(3), 534–545. https://doi.org/10.1080/03075079.2023.2242888

Grainger, P., Carey, M., & Johnston, C. (2024). Why not rubrics in doctoral education? Assessment & Evaluation in Higher Education, 49(8), 1061–1073. https://doi.org/10.1080/02602938.2024.2338924

Krumsvik, R. J. (2024). Chatbots and academic writing for doctoral students. Education and Information Technologies. https://doi.org/10.1007/s10639-024-13177-x

Nature. (2025, March). ChatGPT for students: Learners find creative new uses for chatbots. https://www.nature.com/articles/d41586-025-00621-2

Tensen, D., Grainger, P., & Graham, W. (2025). Using AI to generate formative feedback in doctoral education. Assessment & Evaluation in Higher Education, 1–17. https://doi.org/10.1080/02602938.2025.2536558