The CEO of the largest genAI company announces the new ChatGPT-5 will have PhD-level expertise. You are currently testing if ANY genAI model can review and assess a doctoral confirmation document; a standard task of most tenured (human) academics.

‘How lucky!’ You think to yourself. ‘If CEO Sam Altman’s claims are right, the least ChatGPT-5 could do is provide some critical and formative feedback to a wannabe PhD student.’ So, you add the new ChatGPT-5 model to the testing lineup, and you even pay a few extra dollars to access its smart ‘Thinking’ cousin.

In this study, we compare ChatGPT-o3, ChatGPT-5, ChatGPT-5 Thinking, Claude Opus 4.1, and Gemini 2.5 Pro, using five controlled prompts, including persona-based roles, to evaluate a doctoral confirmation thesis. Oh, and we considered human and student feedback too.

The results were fascinating, including Google Gemini’s inability to grant a mark below perfect (14/14), regardless of the prompt 🥴. But you may wonder, ‘How did ChatGPT-5 and ‘Thinking’ go?’ Ummm, how do I put this… Academically speaking, I’d say ‘💩’.

Actually, ‘💩’ rarely makes it into a Q1 journal, and my seasoned collaborators advised to be more polished and professional. I wanted to talk about the irony behind the ‘polishing a turd’ idiom but chose to be professional instead.

So what are the implications of using genAI to assess doctoral confirmation work?
In the paper, we ‘caution against uncritical use of genAI for formative doctoral feedback. Without supervisory mediation and feedback literacy, students may over-trust flattering or superficially authoritative outputs, particularly under persona prompts. Risks are heightened for candidates working across languages or academic cultures, where subtle meaning-making, epistemic stance, and disciplinary conventions matter (Adel and Alani 2025; Phan 2024b; Rebera et al. 2025). While variability also occurs in human feedback, it is grounded in expertise, context, and socialisation—qualities genAIs cannot replicate. Claims of PhD-level expertise should therefore be treated with caution; current models remain prone to patterned imitation rather than genuine scholarly or sober judgement (Bender et al. 2021; Nadin, 2019; Bentivegna 2025).’

We invite you to view the paper and consider similar testing of your own. Heck, just clicking the link below is a gift of it’s own – let alone reading it or citing it. 😊

The paper can be accessed through the Assessment and Evaluation in Higher Education journal here: https://www.tandfonline.com/doi/full/10.1080/02602938.2026.2619902
or via the DOI link here: https://doi.org/10.1080/02602938.2026.2619902

If you are happy to read the pre-print version, I can, by law, allow you to download it here. It has a few errors and is missing a few references. I recommend that you access the full peer-reviewed version above, if possible.

My deepest gratitude to Dr. Michael D. Carey and Dr. Peter Grainger for their collaborative efforts with this study and submission.

P.S. This has reminded me: Someone should write a paper comparing genAI claims and output to the Bristol Stool Chart. (Don’t look it up)
P.P.S. Type 7 is my new word for ‘AI slop’.


Cite this article is APA here:
Tensen, D., Carey, M. D., & Grainger, P. (2026). Comparing ChatGPT-5 and other GenAI tools in doctoral confirmation: variability, persona effects, and alignment with human feedback. Assessment & Evaluation in Higher Education, 1–17. https://doi.org/10.1080/02602938.2026.2619902


REFERENCES: (in this post)

Adel, A., & Alani, N. (2025). Can generative AI reliably synthesise literature? Exploring hallucination issues in ChatGPT. AI & SOCIETY. https://doi.org/10.1007/s00146-025-02406-7

Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–623. https://doi.org/10.1145/3442188.3445922

Bentivegna, F. (2025). Artificially voiced intelligences: Voice and the myth of AI. AI & SOCIETY. https://doi.org/10.1007/s00146-025-02383-x

Dickieson-Léger, D. (2025). Can Churches Exist in the Deconstruction Space? The Deconstruction Movement and the Decline of the White Protestant Church in North. Consensus, 46(1). https://doi.org/10.51644/RJEK7701

Nadin, M. (2019). Machine intelligence: A chimera. AI & SOCIETY, 34(2), 215–242. https://doi.org/10.1007/s00146-018-0842-8

Phan, H. P. (2024). Narratives of ‘delayed success’: A life course perspective on understanding Vietnamese international students’ decisions to drop out of PhD programmes. Higher Education, 87(1), 51–67. https://doi.org/10.1007/s10734-022-00992-9

Rebera, A. P., Lauwaert, L., & Oimann, A.-K. (2025). Hidden Risks: Artificial Intelligence and Hermeneutic Harm. Minds and Machines, 35(3). https://doi.org/10.1007/s11023-025-09733-0