Artifical Intelligence, Assessment and the Future of Training

Introduction

Since the release of generative AI tools such as ChatGPT in late 2022, educators across the world have been asking an uncomfortable question: if AI can complete many of the assessments we give to students, what are we actually measuring? Newton and Jones tackle this issue head on in a pragmatic review of assessment in the age of generative AI. Although their focus is biomedical science education, the issues they raise resonate across professional education, and particularly in medicine. The paper explores how well AI performs on traditional assessment formats and considers what educators might need to do differently in response.

For those involved in emergency medicine education, which is most of us to be honest, the implications are significant. AI does not just challenge how we teach, it challenges whether our current assessment processes are fit for purpose. The abstract is below, but as always please read the full paper yourself and come to your own conclusions. I’ve had the pleasure of listening to and meeting Philip Newton, and he is a really interesting chap (I’m sure Sue Jones is too, but I’ve not met her yet). As someone involved in designing and overseeing high-stakes emergency medicine assessments, I read this paper with a mixture of fascination and concern.

As disclosure I am currently Dean of RCEM and at college we think about assessments and AI a lot. However, the review below is my own work and does not represent the college view, which is articulated in the links below. That said, this is a fast moving field, so let’s get into what the article is about and why it’s got us so interested.


Abstract

The emergence of ChatGPT and similar new Generative AI tools has created concern about the validity of many current assessment methods in higher education, since learners might use these tools to complete those assessments. Here we review the current evidence on this issue and show that for assessments like essays and multiple-choice exams, these concerns are legitimate: ChatGPT can complete them to a very high standard, quickly and cheaply. We consider how to assess learning in alternative ways, and the importance of retaining assessments of foundational core knowledge. This evidence is considered from the perspective of current professional regulations covering the professional registration of Biomedical Scientists and their Health and Care Professions Council (HCPC) approved education providers, although it should be broadly relevant across higher education.


Newton PM, Jones S. Education and Training Assessment and Artificial Intelligence: A Pragmatic Guide for Educators. British Journal of Biomedical Science. 2025.

What is assessment actually for?

The authors begin with a reminder that assessment sits at the heart of higher education. Universities are defined, in essence, by their ability to award degrees, and those degrees are based almost entirely on the results of assessment, (and since we are an EM blog we can talk about Royal Colleges here too). In professional programmes, that certification carries an additional responsibility: it signals to patients, employers and regulators that graduates are safe to practise. Think about the MRCEM and FRCEM exams and what they mean for progression. MRCEM is often a ticket to the senior resident rota, and FRCEM to consultant positions. That is career progress for the individual, but it also means that those individuals get to make higher stakes decisions that affect patients. Exams are therefore a patient safety issue, and they must therefore be as safe, reliable and as effective as they can be.

Assessment also has an important role in learning itself. Educational research has long demonstrated the “testing effect”, where retrieving information during tests strengthens long-term learning. Regular assessment is therefore not just a way of measuring learning but also a powerful way of promoting it. In other words ‘ Assessment drives learning’, something that every educator knows form the inumerable times we have been asked ‘Is this going to be in the exam?’ by learners (frustrating though that is).

A key point in the paper is that knowledge acquisition is cumulative. Higher order thinking, critical analysis and problem solving depend on a solid foundation of factual knowledge. This is particularly true in scientific and clinical disciplines where specialist terminology and conceptual frameworks underpin practice. The authors caution against a common reaction to AI, the temptation to abandon knowledge testing entirely. For the reasons above, that’s not an option in medicine, and also a reason why primary examinations that test basic science knowledge (in our case subjects like anatomy, physiology, pathology, pharmacology etc.) are so important.

When AI takes the exam

The data on how well AI can take, and pass exams was really interesting, particularly traditional assessment formats such as MCQs and essays. Spoiler – it’s probably better than you are (or myself to be honest).

Multiple choice examinations, long used in medicine and biomedical science, are pretty ubiquitous. The evidence suggests that modern language models perform extremely well on these tests. Early versions of ChatGPT achieved average scores around the mid-50 percent range across higher education MCQ exams. Newer models have improved dramatically, with GPT-4 averaging around 75%, and more recent versions achieving scores of around 94% on the UK Medical Licensing Applied Knowledge Test. Importantly, these tools can now interpret images and analyse data as part of exam questions. That means even visually based or clinically framed MCQs are no longer protected from AI assistance. I’ve used chatGPT to diagnose skin rashes (personal not professional use) and it’s pretty good, and the data suggests it can manage ECGs, radiology, clinical images and more really well. I’m sure many of you are already using AI interpretation tools for ECGs and radiology, and yes, in my experience it performs really well.

The implication is not that MCQs are obsolete, but that the context in which they are used matters enormously. Unsupervised online exams, particularly those introduced during the COVID-19 pandemic, are now highly vulnerable to AI-assisted cheating. If MCQs are to remain part of summative assessment, the authors argue they must take place under secure, supervised conditions.

The problem with essays and other long form assessment

If MCQs are vulnerable, longer responses such as essays, reflections, reviews, project reports etc. appear even more problematic. Generative AI systems now produce essays that markers frequently struggle to distinguish from human writing. In one study researchers inserted AI-generated answers into real exam marking pools. Academics marking the scripts not only failed to identify the AI-generated work but often awarded it marks equal to or higher than student submissions.

This presents a deeper problem than simple plagiarism. Traditional plagiarism detection tools work by identifying copied text from known sources. AI-generated writing is original in the sense that it is newly generated text, so there is no external source to detect. Although AI-detection software exists, it provides probabilistic estimates rather than definitive proof and is increasingly easy to circumvent. I’ve used this on work that I know is AI generated, and also work that I know is not (because I did both), and the assessment tools were not definitive for either, so I cannot rely on AI detection tools either in my role as an educator/assessor. I believe that this means that asynchronous essay/reflection/report assignments may no longer represent a valid assessment method unless academic writing itself is the skill being assessed. In medicine, the most obvious implication is in reflective pieces which I now believe are regularly created using AI. There are even dedicated websites to support AI reflective writing (check out fourteendfishermen website related to the GP portfolio). Now, it has to be said that cheating on written assessments is nothing new. For decades it has been possible to get someone to write an essay for you, for money or other favours, but that was expensive and risky so not everyone did it. So the principle that cheating happens has not changed, but it is now easier and faster, or as our colleague Greg Yates described it: ‘AI may democratise cheating’ (good quote that – a keeper).

The conclusion must be that if our aim is to assess whether a learner can practise a subject (in our case EM), then assessment should focus on their ability to do it, rather than write about it.

Assessing performance rather than prose

Where AI struggles most is in assessments that require real-time demonstration of knowledge or skill, but even that is changing. Practical examinations, laboratory competency assessments, viva examinations and direct observation of practice are all examples of formats that require learners to demonstrate understanding in the moment. In these situations, it becomes much harder to rely on AI tools without the lack of underlying understanding becoming apparent. In medicine, assessments such as OSCEs are probably the least likely to be affected by AI. In the past we used viva-voce exams to test candidates, and they too are largely immune to AI influence, but they were largely abandoned owing to cost and concerns around EDI issues. Viva-voce exams area a face to face interview and if not carefully designed and monitored can be very variable and open to accusations of bias, EDI issues, lack of consistency and more. That’s not to say they cannot be used, but they are by no means an easy fix. There are increasingly calls to return to face to face assessments, but there were valid reasons why we moved away from them, and perhaps something different altogether will be required (though what that may be is yet to be determined). There is also the cost implications of running more face to face exams (they are expensive for candidates and organisations, and on a volunteer model of examiners, difficult to sustain).

In laboratory training environments, trainees often learn by observing techniques, performing them under supervision and then explaining the underlying concepts back to their trainer. These interactive discussions, combined with direct observation of practice, provide a robust way of assessing competence. We can do similar in EM, and the ESLE assessments are probably the closest to the direct laboratory observations, although there is probably work to be done on the calibration of assessors in these and other assessments. Basically, whenever we introduce physical examiners into the assessment process it prevents AI use, but opens up other concerns that will need to be mitigated (if they are possible at all).

The authors suggest that this kind of assessment may become increasingly important as educators seek formats that remain valid in an AI-rich environment. In medicine, I can mostly agree, but whether we have the capabilitry and resources to do so is as yet to be determined.

Professional standards and patient safety

As I said at the beginning medical (including RCEM) exams are really important as they are related to professional progression and regulation.

Biomedical scientists, doctors and other healthcare professionals, must demonstrate competence before entering practice. Allowing learners to progress without genuine understanding would carry clear risks to patient safety. At the moment regulatory standards are linked to exams that require professionals to practise only within their competence and to maintain honesty in their training and qualifications. Submitting AI-generated work without understanding it would therefore raise ethical concerns that extend beyond academic misconduct. This is something we are seeing at all levels of medical education, but how we spot it and what we do about are still difficult. RCEM have published guidelines as have other colleges and universities, but this is a fast moving area and it can be tricky for regulation to keep up with innovation. Remember though that AI in education is not simply about cheating, and it’s not all about reporting people. At the heart of this concern is whether training programmes can still assure competence.

What might this mean for emergency medicine education?

Although the paper focuses on biomedical science, the themes translate well to emergency medicine training. Emergency medicine has long relied less on written coursework and more on direct assessment of clinical performance. Workplace-based assessments, simulation training, procedural sign-offs and case-based discussions all require learners to demonstrate their reasoning and skills in real time. In many ways, emergency medicine training may already resemble the type of assessment ecosystem the authors argue we need. So we have some protection, but there are concerns.

The challenge of knowledge assessment remains. Emergency medicine practice relies heavily on rapid retrieval of core knowledge in time-critical situations. The idea that clinicians can simply “look things up” or rely on AI assistance ignores the cognitive demands of acute and emergency care. Clinical reasoning under pressure depends on well-developed mental models and accessible factual knowledge, plus the ability to synthesise that information and come to a decision (which is often probabalistic rather than definitive).

AI therefore does not remove the need for knowledge, if anything, it highlights why knowledge and recall remains essential for a lot of what we do.

Another challenge is AI literacy itself. As AI tools become integrated into healthcare systems, clinicians will need to understand how these tools work, when they can be trusted and when they cannot. Future emergency physicians may need to interpret AI-generated diagnostic suggestions or risk predictions, much as we currently interpret imaging reports or laboratory results. The difficulty is that we still do not know exactly how AI will be integrated into everyday clinical practice. Designing assessments around AI use is really dynamic at the moment and may therefore be premature.

Further Reading: AI and Medical Examinations

A growing body of literature has examined how well large language models perform on medical and scientific examinations. Taken together, these studies paint a fairly consistent picture: modern AI systems are increasingly capable of passing many of the knowledge-based assessments traditionally used in professional education. To my knowledge this has not been specifically tested in RCEM exams, but as the information below suggests, I suspect it would do pretty well if we allowed it a go.

One of the earliest widely cited studies was published in PLOS Digital Health in 2023. Kung and colleagues tested ChatGPT on questions from the United States Medical Licensing Examination (USMLE). The model performed at or near the passing threshold across all three USMLE steps and was able to generate explanations demonstrating elements of clinical reasoning rather than simple factual recall. Subsequent work examining newer models has demonstrated rapid improvements. Analyses of GPT-4 on USMLE-style questions have reported accuracies approaching or exceeding 85–90%, comfortably above typical passing standards for candidates.

Similar findings have been reported in biomedical science education. Stribling and colleagues evaluated GPT-4 using real graduate-level examination questions and found that the model outperformed the average student in most of the exams tested. Importantly, these exams required data interpretation and applied reasoning rather than simple recall.

Systematic reviews of the literature confirm the trend. Across multiple studies and examination formats, GPT-4 level models demonstrate accuracy rates around 80% on medical licensing-style questions, consistently outperforming earlier generations of language models.

None of this means that AI can practise medicine. Clinical competence involves communication, teamwork, judgement and decision-making under pressure. However, these findings highlight something important: many of the assessments used in professional education measure knowledge in formats that are increasingly easy for machines to replicate.

Which brings us back to the central question raised by Newton and Jones: if AI can pass the exam, what exactly are we assessing?

Final thoughts

Generative AI has exposed weaknesses that already existed in many traditional forms of assessment. Essays and unsupervised online exams were vulnerable long before ChatGPT appeared, but AI has dramatically amplified those vulnerabilities.

Rather than abandoning assessment altogether, the authors argue for a pragmatic response. Foundational knowledge still matters. Secure assessment environments still matter. And authentic demonstrations of competence, whether through practical exams, discussion or observation matters more than ever.

For emergency medicine educators, this may not represent a revolution so much as a reaffirmation of what we already value: assessing clinicians not just on what they can write, but on what they can actually do when it counts. How we develop our exams and assessment tools to make sure that this continues to happen, and that it remains robust is going to need some considerable effort.

vb

S

References

  1. Newton PM, Jones S. Education and Training Assessment and Artificial Intelligence: A Pragmatic Guide for Educators. Br J Biomed Sci. 2025.
  2. Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on the United States Medical Licensing Examination: Potential for AI-assisted medical education. PLOS Digital Health. 2023.
  3. Nori H, King N, McKinney S, et al. Capabilities of GPT-4 on medical challenge problems and USMLE-style exams. arXiv. 2023.
  4. Garabet R, et al. GPT-4 performance on USMLE Step 1 style questions. Medical Science Educator. 2023.
  5. Stribling D, Xia Y, Amer MK, et al. The model student: GPT-4 performance on graduate biomedical science exams. Scientific Reports. 2024.
  6. Liu M, et al. Performance of ChatGPT across medical licensing examinations: systematic review. JMIR Medical Education. 2024.
  7. RCEM guideline on AI use. https://rcem.ac.uk/position-statement/rcem-position-statement-artificial-intelligence/

Lastly…..did I use AI to help develop this blog?

Yes, yes I did! Primarily to somewhat prove the point of the article, but not as much as you might think. The vast majority of what you have read has been typed by me. How much did you use exactly you might ask? Well……, more than a smidge, but not a bunch. I’m happy it’s my work, but I would not submit it for an exam.

Does that answer your question? Probably not, but then who can tell these days 😉

Cite this article as: Simon Carley, "Artifical Intelligence, Assessment and the Future of Training," in St.Emlyn's, July 28, 2026, https://www.stemlynsblog.org/artifical-intelligence-assessment-and-the-future-of-training/.

Thanks so much for following. Viva la #FOAMed

Scroll to Top