The Quiet Reinvention of Assessment
Aka, how AI dismantled then rebuilt how we assess learning in 2026
Hey folks 👋
Two years ago we acknowledged AI had broken assessment. Initially, the reaction was defensive. In higher education, AI bans, huge investment in AI detection tools and a return to in-person examination were standard in 2024.
Meanwhile, in corporate L&D, many of us worried that employees were quietly outsourcing our carefully designed training to AI tools—yet we also suspected there was little we could realistically do to police or prevent it. In the vast majority of organisations, assessments (primarily quizzes) were left unchanged, in the hope that the tension would resolve itself without us having to rethink how we designed, delivered, or measured learning.
Two years on, the publicity of this debate has died down but in the background things have changed a lot. AI didn’t just break the old assessment system; it quietly rebuilt a new one which is objectively better than the system it replaced.
In 2026, 65% of UK students report that their assessment has already changed significantly in response to AI (Higher Education Policy Institute, 2026). This isn’t a local anomaly but the leading edge of a global shift: across major systems in the US, Europe, and Australia, roughly half of students report noticeable changes to assessment expectations or formats, while over 70% of universities have rewritten assessment policies or guidance to account for AI use (EDUCAUSE, 2025).

In corporate L&D, the change is less visible but more structural — nearly half of organisations report being in the “emerging” or more advanced stages of generative AI adoption for learning and talent development, with leading organisations explicitly using AI to drive skills assessment and impact measurement (LinkedIn, 2025). In the workplace, assessment is already moving out of standalone quizzes and into workflow data, simulations, and real performance signals, as conversational intelligence platforms and skills-focused learning ecosystems increasingly treat on-the-job behaviour and outcomes as the primary evidence of capability (LinkedIn, 2025; Crowdmark, 2025).
This month, as part of a project I’m doing with Arizona State University, I’ve taken a deep dive into how assessment in changing in the AI age. The TLDR is: assessment is changing a lot. AI didn’t just dismantle old methods of assessment - it also provided the tools to build new, improved approaches.
Here’s what’s happening right now, and some thought on how it changes your design work.
Let’s dive in!
1. Oral assessment of human understanding
The oldest assessment method in the world is having its biggest comeback in a century.
Oral examination used to be how assessment was done. Medieval universities ran on it — students defended their understanding directly to a master in front of witnesses, and the credential was awarded on the basis of what the student could say, think, and argue under live questioning. The method phased out for one reason: it didn’t scale. As universities grew from hundreds of students to thousands and now millions, one-to-one oral defence became economically impossible. The written essay and the multiple-choice test replaced it not because they were better measures of learning, but because they were the only things you could grade at scale — we’ve spent a century with a less valid assessment method because it was the only one we could afford.
AI changes that economics completely. A structured AI oral exam now costs around forty cents per student — versus thousands of pounds per student for human-led oral assessment at scale. And it isn’t just cheaper — it’s measurably more reliable too. The most rigorous deployment to date, by Panos Ipeirotis et al at NYU Stern, achieved Cronbach’s alpha of 0.75 to 0.80 — meaningfully higher than the 0.50 that essays typically score on the same measure. The thing that was supposed to break assessment has produced a more valid assessment than the one it replaced.

Meanwhile, Droulers, Krautloher and Shaeri’s three-year study across 290 students and 25 subjects, published in Teaching in Higher Education in 2026, found oral assessment consistently outperformed written work on validity, reliability, and student preference.
From a learner perspective, the experience now looks something like this. The student submits a written piece of work as usual; two or three days later, they sit a 10–15 minute oral examination with an AI voice agent. The agent has read their submission, has been briefed on the rubric, and asks targeted questions about the arguments the student made — probing where answers are thin, following up where reasoning is unclear, and adapting in real time to the student’s responses.
The full transcript is then graded by an ensemble of language models that cross-check each other’s judgements, with the human instructor moderating edge cases; the student gets detailed written feedback within hours, not weeks.
What this means for you.
You probably don’t need to replace your written assessments — you need to pair them with an oral check that’s now economically viable for the first time.
Three practical moves. First, when you write learning outcomes, write them as observable verbs — defends, conducts, diagnoses, negotiates, prioritises — rather than knowledge verbs like understands, knows, identifies. If the outcome can be demonstrated by recalling something, AI can fake it; if it can only be demonstrated by doing something — including speaking it under follow-up — AI can’t.
Second, take any written assignment in your current portfolio and ask: what would I want to follow up on if I could? Whatever you’d ask is the start of your oral assessment design — the student writes the essay, defends two or three specific arguments in a 10-minute oral check, and the grade is moderated by both signals.
Third, treat the oral check as the source of truth. The written piece becomes the artefact the conversation is about, rather than the artefact being assessed — that re-orders the academic integrity question from “did they use AI to write this?” (mostly unanswerable) to “can they defend this argument live?” (entirely answerable, and the answer is the assessment).
2. AI personas for real-world skill assessment
For decades, the highest-quality skill assessment in any field has been simulation with a trained human counterpart — a standardised patient in medical training, a role-played client in law school clinical programmes, an actor playing a difficult employee in management development. The method is the gold standard because it tests behaviour in context: not what you know, but what you actually do when a real-feeling person is in the room with you. The problem has always been the same as oral assessment — it’s prohibitively expensive at scale. Recruiting, training, and deploying human standardised patients is so costly that even medical schools could only afford to use them sparingly, and most other professions couldn’t use the method at all.
AI personas have removed that constraint. A trained AI counterpart can be deployed at any time, runs at a fraction of the cost of a human actor, and — counterintuitively — produces a more consistent assessment because the persona’s behaviour is reproducible across students in a way human actors can’t quite be. The methodology that took fifty years to mature in medical education is now spreading to every field where behaviour in conversation matters.
The empirical case is now substantial. Jacobs et al’s multi-institutional deployment of MedSimAI across three medical schools — 410 students and 1,024 encounters — produced OSCE history-taking score improvements from 82.8 to 88.8 (p<0.001, d=0.75).

Foster et al’s randomised crossover trial of SimConverse across two UK medical schools, with 378 participants, produced communication-skills gains 76% as large as actor-based simulation — at a fraction of the cost. Together, these are some of the first studies suggesting that AI-mediated simulation isn’t a poor cousin to the human-actor version — it’s a credible methodology in its own right, and one that travels beyond the medical school budget.

From a learner perspective, the experience is closer to deliberate practice than to traditional assessment. A nursing student conducting a motivational interview, for example, doesn’t sit a multiple-choice test about MI theory — they conduct a 20-minute conversation with an AI patient who presents with substance misuse, exhibits resistance, ambivalence, and eventually change talk.
The student has to navigate the conversation in real time, and their behaviour is scored against a rubric: empathic responding, reflective listening, evocation of change talk, autonomy support. The persona evolves over time so each encounter is genuinely novel, and the student can re-attempt the assessment with a different version of the same case until they’ve demonstrated capability.
What this means for you.
If your subject involves talking to a person — clients, patients, colleagues, customers, students — you can now build a behavioural assessment around that conversation, at scale, for the first time.
The design move is to specify the persona explicitly. Not “a difficult client” — but a particular client, with a particular history, particular emotional state, particular goals and resistances. The more specific the persona, the more your assessment measures the specific behaviours you want to develop; the rubric then becomes a list of observable behaviours rather than abstract competencies. Instead of grading “demonstrates empathy,” you’re grading whether the learner used reflective listening at the three moments in the conversation where it mattered most.
The hardest part isn’t the technology — it’s the rubric design. Spend most of your effort there. Assess the transcript, not the artefact; assess what the learner did in the moment, not what they said about it afterwards.
3. Process assessment to measure thinking, not output
Shifts 1 and 2 are about new methods of summative assessment — better ways to measure what a learner can do at the end of a course or programme. This one is different: it’s about what changes at the summative assessment moment when the artefact is what’s being submitted.
The traditional summative submission has always been the artefact alone — the essay, the report, the analysis, the deck. The grader receives the final product, judges it against the rubric, returns a mark and three sentences of feedback; the process by which the artefact was produced is invisible, and the grade reflects only what the learner can produce, not what they can do.
This worked when producing the artefact required the cognition the assessment was supposed to measure — if you could write the essay, you had probably thought through the argument; if you could produce the analysis, you had probably understood the data. AI has decoupled the two: the artefact and the cognition are no longer linked. A polished essay no longer proves the learner can think; a structured analysis no longer proves they can reason. The artefact alone is no longer enough.
The methodological response — already mature in fields like postgraduate research, professional practice, and clinical training — is to make the process itself part of the summative submission.
The learner submits the artefact alongside the trail of how it was produced: the AI prompts they used, the drafts that show how the argument evolved, the decisions they made about what to keep and discard, the moments where they pushed back on what the AI suggested. AI can produce any artefact, but it cannot — yet — fake the process; the process trail becomes the evidence of cognition that the final product no longer carries.
The institutional adoption of this approach is moving fast. Monash University’s TeachHQ framework specifies process reflection as a required submission alongside any AI-assisted work — what prompts were used, what was kept, what was discarded, why. The University of Sydney mandates AI disclosure on every assignment, with penalties for non-compliance.
The empirical case is less mature here than for oral assessment or AI personas — partly because process documentation is a methodological move rather than a technological one — but the institutional pattern is clear: where universities have moved to process-aware submission, academic integrity disputes have measurably declined, and the conversation about AI use has shifted from adversarial to developmental.
From a learner perspective, this reframes the academic integrity question in a way that’s much easier to live with. The question shifts from “did you use AI?” — which is mostly unanswerable in 2026, and which puts both learner and grader into an adversarial relationship — to “can you account for how you used it?” — which the learner can answer with integrity, and the grader can assess with confidence.
What this means for you.
Add a short, structured process document to any AI-assisted summative submission — three or four prompts is usually enough: what prompts did you use; what did you keep and discard and why; where did you push back on the AI; what would you do differently next time. Keep it under 500 words. Mark it as part of the assessment, not as extra paperwork.
Two design notes that matter. First, design the process document as if it’s the more assessable part of the submission — because in many cases it now is; the reflection on prompt choices, the reasoning about what to keep, the moments of pushback, all surface cognition the final artefact hides. Second, make the rubric explicit; “describe your AI use” produces vague responses, but “describe one moment where you disagreed with what the AI suggested, and why” produces an artefact you can actually grade.
4. Continuous assessment of skills & behaviour
The first three shifts are about discrete assessment moments — better methods of oral examination, persona-based simulation, and process-aware summative submission. This one is different. It’s about what AI makes possible for assessment between and around the discrete moments — and ultimately, what it makes possible for measurement once the boundary between learning and working starts to dissolve.
For most of the history of formal education, assessment has been event-based for one reason: every additional data point cost money. Every formative check required an instructor’s time. Every behavioural observation in the workplace required someone shadowing the employee, taking notes, scoring against a rubric. So we settled for endpoints — the mid-term, the final, the post-training survey — and accepted that a student silently drowning in week 4 wouldn’t surface as a problem until the mid-term in week 7, by which point the intervention window had closed. The measurement model was a compromise with the economics, not a methodological choice.
The same constraint produced the dominant measurement model in corporate L&D: a post-training survey, a knowledge-check quiz, and a completion certificate, all standing in for the thing we actually wanted to measure but couldn’t afford to — whether the learning had transferred into behaviour on the job. Research suggests 80% of sales training is forgotten within 30 days, and even when retention holds, the link between knowledge-check performance and real-work behaviour is weak to non-existent. We measured what we could afford to measure, even though we knew it wasn’t what we needed to measure.
AI has changed the economics of measurement completely, and the methodological boundary between assessing learning and assessing performance is dissolving as a result.
In a structured learning experience, every interaction a learner has with an AI tutor now produces continuous assessment-grade data — what the learner attempted, what they got right unaided, what they needed scaffolding for, how their reasoning developed across sessions. The summative final is no longer the only measurement moment; it’s one signal in a stream. A randomised controlled trial across 5,500 students in 74 Indian schools (Oreopoulos et al. 2026) found high-usage students gained 0.44 to 0.47 standard deviations on end-of-year maths assessments — a substantial effect comparable to high-impact human tutoring interventions. The mechanism is the continuous assessment signal itself: the tutor knows, in real time, what each student can and can’t do, and adjusts.
In the workplace, the same machinery is now being used to test if and how well AI can capture direct and continuous measurement of real performance. Sales calls, client meetings, internal discussions, coaching sessions — all are machine-readable, and the patterns that distinguish high performers from low performers can be extracted and compared against any individual employee’s actual behaviour. Transfer is no longer a thing we assess by proxy — it’s the thing we assess directly.
The single clearest articulation of where this is heading came from Heather Stefanski, CLO at McKinsey, on the Workplace Stories podcast in May 2026. In Stefanski’s vision, assessment of workplace training isn’t an event scheduled around the work any more; it’s an emergent property of the AI “colleague” who works alongside the learner and both supports and tracks their performance. Her position is that L&D specialists should be in the build of every AI agent the organisation deploys, designing the developmental and measurement layer of the work itself rather than building courses adjacent to it.

From a learner perspective, this means the assessment moments are no longer scheduled. They’re continuous, embedded in the learning or the work itself. The student practices a problem with the AI tutor; the system notices when they can answer the next problem without scaffolding and when they can’t. The employee takes a real call; the AI analyses the transcript against the behavioural patterns of top performers; the employee receives feedback within hours on specific moments — the missed discovery question, the objection that wasn’t surfaced, the close that landed too early. The assessment loop closes from weeks or months to a single working day.
What this means for you.
The design move is to define one or two leading capability indicators you’ll watch across the learning experience and into the work itself, not just at endpoints. Pick something concrete: cognitive engagement quality, scaffolding decay, unaided capability checks at intervals, specific observable behaviours on real calls or in real interactions. Decide what threshold triggers an intervention before you deploy the experience, so you don’t end up staring at a dashboard wondering what to do with the data.
Start with the behaviour you want to measure, not the training around it. In a structured learning context, that means identifying the specific capability the learner is building and the leading indicators that capability is forming. In a real-work context, it means identifying what your top performers actually do — the specific moments, the specific moves, the specific patterns of language and timing — and designing your continuous signal to surface those moments. In both cases, the rubric is no longer abstract competencies; it’s a specific list of observable indicators you’d recognise across any session or any call.
And remember — the assessment question and the design question are inseparable here. If the AI working alongside the learner does too much of the cognitive work, the continuous signal you capture is meaningless: you’re measuring what the AI can do, not what the learner can.
Conclusion: What AI Actually Did to Assessment
Look at what just happened. Two of these four shifts only exist because AI exists — oral assessment at scale, persona simulation outside medical schools. The third — process documentation — pre-dates AI but became urgent because AI made the alternative untenable. The fourth — continuous capability signal — is the most radical, and the one where AI hasn’t just changed the method, it’s changed what counts as an assessment moment at all.
AI has played two important roles in the evolution of assessment. It was the destabiliser — making the old methods visibly inadequate, fast — and it was the enabler — making the better methods affordable for the first time. The same technology that broke the essay made oral assessment cheap; the same technology that let students delegate courses powered continuous formative signal; the same technology that broke the post-training survey powered real-work behaviour tracking.
AI didn’t break assessment — it broke the economics of bad assessment, and in doing so, it cleared the ground for the methods we should have been using all along.
If there’s one design principle that connects all four shifts, it’s this. The L&D field spent two years treating AI as an adversary in assessment design; the teams doing the most interesting work are now treating it as a collaborator.
Use AI to draft assessment briefs and critique them; generate question banks and distractors; build persona-based stress tests of your designs before you deploy them — see how a struggling learner, a confident learner, and an edge-case learner would each experience the assessment, then fix the design before any real student touches it. Mark formative work and surface patterns for your attention. Then spend your time on the design judgement, the SME relationships, the strategic decisions about what the assessment is actually for.
The pretence that AI isn’t in the room has wasted two years and produced almost nothing. Assume AI is in the room, and design accordingly. To get started:
Take one assessment you currently run. Ask yourself two questions: what’s the artefact, and what’s the behaviour?
If you’re assessing the artefact, you’re using a methodology AI has made obsolete — pick one of the four shifts and apply it to that one assessment.
You don’t need to redesign the whole assessment system from the outset — you need to move one assessment one rung, then another, then another. That’s how the quiet revolution actually scaled — one assessment at a time, by people who stopped arguing (or stopped pretending) and started building.
My closing take is this: AI didn’t break assessment - it broke the economics that kept low quality assessment methods in place. Replacing those methods with more viable and valuable alternatives is a new and exciting part of our responsibilities with the potential to make us more, not less, effective in our roles.
Happy innovating!
Phil 👋
PS: If you want to explore how to get the most from AI in your day to day work, check out my AI Bootcamp for L&D.




This topic is top of our minds, especially the shift to measure thinking.
On AI-voice agents for oral assessment, are there any vendors developing commercial solutions for this?