AI Can Build Your Course, but Can it Design the Learning?
Aka, the current state of AI's instructional design ability & what it means for L&D
Hey folks 👋
There’s a problem in our field that’s been around for as long as the field has, and it goes like this: people who are expert in a topic tend to believe that knowing something means being able to teach it, so they design the course themselves.
The subject matter experts who don’t “design their own” aren’t usually deferring to our expertise; they’re usually avoiding the labour. Storyboards, documents, images, diagrams, review cycles, SCORM packaging: the whole process took weeks, so they delegated it, we took it on, and the diagnostic conversation got twenty minutes at kick-off if it happened at all.
Enter AI. In a world where production is commoditised, the time constraint that used to protect L&D is gone. As a result, a decades-old wicked problem has — in the last year especially — got considerably worse. In 2026, it’s faster and easier to build a course without L&D than ever before.
This week, I spent a chunk of time in a room with leaders from a large multinational discussing exactly this. The question their L&D team was being asked, more or less directly, was why the company needs an L&D function at all when a domain expert can produce a course in an afternoon.
Put another way: what’s the value of the L&D team in the post-AI workplace? It’s a tough question, and I think we’ve mostly been answering it badly by vaguely asserting that AI output is shallow, or generic, or that it “lacks the human touch.”
So I ran a test to dig deeper into the specific risks we face when we replace L&D professionals with SMEs plus LLMs. The results were uncomfortable reading about what AI can now do, and reassuring about what it still can’t — and not in the ways I expected.
Let’s dive in!
The Case for SMEs Designing their Own Courses
Nobody has run a census of what share of course design is completed by subject experts rather than L&D teams, so I can’t give you a number. But the signals are there, and they point in a very clear direction.
The shift was already underway before AI arrived. In 2023, CIPD’s Learning at Work survey of over 1,100 UK L&D practitioners found a quarter of L&D functions reporting increases in SMEs designing learning experiences over the previous twelve months. In a good number of cases, this trend is driven by necessity: there simply aren’t enough instructional designers to route the design work to. A 2025 survey of L&D leaders, mostly in organisations of 250+ people, found dedicated IDs on only around seven in ten L&D teams.
Then there’s the tooling. Tools like Easygenerator are being built to actively enable SME design. The company reports 50,000+ authors across 2,000+ companies on a platform whose entire proposition is that you don’t need instructional design skills to build a course, and it isn’t alone in that category. Those licences aren’t being bought for L&D teams — they’re bought for the people who used to raise a ticket.
But the clearest signal I have is what I hear. Over the last six months, across bootcamp cohorts and client engagements, a big recurring question has been: the SMEs have started building their own courses and nobody asked us — what do we do about that?
So, the tests I decided to run aren’t hypothetical: they test something real that’s happening right now, with the goal of answering a question that’s asked of L&D more than ever before: why do we need L&D when we have AI?
The Test
From my observations on the ground, non-expert designers tend to brief AI in a very similar way: they tell it it’s an instructional designer, upload a source document, pre-specify the solution, pre-choose the form and name the audience.
So I recreated that behaviour and gave four AI models the same prompt: “You’re an expert in instructional design. Read the document attached. Then help me turn it into a short elearning for new starters.”
The source document I picked was similar to the ones I see a lot of SMEs work with: a real, publicly available corporate Standard Operate Procedure on reporting misconduct — ten sections, dense, procedural.
The models four I tested were the ones used most in both corporate and higher-ed settings: ChatGPT-5.5 High, Claude Sonnet 5 High, Copilot Think Deeper, and — definitely an outlier, but interesting for comparison — Gemini 3.1 Pro.
I scored every output against a 36-point instructional design standard, which came to 288 individual judgements in total.
The standard is the working QA checklist I use on real client design reviews and bootcamp cohort work, assembled from roughly 200 peer-reviewed studies and meta-analyses across performance improvement, cognitive science and evaluation — Gilbert on where performance deficits actually sit, Mager on objectives, Biggs on constructive alignment, Sweller and Mayer on load and multimedia design, Bjork on desirable difficulties, Kirkpatrick and Brinkerhoff on evaluation — plus a set of criteria that exist purely to catch the errors I see most often in practice.
The 36 criteria split into two types of quality measure.
Design craft — 22 criteria. Is the learning well designed?
These are the things you can inspect in a finished course. Do the objectives describe observable performance, or a state of mind? Do objectives, practice and assessment sit at the same cognitive level, or does an apply-level objective get assessed with recall questions? Does practice outweigh exposition? Does feedback address the decision the learner made, or just mark it right or wrong? Do the wrong answers encode misconceptions people actually hold? Are load and multimedia principles operationalised, or named and then violated? Is accessibility designed in or bolted on?
Every one of these leaves a trace. If a design gets them wrong, the evidence is sitting there in the artefact, and someone who knows what to look for can find it.
Design judgement — 14 criteria. Should this learning experience exist, in this form?
These only ever exist in the process. Was the performance problem established before a solution was chosen? Was “this isn’t a training problem” ever a live possible answer? Were the constraints elicited, or invented? Was success defined as something that changes in the workplace, before design started? Does every parameter — seat time, pass mark, item count — have a stated basis, or is it a plausible number? Is there a plan to test it, with a threshold at which you’d stop?
None of these leave a trace. There’s no screen where the missing needs analysis would have been, and nobody can review an omission.

Each criterion scores 0, 1 or 2 — absent or actively wrong, gestured at but not operationalised, or met with evidence in the artefact. And every criterion has a stated failure signature, written in advance, so I’m scoring against a description rather than a feeling. For “objectives describe observable performance,” the signature is uses “understand”, “be aware of”, “appreciate the importance of.” For “practice outweighs exposition,” it’s content, content, content, knowledge check.
Four criteria are automatic fails regardless of the total: any zombie theory used as design rationale — learning styles, the learning pyramid, the ten-minute attention span, 70:20:10 as empirical law — and any case where a non-learning problem gets solved with a course.
Finding #1: Some AI models have genuinely impressive design expertise
Based on previous tests, part of me expected wall-to-wall slop, but I didn’t get it.
What surprised me wasn’t just the impressive level of some models’ design expertise, but also the inconsistency of expertise across models. The craft scores range wildly, from an impressive 77% to a scarily low 9%. This range matters on its own, and I’ll come back to it, because it means “can AI design learning?” is a far more complex question than we tend to acknowledge.
But first, the top of the range is the part worth focusing on.
What the strongest models did well
ChatGPT-5.5 and Claude Sonnet 5 High demonstrated five instructional design competencies between them.
Asked clarifying questions of the user. Claude Sonnet 5 High was the only model that stopped before building. It outlined a five-screen structure, then asked: “Want me to build this out as an actual interactive module, or as a script/storyboard document you can hand to a designer? What format would be most useful for you right now?”. Not the questions an expert Instructional Designer would ask at this stage (more on this later) but a real fork with real consequences, offered before committing to either.

Claude then versioned the deliverable as “Draft v0.1 — for design and compliance review” and flagged two decisions back to me at handoff rather than presenting them as settled: that the contact details needed sign-off from Corporate Ethics & Compliance, and that three sections of the SOP had been deliberately excluded and should be confirmed as out of scope. It also deferred the LMS specifics — SCORM version, completion criteria — to the L&D team rather than inventing them.
Scoped by exclusion and managed cognitive load. Instead of including everything from the document, ChatGPT-5.5 started by asking itself: does this help a new starter act? If the answer was no, it cut the content — in this case the waivers process, approval dates, document history and the external auditor clause. Deciding what a course is not about is the harder half of scoping, and it’s what separates a designed module from a summarised document.

Sequenced for learning rather than in line with the document. ChatGPT-5.5 opened the learning experience with a decision before it taught anything: it presented a scenario, gave the learner a choice, then delivered the content. That’s the generation effect operationalised, and it runs directly against the grain of a source document that starts with purpose and scope.
Designed smart distractors and feedback. The knowledge checks generated by ChatGPT-5.5 included three distractors that encoded real misconceptions and required genuine thinking — wait until you have proof, you’re new, you probably misunderstood — rather than three absurdities and an obvious answer. Its feedback addressed each wrong answer on its own terms, explaining why that specific error is tempting. Claude Sonnet 5 High also performed well here; its opening feedback line was better than most human-written feedback I see: “It’s understandable to hesitate, but GSK asks every one of us to raise genuine concerns — not sit on them or pass them on informally.”

Flagged risks. When ChatGPT-5.5 proposed a drag-and-drop activity, it also flagged an accessibility risk and suggested an accessible alternative. It refused to hard-code a phone number from a thirteen-year-old document, flagging that it must be verified.
In all of these examples, AI showed value in two ways: it sometimes did competent instructional design work — e.g. building distractors that encode real misconceptions rather than three absurdities and an obvious answer — and it sometimes behaved like a thought partner in getting to a better output — e.g. asking me whether I wanted a working module or a storyboard before it built either.
But look closely at what triggered those behaviours and a pattern emerges. Several of them fired because of something that I included in the prompt. I named the audience, so ChatGPT-5.5 had a learner profile to scope against. I said “short,” so it had a reason to cut. Take those words out and the behaviours have nothing to attach to, which means the design quality declines.
In practice, this means the design ability of AI isn’t consistent or guaranteed; it’s hugely contingent on the brief already containing the right information — and the whole premise of a non-expert briefing an LLM is that they don’t know which information matters. If they don’t include the necessary context, the model won’t tell them it’s missing: not one of the four models I tried asked me anything about the learners, the context, or the problem. Claude Sonnet asked me one question, and it was about which file format I wanted.
So what LLMs offer isn’t design capability so much as design responsiveness. Give these models the right inputs and you might get OK design quality out. Give them a different brief and you get whatever the source document happened to contain, arranged competently with quizzes attached.
Finding #2: All four models chose the same design approach — and it was the wrong one
I asked AI for “a short elearning for new starters”. Every one of the four models accepted that as settled. Not one asked whether an asynchronous online course was the right vehicle for my goal. Not one named an alternative: a facilitated session, a manager-led conversation, a job aid, a bot, a change to the reporting system itself.
All four did exactly what AI tools are built to do: go from brief to execution as rapidly as possible. Rewarded for compliance and speed rather than quality or challenge, the LLM takes a design spec and produces the thing requested. If the request lacks details, AI fills in the blanks by doing the equivalent of an internet search.
In my case, because the brief asked for a short elearning course, I got four short elearning courses — regardless of whether that was the right solution. The design of those courses was, in every meaningful respect, exactly the same design. Here’s the headlines:
Every model used categorisation practice. This is the activity where you give the learner a set of examples and ask them to sort each one into a named category — bribery, fraud, discrimination, conflict of interest, safety risk. Its legitimate purpose is building discrimination: helping someone tell apart two things they’d otherwise confuse, which is why it works well for skills like triaging support tickets or spotting the difference between a near-miss and an incident. All four models used it the same way here — deliver the SOP’s list of reportable conduct, then ask the learner to file examples back into that list. GPT-5.5 built interactive category cards, Claude Sonnet a five-card match-to-category exercise, Copilot a click-each-item hotspot, Gemini four clickable tiles. But in all we had the same underlying design decision: deliver information, then check and reinforce it via categorisation.
Every model used single-turn MCQs. In all four designs, the learner is presented with a series of “single-turn” knowledge checks: the learner makes one decision, receives feedback, and the story ends. Nothing carries forward, no choice changes what happens next, and there’s no cost attached to getting it wrong — which is what separates a scenario from a scenario-shaped multiple-choice question.
Every model used click-to-reveal exposition. The learner clicks a thing, text appears, and the text is the same text that could have been on the screen from the start. Some models used cards, some tiles, accordions and hover states. This is usually justified as “interactivity” and “active learning”, but it isn’t: interaction that only controls when content appears is just navigation. It adds a click, not a decision. Copilot and Gemini both went further and gated progress based on this — the Next button stays disabled until every tile has been opened, which converts a design element into a compliance tracker and creates friction which we know leads to frustration and disengagement, not learning.
Every model ended on a summative knowledge check. In all four designs, after every piece of content has been delivered, there is a final MCQ with between one and five scored items, and a pass mark of 80%. Nobody justified that threshold and nobody could have: when we only have four or five items there's no psychometric basis for any pass mark at all, because a single lucky guess moves the result by twenty or twenty-five percentage points. What's really being measured is whether the learner clicked through and finished. Copilot's four-item check makes this plain, since it includes a True/False item on whether GSK protects people who report — an item every learner will answer correctly, which means a pass is available on two genuine questions out of three.

So, how does this compare to what an expert ID would have designed? There’s two parts to this answer - the ideal one and the helpful one.
Answer 1: Not a Course
An experienced designer, given this brief, would spend the first hour establishing what's actually happening. What are the current reporting rates? Where do reports come from — which functions, which regions, which channels? What proportion get substantiated? What happens to the people who file them, and does anyone know? Are we solving low reporting, or slow reporting, or reports going to the wrong place?
On the published evidence about whistleblowing, the most probable answer here is that this isn’t a knowledge problem at all. If reporting rates are low, the barrier is almost certainly perceived retaliation risk plus doubt about organisational responsiveness — and the interventions that move those are:
manager capability work - because the immediate manager is the single biggest variable in whether someone speaks up
visible responsiveness - publishing anonymised case outcomes to model the safe environment and build a reporting culture.
Answer 2: A Better Course
Every strategy the models chose treats this as a knowledge problem. Every strategy a professional would choose treats it as a willingness (behavioural) problem. That single diagnostic call — knowledge or willingness — determines every downstream decision, but it was never made.
Treated as a behavioural problem, the design would look more like this:
1. Worked example of the judgement, not the procedure — a real employee narrating an ambiguous situation and how they weighed it, because expert reasoning has to be made visible before a novice can imitate it (Sweller; Atkinson et al.).
2. Contrasting cases — pairs of near-identical situations differing on one feature, because discrimination is learned by comparison at the boundary, not by sorting clear examples into named boxes (Kellman, Massey & Son).
3. Behavioural rehearsal — the learner practises actually opening the conversation, because willingness responds to rehearsing the moment rather than to being told what the moment requires.
4. Social proof — real, anonymised accounts of what happened to people who reported, because risk is calibrated by observing outcomes and no amount of asserting “you are protected” substitutes for evidence that the system acts.
5. Performance support at the point of need — a one-page decision aid rather than content to memorise, because the decision happens six months after induction and nobody reopens a module.
6. Spaced return at 30 and 90 days — one scenario, one decision, because spacing is the legitimate evidence on sizing and a new starter’s first ambiguous situation rarely lands in week one (Cepeda et al.).
7. A measure that can fail — reporting volume, channel mix and time-to-report, agreed before the build with a pre-committed threshold, because completion cannot fail and a measure that cannot fail isn’t one.
Why All of the AI Models Landed in the Same Place
The convergence that we see in AI designs isn’t coincidence, and it isn’t a ceiling on what these models can produce. It comes down to what they learned from.
There is an enormous public corpus of finished e-learning. Storyboards, screen specifications, module structures, published courses, vendor templates, conference showcases, agency portfolios, LinkedIn posts showing off a nice interaction. Decades of it, all text and images, all of it scrape-able.
There is almost no public record of the thinking that produced any of it.
The conversation where a designer says “this is a willingness problem, not a knowledge problem, so a module won’t move it” happens in a room, with no minutes. The reasoning that leads someone to choose behavioural rehearsal over categorisation practice sits in their head, or at best in a client deck nobody publishes. The needs analysis lives on a shared drive. The pilot that failed and got quietly shelved is documented nowhere at all, because nobody writes up the thing they decided not to build.
So these models have learned design artefacts in extraordinary volume and design method barely at all. They know what e-learning looks like. They have very little access to why any of it looks that way and whether it actually works or not.
That produces a completely consistent behaviour: ask an LLM for an e-learning module and it generates the statistical centre of every e-learning module it has ever seen. What AI produces isn't the output of a design process. It's the output of a sampling process — the statistical centre of every e-learning module it has ever seen.
Finding #3: Design judgement in missing from every model
The craft scores (how well AI executes design tasks) span sixty-eight points, which tells us that the four models are doing some genuinely different things when they build a course. The judgement scores (how well AI makes design decisions) have far less variation.
Not one model asked what new starters currently do wrong. Not one asked whether anyone had evidence they do anything wrong at all. Every model invented a persona and used it to designed for the invention without question. Every output ended at LMS completion, with no behavioural measure anywhere in any of them. Not one proposed a pilot, a success threshold, or any condition under which the design would be abandoned.
Claude Sonnet’s 36% is the only figure that breaks the pattern, and it breaks it for a reason worth understanding.
It earned almost all of that on process behaviour: it asked questions, offered a fork, versioned its draft, flagged risks at handoff and deferred to the human on some of the details.
But the one model that behaved like a genuine thought partner still never asked the question that determines whether the work should exist. Which is worth sitting with, because collaboration and diagnosis look almost identical from the outside and do entirely different things. Asking how would you like this? is collaboration. Asking are you sure you want this? is diagnosis. Claude Sonnet did a great deal of the first and none of the second.
Why this gap can’t close by accident
The explanation from Finding #2 applies here too, one level up.
Craft has a public record, because every finished course is a worked example of the craft. Diagnosis doesn’t have the same record. There is no scrape-able corpus of needs analyses, no archive of the conversation where someone established that reporting rates were low before deciding what to do about it, no published record of the success metric a designer agreed with a stakeholder six months before launch.
So the diagnostic half of the discipline isn’t just underrepresented in AI’s training data: it’s essentially absent from it.
There’s a second reason too, and it’s about incentives rather than data. Nothing in how these models are built rewards reliability or quality. Instead, they’re rewarded for speed and task completion. In practice, this means that the models default to whatever appears most frequently in the documents they learned from — a very large, very mixed pile of existing course documentation — good, bad and indifferent.
Finding #4: Three of the four made errors that would compromise compliance
The SOP I used as my source document requires every employee to report suspected fraud, regardless of materiality, to Corporate Security & Investigations and the local Finance Director.
Claude Sonnet cut a mandatory duty — and documented the decision.
Sonnet cut this, reasoning that “that’s for specific functions, not new-starter onboarding.” It read the report’s recipients as the obligation’s audience. It then recorded the exclusion in an appendix headed, “Content not carried into this new-starter module” — producing a compliance sign-off artefact stating that a universal duty is out of scope for induction.
Copilot taught the wrong procedure in its only practice activity.
In Copilot’s design there’s a simulation that goes like this: “You notice a colleague submitting duplicate invoices for the same supplier. You suspect fraud.” The channels offered in response are line manager, Corporate Ethics & Compliance, and the Speak Up portal.
When compared to the SOP, these channels are incomplete, some are hallucinated and the mandated channel isn’t among them. The rubric then scores “correct channel choice” as a pass condition for a scenario which had no correct option.
Gemini deleted an escalation right. Item §6.2 in the SOP allows you to go outside your line manager if they’re implicated or if you reported it and nothing happened. Gemini kept the first condition and dropped the second, so anyone trained on that module believes the right to escalate exists only when their manager is personally involved.
What links all four of these errors is that none of them is a mistake about the content. Every model read the SOP accurately. What they got wrong was the relationship between the source and the design — which duty applies to which audience, what a practice activity implicitly teaches, what a marked-correct answer certifies, and what quietly failed to make it across.
A SME reviews what’s on the screen, and everything on the screen is faithful to the policy. These errors live in what the screen does with the policy, and in what the screen leaves out — and nobody reviewing for accuracy is looking there, because accuracy review reads forwards from the module. It never reads backwards from the source, asking what this document requires of every reader and whether all of it arrived.
Finding #5: All four wrote in a register that quietly signals “nobody spent long on this”
There’s a well-grounded finding in multimedia learning research called the personalisation principle: learners do measurably better when content addresses them directly and in a conversational register than when it addresses them formally and impersonally (Mayer). This isn’t because friendliness is pleasant, but because direct address and relevance prompts engagement, deeper processing and therefore learning.
And here’s the thing these four outputs demonstrate: they all get the form of personalisation right and the substance of it completely wrong.
Every one of them uses “you.” Every one addresses the learner directly, in a warm, conversational tone. Copilot specifies a “friendly narrator voiceover.” Gemini opens with “Welcome to GSK. As a new team member, you play a vital role.” On the surface, that’s the principle applied.
But here’s the problem…
Not one model asked for learner demographics or psychographics
All four reproduced a US toll-free number and a Philadelphia PO box as the primary reporting route — for a company headquartered in Brentford, with employees across dozens of jurisdictions. That’s partly faithful behaviour, since those details are in the source. But the source is written in British English throughout, and two models Americanised it anyway: “local labor laws,” “ethical behavior,” “Zero Tolerance.” Copilot’s production note reads “replace US phone example with local Speak Up numbers for learners outside the US” — framing the global majority of the workforce as the exception case.
So the module says you in a friendly voice, and then hands you a phone number you can’t dial, in a spelling that isn’t yours. Whatever that communicates, it isn’t this was made for you.
All four models spoke in “AI voice”
When AI thinks “learning” it also thinks “sound professional”. The result is language which feels impersonal and at time inapproropate for the learning context.
In my tests, it showed up in four distinguishable ways:
Unfalsifiable affirmations. “Your voice matters.” “You are protected.” “Zero tolerance for retaliation.” “Speaking up is part of doing the right thing.” “Early reporting prevents harm.” Every one is a statement no learner can test, and none of them is instruction. Gemini’s closing narration names the beneficiary outright — “help protect our company and patients” — and the learner isn’t in that sentence. For a module whose whole job is convincing someone that reporting is safe for them, this is the register least likely to work: it reads as the organisation reassuring itself in the second person.
Elevated abstractions where a concrete detail was needed. “Doing the right thing.” “Our Speak Up culture.” “Ethical behaviour.” These sound like values statements because they are — the models pulled the SOP’s aspirational language forward and used it as instructional content. But you can’t practise “doing the right thing.” A new starter needs to know what to say when their manager asks them to leave something out of a report, and abstractions at that altitude are un-actionable by design.
Chatbot register bleeding into learner-facing copy. “Let’s dive in.” “Let’s practise how to...”. This is the voice these models use to talk to you, the person prompting them in a chat, appearing verbatim in the narration script a learner is supposed to hear.
Compulsive triads. Notice it → Raise it → Use the right route. Recognise / Report / Understand / Practise. Report → Triage → Investigation → Outcome → Feedback. Rule-of-three structures in all four outputs, imposed on content that doesn’t have three parts. The reporting decision isn’t a three-stage process; it’s one judgement call under uncertainty. Forcing it into a triad makes it look procedural, which is precisely the misreading the whole module needs to avoid.
None of these things are a defect on their own, but cumulatively they produce copy that anyone who has spent an hour with AI tools will recognise on sight. The conclusion a learner draws isn’t this was written by AI. It’s nobody spent long on this or cares about this, which is the opposite of what’s needed to engage and learn.
A SME is the last person who’d catch these sorts of problems because they review for factual accuracy not tone. Every line above is accurate — “zero tolerance for retaliation” is a fair summary of §9, “doing the right thing” is lifted from the SOP’s own purpose statement. So an accuracy review finds nothing, because nothing is false. The problem is that none of it is instruction, and spotting the difference between a true statement and a teachable one is a design skill, not a subject-matter one.
What This Means in Practice
AI can do competent learning design work — if. That conditional “if” is the key finding from my tests, and it has two halves:
If the model is good enough. One reached 77% on craft, one 59%, the other two 45% and 9%. The bottom output wasn’t a weaker version of the same thing: no correspondence between objectives, practice and assessment, no feedback written anywhere, an umbrella cartoon standing in for psychological safety. So competent design was available from one model in four — and the two most likely to be sitting in an enterprise licence by default were the two at the bottom.
If the brief already contains what the model needs. Almost everything I praised in Finding #1 fired because of words I happened to include. I named the audience, so ChatGPT-5.5 had a profile to scope against. I said “short,” so it had a reason to cut. Take those out and the behaviours have nothing to attach to — and the whole premise of a non-expert briefing an LLM is that they don’t know which words matter. None of the four told me.
So the honest conclusion isn’t that AI can design learning; it’s that AI executes design well when someone who already knows what to ask for is doing the asking and the checking.
This finding gives us a clearer division of labour than we’ve had before and with it a clearer picture of what the AI-enhanced workflow might look like:
Step 1: Before anything gets built: L&D runs problem
Three questions, in this order, and none of them is what should the course cover?
What’s the problem? Not the topic — the performance gap. What are people doing now that they shouldn’t be, or not doing that they should? And how do we know: incident data, audit findings, manager reports, exit interviews, anything other than someone’s impression that this feels like a risk. If nobody can point to evidence, that’s the first finding.
What’s causing it? This is the question that determines everything downstream, and it’s the one that got skipped in all four of my tests. Is this a knowledge gap — people genuinely don’t know what to do? Or is it fear, incentives, tooling, process, or a manager population who don’t handle it well when someone raises something? Gilbert’s work says the environmental causes outnumber the repertoire ones by a wide margin, so probably not knowledge is the right starting prior.
The SME holds most of the answers to the first question. The organisation and external research data holds the answer to the second. L&D’s job is analysing all three and refusing to move on until they’re answered — because nobody else in the room has a reason to ask them, and AI structurally won’t.
The output is one thing: an evidence-based problem statement.
Step 2: L&D define what do we need to build to solve the problem (if anything)?
Once we have a problem, L&D must ask:
What’s the optimal response? Only now. And it’s a genuinely open question, with “nothing” and “not a learning intervention” both live. If the cause is fear, a course can’t move it. If the cause is a manager population, the intervention is manager work. If it’s a knowledge gap that appears once every eighteen months, it’s a job aid rather than a module.
If it is a learning problem, a second call follows immediately: is it a knowledge, skills or behaviour problem (or a combination of the three)? That single decision determines every strategy downstream — and my test showed all four models defaulting to the same four techniques regardless of the answer.
Neither the SME nor AI can make this call, for different reasons.
The SME can’t, because they’ve usually already decided. They arrive with a solution — we need a course on this — and asking them whether a course is the right answer is asking someone to argue against the thing they came in for. That isn’t a failure of expertise; it’s a structural position. It also requires fluency in the alternatives: performance support, process change, incentive design, manager capability. A compliance expert knows compliance. They have no particular reason to know what shifts a willingness gap, or that willingness gaps and knowledge gaps need different treatment at all.
AI can’t, for three reasons that are related:
1. It has no access to the right sort of context: it can’t see the data needed to make the right design decisions — every model in my test invented a persona and designed for the invention.
2. It’s trained toward helpfulness: helpfulness reads as doing the thing asked, so the brief becomes the specification rather than a hypothesis.
3. Its training data is biased: AI’s “brain” is full of courses designed with knowledge transfer strategy (content > quiz), so unless we work very hard to ban it, it defaults to that solution.
This is why all four of the models I tested correctly identified fear of retaliation as the barrier and then built a course anyway. The analysis was there, the ability to act on it wasn’t, because nothing in the workflow asked for it and nothing in the model’s incentives raises it unprompted.
This stage is where L&D earns its place in a post-AI world. L&D’s role isn’t building better than AI can — AI demonstrably builds well. Instead, our value comes from being the only party in the process with the expertise to say this isn’t a learning problem, or it’s a learning problem but not that kind.
Step 3: Hand over to SME + AI
Once the brief and constraints are in place, the SME + AI can go with one caveat: the brief they’re working from must contain all of the things my test showed that SMEs + AI can’t supply themselves: the audience, the constraints, the strategy, the delivery context, the success measures, the tone of voice .
With these guardrails in place, AI can execute and collaborate with SMEs less like an apprentice and more like an expert team mate. Meanwhile, L&D can focus elsewhere, until the design is complete.
Step 4: Expert L&D review
In the post-AI world, expert review is everything. In our work, we now have two sorts of QA to run:
Pedagogical QA — did it design what I asked for?
In this process, you start with the brief, not the module. You specified a strategy, an audience, a form, a duration and a success measure. L&D’s job is to read the output against each one and check it was delivered:
Did it use the strategy you named? If the brief said behavioural rehearsal and the module contains a sorting exercise and a quiz, it substituted the common technique for the specified one — and it won’t have flagged the substitution.
Did it hold the constraints? Duration, screen count, delivery context, device mix, accessibility requirements, localisation. These get quietly renegotiated: “short” becomes twenty-seven minutes, and the model presents the new number as though you’d asked for it.
Did it design for the audience you gave it, or for a persona it invented? Look for the composite learner who doesn’t exist — the one with no prior knowledge, no time pressure, and no relationship with the person they’d have to report.
Is the assessment still measuring your objective? Objectives, practice and assessment in three columns, read across. An apply-level objective tested with recall means the assessment drifted to what’s easy to score.
And did it invent anything you didn’t ask for? Pass marks, item counts, seat time, screen counts. Every number should trace back to your brief, design arithmetic, or a citation. If it can’t, it was filled in — confidently, and without being mentioned.
The review isn’t asking is this good? It’s asking is this what I specified? — and the gap between the two is where the design quality is assured.
AI QA — does it avoid common AI errors and tells?
These are the errors that come from how these models work rather than from a lack of skill. They’re predictable, which makes them checkable — and they show up in five clusters.
Silent omission. The model drops something and doesn’t tell you. It cut a section, mis-scoped an obligation, or collapsed a two-part condition into one — and there’s no gap in the output where it used to be. This is why the check has to run backwards from the source, listing what the document requires and confirming each item arrived. Reading forwards from the module can only ever verify what’s there.
Confident fabrication. Any value the model doesn’t have, it invents — and it invents at the same confidence as a value it derived. Pass marks, seat time, item counts, screen counts, per-section time budgets that add up perfectly over numbers from nowhere. In the worst case it fabricates procedure: Copilot offered three reporting channels for a fraud scenario and the mandated one wasn’t among them.
Named but not built. The vocabulary of good design arrives without the design. “Branching scenario,” “simulated report,” “scoring rubric,” “tailored feedback,” “remediation pages.” Every term is correct and none describes what’s in the file. Find the branch. If you can’t, it isn’t there.
Strategic monoculture. The model reaches for the statistical centre of everything it has seen — categorisation practice, click-to-reveal, single-turn MCQ, terminal quiz. If all four appear regardless of what the behaviour needs, the format chose the pedagogy rather than the other way round.
AI voice. Unfalsifiable affirmations, values language standing in for instruction, chatbot register leaking into narration, rule-of-three structures imposed on content that hasn’t got three parts.
What links all five is that each one reads as correct. Nothing is false, nothing is missing from view, nothing looks unfinished. Which is why an accuracy review — the only review most of this content gets — clears every one of them.
Every one of those reads as accurate. None of them teaches anything. And an accuracy review — which is the only review most of this content gets — finds nothing wrong with any of it.
Conclusion
The wicked problem was never that subject experts write bad courses: it’s that knowing a thing and teaching a thing are two very different capabilities.
For decades, slow and specialised production time was an accidental safeguard against SME design: a SME could design a course, but they usually couldn’t build it without pulling in L&D. AI broke this constraint.
What's emerging as a result a rising volume of learning experiences that are well-formatted, accessible, professionally structured, but quietly and deeply broken. The defects are real but they don't surface: a missing needs analysis leaves no gap on the screen, a substituted strategy looks like a strategy, a mandatory duty that never made it across is invisible to anyone reading forwards from the module. Nothing looks wrong, so nothing gets sent back — and the volume is growing faster than anyone's capacity to check it.
For L&D, this isn’t a story about being replaced. Production was the bottleneck everyone complained about, but diagnosis & solution definition was the bottleneck that always mattered. The real challenge is helping those around us to understand this, and presenting them with a revised approach to L&D which delivers not just speed and efficiency but also business impact.
Happy innovating!
Phil 👋
PS: What to explore how to lead in the world of AI & L&D with me? Apply for a place on my bootcamp.






What a fantastic article! @Dr. Philippa Hardman, could you share your 36-point instructional design standard? Thank you very much in advance! Best regards, Morten