When anyone can build a course, the real job is deciding which ones shouldn’t exist
Why deciding is the only L&D skill AI can't replace.
Hello folks 👋
Last week, the investor Andrew Chen put a one-line thesis on his X timeline:
“When anyone can build, the person who decides WHAT to build becomes the bottleneck.”
Reading the post and the responses, something struck me : Chen is diagnosing a tech problem, but he’s also describing an L&D problem.
If you work in L&D right now, you already know: “time to build” isn’t the #1 bottleneck anymore. AI did to course production what it did to code production: a course that used to take six weeks now takes an afternoon. Storyboards generate themselves. Scripts write themselves. Assessment items, micro-learning, scenario branching — all of it, all of the time, on tap.
The biggest AI risk that L&D faces isn’t that it gets left behind: it’s that we build more — and flood the organisation with meh-quality content nobody needed in the first place.
In this post, I’ll make the case that:
→ The L&D job has just split in two — and most of us are still working on the wrong half.
→ There’s a new operating model coming for the role, and it’s already running inside a lot of the companies you’ve heard of.
→ The smartest critique of everything I’m about to argue comes from Ethan Mollick — and I think he’s half right.
Let’s dive in!
The Reframe: AI is Commoditising Production (and elevating decision-making)
Perhaps clearest articulation of how L&D is changing came last year from Tomer Cohen, LinkedIn’s Chief Product Officer, on Lenny’s Podcast last December (Cohen, 2025).
Cohen’s framing was blunt: AI is collapsing the cost of execution (in our case, content & course production). What stays durably human in the age of AI, he argues, is judgement — the deciding — and the ability to rally people around the decisions we make.
If Cohen names the principle, Boris Cherny — who leads the Claude Code team at Anthropic — has shown what it looks like operationally. Two insights from a recent in-depth interview resonante here (Pragmatic Engineer, 2026):
→ Product Requirement Documents are dead on his team. The product requirements document — the spec teams used to write before building — has been replaced by working prototypes. They generate the thing, look at it, decide whether it deserves to exist. Judgement-on-the-artifact, not judgement-on-the-spec.
→ Code output per Anthropic engineer is up 200% in a year — and that surge has created a new bottleneck around code review, which Anthropic now hears about from customers weekly (VentureBeat, 2026).
If we read those two findings again with L&D goggles on, the parallels are clear and uncomfortable:
The course design document is the L&D PRD. Storyboards, scripts, learning objectives, assessment blueprints, design rationales — the artefacts most teams still treat as foundational deliverables — are functionally the same as the PRD: a written specification of an artefact you’re about to build. Tech is realising the spec is no longer the highest-leverage artefact, because AI can now generate the working thing in less time than it takes to spec it. The same is now true in L&D. Why write a forty-page design document for something AI can prototype in an afternoon? The unit of judgement has shifted from spec to artefact — and most L&D processes haven’t caught up.
The 200% output / review bottleneck is our future, not theirs. Anthropic produced more code than its review capacity could absorb, and their customers feel it weekly. Now imagine the L&D equivalent. AI lets one designer generate ten times more content than they could a year ago. The bottleneck is no longer authoring tools, SME availability, or sign-off cycles. It’s whether anyone’s expert eye has actually evaluated whether what’s being shipped will produce behaviour change. Most L&D functions don’t yet have a review layer that scales with AI-driven output — and the ones that do are mostly checking compliance, not craft.
→ The role splits along the same fault line. Cherny’s team didn’t make engineers redundant — it elevated them out of writing specs and into judging artefacts. That’s the same move L&D is being asked to make right now. The designer who keeps doing the production work AI can do faster will be measured against AI’s speed and lose. The designer who moves up the stack to judging — is this the right intervention, is this draft pedagogically sound, is this assessment actually going to discriminate, should this even ship — becomes more valuable, not less.
The pattern is consistent across both fields: production capacity exploded, the judgement layer didn’t scale, and the role and value of L&D is now increasingly defined by which side of that gap we choose to operate on.
The Emerging L&D Role, aka the 3Ds
Here’s the emerging model of L&D that I see on the ground, presented as simply as I can put it: three streams of work, two of which belong to AI and one of which belongs to the human in the loop.
Stream 1: Data — AI’s job. Learner research synthesis, capability gap analysis, content drafts, storyboards, scripts, assessment items, scenarios, transcripts of synthetic users working through the design. All of it, in minutes.
Stream 2: Doing — AI’s job. Production, formatting, building, deploying. Multi-agent workflows where one agent summarises sources while another drafts the script while another generates assessment items while another analyses learner data from the last iteration.
Stream 3: Deciding — the human’s job. Is this the right capability gap to solve? Is this the right intervention — or should we not be building a course at all? Is this draft pedagogically sound, or just plausible? Will this assessment discriminate? Will this scenario produce behaviour change, or just look like it will? Of these fifteen AI-generated variants, which three are worth scaling?
The emerging “post-AI” L&D professional isn’t a deskilled functionalist or a generalist — they’re a specialist judge. The job is no longer about executing and building, but deciding well.
Deciding well requires three things that structurally AI doesn’t have:
→ Deep learning expertise. Not familiarity with frameworks. Expertise — the kind that comes from years of designing, testing, watching things fail, and pattern-matching across hundreds of cases. The expertise to look at an AI-generated draft and know it’s plausibly good but pedagogically wrong. To spot the assessment item that won’t discriminate. The scenario that won’t transfer. The objective that’s verb-rich but unmeasurable. AI can generate fluent learning content all day. It can’t tell you which version will actually change what people do at work — because that requires understanding cognitive load, retrieval practice, transfer conditions, motivation, and the difference between performance and learning. That knowledge sits in the practitioner, not the model.
→ Specific business context. Generic best practice is what AI is fastest at. This business, this team, this manager culture, this set of competing priorities, this history of failed initiatives — that’s where judgement actually lives. The same training intervention that works at one company will fail at the next, because the context around it is different. Knowing which capability gap is real, which intervention fits the system it’s landing in, and which design choices will survive contact with reality is contextual work. AI can recommend in general. It can’t decide in particular.
→ Human accountability for impact. This is the dimension most easily overlooked, and the most structurally important. Professional judgement isn’t just making the right call. It’s making the call and being answerable for it. When a learning designer signs off on a programme that doesn’t transfer, or kills a request that the business later regrets killing, or refuses to build something the CEO wanted — there’s a name attached. A reputation. A career. AI can produce a recommendation. It cannot stake anything on it. And the more AI handles the production work, the more visible — and the more valuable — that human accountability becomes.
These three things compound when they combine: Deep expertise without business context is academic. Business context without expertise is opinion. Both without accountability is theatre. The L&D professional who holds all three is the one organisations will pay for in 2026 and beyond.
Building this for Real: an experiment in progress
This week I started an experiment with a Fortune 500 to test exactly this architecture — the three Ds — in practice.
The setup: one learning designer, working with a chain of four AI agents, one for each stage of the instructional design lifecycle.
→ The Problem Framing agent runs analysis. It interrogates the brief, scans the evidence, and — critically — assesses whether learning is even the right intervention. It can recommend “this isn’t a learning problem” and refuse to hand off downstream.
→ The Behaviour Change Design agent runs design. It defines target behaviours in observable terms, surfaces cues and reinforcement loops in the actual workflow, and produces options with trade-offs — never a single verdict.
→ The Quality agent runs review. It checks objective-to-assessment alignment, accessibility, and consistency. Bounded, deliberate, not “review everything.”
→ The KPI Tracking agent runs evaluation. It splits planning from measurement, and refuses to fabricate a status when the evidence isn’t there.
The agents handle the Data and the Doing. Between every stage sits a Deciding gate. Each agent gathers, drafts, surfaces options, flags assumptions and open questions — then stops. Nothing moves forward until the human in the loop makes the judgement call. When they do, the next agent executes on command.
This is the emerging future of both the L&D operating model and the role of the humans within it: a return to the part of the job that always mattered most — and the part most of us spent the last decade unable to fully do because production work ate our weeks.
Day-to-day, the work shifts upward. Less time in authoring tools, more time in stakeholder rooms making the case for or against an intervention. Less time formatting modules, more time reading what the agents have produced and judging it against what the business actually needs. Less time managing project plans, more time orchestrating an agent stack and gatekeeping its outputs.
The skills that matter compound: learning science expertise applied to specific business problems, the orchestration fluency to run agentic workflows well, and the professional standing to put your name to a decision and defend it.
With the right operating model in place, the L&D role doesn’t shrink in this future. It returns to its proper shape: judgement-led, expertise-anchored & accountable.
But wait — isn’t AI getting good at judgement too?
Something I’ve heard this a lot recently is that AI is rapidly getting better at judgement.
This argument is right, narrowly. Modern agents are getting good at a certain kind of judgement. Things like sequencing tasks, prioritising sub-goals, deciding when something’s “done enough” to move forward. Boris Cherny described a team of agents at Anthropic recently which are doing this constantly, and the better ones really are better at it than humans.
But that’s not the kind of judgement L&D — or medicine, law, or strategy — actually trades in. Professional judgement is structurally different on three dimensions:
→ Routine procedural judgement. Sequencing, prioritisation, “is this draft good enough to ship to the next stage?” AI can do this, and increasingly will.
→ Domain-expert judgement. “Is this assessment item pedagogically sound? Will this scenario actually transfer? Does this design align to the objective?” AI can simulate this and is improving fast — but it’s still the contested zone, where the tacit case-base of an experienced learning designer outperforms the model. Precedent from the world of medicine shows us that AI’s domain-expertise will increase in the next two to three years, but not remove the need for a human expert entirely.
→ Professional judgement proper. “Should this exist at all? Am I willing to put my name to it? Will I defend this decision in five years’ time when someone asks why we built it?” This is where accountability, refusal, and stakes-bearing live. AI structurally cannot do this — not because it isn’t smart enough, but because it isn’t a party to the situation.
AI can recommend, but it cannot be accountable. Accountability is the structural reason professional judgement carries weight in organisations: a real human, with a name and a track record, is staking their professional credibility on the call. When a learning designer signs off on a programme that doesn’t transfer, or a compliance refresh that gets the company sued, or a leadership intervention that flames out — somebody answers for it. Liability lives somewhere. Regulation lives somewhere. The implicit social contract of expertise lives somewhere.
That somewhere can never be a model.
The same logic applies to refusal. A learning designer telling a stakeholder “training won’t fix this — your incentive structure is wrong” is exercising a kind of judgement AI can mimic but cannot inhabit. The refusal only counts when it carries professional credibility, organisational standing, and the implicit threat of professional consequences if it’s ignored. AI can flag concerns. Only humans can refuse on behalf of a profession.
So Mollick and others are right that the procedural layer of judgement is leaking to AI, and right that it’s doing so faster than people realise. He’s missing that the more sophisticated AI gets at that layer, the more clearly the human-only zone clarifies itself: it’s not the cognitive work, it’s the answerability.
The smarter the agents get, the more - not less - the valuable the accountable human becomes.
Concluding Thoughts
The question we’ve been asking for the last two years — “how do I get faster at building?” — was the wrong one.
The real question is: can I look at fifteen AI-generated learning assets and decide which three are worth scaling — and put my name to that decision?
That’s the job. And it’s exposed in a way it hasn’t been before.
The compliance refresh that won’t change behaviour. The onboarding course that should be a checklist. The leadership programme that should be coaching. The skills gap that’s actually a manager problem. The performance issue that won’t be fixed by training.
Anyone with AI can generate a course in an afternoon. Almost nobody can look at a business problem and decide that “training isn’t the answer here” — and be right. And almost nobody can stake their professional reputation on that refusal, year after year, in front of stakeholders who’d prefer to hear yes.
That decision is the job. It always was. Production work was just blocking the view.
The first step isn’t to fire up an agent. It’s to look honestly at your own workflow and decide what belongs where.
Take a single project you’re working on right now — a course, a programme, an intervention — and walk through every task it involves, from the initial intake conversation to the final post-launch review. For each task, ask one question:
→ Is this Data? Information-gathering, synthesis, drafting, generating, analysing, formatting. Anything where the work is producing or transforming material. AI’s job.
→ Is this Doing? Building, deploying, updating, distributing, running. Anything where the work is execution. AI’s job.
→ Is this Deciding? A judgement call that requires expertise, context, stakes, and accountability. Is this the right gap? Is this the right intervention? Is this draft pedagogically sound? Will this actually transfer? Should we be building this at all? The human’s job.
Once you’ve mapped one project, you have the start of an operating model. Decide what to delegate. Decide what to protect. Decide where the gates sit between stages — the moments where AI has to stop and wait for you to make the call.
That mapping exercise is the cheapest, fastest, highest-leverage move any L&D professional can make this month. It costs nothing, it requires no tools and it tells you, concretely, where your job is heading.
Happy deciding,
Phil 👋
PS — If you want to learn how to build the agents that handle the Data and Doing once you’ve decided what to delegate, apply for a place on my AI for L&D Bootcamp





