What Anthropic's New"Record a Skill" Feature Means for L&D
Aka, how Anthropic's new screen capture feature could change how we design & deliver learning experiences
Hey folks 👋
Anthropic just launched a new feature called Record a skill inside Claude Cowork. Most people — including most L&D folks — haven’t heard of it yet (and still don’t have access to it yet). But I think this could be one of the most consequential AI feature releases for our field this year, both for what it makes possible and for a new risk it introduces that I haven’t seen anyone name.
In this week’s post, I’ll give you lowdown on what the tool actually is, where it fits in tour day to day work, what it genuinely solves, and where it might make our oldest problem worse while appearing to fix it.
Let’s dive in!
What Is “Record a Skill”?
The concept is simple. Instead of writing instructions for an AI by hand, you record yourself doing a task — your screen, your clicks, your typing, and (if you leave the mic on) your voice as you narrate. When you’re done, Claude analyses the recording and proposes a “skill”: a structured, reusable description of the task that the AI can then follow to do the work itself, or that you can review, edit and share.

Some practical details:
→ Record a Skill lives inside Cowork in Claude (Pro, Max and Team plans at the time of writing)
→ A single recording can run for up to ~10 minutes.
→ Everything on screen is captured, along with anything you say — so you can add context as you go, e.g. explain why you're doing each step, what you're deliberately not doing, and when you'd do it differently
→ When you press "done", Claude analyses the recording and drafts the skill — a step-by-step file which enables you to repeat the task infinitely. You can review and edit it before saving, so nothing is created without your sign-off
Anthropic are selling this feature as a productivity hack, but when I read about Record a Skill through an L&D lens something jumped out: this isn’t just a productivity feature — it’s potentially a powerful knowledge extraction machine.
Record a Skill is the first mainstream tool that captures expertise from behaviour and action rather than from conversation about behaviour — and that has implications for every stage of our workflow.
A ten-minute recording could double as a cognitive task analysis. A narrated expert performance could become the source for every derivative asset — job aid, storyboard, assessment checklist — from a single capture. And the same mechanism could automate L&D’s own repetitive admin, buying back hours for the judgement work only we can do.
But — as ever — there’s a catch, and it’s a big one….
From Asking About the Work to Watching It
In a recent post I explored the wicked problem of extracting knowledge from Subject Matter Eexperts (SMEs). In it, I mapped four ways teams are using AI to extract expert knowledge today:
→ AI mines the conversation — L&D folks record SME interviews, then transcribe and code them with help from AI.
→ AI interviews the SME — L&D folks build GPTs which can interview SMEs on their behalf.
→ AI as junior SME — L&D folks use research-grounded AI tools like Perplexity and Elicit to “become expert” in topics without over-reliance on SMEs.
→ AI analyses the SME’s outputs — L&D folks gather the work / outputs of SMEs and use AI to recreate their brain and expertise, e.g. using NotebookLM.
Notice what all four have in common: the knowledge comes either from talking about the work or from documents produced by the work. None of them observes the work itself.
Record a Skill is the fifth approach: AI watches the SME work. Until now, this category — sometimes called ambient or observant AI — lived almost entirely in medicine and software engineering research labs: eye-tracking radiologists, logging analyst click-streams. Very soon, it will be available to anyone with a Claude license.
The question that sprung to my mind was: does observation of someone at work solves the problem the other four approaches couldn’t? I think the honest answer is: partly — in a way that’s both genuinely valuable but also genuinely dangerous.
What “Observant AI” Solves: The Action Layer
Cast your mind back to the cognitive task analysis research I cited in my recent post on working with SMEs. When expert surgeons verbally described procedures they’d performed hundreds of times, they omitted an average of 51% of action steps and 73% of decision steps from their accounts (Sullivan et al., 2014).
Shifting from narration of skills to recording of skills has the potential to improve knowledge elicitation and solve “the curse of expertise” in two ways:
The screen captures what the expert never says. In an interview, if the expert doesn’t mention a step, it’s gone — that’s the 51% of action steps lost to free recall. In a recording, the step happens on screen whether or not it’s narrated. The tab they switch to, the field they check first, the report they open before the meeting: the behaviour can’t omit itself. For desk-based action knowledge (i.e. how people do a task on a computer), watching is close to a complete fix for what asking loses.
The narration itself is better narration. In a recording, the parts that the expert does explain are also captured differently. An interview is retrospective; the expert reconstructs the task from memory, days or weeks after last performing it. A recording captures them talking while doing: think-aloud rather than recall. This is the foundation of protocol analysis — concurrent verbalisation accesses reasoning while it’s still in working memory, instead of asking the expert to remember what they were thinking afterwards (Ericsson & Simon, 1993). Same expert, same willingness to explain — and the explanation is more accurate simply because of when it’s given.
Put the two together and the comparison with the established process of knowledge elicitation is pretty stark. An interview asks the expert to reconstruct the task from memory, then asks the designer to reconstruct the expert’s reconstruction from notes. A recording removes both of the steps where detail is lost: the task itself becomes the source document.
Of course, not all SME expertise happens on a desktop. But the potential use cases of a tool like within L&D are still huge. Systems training, tool operation, admin and compliance workflows — procedural, software-mediated work — is the territory where much of corporate learning content lives, and exactly the territory where a screen recording captures the whole performance.
A tool like this also has the potential to change the L&D maintenance model. Captured knowledge has always decayed the moment the process changed and keeping up with these changes has distracted L&D teams for decades. AI tools which record and then scale skills could reset this: when the system UI changes, you don’t revise the documentation, you re-record and regenerate.
What This Doesn’t Solve: The Decision Layer
Cast your mind back again to my recent post on working with SMEs. In it, I cited research that shows that when experts are interviewed verbally, as well as losing ~50% of action steps, they also omit 73% of decision steps from their accounts (Sullivan et al., 2014). Here, tools like Record a Skill can’t help, for tww reasons:
The screen can’t capture decisions, because decisions leave no behavioural trace. Recording experts at work fixes the “capturing action steps” problem because the behaviour can’t omit itself — the click happens on screen whether it’s narrated or not. But a decision is not a behaviour. A rejected alternative was never enacted: the screen shows the path the expert took, and can never show the three paths they considered and dismissed in the same moment. TLDR: where action knowledge can’t hide from the camera, decision knowledge is invisible to it by construction.
The expert’s narration still has gaps. Think-aloud beats retrospective recall, but it only surfaces reasoning the expert can consciously reach. The curse of expertise says the most important reasoning isn’t reachable: with practice, judgement becomes compiled and automatic, running below conscious access. The expert doesn’t experience the decision as a decision — it takes a fraction of a second and feels like simply continuing. There’s no pause to narrate, because from the inside, nothing happened. This is why even the best structured interview methods still leave roughly a third of expert knowledge unspoken (Sullivan et al., 2014): the ceiling isn’t the elicitation method, it’s the expert’s own access to their expertise.
So, while recording fixes the omission of actions on both fronts it can’t solve the articulation problem, aka the curse of knowledge and the deepest layer of the SME bottleneck. This matters, because decision knowledge isn’t a nice-to-have layer on top of the procedure; it’s the layer that makes an expert worth capturing in the first place.
If the story ended there, we’d have a useful-but-partial tool. But I have another thought here too…
A New Risk: False Completeness
As we know all too well by now, large language models do not tolerate incompleteness. When recording a skill, an expert who acts without explaining will have their explanation plausibly inferred by the model - i.e. AI will fill in the blanks in order to “package the skill”.

In practice, this means that the resulting artefact contains both reliable, observed steps and AI-generated reasoning, both seamlessly interleaved with nothing to differentiate which is which.
Four Ways to Use Record a Skill in L&D (and Where Each Breaks)
So how do we use a tool like this well, given what we know it captures and what we know it misses? Here are the five use cases I’d start with — each with the benefit and the failure mode stated, because with this tool they always come as a pair.
Use Case 1 — The ten-minute task analysis.
Record an expert performing the task, narrating as they go, and the AI’s decomposition is — functionally — a cognitive task analysis that costs ten minutes instead of ten hours. Two hacks can multiply value here: record a novice and an expert doing the same task, and the delta between the two files is your curriculum — a behaviourally-defined gap, not an asserted one. Record three experts, and where they diverge you’ve found either legitimate style variation (don’t train it) or an unresolved process question (fix that before building anything).
Where it breaks: the file contains observed steps and inferred reasoning, seamlessly interleaved. Used it as a draft to be corrected with the SME, not a source of truth.
Use Case 2 — Design practice activities from real decision points.
Every pause, choice and narrated “because” in a recording is a candidate practice activity, drawn from real system states and real inputs — rather than the sanitised examples SMEs invent in interviews, which are systematically cleaner than reality.
Where it breaks: a recording captures one path through the task. The edge cases, exceptions and rejected alternatives — the raw material of good practice design — only make it into the capture if the narration deliberately surfaces them. Record the hard case, not just the clean run.
Use case 3 — Design evaluation against a real standard.
Most workplace assessment measures recall or scenario judgement as proxies for performance. An expert recording gives you a criterion performance to assess against — the standard defined behaviourally, not rhetorically.
For example, instead of asking trainee analysts to describe a supplier risk check, compare their recorded attempt against your senior analyst's: did they open the payment history first, cross-check both registries, pause on the young accounts? The question shifts from "can they talk about good practice?" to "does their performance match one?"
Where it breaks: one expert’s path is only ever one expert’s path. Before it becomes “the standard,” you need to know whether other competent performers would do it differently (see use case 1’s three-expert test).
Use case 4 — Automate software-based admin (start here!).
Enrolment chasing, LMS report pulls, comms scheduling, certificate processing — the function’s most repetitive work is exactly the low-judgement, software-mediated sort of task set where this tool thrives. Early testing suggest that record a skill can reliably capture and automate unattended these sort of basic, functional tasks. So maybe the best use of a tool like this is to buy back hours from these functional, low risk tasks so you can do higher value judgement work without AI.
Where it breaks: it doesn’t, really — which is itself the finding here. If a ten-minute recording produces a skill that reliably executes a clearly defined task, AI can repeat it as reliably as a junior colleague (with similar caveats re. checking their work, especially at the start of your work together).
The Designer’s New Job: Designing Narration?
Early adopters of record a skill have already worked out that the quality of a captured skill is entirely dependent on the quality of the narration: “state the why - and why not something else - for every step” is the advice circulating in the community. What they've landed on is something our field has a name for: a cognitive task interview (CTA)— except here, the expert is interviewing themselves while they work.
In practice, this means that tools like record a skill don’t remove but rather move the knowledge extraction bottleneck. The knowledge extraction bottleneck is no longer who conducts the elicitation — the recording does that. It’s who designs the script the expert follows while working. This strikes me as an instructional design competence which requires understanding of ~40 forty years of of CTA research on how to surface tacit knowledge.
In practice, these expertly designed templates would require the narrator to surface task, action and reason, e.g.
→ “I’m doing [task] so that [outcome]” — intent before action
→ “Before I decide here, what I’m looking at is [cue]” — the perceptual layer; what the expert checks before they choose
→ “I’m NOT doing [alternative] here, because [reason]” — rejected options; the knowledge that leaves no behavioural trace
→ “This changes when [condition] — in that case I’d [variation]” — conditionality; the decision tree beyond the path being recorded
→ “If [signal] looked off, I’d stop and [action]” — failure detection; what “wrong” looks like before it becomes an error
→ “That was the tricky part — everything else is routine” — salience; telling the model where the judgement actually lives
No script breaks the articulation ceiling — the truly compiled knowledge stays compiled. But the gap between an ad-lib narration and a designed one is the gap between capturing what the expert happens to mention and capturing what a trained interviewer would have asked.
Maybe as L&D folks we won’t conduct the elicitation any more, but spend the time we save there designing it optimally?
Closing Thoughts
Every wave of knowledge extraction technology has promised to solve the SME bottleneck, and every wave has solved one layer while leaving the others standing:
→ Expert systems (1980s) captured the rules — and hit the articulation wall Feigenbaum himself predicted
→ Knowledge management (1990s–2000s) captured the documents — which went stale the moment the process changed
→ Rapid authoring (2000s) put the tools in SMEs’ hands — and the pedagogy walked out of the room
Anthropic's new record a skill feature (and the inevitable wave of other "observant AI" features that will emerge in 2026) fits the same pattern: it solves the access problem but not the articulation (quality and impact) problem.
Zoom out and there’s a pattern emerging for how AI is impacting L&D work. Every time AI absorbs a layer of our work, the bottleneck doesn’t disappear; it shifts, from production to expertise.
When AI could suddenly generate courses, the scarce skill stopped being authoring and became knowing what to build and whether to build at all. When AI could suddenly generate feedback, the scarce skill stopped being writing comments and became defining what great feedback looks like. And now that AI can conduct the elicitation — watching, transcribing, decomposing — the scarce skill stops being running the interview and becomes designing it: knowing which prompts surface tacit knowledge and which produce polite surface description, knowing when a capture is close to complete and when it’s the visible shell of an expertise the screen can’t see, knowing that a fluent file is a hypothesis to be falsified rather than a record to be trusted.
And this is the part of this story I find genuinely encouraging and worrying. AI tools keep getting better at the doing, which raises the premium on the judgement that directs the doing.
Record a skill doesn’t need less instructional design expertise from us — it needs more, and a more sophisticated kind: less time producing, more time deciding what the machine should capture, how it should be corrected, and what a human should do with the result.
The tool can now watch the expert work. Deciding what to record, what the expert should say while it watches — and what happens to the record afterwards — remains the deeply specialised and expert work of L&D folks.
Happy recording!
Phil 👋
PS: What to explore how AI is impacting L&D with me and a cohort of L&D pros like you? Apply for a place on my bootcamp.



