Four ways to judge whether training worked
Techniques for evaluating training programs
In one line: you can judge training at four levels: did people like it, did they learn it, do they now do it at work, and did it make a difference? A lot of checking stops at the first.
What Kirkpatrick did
This isn’t an experiment. It’s a practical framework, set out by a university teacher for people who ran training in companies.
Donald Kirkpatrick had studied how to evaluate a training programme for supervisors in his PhD at the University of Wisconsin (1954). Training departments wanted to show they were worth the money, but had no shared way to talk about evaluation beyond “people seemed to like it”. So in 1959 and 1960 he wrote four articles for a training managers’ journal, one for each of four “steps”. He later expanded them into a book, Evaluating Training Programs: The Four Levels (1994), which became the standard reference.
What he argued
- Level 1, Reaction: did they like it? Usually a feedback form at the end. The easiest level to check, and the most used. It tells you about the room, the pace and the trainer, but not whether anyone learned anything.
- Level 2, Learning: did they learn it? Checked with a quiz, a demonstration or watching people try it. Harder to set up, and more meaningful.
- Level 3, Behaviour: do they do it at work? This needs a follow-up weeks or months later, by watching people or asking them. It’s arguably where the value is, and it’s checked far less often than Level 1.
- Level 4, Results: did it make a difference? For example, fewer medication errors, or fewer patients’ warning signs being missed. Very hard to pin on one session, because so much else changes at the same time.
How much should you trust it?
Useful shared language; not a tested theory, and the levels don’t lead to each other automatically.
- Kirkpatrick built the levels from experience, not from data. He didn’t test how often doing well at one level leads to doing well at the next.
- When others did test it, the links were weak. A 1997 review by Alliger and colleagues pooled 34 studies. It found that how much people enjoyed training was barely related to how much they learned or changed. Ratings of how useful it was did a little better, but were still only weak signals.
- Critics say the model makes training look like a neat ladder, when change at work also depends on the team, the manager and the system people return to.
- It was designed for single training events, not for the slow learning that happens by working alongside others every day.
If you need to convince someone
In our words (a paraphrase, not a quotation): “Did they like it?” is a different question from “Did it change anything?” A feedback form can only answer the first.
What it means for you
A feedback form asking “Was the session well organised?” and “Was the trainer knowledgeable?” mostly tells you about the trainer and first impressions. It says nothing about whether anyone works differently three months later. A small, realistic plan for any session:
- Level 1: a short form or paper slip at the end, used mainly to fix practical things (the room, the timing, the venue nobody could find).
- Level 2: one quick task or question that checks the main thing people should be able to do.
- Level 3: one follow-up at about 30 days, by email or a message: “What have you changed since the session? What got in the way?”
Start with what people need to do differently, and you already know what to look for at Level 3.
How it appears in Teaching That Lands
Kirkpatrick is the backbone of Session 3. Each table sorts 18 evidence cards from a made-up four-day programme, from delivered to impact. A staircase slide then puts a Kirkpatrick level on each step, and shows how little of the usual evidence says anything changed. Participants then rewrite their own outcomes, choose evidence that matches each one, and plan their follow-through: what their learners will revisit at 30, 60 and 90 days, who will hold them to it, and what support they need from that person. The programme judges itself the same way, with three short questions at 30 and 90 days: see the Evaluation Toolkit.
The small print
- Dates. Kirkpatrick began this work in his 1954 PhD. The articles summarised here came out in 1959 and 1960. Kirkpatrick called the levels “steps” at first.
- Level 5. Jack Phillips later added a fifth level, return on investment: did the benefits outweigh the cost? It isn’t part of Kirkpatrick’s original model.
- The New World model. In the 2010s, James and Wendy Kirkpatrick added “required drivers”: the things at work, such as reminders, support and a manager who asks, that help learning turn into doing.
- Other approaches. Brinkerhoff’s Success Case Method argues that average scores hide the story. It looks closely at the people who did and did not use the training, to find out what helped or blocked them. For learning that happens by joining in everyday work rather than on a course, see Lave and Wenger.
Kirkpatrick, D. L. (1959). Techniques for evaluating training programs. Journal of the American Society of Training Directors.