Theory Library Kirkpatrick Model

Did the training work? Four levels

Kirkpatrick Model

Judge training by what people do differently afterwards, not just by whether they liked it or passed a quiz

What it says, in plain language

The idea in one sentence: there are four ways to judge whether training worked: did people like it, did they learn it, do they now do it at work, and did it make a difference? Each one is harder to check than the last, and gets closer to whether anything really changed.

You’ve seen this already

Imagine you go to an evening cookery class to learn a quick, healthy curry. Four questions could tell you whether it worked:

  1. Did you enjoy the evening?
  2. By the end, could you make the dish while the teacher watched?
  3. A month later, do you actually cook it at home on a busy weeknight?
  4. Is your family eating better?

You could love the evening and never cook the curry again. Plenty of us have a shelf of recipe books to prove it. Only questions 3 and 4 tell you the class changed anything.

Why it matters for your teaching

Healthcare training is often judged mainly by the first question: the feedback form at the end. Here are the four levels for a session teaching a team a new way to escalate a patient who is getting worse:

LevelThe questionWhat you might check
1. ReactionDid they like it?The feedback form: “useful, well run”
2. LearningDid they learn it?Each person talks through a case, or gets a short quiz right
3. BehaviourDo they do it at work?A month later: are the healthcare assistants, nurses and doctors using the new escalation steps?
4. ResultsDid it make a difference?Are fewer patients’ warning signs being missed?

Level 3 is where a lot of training quietly falls down. People learn the thing in the room, mean to use it, and then the old routine wins. That gap rarely closes on its own. It needs things at work to support it: reminders, a manager who asks, time to practise, and a team where it feels safe to try something new (see psychological safety).

This is why Teaching That Lands starts with what people need to do differently. If you know that from the start, you already know what Level 3 looks like, and you can plan how you’ll check it.

Try this

  • Decide what Level 3 looks like before you teach. Finish this sentence: “A month from now, I’ll know it worked if I see people…”
  • Plan one follow-up. At about 30 days, ask three short questions: what did you change, what happened, what got in the way? Decide now who will send them.
  • Use the feedback form for what it’s good at. It tells you about the room, the timing and the venue nobody could find. Whether people enjoyed it says very little about whether they’ll change. Asking if they found it useful tells you a bit more, but it’s no substitute for checking what they do.
  • Don’t panic if a harder session scores lower. Active sessions that make people think can feel like harder work, and get less glowing feedback, even when people learn more.

How sure are we?

The four levels are a way of organising evaluation, not a tested theory. They are widely used because they are simple and give people shared words. They have also been widely criticised (for example, Reio and colleagues, 2017).

The main criticism is that the levels don’t lead to each other automatically. A 1997 review that pooled many studies (Alliger and colleagues) found that how much people enjoyed training was barely linked to how much they learned or changed. Ratings of how useful it was did a little better, but were still weak signals. Level 4 is hard to pin on any one session, because too much else changes at the same time.

There is also evidence that liking can point the wrong way. In a study of university physics classes, students in active classes learned more, but felt they had learned less than students given polished lectures (Deslauriers and colleagues, 2019).

A common misreading is that every session needs Level 4 evidence. Use the level that answers the question you care about.

Where it comes from

Donald Kirkpatrick, an American professor, began working on how to evaluate training in his 1954 PhD. He set out the four levels (he called them “steps”) in four articles in 1959 and 1960, and in a 1994 book. Later, James and Wendy Kirkpatrick developed the New World Kirkpatrick Model (set out in their 2016 book), which adds what the workplace must do to support change. They call these “required drivers”.

The evidence

Kirkpatrick, D. L. (1959–1960). Techniques for evaluating training programs (parts 1–4). Journal of the American Society of Training Directors, 13(11)–14(2).
Kirkpatrick, D. L., & Kirkpatrick, J. D. (2006). Evaluating Training Programs: The Four Levels (3rd ed.). Berrett-Koehler.
Kirkpatrick, J. D., & Kirkpatrick, W. K. (2016). Kirkpatrick's Four Levels of Training Evaluation. ATD Press.
Alliger, G. M., Tannenbaum, S. I., Bennett, W., Traver, H., & Shotland, A. (1997). A meta-analysis of the relations among training criteria. Personnel Psychology, 50(2), 341–358.
Deslauriers, L., McCarty, L. S., Miller, K., Callaghan, K., & Kestin, G. (2019). Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom. Proceedings of the National Academy of Sciences, 116(39), 19251–19257.
Reio, T. G., Rocco, T. S., Smith, D. H., & Chang, E. (2017). A critique of Kirkpatrick's evaluation model. New Horizons in Adult Education and Human Resource Development, 29(2), 35–53.