Every year, organisations spend significant sums on training - leadership programmes, technical upskilling, compliance training, management development. And every year, when asked what the return on that investment has been, most training functions produce the same answer: participation numbers, satisfaction scores, and hours logged. These are not measures of value. They are measures of activity.
Why Smile Sheets Persist
The post-training satisfaction survey - commonly called the "smile sheet" in the learning profession - has been the dominant measure of training effectiveness for decades, despite persistent evidence that satisfaction has almost no correlation with learning transfer or business impact. The reason it persists is straightforward: it is easy to administer, and it consistently produces scores that make training programmes look successful.
Participants who have attended a well-facilitated, engaging programme with good content almost always rate it positively. The problem is that positive satisfaction ratings do not predict whether participants will apply what they learned, whether their behaviour will change, or whether that change will produce any measurable business outcome. Organisations that optimise their training programmes for satisfaction scores may be building very enjoyable learning experiences with very limited business impact.
The Kirkpatrick Model: Applied Seriously
The Kirkpatrick model - developed by Donald Kirkpatrick in the 1950s and refined since - provides a four-level framework for evaluating training effectiveness that, applied rigorously, changes what organisations know about their learning investments.
Level 1: Reaction. Did participants find the training relevant, engaging, and well-facilitated? This is the smile sheet level. It is worth measuring but should not be confused with effectiveness. The Kirkpatrick model explicitly treats Level 1 as the lowest level of evidence about programme value.
Level 2: Learning. Did participants actually acquire the knowledge, skills, or attitudes the programme intended to develop? Level 2 evaluation requires pre- and post-assessments that measure what was learned - not whether participants enjoyed the programme. Knowledge assessments, skills demonstrations, and attitude surveys before and after the programme provide evidence of whether learning actually occurred.
Level 3: Behaviour. Did participants apply what they learned on the job? This is where most training measurement falls short. Behaviour change requires follow-up assessment weeks or months after the programme, involving not just the participants but their managers and peers. Did the sales manager who attended the negotiation programme actually negotiate differently? Did the finance team member who completed the financial modelling course produce better models? Level 3 evaluation requires investment in follow-up processes that most organisations have not built.
Level 4: Results. Did the behaviour change produce measurable business outcomes? This is the level that justifies training investment to business leaders - and the level that training functions most consistently fail to reach. The challenge is attribution: business outcomes are produced by many factors, and isolating the contribution of a training programme requires experimental rigour that is difficult to achieve in operational environments. But the difficulty of rigorous attribution does not excuse the absence of any attempt to connect training investment to business results.
Making Kirkpatrick Practical
Applying Kirkpatrick rigorously to every training programme is not practical and not necessary. A sensible approach prioritises evaluation investment in proportion to training investment and business criticality.
For high-investment, strategically important programmes - leadership development for senior managers, technical upskilling linked to a business transformation, capability building for a new market entry - the investment in Level 3 and Level 4 evaluation is justified and should be built into the programme design from the outset. The evaluation methodology should be designed before the programme is delivered, not added retrospectively.
For volume training - compliance programmes, onboarding, product knowledge refreshers - Level 1 and Level 2 evaluation is typically sufficient, with periodic spot-checking at Level 3 to ensure behaviour transfer is occurring.
The discipline that drives improvement is not the evaluation itself but the accountability structure it creates. When training programme designers know their programmes will be evaluated at Level 3 and Level 4, they design them differently - with application exercises, job aids, manager involvement, and follow-up reinforcement that are absent from programmes designed only to produce good satisfaction scores.
The Training Needs Analysis Connection
Training ROI measurement is most effective when it is connected to a robust Training Needs Analysis that defined the capability gap and the business problem the training was intended to address. A programme designed to close a specific, measurable capability gap can be evaluated against whether it closed that gap. A programme designed to "develop leadership capability" in a general sense cannot be meaningfully evaluated at Level 3 or Level 4 because success was never defined specifically enough to measure.
Organisations that invest in rigorous TNA before designing training programmes find that ROI measurement is substantially easier - because the TNA defined what success looks like before the training began.