The end-of-programme survey is one of the most trusted instruments in leadership development. It is fast, cheap, easy to administer and produces a clean number that can travel up the organisation to whoever commissioned the work. It carries an implied claim: participants who report a positive experience have learned something of value.
Evaluation research building on Kirkpatrick's framework distinguishes participant reaction, learning, behaviour and organisational results. [1]
Donald Kirkpatrick set out his four-level model in 1959. It distinguishes reaction, learning, behaviour and results. [1] Reaction, the immediate response of participants, sits at level one. Learning, the acquisition of knowledge or skill, sits at level two. Behaviour, the transfer of that learning into daily practice, sits at level three. Results, the organisational outcomes that follow, sit at level four.
Post-programme surveys are level one instruments. They were built to capture reaction, and they capture it well. A participant leaving a programme knows whether the room was warm, whether the facilitator held their attention, whether the material felt relevant, whether they enjoyed the two days. These are legitimate things to measure.
An exit survey alone cannot establish level-two learning or level-three behaviour change. It can record participants' immediate reactions and their perceptions of learning. Evidence of retention and transfer requires later assessment. [1] The learning that a programme is intended to produce forms slowly, becomes visible in practice, and reveals itself in the moments when a leader chooses differently under pressure. That evidence sits well beyond the closing session.
Wisdom research has examined the reflective capacities associated with mature judgement. Coaching research has examined the practices that sustain goal-directed behaviour change over time. Monika Ardelt's three-dimensional wisdom scale comprises cognitive, reflective and affective dimensions of wisdom. Its reflective dimension concerns the capacity to examine experience from multiple perspectives. [2] In a ten-week study of life coaching, Anthony Grant found significant improvements in participants' goal attainment and metacognition. [3]
Both bodies of work point in the same direction. The reflective and behavioural changes a leadership programme aims to produce form through a feedback loop that runs well beyond the closing session. Immediate satisfaction data sits outside that loop, and any evaluation architecture that wants to speak to development has to reach into the weeks and months that follow.
Schedule outcome assessment around the opportunities participants have to apply the intended behaviour. Evaluation guidance distinguishes short-, intermediate- and long-term outcomes, with each measured at a point when the relevant change can be observed. [4] A ninety-day review can examine early application. A six-month review can examine whether the practice has become established.
An immediate post-programme survey can capture participants' self-reported reactions to the experience, including perceived relevance, satisfaction and views of the facilitation. These measures inform programme design and delivery. [1]
These are useful to the design and delivery team, and together they describe the quality of the experience the participant has just had. The commissioning organisation still needs separate evidence to answer the learning question, the behaviour question and the return question. High satisfaction scores are compatible with programmes that produced deep behavioural change. They are equally compatible with programmes that produced none. The instrument cannot distinguish between the two.
Organisations that want to know what their programmes produce can build an assessment architecture that matches the claim they are trying to test. Four practices carry most of the weight.
Assess reaction at the point of exit, and treat that data as reaction data. It answers a useful question about the participant experience and it earns its place in the evaluation. It does not answer the learning question.
Assess learning at a delay of six to twelve weeks, once the material has been tested against real work. Ask participants to describe a specific situation in which the programme content shaped a decision, and what the outcome was. Narrative evidence gathered through a consistent prompt can show how participants applied programme content in a defined work situation. Code responses against the programme's intended behaviours and review them alongside manager and peer observations. Establish a pre-programme baseline for the same behaviours, so the evaluation can examine change over time. [5]
Assess behaviour at ninety days and again at six months, drawing on the perspective of colleagues, direct reports and clients where appropriate. Colleagues, direct reports and clients can contribute evidence of behaviour change because they experience the leader's practice in live working relationships. Set the purpose, respondents, access arrangements and feedback process before collecting multi-rater evidence. Participants should know how observations will be used in the evaluation and development process.
Assess results against the specific organisational outcome the programme was commissioned to influence, and accept that the causal chain is long. Honest evaluation names the uncertainty in that chain and sets out the evidence at each link.
Organisations spend significant sums on leadership development, and the people receiving that development invest their attention and their trust. A measurement system earns trust when it can identify participant experience, learning in practice, behaviour change and organisational results with appropriate evidence. Assessment built to that standard gives organisations a defensible account of participant experience, learning in practice, behaviour change and organisational results.
Bibliography
[1] Bates, Reid. "A Critical Analysis of Evaluation Practice: The Kirkpatrick Model and the Principle of Beneficence." Evaluation and Program Planning, 2004.
[2] Ardelt, Monika. "Empirical Assessment of a Three-Dimensional Wisdom Scale." Research on Aging, 2003.
[3] Grant, Anthony M. "The Impact of Life Coaching on Goal Attainment, Metacognition and Mental Health." Social Behavior and Personality, 2003.
[4] W. K. Kellogg Foundation. "Logic Model Development Guide." W. K. Kellogg Foundation, 2004.
[5] Shadish, William R., Thomas D. Cook, and Donald T. Campbell. Experimental and Quasi-Experimental Designs for Generalized Causal Inference. Houghton Mifflin, 2002.