KS Insight
Home/Program Evaluation

Program
Evaluation

We measure whether our programs change how people lead, with the same discipline we bring to assessing leaders.

How we measure

We hold our own programs to the same standard

Participants rate themselves on a set of concrete leadership behaviors before a program begins and again after it ends. Each behavior is rated twice: how often they do it, and how well it works when they do. That pairing separates capability from deployment.

At the start, each participant names the thing they came to fix. At the end, we test whether they moved on it. Where the client has team-rated feedback, we check the self-ratings against how each leader’s team actually experiences them. The client receives an impact report covering delivery, engagement, and skill movement.

We have run this evaluation work across our programs, including a graduate leadership course at Columbia SIPA, adaptive leadership cohorts at the Broad Institute, and the Manager to Leader programs at KAUST.

Learnings

What the data has taught us

Capability moves before behavior does.

In a recent senior cohort, self-rated effectiveness rose on every behavior we measured while frequency barely moved. That is the right shape for a skills program: people get better at the hard conversation before they seek it out more often. When frequency stays flat for too long, the constraint is usually the organization, and we say so in the report.

Averages hide the movement that matters.

Rating each behavior twice lets us track individuals between four states, from doing the hard thing often but badly to doing it often and well. In one cohort, the group that was struggling most nearly emptied by the end of the program. That change was invisible in the mean scores.

A falling score can be good news.

Some participants rate themselves lower after a program than before. When we read their reflections, they are usually the ones who engaged most deeply: the program raised their standard for what good looks like. We read every self-rating next to the participant's own reflections before we count it.

Consequences teach what feedback cannot.

In our simulations, the strongest learning we have documented came from teams who made a bad call, watched it cost them the trust of the people around them, owned it, and rebuilt. The behavioral change across the following weeks was visible in their work. No debrief produces that on its own.

Asking for your worst move beats asking what went well.

The single sharpest reflection instrument we have used asks participants to name their worst leadership move and reason through what it cost. The negative framing forces honest, risk-aware thinking.

The same shift keeps appearing.

Across cohorts, from graduate students to senior scientists, the end-of-program reflections describe the same movement: from the leader as the person with the answers to the leader as the person who manages uncertainty and enables others. The most common self-diagnosed growth edge is acting too slowly when the situation is unclear.

Where this is going

The next question is whether the team feels the difference

The next step we are building is a reporting cycle that pairs each leader’s self-assessment with their team’s rating of them, before and after a program. That shows whether the team feels the difference.

Talk to us about measuring your program

If you are commissioning leadership development and want evidence that it worked, book a call. We can walk you through the instruments and a sample impact report.