Insights Whitepaper
Whitepaper

Counting graduates is not measuring leaders

Adeolu Timothy 19 August 202519 Aug 2025 20 min read Measurement · Leadership Development · Evaluation · Qualitative Research · Governance
Summary

Selection and training programmes report how many people they admitted, trained and graduated. Almost none can say what those people went on to build, and the ones that try mostly reach for a bigger number. The opposite of a bad number is not a better number. Some of what these programmes do is only reachable through accounts and mechanisms, and qualitative evidence has standards of its own that testimonial walls fail as badly as graduate counts fail statistical ones. What to decide at intake for both, why an agreement percentage fails both at once, and why the dependants multiplier deletes the best evidence a programme owns.

Summary. Every selection programme publishes admissions, cohorts and graduates. Almost none can say what those people built five years later, and the standard fix, a bigger number, is half an answer at most. This sets out the two kinds of evidence these programmes need, the standards each has to meet, why an agreement percentage fails both at once, and why counting the wrong thing damages the thing being counted. My own programme is audited here on the same terms.

#The number nobody asked for

Every selection programme on earth publishes the same figures. Applications received. Participants selected. Cohorts run. Graduates certified. Countries represented. Those numbers are true, they are cheap to produce, and they answer a question nobody was asking.

The question people are asking is what happened afterwards. Not what participants thought of the training while it was fresh, but what they built, how it happened, and whether it would have happened anyway.

Two failures are stacked here, and almost everyone who notices the first walks into the second. The first is reporting outputs and calling them impact. The second is the correction: a bigger number, a longer survey, a multiplier. That correction assumes the problem was insufficient quantification. Often it was not. Some of what a good programme does to a person is not a small effect that better instruments would detect. It is a different kind of thing, and it needs a different kind of evidence, held to its own standard rather than apologised for as soft.

#The evidence I have

In 2018 I was selected for the Young African Leaders Initiative Regional Leadership Centre for West Africa, hosted at the Ghana Institute of Management and Public Administration in Accra. The four regional centres were announced in 2014, in Ghana, Kenya, Senegal and South Africa; the West Africa centre serves nine countries, and the model has graduated more than thirty thousand people.

The training was good and the cohort was serious. What I noticed then, and have kept noticing, is that the reporting stopped at the graduation ceremony. I know how many of us were in the room. I do not know how many are still running the businesses we described in our applications, how many entered public service, or how many civic initiatives outlived their founders' enthusiasm. I also do not know the more interesting thing: what it was about those weeks that changed the people it changed, and why it slid off the rest of us. Neither question was asked in a form that could be answered later.

Substitute Chevening, Erasmus Mundus, DAAD, Commonwealth Scholarships, a national graduate scheme or a corporate high-potential programme and the shape of the gap is identical.

#Output, outcome, impact

The vocabulary has existed for decades. The OECD Development Assistance Committee distinguishes outputs, the products of an activity, from outcomes, their effects, from impact, the higher-order and longer-term effects. The criteria were revised in 2018-19, treating effectiveness as achievement of stated objectives and reserving impact for the "so what" question evaluations routinely skip.

Training a thousand people is an output, entirely within the programme's control and therefore always achieved. What those thousand do afterwards is an outcome, only partly attributable, and therefore the thing nobody wants to be measured on.

An output is a number you control. An outcome is a number that can go badly. Programmes report the first and describe the second in adjectives.

Note what the framework does not say: nothing in it requires impact to arrive as a quantity. Some consequences are countable, some are only describable, and the vocabulary is silent on which is which. Most programmes read the silence as permission to count.

#What an agreement percentage actually is

This is not a case of nobody trying. The US State Department's Bureau of Educational and Cultural Affairs evaluated the Mandela Washington Fellowship across the 2014 to 2018 cohorts, published August 2020, surveying alumni, institute staff, host organisations and US community members. It is a serious piece of work.

Read the headline figures. Ninety-six per cent of Fellows "believe participation in the Fellowship helped with achieving professional and personal goals". Eighty-one per cent "agreed that the Fellowship has given them valuable job skills to use back home". Sixty-seven per cent "feel that the Fellowship helped in using the new skills/knowledge to manage a business or organization or team".

The usual criticism is that these are self-reported and therefore soft. That misses the actual defect. Look at what happened to the underlying material. Somebody had a real experience, with a shape and a cause and a turning point. They were handed a scale. Everything specific was discarded and what survived was a tick. Then the ticks were counted and given a decimal point.

That is a story with the story removed and a percentage sign attached to the residue. It inherits the unfalsifiability of a testimonial, since nobody can disagree with what an alumnus believes, and borrows the authority of a statistic, since it has a number in it. It is not a weak quantitative finding. It is qualitative data destroyed in transit, then reported as though the destruction were a method.

The tell is that no version of the world produces a low figure. People who disliked the programme did not respond; people who responded like it. Ninety-six per cent was structurally guaranteed before the first form went out, which is why a longer scale or a bigger sample would not help.

The quantitative side fails in parallel. McKenzie and Woodruff's review of business training evaluations across the developing world found studies commonly suffer small samples, measure effects within a year, and lose participants to attrition. Over those windows the effects were modest: owners adopted some taught practices in small magnitudes, and few studies detected significant effects on profits or sales. A programme that surveys at week twelve and never again guarantees itself a flattering result on one side and a null on the other.

#Two kinds of evidence, two sets of standards

The useful distinction is not hard versus soft. It is between evidence answering how many and how much and evidence answering what happened, to whom, by what route. Both are answerable rigorously. Both are usually done badly.

Quantitative rigour is familiar enough that its absence embarrasses: a denominator, a response rate, a comparison group, a pre-specified outcome, attrition reported alongside the result.

Qualitative rigour is equally well developed and almost never applied by the programmes leaning hardest on stories. Guba and Lincoln set out credibility, transferability, dependability and confirmability as the qualitative analogues of validity and reliability. For a programme that means:

  • Purposive sampling, stated. You choose whose account you collect on a declared logic: maximum variation, typical case, extreme case. What you do not do is let the sample select itself by asking who wants to write in.
  • Negative cases, deliberately sought. The person who left in week three. If no account makes the programme look bad, you have not done research, you have done recruitment.
  • Saturation, not a target count. You interview until new accounts stop producing new themes. Fifteen is not a number to hit, it is a stopping condition you reached or did not.
  • Verbatim, not paraphrase. The words are the data. Do not tidy the awkward ones; those are usually the true ones.
  • Analysis someone else could repeat. Code the accounts, define the themes, have a second person code a subset. Themes only you can see are yours, not the cohort's.
  • Member checking. Take the finding back to the people it came from and ask whether they recognise themselves in it.

Packaged methods exist for exactly this: Most Significant Change, which puts the selection of which accounts matter through documented deliberation rather than letting someone pick the best quotes; Outcome Harvesting, which starts from observed outcomes and works back to contribution; contribution analysis, which tests a theory of change against rival explanations; realist evaluation, which asks what works, for whom, in what circumstances.

None of that is softer than a regression. The failure mode of nearly every leadership programme is to skip the discipline and keep the vocabulary.

A testimonial wall is not qualitative research. It is qualitative research's clothing worn by a marketing exercise, and the difference is entirely in the sampling.

#What no number reaches

Some of what these programmes do is not a small effect awaiting a better instrument. It is not a quantity.

Sen's capability approach makes the distinction usable. A functioning is what a person actually does: income earned, job held, firm registered. A capability is the set of things they could do, the real freedom available. Two people with identical incomes can hold very different capability sets, and the capability set governs what happens when circumstances change. Programmes measure functionings, because functionings fit in a column, then claim to have changed capabilities.

Take a recognisable example. A participant is quiet in team meetings. Someone takes them aside early, privately, and explains why contributing something however small is not optional. The habit sticks and travels into every organisation they work in for a decade.

"The programme built confidence" is so thin it would sit identically under any programme on earth. One-to-one correction on a single concrete behaviour, delivered privately and early, transferring to later workplaces, is something another programme could act on tomorrow. Only the second is usable, and no instrument produces it, because the first thing every instrument does is discard the behaviour, the timing and the delivery and keep the adjective. Precision here comes from thickness of description, not decimal places.

Scott's argument about legibility is the wider version. Institutions simplify what they administer into countable categories, and the simplification becomes the official reality. What did not fit stops being administered, then stops being noticed, then stops being funded. A programme that can only report what it counts will, within a few budget cycles, only do what it can count.

#Numbers cannot be read without accounts

There is a practical argument that does not depend on valuing stories intrinsically.

The programme I run graduates on one condition: income earned from the trained skill, verified by an outside party, before day ninety. Set aside whether that is a good rule and take it as a metric with the properties evaluators ask for, since it is externally verified, defined before intake, and cannot be satisfied by attendance.

Suppose the number comes back at sixty-two per cent. What do you do with it?

Nothing, on its own. You cannot tell whether sixty-two is good, what separated the sixty-two from the thirty-eight, or whether the binding constraint was skill, clients, equipment, electricity, household pressure or nerve. You cannot tell which part of the ninety days did the work, so you cannot tell what to cut when money is short, and you cannot tell anyone else how to reproduce it.

Every one of those is answerable only through accounts. The number tells you whether to keep going; the accounts tell you what you are actually doing, which is the only thing that lets you do it again elsewhere. A programme with metrics and no mechanism knows that it works and has no idea why, so it cannot survive its founder, cannot be franchised, and cannot be defended when a funder asks what specifically they are buying.

It holds the other way too. The accounts alone would let me believe anything I wanted; the number is what stops the stories being flattering. Each is load-bearing for the other.

#Counting deforms what it counts

The last argument for keeping some things qualitative is protective.

Campbell put it plainly in 1979: the more any quantitative social indicator is used for social decision making, the more it will be subject to corruption pressures and the more apt it will be to distort and corrupt the processes it was intended to monitor. Strathern's compression is the version people quote: when a measure becomes a target, it ceases to be a good measure.

This is not abstract for me. Earn before day ninety is a target and I chose it. Consider what it rewards under pressure. A fast, low-value gig satisfies it. A participant who should spend six months building a durable skill can be steered toward whatever earns soonest, because that is what closes the cohort. Selecting people already close to earning improves the rate without the programme doing anything. None of that requires anyone to cheat, only that people respond to what is counted, which they always do.

The defence is not a better metric, because the same logic eats that one too. The defence is a second stream of evidence the target cannot reach: what the work was, whether it lasted, whether the person could do it again unassisted, what they turned down. Those are questions for accounts. Their function in the system is to catch the metric lying.

Anything you convert into a target, you deform. Keeping the mechanism in prose keeps it honest, because prose is much harder to optimise for.

#Where my own programme fails this test

It would be cheap to run this argument only against other people's programmes, so here is mine on the same terms. Cycle 28 publishes an alumni page, and it fails in both columns.

Quantitatively there is no denominator. Trajectories with no cohort total and no intake year are an illustration, not a retention figure, and the total is the second question any serious funder asks.

The qualitative side is the real indictment, more embarrassing precisely because this piece argues that is where the useful knowledge lives. Against the standards above the page meets none. The sampling is self-selection by people who felt warmly enough to write in. There are no negative cases, because nobody who left in week three is on that page and I have not asked them. No stopping rule was declared, so the number of accounts is the number that arrived. Nothing was coded, so any theme I see is my impression, not a finding. Nobody has taken a summary back to participants to check they recognise themselves in it.

That is a marketing page where there should be a study, and I built it. The fix is not more stories, and certainly not better-written ones. It is roughly twelve interviews on a stated logic, a third with people for whom the programme did not work, kept verbatim and coded by two people. That is a week of work I have not done.

The failure mode is not incompetence. It is what happens by default when nobody decides the evidence design before the cohort starts, which is the whole argument.

#What to decide at intake

Name the outcome per track, and the mechanism you believe produces it. Two or three verifiable claims, not twenty. Business track: firm still trading at year five, employees on payroll, revenue band. Public policy: employment in a public institution, seniority band, named instrument worked on. Then write down, before the cohort starts, why you think the programme will produce that. The written mechanism is what accounts later confirm, complicate or destroy; a theory of change you never test is decoration.

Decide the qualitative sample at intake too. Who you will interview, on what logic, at what intervals, and how you will keep reaching the ones it did not work for. Retrospective qualitative sampling is not sampling, it is whoever is still fond of you.

Take consent at intake, not at year five. Long-term follow-up is a processing purpose and must be specified up front. GDPR Article 5 requires personal data to be collected for specified, explicit and legitimate purposes, limited to what is necessary, and kept in identifiable form no longer than needed, and comparable duties apply in most jurisdictions these programmes recruit from. In practice: a clause covering decennial follow-up, a named lawful basis, a retention rule and a withdrawal path, agreed while the participant is still reachable. Consent to be quoted is a separate question from consent to be counted.

Prefer records over recollection for facts, and recollection over records for mechanism. The UK's Longitudinal Education Outcomes dataset links education records to tax, benefits and self-assessment data, letting the Department for Education report employment and earnings three and five years out without a survey. Most programmes will not get tax records, but company registries, procurement records, professional bodies, published appointments and grant databases do not depend on an alumnus answering an email in 2035. Use them for what happened, never for why, because they cannot say.

Keep a comparison group, free, at selection. Applicants scoring just below the threshold are on average near identical to those just above. Comparing them years later isolates something much closer to the programme's contribution than any before-and-after chart. It costs only the discipline to retain scored applications and take the same consent from near-misses, and interviewing a few of them is the cheapest counterfactual account you will get.

Publish attrition and sampling with every figure and every quote. A number from a thirty per cent response rate where responders are disproportionately the successful is not a finding, and a quote with no statement of how it was chosen is not either. The LEO commentary is careful that the data is not a measure of institutional value added; state limits in the same breath as the claim.

#The multiplier temptation

Once you have a wall of alumni profiles the obvious move is to multiply. If each supports on average three dependants, a hundred graduates becomes four hundred lives changed and the deck gets a bigger number without anyone collecting anything. I have felt the pull, which is why it gets a section.

The instinct is right. Household effect is real and larger than the individual effect. The arithmetic is what fails:

  • Wrong population. Applied to everyone who graduated, not everyone still earning. If a third have stopped by month twelve, a third of the dependants are imaginary and nothing in the calculation knows.
  • Undefined term. School fees and a contribution to groceries both count as one household at the same weight, forever.
  • Double counting. Two graduates from one household bring six dependants to the total; the household has three.
  • No counterfactual. Most were being supported, precariously, before the programme existed. The claim is a change; the multiplier reports a level.
  • Unfalsifiable. No result can contradict it, because the three was assumed. A number that cannot come out lower is the graduate count in a more flattering unit.

The deeper objection is the one this piece has been building to. What happens in a household when someone starts earning is not a coefficient. It is a sequence: which bill got paid first, who went back to school, which relative stopped being asked for money, what happened when the income dropped again. That is thick, causal, highly variable, and the most persuasive material a programme of this kind owns. Multiplying by three does not summarise it. It deletes it and leaves a number no funder believes anyway.

The replacement runs in both columns and is cheap. Quantitatively, "how many people depend on the income you earn" is one field at intake and one at follow-up, giving a distribution and a change rather than a coefficient. Qualitatively, six household accounts collected properly, two of them where the earning did not hold, will do more in a funding conversation than any multiplied figure has. Shape first, figures illustrative:

Of 100 participants admitted, 62 reached verified income before day 90. At intake they reported a median of 3 people dependent on their earnings. At 12 months we reached 46 of the 62, of whom 38 were still earning from the skill. Everything after "at 12 months" rests on a 74 per cent response rate. Alongside it we hold 12 interviews sampled for maximum variation, 4 with participants who did not reach verified income, coded by two people.

Nobody can argue with that, which is the point. The multiplied version invites the first competent funder in the room to ask where the three came from, and you lose them over a number you never needed to invent.

#Where automated systems help, and where they do not

The temptation is to describe an AI system that measures leadership impact. That is not a thing, and claiming it is will discredit the underlying work. What machines do well here is narrow and real: record linkage and entity resolution, matching an alumna on a 2018 intake form to a company director in a registry and a grantee in a funder database; structuring free text into countable fields; proposing a first-pass coding frame across a hundred transcripts for a human who keeps the codebook; flagging improbable claims and contradictions for a phone call.

What such a system must not decide: whether the programme caused an outcome, whether a participant counts as a success, or what an unverified claim means. One prohibition belongs on the list precisely because it looks so useful: it may not score sentiment on participant accounts and report the average. That is the ninety-six per cent failure rebuilt with a model, at higher volume and better branding. Attribution is a research design question, not an inference task. The model classifies and flags; a human, and a stated method, judge. Write the prohibitions down before building, because the first request after the dashboard ships is always to let it rank people.

#When not to do this

A cohort of twenty will never produce a meaningful statistical comparison however carefully you retain near-misses. Run the consent clause and the record checks anyway, since they cost little and compound, but report a case series and say so. Twenty people described honestly beats twenty dressed up as a finding, and this is exactly where the qualitative arm should carry the argument.

A programme whose participants are at risk from being identified, in a repressive setting or a regulated profession, should collect less rather than more. Consent for decennial follow-up is not meaningful when the follow-up is itself the exposure, and a named quotation is a larger exposure than a row in a count.

A first-year programme has better uses for the effort than an evaluation framework. It needs only the six things below, which are the ones that cannot be added later.

#The minimum viable version

  1. Five structured fields at intake: a stable identifier, the cohort or intake year, two contact routes that survive a job change, and the number of people dependent on the participant's income.
  2. A consent clause covering ten years of follow-up, with a named basis, a retention rule, and separate permission to be quoted and named.
  3. Scored applications retained, including rejections near the threshold.
  4. One annual check against public records, not a questionnaire.
  5. Twelve interviews a year on a written sampling logic, at least a third with people it did not work for, kept verbatim.
  6. One number published each year with its denominator and response rate, and one finding from the interviews published with its sampling logic.

The last two are what make the other four mean anything.

#Key takeaways

  • Decide both evidence streams before selecting the first cohort. Afterwards it is archaeology, not measurement.
  • Treat an agreement percentage as a defect rather than a finding. Collect the account properly or collect a verifiable fact, but do not tick-and-total the difference.
  • Hold stories to qualitative standards instead of excusing them from all standards: stated sampling logic, negative cases sought, a stopping rule, verbatim records, analysis someone else could repeat.
  • Never multiply a count by an assumed coefficient. If the household effect matters enough for the deck, it matters enough to be a field at intake and a handful of properly sampled accounts.
  • Keep a stream of evidence your target cannot reach, because whatever you turn into a target you will eventually deform. Its job is to catch the metric lying.
  • Let machines link records, structure text and propose codes. Never let them score sentiment and report the average, or decide attribution, success, or what an unverified claim means.

#The ask

Training a thousand people is an output. Understanding what those thousand went on to build, and how, is impact. The distinction is old and the tooling is ordinary. What is missing is the decision, taken before selection rather than after the funding is cut, that you will be able to say both what happened and why, and that neither half will be allowed to stand in for the other.

Written from both sides: as a 2018 participant in the YALI Regional Leadership Centre for West Africa, and as the founder of a training programme whose own published evidence currently has neither a denominator nor a sampling logic attached to it.

#References

Books and journal articles are cited bibliographically; everything else links to source.

  1. "YALI Regional Leadership Center West Africa-Accra completes training of 108 young Africans in leadership skills." Ghana Institute of Management and Public Administration, 16 August 2023. Centre hosted at GIMPA, four centres announced July 2014, nine countries in the sub-region. https://gimpa.edu.gh/2023/08/16/yali-regional-leadership-center-west-africa-accra-completes-training-of-108-young-africans-in-leadership-skills/
  2. "Young African Leadership Initiative Legacy and Localization." Arizona State University, International Development Initiative. Four centres in Ghana, Kenya, Senegal and South Africa, more than 30,000 graduates. https://internationaldevelopment.asu.edu/projects/young-african-leadership-initiative-legacy-and-localization/
  3. Evaluation of the Mandela Washington Fellowship (2014-2018), infographic report, August 2020. Bureau of Educational and Cultural Affairs, US Department of State. Source of the 96, 81 and 67 per cent figures, quoted verbatim above. https://2012-2025.eca.state.gov/files/bureau/05b_mandela_washington_fellowship_infographic_report_vf.pdf
  4. Applying Evaluation Criteria Thoughtfully. OECD Publishing, 2021. The six DAC criteria; criteria adapted 2018-19 with definitions published December 2019. https://www.oecd.org/content/dam/oecd/en/publications/reports/2021/03/applying-evaluation-criteria-thoughtfully_45a54ea7/543e84ed-en.pdf
  5. McKenzie, D. and Woodruff, C. "What Are We Learning from Business Training and Entrepreneurship Evaluations around the Developing World?" The World Bank Research Observer, 29(1), 2014, pages 48-82. doi:10.1093/wbro/lkt007 Open working paper version: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2149390
  6. Guba, E. G. and Lincoln, Y. S. Naturalistic Inquiry. Sage, 1985. Credibility, transferability, dependability and confirmability.
  7. Davies, R. and Dart, J. The "Most Significant Change" Technique: A Guide to Its Use. 2005.
  8. Wilson-Grau, R. Outcome Harvesting: Principles, Steps, and Evaluation Applications. Information Age Publishing, 2018.
  9. Mayne, J. "Contribution analysis: An approach to exploring cause and effect." ILAC Brief 16, 2008.
  10. Pawson, R. and Tilley, N. Realistic Evaluation. Sage, 1997.
  11. Sen, A. Development as Freedom. Oxford University Press, 1999. Capabilities and functionings.
  12. Scott, J. C. Seeing Like a State. Yale University Press, 1998. Legibility and administrative simplification.
  13. Campbell, D. T. "Assessing the impact of planned social change." Evaluation and Program Planning, 2(1), 1979, pages 67-90.
  14. Strathern, M. "'Improving ratings': audit in the British University system." European Review, 5(3), 1997, pages 305-321. Source of the common formulation of Goodhart's law.
  15. Geertz, C. "Thick Description: Toward an Interpretive Theory of Culture." In The Interpretation of Cultures. Basic Books, 1973.
  16. "Graduate labour market outcomes (LEO): methodology." Department for Education, Explore Education Statistics. Linkage of National Pupil Database, HESA, further education records, HMRC and DWP data; stated limits on causal interpretation. https://explore-education-statistics.service.gov.uk/methodology/leo-graduate-and-postgraduate-outcomes
  17. "Longitudinal Education Outcomes." ADR UK. Scope of roughly 38 million individuals, 1995 to 2021, England only, and the accreditation required for access. https://www.adruk.org/data-access/flagship-datasets/longitudinal-education-outcomes/
  18. "A beginner's guide to Longitudinal Education Outcomes (LEO) data." Wonkhe. On LEO not being a measure of value added or a performance indicator. https://wonkhe.com/blogs/a-beginners-guide-to-longitudinal-education-outcomes-leo-data/
  19. Regulation (EU) 2016/679 (General Data Protection Regulation), Article 5. Principles relating to processing of personal data. https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:32016R0679
  20. Cycle 28. Programme structure and the earn-before-day-90 graduation condition, referenced here as the author's own worked example. https://cycle28.org/

Shipping something like this?

Everything above came out of building and running the thing, not researching it. If you are hitting the same problems, I work with product and engineering teams on exactly this, architecture reviews, technical strategy, and staying on through implementation.