Summary. Every selection programme publishes admissions, cohorts and graduates. Almost none can say what those people built five years later, and the standard fix, a bigger number, is half an answer at most. Two kinds of evidence are needed, each with its own standards. My own programme is audited here on the same terms.
#The number nobody asked for
Applications received. Participants selected. Cohorts run. Graduates certified. Those figures are true, cheap to produce, and answer a question nobody was asking.
The question people are asking is what happened afterwards, how, and whether it would have happened anyway. Two failures stack here, and most people who notice the first walk into the second. The first is reporting outputs and calling them impact. The second is the correction: a longer survey, a bigger sample, a multiplier. That assumes the problem was insufficient quantification. Often it was not.
I went through the YALI Regional Leadership Centre for West Africa in 2018. Good training, serious cohort, reporting that stopped at the graduation ceremony. Substitute Chevening, Erasmus Mundus or any corporate high-potential scheme and the shape is identical.
#What an agreement percentage actually is
The US State Department evaluated the Mandela Washington Fellowship across its 2014 to 2018 cohorts. Ninety-six per cent of Fellows "believe participation in the Fellowship helped with achieving professional and personal goals".
The usual criticism is that this is self-reported and therefore soft. That misses the defect. Somebody had a real experience, with a cause and a turning point. They were handed a scale. Everything specific was discarded, what survived was a tick, and the ticks were counted and given a decimal point.
That is a story with the story removed and a percentage sign attached to the residue. It has the unfalsifiability of a testimonial and borrows the authority of a statistic. The tell is that no version of the world produces a low figure: people who disliked the programme did not respond. Ninety-six per cent was guaranteed before the first form went out.
An output is a number you control. An outcome is a number that can go badly. Programmes report the first and describe the second in adjectives.
#Stories have standards too
The useful split is not hard versus soft. It is between evidence answering how many and evidence answering what happened, and by what route. Both can be done rigorously. Both are usually done badly.
Qualitative rigour is well developed and almost never applied by the programmes leaning hardest on stories: purposive sampling on a stated logic, negative cases deliberately sought, interviewing until themes stop appearing rather than until you hit a count, verbatim records, coding a second person could repeat, and taking the finding back to the people it came from.
A testimonial wall meets none of those. It is qualitative research's clothing worn by a marketing exercise, and the difference is entirely in the sampling.
#What no number reaches
Sen's distinction helps. A functioning is what someone does: income earned, job held. A capability is what they could do. Programmes measure functionings, because functionings fit in a column, then claim to have changed capabilities.
A participant is quiet in meetings. Someone takes them aside early and explains why contributing something, however small, is not optional. It sticks for a decade. "The programme built confidence" would sit identically under any programme on earth. The specific version is something another programme could act on tomorrow, and no instrument produces it, because the first thing every instrument does is discard the behaviour and keep the adjective.
#Numbers cannot be read without accounts
My own programme graduates on verified income before day ninety. Say it comes back at sixty-two per cent. What do you do with it?
Nothing, alone. You cannot tell what separated the sixty-two from the thirty-eight, whether the constraint was skill, clients, electricity or nerve, or which part of the ninety days did the work. So you cannot tell what to cut when money is short, and cannot tell anyone else how to reproduce it. A programme with metrics and no mechanism knows that it works and has no idea why.
It runs both ways. The accounts alone would let me believe anything; the number is what stops the stories flattering me.
#Counting deforms what it counts
Campbell, 1979: the more an indicator is used for decision making, the more it distorts the process it was meant to monitor.
Earn before day ninety is a target and I chose it. A fast, low-value gig satisfies it. Someone who should spend six months on a durable skill can be steered toward whatever earns soonest. Selecting people already close to earning improves the rate without the programme doing anything. No cheating required, only that people respond to what is counted.
The defence is not a better metric. It is a second stream the target cannot reach, whose job is to catch the metric lying.
#Where my own programme fails
Cycle 28 publishes an alumni page with no denominator: trajectories with no cohort total and no intake year are an illustration, not a retention figure.
The qualitative side is worse, and more embarrassing given the argument above. The sampling is self-selection by people who felt warm enough to write in. No negative cases, because nobody who left in week three is on that page and I have not asked them. Nothing coded. Nobody has checked the summary back with participants. That is a marketing page where there should be a study, and I built it.
#Do not multiply
The tempting move is a coefficient: three dependants each, so a hundred graduates becomes four hundred lives changed. It is applied to everyone who graduated rather than everyone still earning, double counts shared households, and cannot come out lower, because the three was assumed.
Worse, it deletes the best material you own. What happens in a household when someone starts earning is a sequence: which bill got paid first, who went back to school, what happened when the income stopped. Ask "how many people depend on your income" at intake and again at follow-up, and collect six household accounts properly, two where the earning did not hold.
#The ask
Decide both evidence streams before selecting the first cohort; afterwards it is archaeology. Publish the denominator and the sampling logic in the same breath as the claim. And keep something the target cannot reach, because whatever you turn into a target, you will eventually deform.
Written as a 2018 YALI participant, and as the founder of a programme whose own published evidence has neither a denominator nor a sampling logic attached to it.
#References
- Evaluation of the Mandela Washington Fellowship (2014-2018), August 2020, US Department of State. Source of the 96 per cent. https://2012-2025.eca.state.gov/files/bureau/05b_mandela_washington_fellowship_infographic_report_vf.pdf
- Applying Evaluation Criteria Thoughtfully. OECD Publishing, 2021. https://www.oecd.org/content/dam/oecd/en/publications/reports/2021/03/applying-evaluation-criteria-thoughtfully_45a54ea7/543e84ed-en.pdf
- McKenzie, D. and Woodruff, C. World Bank Research Observer, 29(1), 2014. doi:10.1093/wbro/lkt007
- Guba, E. G. and Lincoln, Y. S. Naturalistic Inquiry. Sage, 1985.
- Sen, A. Development as Freedom. Oxford University Press, 1999.
- Campbell, D. T. "Assessing the impact of planned social change." Evaluation and Program Planning, 2(1), 1979.
- Longitudinal Education Outcomes, ADR UK, on administrative linkage. https://www.adruk.org/data-access/flagship-datasets/longitudinal-education-outcomes/