Excerpt
Development donors increasingly demand exhaustive data collection, mistaking volume for rigor and full coverage for credibility. But abandoning statistical sampling in favor of mass data collection produces unreliable, meaningless numbers, and AI will only accelerate this problem unless the sector relearns the value of well-designed samples. The fix isn’t more technology. It’s statistical literacy and the confidence to push back.
The Quiet Numbers Crisis
There is a quiet crisis running through international development, and it has nothing to do with a lack of data. It has to do with the wrong kind of data, collected under impossible expectations, interpreted by people who have forgotten what statistics are actually for.
A Familiar Scenario
Consider this scenario. A major donor funds a sanitation programme across several cities in West Africa. The results framework is ambitious. The donor wants data on every participant, every household, every transaction.
Sound familiar?
The project team does what they can despite already being stretched extremely thin. They collect what is possible and report what sounds plausible. But somewhere between the project sites and the boardroom, the numbers stop meaning anything.
This is not a hypothetical. It is happening now, across programmes and geographies, driven by a growing disconnect between what funders demand and what is actually measurable in the real world.
How the Sector Lost Sight of Sampling
The core problem is that the sector has drifted away from the logic of sampling and now seems to be in the business of collecting data for data’s sake. For decades, rigorous development practice relied on well-designed representative samples. We look at smaller groups, carefully selected, that could tell you something meaningful about a larger population. This approach is statistically sound, practically achievable, and used routinely in research environments far more sophisticated than most development programmes.
But somewhere along the way, funders began to confuse the scale of data with the quality of insight. The assumption took hold that more data is better data and that full coverage is more credible than a sample. The aesthetic appeal of a well-designed dashboard with thousands of data points tells you more than a carefully analysed survey of one hundred people.
More Data, Less Insight
It does not. In fact, the opposite is often true. When you abandon sampling logic in favour of mass data collection, you typically end up with large volumes of inconsistent, unverifiable information that tells you very little about what is actually happening on the ground. You lose the ability to draw clean conclusions and lose perspective of what the real impact and human story is behind the numbers.
There is also a human and waste cost. Teams in the field spend enormous time collecting data that will never meaningfully inform decisions. In this case, local staff are diverted from implementation, monitoring budgets often balloon. And the credibility of the entire evidence base erodes because everyone knows the numbers are not quite real or meaningful—but no one wants to say so in a donor report.
Why AI Will Make This Worse Before It Makes It Better
Artificial intelligence is about to make this worse before it makes it better and we see it already happening. A quick shift again from qualitative analysis back to data-driven reporting demands. As AI-assisted data collection and analysis tools enter the sector, there will be enormous pressure to generate even more data, faster, at lower cost. This could be genuinely transformative if paired with statistical rigour and a comprehensive knowledge of statistics and the new technology tools paired with a nuanced understanding of the socio-cultural stories behind the numbers to create a complete picture of impact. But if the underlying logic remains broken—if funders still equate volume with validity—AI will simply automate the production of impressive-looking nonsense.
The Fix is Conceptual, Not Technical
The fix is not technical. It is conceptual. Donors need to invest in their own statistical literacy. Implementing organisations need the confidence to push back. And the sector needs to agree, collectively, that a well-designed sample is not a compromise. It is the gold standard.
The numbers exist and we can find information in them. We just need to stop asking for the wrong ones.






Comments are closed