Method behind Measuring Change

Impact Measurement

What a technology, service, or programme actually changes, for whom, and through what pathway - measured directly with the people it is meant to serve, and reported honestly.

Most writing about a new technology or programme arrives ahead of the evidence. Impact measurement closes that gap: it puts honest questions to the people the work is meant to serve, from the human side first, independently of the intervention, and reports what the evidence actually says, including when that is inconvenient.


What we measure - three dimensions of impact

Every study sits inside the same measurement frame, so a result in one context can be read against another and a one-off study becomes a trajectory over time. Three dimensions, each answering a different question.

Dimension 01

Reach - who is actually being reached, and who is being missed

The most flattering-looking programme can be reaching the wrong people, or the same people repeatedly, or leaving the hardest-to-reach unmoved. We measure coverage, inclusion, and, critically, the gap between the population an intervention was designed for and the one it is landing with. Distribution is central, not a footnote.

Dimension 02

Depth - how much things actually change

Beyond headline satisfaction: the magnitude of change in income, time, health, opportunity, or whatever the intervention is aiming at. And how meaningful the change is in the respondent's own terms, not in the language of the funder's logframe.

Dimension 03

Experience - what it is like to be on the receiving end

Satisfaction, friction, unintended effects, and the qualitative texture that numbers alone cannot carry. What would the user change first? What would they never give up? The answers to those questions are usually where the useful insight lives.

Reach, depth, and experience read together give a project a picture of impact that a self-report or an output count cannot: whether the intervention landed with the right people, whether it made a material difference, and whether the people on the receiving end would recommend or repeat it. See the in-depth interview guide for how the Lab reaches depth in the qualitative work, and the European impact-tracking template for the shape of a full baseline-to-endline design.


How strong is the evidence

Not all evidence carries the same weight. The Lab classifies every finding it reports against a five-tier pyramid, from what was delivered at the base to a causal contribution established at the top. The pyramid is the honest answer to the question every funder eventually asks, which is whether the number in the report is a claim, an observation, or a conclusion.

Evidence Strength Pyramid. From base to apex, five tiers: Activity Data (what was delivered or reached); Self-Reported Outcome (participants report perceived change); Observed Change (behaviour or condition has measurably changed); Triangulated Evidence (multiple sources confirm the change); Attributable Impact (causal contribution established). An arrow on the left runs from 'weaker evidence' at the base to 'stronger evidence' at the top.
Every finding the Lab reports carries an explicit place on this pyramid. Most claims made about programmes sit on the bottom two tiers and are reported as if they sat on the top two. Naming the tier is how the reporting stays honest.

Two working rules follow from the pyramid. Most claims made about programmes sit on the bottom two tiers, and most reports present them as if they sat on the top two, which is where the reporting-loop pressure we describe in Why funders learn from grantees shows up in practice. And moving a study up the pyramid is a design choice made before the fieldwork, not a rhetorical move made in the write-up: a baseline captured, a second source lined up, a control identified. What you can honestly claim at the end is set by what you built in at the start.


The kinds of impact we track

Impact is not one thing. On any given study, we typically measure a mix of the following, calibrated to the specific decision the research is meant to inform.

  • Outcomes, not outputs. What actually changed for the person, not what the project produced. "Ten workshops delivered" is an output; "the participants can now negotiate a fair price" is an outcome. Only the second one earns the word impact.
  • Quality-of-life change. Concrete shifts in income, time, health, safety, learning, autonomy - measured in the units respondents actually use.
  • Distributional effects. Who was reached, who was excluded, and whether inequalities inside the target population widened or narrowed. Aggregate figures routinely hide this.
  • Unintended effects. The consequences a project didn't plan for, in either direction. Some are the most important findings a study produces.
  • Customer experience. How the intervention feels to use, the friction points that predict whether it will be sustained, and the specific improvements respondents ask for.
  • Attribution and additionality. Whether the change happened because of the intervention or alongside it - the distinction that separates evaluation from press release.
  • Sustainability. Whether the benefit persists once the project ends, or fades. Only measurable by revisiting.

How we do it

We start from the decision, not the data. A short design conversation establishes what decision the research must inform, the population that holds the answer, and the specific evidence that would genuinely change the decision. Most research is wasted here, by gathering data that was never going to settle anything. See How It Works for the five-stage method the Lab runs on every engagement.

We build the instrument around the respondent. Short, clear, locally grounded, and structured for repetition and comparison. We use validated instruments where they exist and design bespoke ones where they do not. Every instrument is built for repeat use, so a single study becomes a time series with almost no marginal cost.

We capture a baseline before the intervention. Impact is a change, and a change needs a before. Baselines are captured with the same instrument that will be used again later, so the comparison is valid. Without this, an impact claim can only be asserted, not evidenced.

We reach respondents directly, in-language. Trained local researchers who know the place gather data by phone or in person, with informed, recorded, revocable consent. We reach the people the intervention is meant to serve, not proxies for them. See Field Research for the field methods behind this.

We report the inconvenient finding. Independence is the whole point. We publish what the evidence says, including the parts that contradict a client's hopes. That posture is what makes the finding worth the effort of gathering it.


What you receive

Every engagement ends in something you can use, in the following forms:

  • Clean, well-documented data, ready to be re-analysed or re-run.
  • A written analysis in plain language, stating what the evidence shows and what it cannot conclude.
  • The instruments themselves - surveys, interview guides, indicator sets - so the study can be repeated next quarter or next year.
  • Benchmarking against comparable work where a comparable exists.
  • A direct reading of what the finding means for the decision in front of you, no translator required.

Turnaround is typically measured in weeks. Instruments are designed for repetition, so what starts as a one-off becomes an ongoing measurement capability.


Where impact measurement earns its cost

The Lab is most useful at a few specific moments:

  • Before a decision to scale - where the cost of being wrong about how the intervention lands is high. This is Context Entry work.
  • When a funder or board needs evidence that will survive scrutiny - because independent measurement produces the credibility a self-assessment cannot.
  • When a project is still live - a mid-course reading can still change the outcome, and cost less than fixing it later.
  • When the intervention is technical as well as social - payment systems, e-mobility, energy access, water, AI and digital services - where a purely social or purely technical evaluation would miss what matters. The Lab pairs both perspectives.
  • When you specifically need an outside voice - because an internal one, however honest, will be discounted.

How we differ from what already exists

There are excellent impact-measurement firms. The Lab's difference is a specific combination three organisations rarely hold together.

Peers we respect

3ie IDinsight IPA J-PAL

Each is excellent at what it does. See 3ie on rigorous impact evidence, IDinsight on decision-focused evaluation, Innovations for Poverty Action and J-PAL on the RCT standard. The Lab does not try to be any of them.

  • Social-science fieldcraft and technical literacy in the same team. Policy evaluators understand programmes but not the technologies inside them. Technical evaluators understand systems but not the people around them. The Lab refuses the split. See About.
  • Rooted in Delft, working through partners on the ground. Not a survey firm parachuted in; researchers who already speak the language and know the place. See About.
  • Both a publisher and a measurement partner. The articles the Lab runs openly on its own account are the proof of the method; the contracted measurement uses the same method, made available.

Related reading


To discuss a study or a collaboration, see Contact. For the field research behind the measurement, see Field Research. For the wider reading behind it, see Articles.