← The Water for Health Fund  ·  Technical annex

Cost per averted DALY, shown all the way down

The fund's primary metric, computed to GiveWell standard, with every assumption, adjustment, and cost basis visible.

Most cost-effectiveness claims in this sector are a single number with an invisible denominator. This annex does the opposite: it publishes the machinery. All seven programs in the portfolio are modeled end-to-end, each on an honestly stated cost basis: the flagship worksheet is below, and every program has its own worksheet page and downloadable workbook in the portfolio section.

PRELIMINARY · v1 · AUGUST 2026 · MODELS UNDERGO ADVERSARIAL REVIEW BEFORE FIGURES ARE FINAL
01The Metric

Four numbers per program, not one

Every program in the portfolio is scored on the same four metrics. Together they answer the two questions a funder should insist on: what does a dollar buy, and would the claim survive a referee.

1 · $ / DALY averted
Health value on a full delivery cost basis, GiveWell-comparable, reported both raw and evidence-adjusted. The evidence adjustment is a published four-factor pipeline (internal validity × external validity × publication bias × implementation quality), never a hidden haircut.
2 · Marginal donor $ / DALY
The fund's primary metric: the same DALYs divided by the net philanthropic cost after third-party revenue. Where carbon revenue covers full delivery cost, this falls to zero, and the health benefit is free to philanthropy.
3 · $ / tonne CO₂e vs SCC
Climate value: full delivery cost per credited tonne, against the $185 social cost of carbon (Rennert et al. 2022, Nature). Reported both as credited and integrity-adjusted.
4 · Credit integrity ratio
Real emission reductions divided by credited reductions, decomposed into baseline practice × sensor-verified delivery × leakage. The health model and the credit quantification are required to use the same delivery number, so the model cannot tell one story to the carbon buyer and another to the health funder.

Each model also reports self-sustainability: the ratio of carbon revenue to delivery cost, in the base case and at the contracted downside price floor. Moral-weight sensitivity is run under three published frameworks (GiveWell default, egalitarian, wellbeing-anchored); the GiveWell default is the headline.

02Worked Example

Amazi Meza Rwanda, in its most conservative frame

The flagship school safe-water program: 2,000 schools and 1.98 million students at scale, Gold Standard registered, Article 6 credits contracted and issuing since 2024. The model below is deliberately strict with itself.

What the model counts

All-in delivery cost over 2024–2033: filtration hardware and replacements, rainwater systems, and in-country staffing, $12.25M against 14.7M student-years served ($0.83 per student-year). Health benefit is built bottom-up: diarrheal illness averted among students only, valued as years lived with disability, using the meta-analytic filtration effect (Wolf et al. 2022, Lancet) and sensor-verified delivery performance.

What it deliberately excludes

Under-five mortality and household spillovers (the effects measured in Rwanda's cluster-randomized trial, Kirby et al. 2019), respiratory-infection reductions, and fuel and time savings (the dominant benefits in Barstow et al. 2019). Every exclusion is conservative: each would lower the cost per DALY, some substantially. They enter as the evidence is traced, not before.

$6,822
per DALY averted, raw, on full delivery cost, students-only YLD frame. Evidence-adjusted: $14,026 (composite adjustment 0.49).
106%
of full 2024–2033 delivery cost covered by contracted carbon revenue ($13.0M against $12.25M). Steady-state coverage reaches 2.8–3.4× from 2029.
$0
marginal donor cost per averted DALY at contracted prices: carbon revenue carries delivery, so every averted DALY is free to philanthropy.
$16.43/t
full delivery cost per credited tonne, 11.3× below the $185 social cost of carbon.
$23.47/t
integrity-adjusted cost per tonne (real rather than credited reductions), still 7.9× below the SCC.
0.70
credit integrity ratio: residual baseline factor 0.78 (flagged for a measured baseline survey) × sensor-verified delivery 0.90 × leakage 1.0. The 2026 methodology reconstruction had already removed roughly 45% of the earlier crediting basis.

Why publish a $6,822 figure when the site's headline is under $100? Because they answer different questions on different bases, and this annex exists to keep those bases separate. The $6,822 is the strictest possible gross frame: school-age illness only, all costs in, every unproven benefit out. The stronger gross figures live in the portfolio's community programs, where under-five exposure dominates and the randomized mortality-relevant evidence applies: Asili DR Congo models at $355 per DALY on its philanthropic cost basis, and Burundi reaches $2,908 under the Kremer mortality scenario on full cost; every program's worksheet is linked in the portfolio section. And the fund's primary metric, marginal donor cost, is what carbon revenue drives toward zero regardless of frame.

03The Worksheet

The entire model, cell by cell

Everything below is the complete v1 workbook: every input with its provenance, every calculation in the order it runs, every scenario. Nothing is summarized away.

Provenance key: Proforma input Evidence-anchored Judgment · flagged Sensor-measured Formula
↓ Download the workbook (.xlsx)
A · Program series, 2024–2033 (Virridy carbon proforma, base scenario)
Series2024202520262027202820292030203120322033Total
Students enrolled197,600444,600839,8001,383,2001,976,0001,976,0001,976,0001,976,0001,976,0001,976,00014,721,200 student-yrs
Net credits (tCO₂e)7,28232,21537,44069,88799,83899,83899,83899,83899,83899,838745,852
Carbon revenue ($)0109,230515,440636,4801,257,9661,896,9221,996,7602,096,5982,196,4362,296,27413,002,106
Rainwater capex ($)250,000300,000450,000550,000250,000000001,800,000
Filter capex ($)852,8911,245,221234,423994,4221,084,824000004,411,781
Replacement filters ($)24,70055,575104,975172,900247,000247,000247,000247,000247,000247,0001,840,150
In-country staffing ($)420,000420,000420,000420,000420,000420,000420,000420,000420,000420,0004,200,000

All rows are proforma inputs from the base scenario of the portfolio financial model; net credits carry the ~5% Article 6 host-country retirement already deducted, on the Gold Standard SDWS V2.0 yield basis (0.053 credits per person-year). Credits issued to date: 16,153 tCO₂e (Sep 2024) plus 17,758 pending, matching the 2024–25 rows.

B · Cost basis: one denominator
TOTAL_PROGRAMME_COST, 2024–2033$12,251,931Formula rainwater + filter capex + replacements + staffing
Student-years served14,721,200Formula sum of the enrollment row
Cost per student-year$0.832Formula cost ÷ student-years

Every figure on this page divides by this one cost. The build fails programmatically if any downstream number uses a different denominator. Deliberately excluded from this cost: an allocated share of portfolio-level management, verification, and sensor operations (an optimistic omission, priced in scenario S7), and post-2033 operating years within the 15-year crediting period (a conservative omission: the steady-state years are the cheapest). The organization's published "∼$12M investment" reconciles with this basis as a cross-check; it is not the denominator.

C · Health chain inputs
InputValueProvenanceBasis
Diarrheal incidence, school-age0.60 episodes/child-yrJudgment · flaggedCountry-specific GBD lookup pending; tested 0.30–1.20 in scenarios S1–S2. The widest single uncertainty in the model.
Diarrhea reduction, point-of-use filtration34%Evidence-anchoredWolf et al. 2022 (Lancet), filtration risk ratio 0.66. The program's own randomized result (29%, Kirby et al. 2019) is consistent; the meta-analytic value is used, and a conservative 25% is tested in S5.
School share of daily drinking water30%Judgment · flaggedStudents drink roughly a third of their water at school; no time-use survey yet. Tested 20–40% in S3–S4.
Delivery effectiveness (safe at the point)90%Sensor-measuredPortfolio monitoring record: 90.3% of samples safe under monitored delivery (n=1,685) against 24.3% at baseline (n=4,943).
Disability weight, diarrheal episode0.188Evidence-anchored flaggedGBD moderate diarrhea. The mild/moderate severity mix is unresolved (mild is 0.074), so this choice is flagged.
Episode duration4.3 daysJudgmentTypical acute episode.
YLD per episode0.002215 DALYFormula0.188 × 4.3 ÷ 365
D · Calculation chain
StepValueComputation
Episodes averted, 2024–2033810,84414,721,200 student-yrs × 0.60 × 0.34 × 0.30 × 0.90
DALYs averted (raw)1,796episodes × 0.002215 YLD. Years lived with disability only: school-age diarrheal mortality is excluded (conservative).
$ / DALY averted, raw$6,822$12,251,931 ÷ 1,796
Evidence adjustment (section E)× 0.4861,796 → 874 DALYs adjusted
$ / DALY averted, evidence-adjusted$14,026$12,251,931 ÷ 874
xCash, GiveWell default weights0.101×DALYs × 2.3 units ÷ cost ÷ 0.00335 units/$ (GiveDirectly benchmark)
xCash, egalitarian / wellbeing weights0.044× / 0.158×same chain, DALY moral weight 1.0 / 3.6
xCash, moral-uncertainty weighted0.103×0.45 × GiveWell + 0.25 × egalitarian + 0.30 × wellbeing. Ranking is stable across all three frameworks.
E · Evidence-adjustment pipeline
FactorScoreBasis
Internal validity0.80The effect anchor is meta-analytic, but the chain multiplies it by two judgment inputs (incidence, school share).
External validity0.80Household point-of-use trials transferred to school gravity-filter delivery.
Publication bias0.80Systematic review without an independent funnel-plot confirmation in this model.
Implementation quality0.95Sensor-instrumented delivery, continuous monitoring record, registry issuance track record.
Composite0.486Formula product of the four. Reported alongside the raw figure, never silently multiplied into it.
F · Credit integrity and carbon metrics
Residual baseline-practice factor0.78Judgment · flagged The 2026 methodology reconstruction (V2.0) already cut the crediting basis by ~45%; this prices the residual risk that baseline water-boiling practice is still overstated. A measured school baseline survey replaces it.
Delivery effectiveness0.90Sensor-measured The same cell the health chain uses. One usage number, two consumers.
Leakage / rebound1.00Negligible for water treatment.
Credit integrity ratio0.70Formula 0.78 × 0.90 × 1.00
$ / tonne, credited basis$16.43$12,251,931 ÷ 745,852 t · 11.3× below the $185 social cost of carbon
$ / tonne, integrity-adjusted$23.47$12,251,931 ÷ (745,852 × 0.70) · 7.9× below the SCC. Fails the SCC test only if integrity falls below 0.089.
Implied realized price$17.43/t$13,002,106 revenue ÷ 745,852 t, blended across the Article 6 offtake and post-contract sales
Revenue coverage of full cost106.1%$13,002,106 ÷ $12,251,931 → net philanthropic cost −$750,175: self-financing over the window
Steady-state coverage2.84× → 3.44×2029 and 2033 revenue against the $667,000 steady-state year (replacements + staffing)
G · Scenarios (one-way; a preliminary model carries no Monte Carlo by design)
ScenarioVaried inputDALYs$ / DALY rawxCash GWxCash MUReading
Base case1,796$6,8220.101×0.103×the strict frame, as above
S1 · incidence low0.30898$13,6450.050×0.052×worst single-input case; climate finding unchanged
S2 · incidence high1.203,592$3,4110.201×0.207×
S3 · school water share low20%1,197$10,2340.067×0.069×
S4 · school water share high40%2,394$5,1170.134×0.138×
S5 · conservative effect size25%1,320$9,2780.074×0.076×GiveWell-style conservative reduction
S6 · + fuelwood savings$5.2M / 10 yr1,796$1,308 -equivalent0.525×0.528×13,000 t/yr at $40/t would dominate health value 4:1. Unverified incidence; scenario-only until measured, never central.
S7 · + allocated portfolio overhead+$2.5M1,796$8,2140.084×0.086×cost per tonne rises to $19.78, still 9.4× below the SCC
S8 · downside price floor$10/t1,796$2,669 net basiscoverage 60.9%revenue $7.46M, net philanthropic cost $4.79M; the donor-cost ladder's third rung

Recomputed robustness statement, not asserted: no single input moves the health frame anywhere near the cash-transfer bar on full cost, and none reverses the climate finding. The self-financing conclusion is the fragile one, and its reversal condition is stated in section F and S8.

H · The model's critique of itself
DimensionRatingAssessment
Effect-size qualityWEAKThe meta-analytic anchor is solid, but the chain multiplies it by two judgment inputs with no Rwanda-specific data, and the program's own randomized result is not yet independently traced into the model file. Preliminary status is correct; the burden lookup and source tracing gate the upgrade.
Counterfactual robustnessMODERATEHealth-side funging is low: no other funder provides school point-of-use treatment at this scale in Rwanda. Carbon-side additionality is the sharper question: a registered project with a contracted buyer could plausibly attract another developer, and v2 must model it.
Moral-weight transparencySTRONGAll frameworks reported; ranking stable. The decision-relevant fact is the frame (gross vs marginal donor cost), not the framework.
Attribution and integrityMODERATE-STRONGVirridy develops and operates the program, so attribution near 1.0 is defensible. The integrity ratio is decomposed and shares its delivery term with the health chain; its weak link, the 0.78 baseline residual, is flagged, not hidden.
Excluded factorsFLAGExcluded: under-five and household spillovers, respiratory effects, fuel and time savings (conservative on benefits); allocated overhead and post-2033 years (boundary choices priced in S7 and noted in B). Net direction: conservative on benefits, optimistic on the cost boundary.
Ranking stabilityMODERATEHealth and climate frames are stable across S1–S8. The self-financing finding is the sensitive one: it fails below roughly $16.40/t average realized price or a 6% issuance shortfall.
04Donor Cost

The donor-cost ladder

The same program, the same DALYs, three honest cost bases. The metric that matters to a donor is the third.

05Method

GiveWell's machinery, plus rules GiveWell doesn't need

The protocol adopts the standard machinery of professional cost-effectiveness analysis, then adds discipline specific to carbon-financed delivery.

Standard machinery

Bottom-up disease-burden construction (never a percentage applied to an aggregate burden). A four-factor evidence-adjustment pipeline with published scoring rubrics, reported alongside the raw estimate, never silently multiplied in. Three moral-weight frameworks run in parallel. Counterfactual and funging analysis with named alternative funders. One-way sensitivity on every load-bearing input, with probabilistic analysis reserved for models whose inputs can honestly carry distributions.

Carbon-specific discipline

A single cost denominator per model, asserted in code so the build fails if any figure divides by a different cost than the one that produced the benefit. One delivery number shared by the health chain and the credit quantification. An integrity ratio that decomposes how credited tonnes relate to real ones, informed by the continuous sensor record rather than annual surveys. And a hard labeling rule: revenue-leveraged figures are never presented as gross cost-effectiveness.

Adversarial review before any figure is final

Every model faces a three-seat referee panel before its numbers are cited: an econometrician (recomputes every chain from raw inputs), an uncertainty analyst (attacks the sensitivity and scenario structure), and a skeptical grantmaker briefed to argue against funding, who for this portfolio also takes the carbon-market critic's chair: suppressed-demand baselines, self-reported usage, and additionality are attacked before an outside critic does. The worked example above is v1 and enters that review now; the independent, GiveWell-partnered evaluation in design supersedes all of it.

Delivery inputs draw on the portfolio's continuous verification layer: 1,273 monitored water sites, with pooled water quality moving from 24% of samples testing safe at baseline (n=4,943) to 90% under monitored delivery (n=1,685). See the evidence base on the main page. Sensor credibility is applied only where the sensor actually measures: water safety at the point of delivery, not household behavior, and the models keep that boundary explicit.

06Portfolio

All seven programs, each on its honest basis

Every program is modeled with the same protocol; each row links to that program's full cell-by-cell worksheet and downloadable workbook. The cost bases differ by design, so the rows are labeled rather than averaged: a single portfolio $/DALY would mix denominators, and this annex does not do that.

ProgramCost basis (2024/25–2033)$ / DALY raw → net donor$ / tCO₂e (integrity-adj)IntegrityWorksheet
Amazi Meza · RwandaFull delivery · $12.25M$6,822 → $0 (coverage 106%)$16.43 ($23.47) · 7.9× under SCC0.70on this page
Amazi Water · BurundiFull delivery · $53.4M$5,307 → $2,450 ($15 floor)$27.87 ($44.23) · 4.2×0.63worksheet
Asili · DR CongoPhilanthropic rehab · $1.2M$355 ($1,133 adj); tariffs fund O&M$1.08 ($2.30) · 80×0.47 flaggedworksheet
LifeStraw · KenyaCarbon financing only · $2.35M$7,312 financing basis, labeled$18.21 ($25.29) · 7.3×0.72worksheet
MWA DRIP · KenyaGrant + purchases · $7.49Mnone by design pathway not yet designed$13.79 ($18.03) · 10.3×0.77worksheet
Helvetas · MadagascarCarbon financing only · $2.75M$572 financing basis, labeled$10.31 ($20.21) · 9.2×0.51 flaggedworksheet
Water Mission · TanzaniaOfftake outlay · $10.39Mnone by design another org's delivery$12.83 ($17.82) · 10.4×0.72worksheet

What rolls up honestly across bases: 5.5M tonnes modeled across the seven programs, 2024–2033, and every program's integrity-adjusted cost per tonne sits between 4.2× and 80× below the $185 social cost of carbon on its own stated basis. What does not roll up: a single portfolio $/DALY (mixed denominators), and the strongest per-DALY rows (Asili, Helvetas) are the ones on partial cost bases, which is exactly why the basis column exists. Integrity flags: the three programs still on pre-reconstruction crediting baselines (Asili 0.243, Helvetas 0.21, Burundi 0.170 tonnes per person-year) carry the deepest integrity haircuts here, quantified before any critic asks. All models are v1 preliminary, in adversarial review.