Skip to content

Predictive Modelling DrivenData Pump It Up · MMA 869 Machine Learning & AI

Ranked #1 on DrivenData predicting which Tanzanian water pumps fail

Tanzania needs to know which of 59,400 water pumps to repair. I built a CatBoost-led blend (0.8235) inside a six-person team whose stacked ensemble hit 0.8308 accuracy, #1 on the leaderboard at submission, against a 0.5431 baseline.

6 min read

Result
0.8308 accuracy for the team’s stacked ensemble and 0.8235 for my own blend, against 0.5431 for always guessing the most common class. #1 on the leaderboard at submission; #3 of 8,657 ranked entries as of 3 Oct 2026.
Impact
Of every 100 broken pumps, my blend correctly sends a repair crew to 77. Trading about 0.5 accuracy points would raise that by roughly 6.
My role
Modelling pair: built my own CatBoost blend, ran the model and error analysis, generated the final submissions · Team of 6
Context
DrivenData Pump It Up · MMA 869 Machine Learning & AI
Timeframe
Summer 2026 (submitted Sep 2026)
Tools
CatBoostLightGBMRandom ForestStackingscikit-learnPython
0.8308
team accuracy on the public leaderboard (CV 0.8208)
+2.0 pts
from native categorical encoding, against under 0.5 from all tuning
−2.3 pts
leaderboard cost of 12 features that cross-validation liked
Native categorical encoding did most of the climb
Accuracy at each step of the project Horizontal bars. First CatBoost 0.7906 in cross-validation; adding the high-cardinality columns 0.8109; proper 5-fold CV 0.8133; my three-model blend 0.8154 in CV and 0.8235 on the public leaderboard; the team's stacked ensemble 0.8308 on the public leaderboard. Guessing the most common class scores 0.5431. 0.78 0.79 0.80 0.81 0.82 0.83 First CatBoost 14 simple categories · CV First CatBoost (14 simple categories · CV): 0.7906 0.7906 + ward, subvillage, funder… high-cardinality columns · CV + ward, subvillage, funder… (high-cardinality columns · CV): 0.8109 0.8109 Proper 5-fold CV same model · CV Proper 5-fold CV (same model · CV): 0.8133 0.8133 My blend CatBoost + LightGBM + RF · CV My blend (CatBoost + LightGBM + RF · CV): 0.8154 0.8154 My blend public leaderboard My blend (public leaderboard): 0.8235 0.8235 Team stacked ensemble public leaderboard Team stacked ensemble (public leaderboard): 0.8308 0.8308
One step carries the chart: letting CatBoost use the high-cardinality columns added 2.0 points. All hyperparameter tuning combined added under 0.5. For scale, always guessing "functional" scores 0.5431.

A repair crew can only visit so many pumps

Tanzania’s water ministry surveyed 59,400 water points. 38% were broken and another 7% needed repair. A crew sent to a working pump is a wasted day; a broken pump nobody visits leaves a village walking to the next well. The competition turns that into one question per pump: functional, needs repair, or non-functional? It is scored on plain accuracy, and the bar to beat is 0.5431: what you get by calling every pump functional.

The data hides its gaps as zeros

Each row is one pump at one survey date, with about 40 columns covering location, hardware, funder, installer, water quantity and payment. Three problems shaped everything after:

  • Fake zeros. 34.9% of construction_year and 34.4% of gps_height values are 0, meaning “unknown”. I turned them into missing values and added a flag, because missingness is itself a clue: pumps with a recorded build year work 56.1% of the time, pumps without one 51.0%.
  • Huge categories. ward has 2,092 values and subvillage 19,287.
  • A rare middle class. Only 7.3% of pumps “need repair”, so accuracy rewards ignoring them.

Anything learned from the labels, including ward-level medians and neighbour features, was rebuilt inside each cross-validation fold so no pump ever saw its own answer.

How the model was built

Letting CatBoost read 2,092 wards was worth 2 points

Our first models were stuck at 0.79 because one-hot encoding subvillage would create about 19,000 columns, so we had dropped the location columns. The signal was in exactly those columns. Failure is local, and it gets stronger the smaller the area:

Area level Areas Spread in failure rate
region 21 0.114
lga 125 0.171
ward 2,092 0.239

CatBoost replaces each category with how pumps in that category usually do, computed in a shuffled order so no pump sees its own label. That made ward, subvillage, funder, installer and scheme_name usable and added +2.0 accuracy points. All hyperparameter tuning combined added under 0.5. I rejected MICE imputation (it erased the missingness signal) and one-hot encoding.

Nearby pumps predict each other

Each pump got the outcome mix of its 10, 25 and 50 nearest pumps by GPS, built out-of-fold. Neighbours share a water table, a mechanic and a budget, and these features ended up carrying 22.2% of the model’s importance.

Cross-validation ranked two blends backwards

Cross-validation ranked our two best blends backwards
Cross-validation score versus public leaderboard score for four pipelines Dumbbell chart. Open circle is cross-validation, filled circle is leaderboard. The blend with 12 extra features dropped from 0.8166 in CV to about 0.80 on the leaderboard. The submitted blend rose from 0.8154 to 0.8235. A teammate's XGBoost and LightGBM pipeline went from 0.8125 to 0.8158. The team stacked ensemble went from 0.8208 to 0.8308. 0.80 0.81 0.82 0.83 Blend with 12 extra features Blend with 12 extra features: cross-validation 0.8166 Blend with 12 extra features: leaderboard ~0.80 CV 0.8166 → ~0.80 Blend without them (submitted) Blend without them (submitted): cross-validation 0.8154 Blend without them (submitted): leaderboard 0.8235 CV 0.8154 → 0.8235 Teammate: XGBoost + LightGBM Teammate: XGBoost + LightGBM: cross-validation 0.8125 Teammate: XGBoost + LightGBM: leaderboard 0.8158 CV 0.8125 → 0.8158 Team stacked ensemble Team stacked ensemble: cross-validation 0.8208 Team stacked ensemble: leaderboard 0.8308 CV 0.8208 → 0.8308
Open circle = cross-validation, filled = public leaderboard. CV put the 12-feature blend ahead; the leaderboard put it 2.3 points behind. The two files disagree on 804 of 14,850 pumps, far too many for luck.

I added 12 engineered features: counts per funder and installer, quantity combinations, date parts. Cross-validation said +0.14 points. The leaderboard said −2.3. The two submissions disagree on 804 of 14,850 pumps, far too many for luck. We had made about ten choices against the same CV rows, and the extra features memorised quirks of those rows. The best of them was the #1 feature by importance in the worse model. I deleted all twelve. Cross-validation is good for comparing options, but after many rounds of tuning against it, it overstates how well the model will do on new data.

0.8308, #1 on the leaderboard at submission

  • Team stacked ensemble: 0.8308 on the public leaderboard, cross-validated at 0.8208. It was #1 on the leaderboard when we submitted in September 2026; newer entries have since posted 0.8325 and a tied 0.8308, so it stands #3 of 8,657 ranked entries as of 3 October 2026. (About 20,000 people have joined the competition; most never submitted.)
  • My blend: 0.8235 on the leaderboard and 0.8154 in stratified 5-fold CV (CatBoost 0.50, LightGBM 0.25, Random Forest 0.25, weights chosen on out-of-fold predictions only).
  • Validation spread: the same CatBoost scored 0.8091 to 0.8207 across folds, a 1.1-point swing from luck alone, which is why every comparison used all five folds.

Per-class results for my blend (out-of-fold, all 59,400 pumps):

Actual ↓ / Predicted → Functional Needs repair Non-functional
Functional 91.6% 1.3% 7.1%
Needs repair 57.1% 28.6% 14.4%
Non-functional 21.6% 1.0% 77.4%

Read as a repair plan: of every 100 broken pumps, the model sends a crew to 77 and misses 22. “Needs repair” is the weak class: precision 0.66, recall 0.29.

What drives pump failure

Location and neighbours make up 45.6% of the model (CatBoost importance)
Share of feature importance by theme Horizontal bars. Location 23.4 percent and neighbour features 22.2 percent together make up 45.6 percent. Hardware 15.9, who funded or runs the pump 13.3, water quantity 10.2, payment and water quality 4.8, age 4.5, everything else 5.7. 0% 5% 10% 15% 20% 25% Where the pump is ward, lga, region, basin Where the pump is (ward, lga, region, basin): 23.4% 23.4% How nearby pumps are doing 10 / 25 / 50 nearest How nearby pumps are doing (10 / 25 / 50 nearest): 22.2% 22.2% Pump hardware type, extraction, source Pump hardware (type, extraction, source): 15.9% 15.9% Who funded, built, runs it funder, installer, scheme Who funded, built, runs it (funder, installer, scheme): 13.3% 13.3% Water availability quantity Water availability (quantity): 10.2% 10.2% Payment and water quality Payment and water quality (): 4.8% 4.8% Age pump_age Age (pump_age): 4.5% 4.5% Everything else Everything else (): 5.7% 5.7%
Location plus neighbours is 45.6% of the model. Nearby pumps share a water table, a mechanic and a maintenance budget, so they fail together.
  • A dry pump is almost always broken: 96.9% of pumps recorded as dry are non-functional.
  • Who runs a pump matters as much as its hardware. Funder, installer and scheme (13.3%) nearly match all hardware columns combined (15.9%).
  • Payment signals maintenance. 75% of pumps where users pay an annual fee work, against 45% where they never pay.

These shares come from CatBoost’s built-in importance, which shows what the model uses, not what causes failure.

Recommendation: rank by risk, then pay for recall

A ministry should send crews in order of predicted probability of failure, not by district. If a missed broken pump costs more than a wasted visit, move the decision threshold: our prior-tuning work showed about 0.5 accuracy points buys roughly 6 points of recall on broken pumps. And because the model is blind to one group (below), young hand pumps should stay on a routine inspection rotation.

Limitations and what I’d do next

  • It misses young hand pumps that still have water. Among broken pumps it misses, 57.1% have enough water and the median age is 11 years, against 23 for the ones it catches. They look healthy on every recorded column; two rounds of features built to fix this failed.
  • Accuracy is the wrong objective for a real ministry. A deployment would set the threshold from repair and travel costs, not from a leaderboard.
  • Importance is not causation. Next step: permutation importance or SHAP on a held-out fold.
  • The leaderboard is not a fully independent test. We used leaderboard scores when choosing which blend to submit (the 12-feature decision above), so 0.8308 is somewhat optimistic as an estimate of performance on new data. Cross-validation (0.8208) is the more honest figure.
  • Every strong model was tree-based. Our logistic-regression baseline scored about 0.76.

Team and credits

Team Danforth had six people. I was half of the modelling pair: I built my own CatBoost, LightGBM and Random Forest pipeline (0.8235), ran the model analysis, confusion matrix, feature importance and error analysis, and generated the final submissions. Teammates led data cleaning and feature engineering, built further CatBoost and XGBoost pipelines, and assembled the stacked ensemble with a meta-learner on top that scored 0.8308.

Technical appendix

Score ladder (my pipeline)

Step CV accuracy Change
CatBoost, 14 simple categories 0.7906 single 80/20 split
+ ward, subvillage, funder, installer, scheme_name 0.8109 +2.0
Stratified 5-fold CV 0.8133 measured properly
+ LightGBM and Random Forest blend 0.8154 leaderboard 0.8235

Final settings. CatBoost 700 rounds, learning rate 0.05, depth 8, l2_leaf_reg 3 (early stopping found 710 rounds best). LightGBM 800 rounds, learning rate 0.05, 96 leaves. Random Forest 400 trees, min_samples_leaf 2, max_features="sqrt".

Tuning. RandomizedSearchCV on XGBoost (50 fits, best 0.7449, held back by one-hot encoding); GridSearchCV on CatBoost (24 fits); class-prior multipliers over 2,025 combinations, rejected because the gain (0.8187 → 0.8188) vanished on held-out rows.

What didn’t work. MICE imputation, StandardScaler (no effect on trees), one-hot encoding, grouping rare funders into “other”, the 12 extra features (−2.3 on the leaderboard), conditional neighbour features (−0.10, 4 of 5 folds worse), and class weighting for accuracy (needs-repair recall 0.21 → 0.70 but accuracy 0.79 → 0.74).

Out-of-fold scores, my blend. Accuracy 0.8154 · macro F1 0.6885 · balanced accuracy 0.6584 · Cohen’s κ 0.6497 · log loss 0.4679.