Predictive Modelling DrivenData Pump It Up · MMA 869 Machine Learning & AI
Ranked #1 on DrivenData predicting which Tanzanian water pumps fail
Tanzania needs to know which of 59,400 water pumps to repair. I built a CatBoost-led blend (0.8235) inside a six-person team whose stacked ensemble hit 0.8308 accuracy, #1 on the leaderboard at submission, against a 0.5431 baseline.
6 min read
- Result
- 0.8308 accuracy for the team’s stacked ensemble and 0.8235 for my own blend, against 0.5431 for always guessing the most common class. #1 on the leaderboard at submission; #3 of 8,657 ranked entries as of 3 Oct 2026.
- Impact
- Of every 100 broken pumps, my blend correctly sends a repair crew to 77. Trading about 0.5 accuracy points would raise that by roughly 6.
- My role
- Modelling pair: built my own CatBoost blend, ran the model and error analysis, generated the final submissions · Team of 6
- Context
- DrivenData Pump It Up · MMA 869 Machine Learning & AI
- Timeframe
- Summer 2026 (submitted Sep 2026)
- Tools
- CatBoostLightGBMRandom ForestStackingscikit-learnPython
- 0.8308
- team accuracy on the public leaderboard (CV 0.8208)
- +2.0 pts
- from native categorical encoding, against under 0.5 from all tuning
- −2.3 pts
- leaderboard cost of 12 features that cross-validation liked
A repair crew can only visit so many pumps
Tanzania’s water ministry surveyed 59,400 water points. 38% were broken and another 7% needed repair. A crew sent to a working pump is a wasted day; a broken pump nobody visits leaves a village walking to the next well. The competition turns that into one question per pump: functional, needs repair, or non-functional? It is scored on plain accuracy, and the bar to beat is 0.5431: what you get by calling every pump functional.
The data hides its gaps as zeros
Each row is one pump at one survey date, with about 40 columns covering location, hardware, funder, installer, water quantity and payment. Three problems shaped everything after:
- Fake zeros. 34.9% of
construction_yearand 34.4% ofgps_heightvalues are 0, meaning “unknown”. I turned them into missing values and added a flag, because missingness is itself a clue: pumps with a recorded build year work 56.1% of the time, pumps without one 51.0%. - Huge categories.
wardhas 2,092 values andsubvillage19,287. - A rare middle class. Only 7.3% of pumps “need repair”, so accuracy rewards ignoring them.
Anything learned from the labels, including ward-level medians and neighbour features, was rebuilt inside each cross-validation fold so no pump ever saw its own answer.
How the model was built
Letting CatBoost read 2,092 wards was worth 2 points
Our first models were stuck at 0.79 because one-hot encoding subvillage would create about
19,000 columns, so we had dropped the location columns. The signal was in exactly those columns.
Failure is local, and it gets stronger the smaller the area:
| Area level | Areas | Spread in failure rate |
|---|---|---|
| region | 21 | 0.114 |
| lga | 125 | 0.171 |
| ward | 2,092 | 0.239 |
CatBoost replaces each category with how pumps in that category usually do, computed in a shuffled
order so no pump sees its own label. That made ward, subvillage, funder, installer and
scheme_name usable and added +2.0 accuracy points. All hyperparameter tuning combined added
under 0.5. I rejected MICE imputation (it erased the missingness signal) and one-hot encoding.
Nearby pumps predict each other
Each pump got the outcome mix of its 10, 25 and 50 nearest pumps by GPS, built out-of-fold. Neighbours share a water table, a mechanic and a budget, and these features ended up carrying 22.2% of the model’s importance.
Cross-validation ranked two blends backwards
I added 12 engineered features: counts per funder and installer, quantity combinations, date
parts. Cross-validation said +0.14 points. The leaderboard said −2.3. The two submissions disagree
on 804 of 14,850 pumps, far too many for luck. We had made about ten choices against the same CV
rows, and the extra features memorised quirks of those rows. The best of them was the #1 feature by
importance in the worse model. I deleted all twelve. Cross-validation is good for comparing
options, but after many rounds of tuning against it, it overstates how well the model will do on new
data.
0.8308, #1 on the leaderboard at submission
- Team stacked ensemble: 0.8308 on the public leaderboard, cross-validated at 0.8208. It was #1 on the leaderboard when we submitted in September 2026; newer entries have since posted 0.8325 and a tied 0.8308, so it stands #3 of 8,657 ranked entries as of 3 October 2026. (About 20,000 people have joined the competition; most never submitted.)
- My blend: 0.8235 on the leaderboard and 0.8154 in stratified 5-fold CV (CatBoost 0.50, LightGBM 0.25, Random Forest 0.25, weights chosen on out-of-fold predictions only).
- Validation spread: the same CatBoost scored 0.8091 to 0.8207 across folds, a 1.1-point swing from luck alone, which is why every comparison used all five folds.
Per-class results for my blend (out-of-fold, all 59,400 pumps):
| Actual ↓ / Predicted → | Functional | Needs repair | Non-functional |
|---|---|---|---|
| Functional | 91.6% | 1.3% | 7.1% |
| Needs repair | 57.1% | 28.6% | 14.4% |
| Non-functional | 21.6% | 1.0% | 77.4% |
Read as a repair plan: of every 100 broken pumps, the model sends a crew to 77 and misses 22. “Needs repair” is the weak class: precision 0.66, recall 0.29.
What drives pump failure
- A dry pump is almost always broken: 96.9% of pumps recorded as
dryare non-functional. - Who runs a pump matters as much as its hardware. Funder, installer and scheme (13.3%) nearly match all hardware columns combined (15.9%).
- Payment signals maintenance. 75% of pumps where users pay an annual fee work, against 45% where they never pay.
These shares come from CatBoost’s built-in importance, which shows what the model uses, not what causes failure.
Recommendation: rank by risk, then pay for recall
A ministry should send crews in order of predicted probability of failure, not by district. If a missed broken pump costs more than a wasted visit, move the decision threshold: our prior-tuning work showed about 0.5 accuracy points buys roughly 6 points of recall on broken pumps. And because the model is blind to one group (below), young hand pumps should stay on a routine inspection rotation.
Limitations and what I’d do next
- It misses young hand pumps that still have water. Among broken pumps it misses, 57.1% have enough water and the median age is 11 years, against 23 for the ones it catches. They look healthy on every recorded column; two rounds of features built to fix this failed.
- Accuracy is the wrong objective for a real ministry. A deployment would set the threshold from repair and travel costs, not from a leaderboard.
- Importance is not causation. Next step: permutation importance or SHAP on a held-out fold.
- The leaderboard is not a fully independent test. We used leaderboard scores when choosing which blend to submit (the 12-feature decision above), so 0.8308 is somewhat optimistic as an estimate of performance on new data. Cross-validation (0.8208) is the more honest figure.
- Every strong model was tree-based. Our logistic-regression baseline scored about 0.76.
Team and credits
Team Danforth had six people. I was half of the modelling pair: I built my own CatBoost, LightGBM and Random Forest pipeline (0.8235), ran the model analysis, confusion matrix, feature importance and error analysis, and generated the final submissions. Teammates led data cleaning and feature engineering, built further CatBoost and XGBoost pipelines, and assembled the stacked ensemble with a meta-learner on top that scored 0.8308.
Technical appendix
Score ladder (my pipeline)
| Step | CV accuracy | Change |
|---|---|---|
| CatBoost, 14 simple categories | 0.7906 | single 80/20 split |
+ ward, subvillage, funder, installer, scheme_name |
0.8109 | +2.0 |
| Stratified 5-fold CV | 0.8133 | measured properly |
| + LightGBM and Random Forest blend | 0.8154 | leaderboard 0.8235 |
Final settings. CatBoost 700 rounds, learning rate 0.05, depth 8, l2_leaf_reg 3 (early stopping
found 710 rounds best). LightGBM 800 rounds, learning rate 0.05, 96 leaves. Random Forest 400 trees,
min_samples_leaf 2, max_features="sqrt".
Tuning. RandomizedSearchCV on XGBoost (50 fits, best 0.7449, held back by one-hot encoding); GridSearchCV on CatBoost (24 fits); class-prior multipliers over 2,025 combinations, rejected because the gain (0.8187 → 0.8188) vanished on held-out rows.
What didn’t work. MICE imputation, StandardScaler (no effect on trees), one-hot encoding, grouping rare funders into “other”, the 12 extra features (−2.3 on the leaderboard), conditional neighbour features (−0.10, 4 of 5 folds worse), and class weighting for accuracy (needs-repair recall 0.21 → 0.70 but accuracy 0.79 → 0.74).
Out-of-fold scores, my blend. Accuracy 0.8154 · macro F1 0.6885 · balanced accuracy 0.6584 · Cohen’s κ 0.6497 · log loss 0.4679.