Skip to content

Predictive Modelling MMA 869 Machine Learning & AI · individual assignment

Chose a turbine model on cost, not accuracy, saving $137,500 a year

Four course cases, one theme: pick the model by the decision it drives. A cost matrix showed the “less accurate” random forest beats an RNN by $137,500 a year on turbine maintenance.

3 min read

Result
Random forest costs $1.63M a year against $1.77M for the RNN and $5.12M for no model. Credit-risk pipeline reached 0.928 ROC-AUC on held-out data, up from a 0.819 baseline.
Impact
In the course’s wind-farm case, predictive maintenance cuts costs by about two thirds; picking the right model is worth a further $137,500 a year. Figures use the case’s given costs.
My role
All analysis, code and write-up · Solo
Context
MMA 869 Machine Learning & AI · individual assignment
Timeframe
Summer 2026 (submitted Sep 2026)
Tools
scikit-learnRandom ForestK-MeansDBSCANAssociation rulesPython
Links
Code and data available on request
$137,500
a year: random forest over RNN, same failure data
0.928
held-out ROC-AUC on credit risk (baseline 0.819)
1.000
adjusted Rand index: K-Means and DBSCAN found the same 5 segments
The random forest is $137,500 a year cheaper than the RNN
Annual maintenance cost under each policy Horizontal bars. Waiting for failures costs 5.12 million dollars a year. The RNN costs 1.77 million. The random forest costs 1.63 million, the lowest. $0M $1M $2M $3M $4M $5M No model fix after failure No model: $5.12M $5.12M RNN higher recall RNN: $1.76M $1.76M Random forest higher precision Random forest: $1.63M $1.63M
Annual cost over 255,501 turbine-days: caught failure $2,500 (inspection + service), missed failure $20,000, false alarm $500. Either model cuts costs by about two thirds.

Accuracy could not tell the two turbine models apart

A wind farm with 700 turbines fails about once every two days. A breakdown costs $20,000 to repair; an inspection costs $500, and fixing a turbine caught early costs $2,000 more. Two models predict failures, and both score above 99% accuracy, because only 256 of 255,501 turbine-days were failures. A model that never predicted a failure would score above 99% too.

So I priced every outcome and applied it to each model’s confusion matrix:

Caught Missed False alarms Annual cost Saving vs no model
Random forest 201 55 50 $1,627,500 $3,492,500
RNN 226 30 1,200 $1,765,000 $3,355,000

The RNN’s better recall is real: its 25 extra catches save $437,500. Its 1,150 extra false alarms cost $575,000. Net, the random forest wins by $137,500 a year.

Recommendation: choose models by expected cost, and check the costs first. The answer flips if inspections fall below about $380 or breakdowns rise above about $25,500.

Engineering affordability beat tuning on credit risk

Feature engineering did most of the work on credit risk
Credit-risk ROC-AUC at each pipeline step Horizontal bars. Baseline 0.8189, feature engineering 0.9096, feature selection 0.9082, tuning 0.9188, held-out test 0.9279. 0.80 0.84 0.88 0.92 Baseline Random Forest, raw features Baseline: 0.8189 0.8189 + feature engineering instalment, ratios + feature engineering: 0.9096 0.9096 + feature selection 21 of 41 kept + feature selection: 0.9082 0.9082 + tuning 32-combo grid + tuning: 0.9188 0.9188 Held-out test never seen before Held-out test: 0.9279 0.9279
ROC-AUC, stratified cross-validation on training data, then one score on the untouched test set. Engineering affordability features added 0.09; tuning added 0.01.

The task was to flag bad credit risks without leaking information. Every step sat inside one scikit-learn pipeline, so feature engineering, selection and tuning were refit on each training fold only, and the test set was scored once at the end. The biggest gain came from features a lender would recognise: monthly instalment (the same $20,000 is a very different risk repaid over 6 months than over several years) and amount per previous account. I scored on ROC-AUC because a model that catches no bad loans can still look accurate.

On the held-out set the tuned model reached 0.928 ROC-AUC, catching 80% of bad risks (recall 0.799) at 0.614 precision. I chose class weighting deliberately: missing a bad risk costs a lender more than refusing a good one.

Two clustering methods agreed on five customer segments

For a jewellery store’s customer base, I compared raw, standard-scaled and min-max-scaled data for K-Means at k = 2 to 10. Unscaled data never passed a silhouette of 0.74 because income (in dollars) swamped spending score (0 to 1). Both scaled versions chose k = 5 with silhouette 0.805. DBSCAN, tuned separately, found the same five clusters (adjusted Rand index 1.000), which is strong evidence the segments are real rather than an artefact of one algorithm. I turned each into a persona, from high earners who spend everything and save $4.1k on average to older, low-income savers.

Association rules: lift separates habit from cause

For a grocery store, I worked through when support, confidence and lift each mislead: milk with eggs has high support but lift near 1, because milk is in most baskets anyway; a rare item like cake can have very high confidence and lift with a specific partner. Lift is the one to act on.

Limitations

  • These are course cases. The costs, data and confusion matrices were given, so the dollar figures show method, not money I saved a real company.
  • The wind-farm comparison takes the confusion matrices as fixed. A real team would also tune each model’s threshold against the cost matrix, which could narrow the gap.
  • Credit-risk tuning used a 32-combination grid with 5 folds to keep runtime short; a wider search might add a little.
Technical appendix

Credit-risk pipeline. FunctionTransformer feature engineering → OneHotEncoder → SelectFromModel (Random Forest, threshold tuned) → RandomForestClassifier. Best grid: class_weight=balanced, max_depth=10, max_features=sqrt, min_samples_leaf=5, 300 trees, selection threshold = mean. Test confusion matrix: 886 good correctly passed, 105 good refused, 42 bad missed, 167 bad caught (1,200 applicants).

Segmentation. K-Means k = 5 on standard-scaled data: silhouette 0.805, Calinski-Harabasz 3,671. DBSCAN grid over eps and min_samples. The highest-silhouette settings (min_samples 24) only scored well by labelling about 26 customers as noise, and those customers formed a coherent segment of their own. Chosen: eps 0.5, min_samples 8, in the middle of the stable region, with every customer assigned.