Final Model
Why M8 was selected, how it performs on the untouched 5,000-row test set, and what that is worth in avoided maintenance cost.
| Candidate | Val recall | Val F1 | Val precision | Verdict |
|---|---|---|---|---|
| M5 Adam + dropout | 90.99% | 93.95% | 97.12% | Best F1, but misses more failures |
| M6 SGD + class weights | 91.44% | 75.75% | 64.65% | Recall bought with 1-in-3 false alarms |
| M10 Adam + dropout + OS | 91.44% | 88.26% | 85.29% | Ties on recall, overfits the duplicated rows |
| M8 Adam + dropout + CW | 91.89% | 91.48% | 91.07% | Selected — highest recall, precision held |
- Recall is the primary metric, and M8 leads it while keeping precision above 91% — every extra failure caught does not come at the cost of a flood of inspections.
- Its training and validation recall are identical at 91.89%, so the model generalises rather than memorises.
- Class weights are preferred over oversampling because they reweight the loss without duplicating rows.
Adam (lr 0.001) · Binary cross-entropy with class weights · 30 epochs · batch 128 · seed 42 · threshold 0.5
| Predicted healthy | Predicted failure | |
|---|---|---|
| Actual healthy | 4,690 | 28 |
| Actual failure | 35 | 247 |
- 247 failures caught, 35 missed
- 28 false alarms, 4,690 healthy units cleared
- ROC-AUC 0.939
- Validation recall was 91.89% and test recall is 87.59% — a small, expected drop on genuinely unseen data with no sign of overfitting.
Unit costs are assumed ($5k inspection / $15k repair / $60k replacement); the project statement supplies only the ordering replacement > repair > inspection
Unit costs are illustrative and preserve the stated ordering; the percentage saving is what should be read, not the absolute dollars.
Real M8 test-set probabilities on 5,000 unseen generators. Move the threshold to see the operating point change live.
Recall
87.59%
247 caught · 35 missed
baseline
Precision
89.82%
28 false alarms
baseline
F1 score
88.69%
accuracy 98.74%
baseline
Expected cost
$5.95M
64.9% below the no-model baseline
baseline
At 0.50 the model catches 247 of 282 failures, raises 28 false alarms and costs $5.95M versus $16.92M with no model.
Expected cost = 247 repairs × $15,000 + 35 replacements × $60,000 + 28 inspections × $5,000.
Cost is flat between roughly 0.35 and 0.70 — the operating point is robust, so 0.50 is kept for interpretability while a lower threshold stays available whenever missed failures become more expensive.
Click Generate to produce an evidence-based observation, insight, recommendation, and business benefit for this section.