PM

System Initializing

Waking up industrial sensors...

Render Free Tier is spinning up the server. This usually takes 30-60 seconds.

Home Dataset EDA Bias SMOTE Balanced Normalization Models Eval Hybrid Result Predict ๐Ÿ“„ Report

Predictive
Maintenance
Using ML

In industrial environments, unexpected machine failures lead to significant financial losses, production downtime, and safety risks. Traditional maintenance is either too early (wasteful) or too late (costly).

This project builds a data-driven predictive system that analyses real machine sensor readings to forecast failures before they happen โ€” enabling timely, targeted maintenance.

10,000
Data samples
5
Features
5
ML models
97.5%
Best accuracy
sensor_feed.live
Air Temp298K
Process Temp308K
Rot. Speed1551rpm
Torque42Nm
Tool Wear198min

What Data We Have

The AI4I Predictive Maintenance Dataset contains 10,000 operational snapshots from a simulated milling machine. Each row captures the machine's state at one point in time, labelled as failure or no-failure.

๐ŸŒก
Air Temperature [K]
Ambient air temperature around the machine. Generated using a random walk process around 300 K. Affects cooling efficiency and thermal environment of the machine.
295 โ€“ 304 K
๐Ÿ”ฅ
Process Temperature [K]
Internal machine operating temperature. Always 8โ€“10 K higher than air temperature due to heat generation. Thermal stress on components rises sharply above threshold.
305 โ€“ 314 K
โš™๏ธ
Rotational Speed [rpm]
Spindle speed in revolutions per minute. Very high or very low speeds indicate operational stress or tool wear failure modes.
1168 โ€“ 2886 rpm
๐Ÿ”ฉ
Torque [Nm]
Rotational force applied to the cutting tool. High torque combined with low rotational speed is a strong indicator of mechanical overload and overstrain failure.
3.8 โ€“ 76.6 Nm
๐Ÿช›
Tool Wear [min]
Cumulative minutes the cutting tool has been in use. The primary aging indicator โ€” tools beyond 200 min fail more frequently across all failure types.
0 โ€“ 253 min
๐ŸŽฏ
Target (Label)
Binary classification label. 0 = No Failure, 1 = Failure. Derived from five sub-failure types: tool wear, heat dissipation, power, overstrain, and random failures.
0 or 1

Understanding the Raw Data

Before applying any ML technique, we performed a thorough EDA on the original dataset to understand the shape of each feature, detect outliers, and discover relationships between variables.

Feature Distributions (Histogram + KDE)
Mean300.0 K
Std Dev2.0 K
Min295.3 K
Max304.5 K
Near-normal distribution. Random walk generation keeps values in a narrow band around 300 K. No major outliers detected.
Outlier Detection โ€” Boxplots (All Features)
Whiskers = 1.5ร—IQR ยท Points beyond whiskers = outliers
What the boxplots reveal

Rotational speed has a right skew โ€” most machines operate in the 1200โ€“1800 rpm range but a few run much faster. Torque shows high-end outliers beyond 65 Nm, which often correspond to overstrain failures. Tool wear is near-uniform since it increments linearly with usage time. These outliers are intentionally kept โ€” they carry critical failure signal and removing them would reduce model sensitivity.

Correlation Heatmap โ€” Original Dataset
Pearson correlation between all numerical features
Key relationships discovered

Air Temp โ†” Process Temp (r = 0.88): Strong positive correlation. When ambient temperature rises, internal machine temperature follows โ€” increasing thermal stress risk.

Rotational Speed โ†” Torque (r = โˆ’0.88): Strong negative correlation. This is mechanical law โ€” power = speed ร— torque. Extreme combinations (very low speed + very high torque) trigger overstrain failures.

Tool Wear โ†” Target (r = 0.11): Weak but real. Tool wear amplifies other failure modes rather than causing failures on its own.

Class Imbalance โ€” The Critical Problem

During EDA, we discovered that the dataset is severely imbalanced. This is one of the most dangerous problems in real-world ML classification tasks.

Class Distribution โ€” Original Dataset
9,661
No Failure โ€” 96.6%
339
Failure โ€” 3.4%
28 : 1 imbalance ratio
The accuracy paradox

A model that always predicts "No Failure" achieves 96.6% accuracy without learning anything. This is why we must use F1-score and recall as primary metrics โ€” not accuracy.

๐Ÿ’ฅ Real-world consequence

In predictive maintenance, a missed failure (false negative) costs far more than a false alarm. An undetected failure can cause catastrophic equipment damage, production halt, or even worker injury. A biased model optimising for accuracy is useless โ€” it would miss 100% of failures. We must fix the imbalance before training.

SMOTE โ€” Fixing the Imbalance

Synthetic Minority Oversampling Technique (SMOTE) generates new synthetic samples for the minority class by interpolating between existing examples โ€” preserving distribution while balancing classes. We used random_state=42 for reproducibility.

1
Select a minority class sample
โ†’
2
Find K nearest minority neighbors
โ†’
3
Interpolate a new synthetic point
โ†’
4
Repeat until classes are equal
Class Balance โ€” Before SMOTE
Class Balance โ€” After SMOTE
Feature Distribution โ€” Before vs After SMOTE
Validating SMOTE quality

Feature distributions remain structurally similar before and after SMOTE โ€” confirming that synthetic samples follow the same statistical patterns as real data. The distributions overlap closely, meaning the synthetic failure samples are realistic interpolations rather than noise. This is the key advantage of SMOTE over simple oversampling (duplication), which would cause severe overfitting.

Validating the Balanced Dataset

After SMOTE, the dataset grew from 10,000 to 19,322 samples (9,661 per class). We ran a fresh EDA pass to confirm the balance and verify statistical integrity.

Class Distribution โ€” Balanced
Feature Distributions โ€” Balanced Dataset
Correlation Heatmap โ€” Balanced Dataset
Compare with the original heatmap in Section 03
SMOTE preserves correlation structure

The correlation matrix is nearly identical to the original dataset. Air-Process temperature correlation remains at 0.86, and the Speed-Torque negative correlation stays at โˆ’0.85. This is critical evidence that SMOTE generated statistically consistent samples โ€” maintaining the underlying mechanical relationships of the machine rather than introducing artificial patterns.

StandardScaler โ€” Preparing for Training

Features exist at vastly different numerical scales. StandardScaler transforms each feature to zero mean and unit standard deviation so no feature dominates others due to magnitude alone.

Feature Ranges โ€” Before Normalization
Feature Ranges โ€” After Normalization
Normalized Feature Distributions
Why normalization matters here

KNN computes similarity using Euclidean distance. Without normalization, rotational speed (range ~1700) dominates over torque (range ~73) purely due to scale. After StandardScaler, all five features contribute equally to distance calculations.

Logistic Regression uses gradient descent to fit coefficients โ€” unnormalized features cause slow convergence and biased weights.

Random Forest and Decision Tree are scale-invariant by nature but normalization ensures consistent preprocessing across the entire pipeline.

โœ“ 19,322 samples โœ“ Balanced 1:1 ratio โœ“ Mean = 0, Std = 1 โœ“ 80/20 train-test split โœ“ Ready for training

Training & Comparing Models

We trained four base classifiers with a stratified 80/20 train-test split (random_state=42). Stratification ensures both train and test sets maintain the balanced 50/50 class ratio.

Model Accuracy Comparison
All Metrics โ€” Grouped Comparison
Accuracy ยท Precision ยท Recall ยท F1 Score
83.5%
Logistic Regression

Linear decision boundary fails to capture non-linear interactions between features like speed-torque combinations. Useful as a baseline but insufficient for complex failure patterns in sensor data.

95.7%
KNN (k=5)

Finds 5 nearest neighbors in normalized feature space. Achieves the highest recall (98.7%) โ€” misses very few failures. Lower precision (93.1%) means more false alarms. Performance entirely depends on normalization.

95.0%
Decision Tree

Learns explicit IF-THEN rules from feature thresholds. Highly interpretable. Slightly prone to overfitting which the Random Forest ensemble approach corrects. Good balance between precision and recall.

97.4%
Random Forest

Builds 100 trees on random feature subsets, aggregates votes. Best balance of precision (96.4%) and recall (98.3%) among base models. Ensemble approach reduces variance and avoids overfitting compared to a single decision tree.

Confusion Matrix & ROC Curves

Deep-dive evaluation of every model. The confusion matrix shows exactly what each model got right and wrong. The ROC curve shows performance across all decision thresholds โ€” AUC closer to 1.0 is better.

Confusion Matrix
Actual vs Predicted on test set (3,865 samples)
Predicted
No Failure
Predicted
Failure
Actual No Failure
โ€”
โ€”
Actual Failure
โ€”
โ€”
True Negatives (TN)
Correctly said No Failure
False Positives (FP)
False alarms raised
False Negatives (FN)
Missed failures โš 
True Positives (TP)
Correctly caught failures
Evaluation Summary

The Matrix breaks down the exact number of correct and incorrect predictions. The goal is to maximize the Diagonal (Blue/Green) and minimize the others.

ROC Curve โ€” Logistic Regression
AUC: โ€”
ROC Interpretation

The red area represents the model's ability to distinguish between classes. A larger area (AUC) means the model is better at catching failures without raising too many false alarms.

All Models โ€” ROC Curve Comparison
Higher and more left-bowed = better. AUC = area under the curve.
How to read this chart

The diagonal dashed line is a random classifier (AUC = 0.5). Any model above this line is better than guessing. The closer the curve hugs the top-left corner, the better. AUC = 1.0 is perfect. Our Hybrid achieves the highest AUC โ€” confirming it as the strongest model at every threshold.

Hybrid Model โ€” Best of Both Worlds

We combined the top two models using a Hard Voting Classifier (sklearn VotingClassifier). Each model independently predicts the class and the majority vote determines the final output โ€” reducing the individual weaknesses of each model.

Normalized test data (3,865 samples)
๐ŸŒฒ
Random Forest
100 trees ยท feature subsets ยท bagging
97.4% acc
P: 96.4% ยท R: 98.3% ยท F1: 97.4%
๐Ÿ“
KNN (k=5)
Euclidean distance ยท normalized space
95.7% acc
P: 93.1% ยท R: 98.7% ยท F1: 95.8%
Hard Voting โ†’ Majority class wins
Final Output Hybrid (RF + KNN) 97.5% accuracy ยท 97.5% F1
All 5 Models โ€” Final Performance Comparison
Including Hybrid โ€” best performer across all metrics
Why does combining models help?

RF is strong at precision โ€” it avoids false alarms by being conservative. KNN is strong at recall โ€” it catches almost all true failures by considering local feature space proximity. When RF misclassifies a borderline failure case, KNN's proximity-based vote often corrects it, and vice versa. The combination produces a classifier that is simultaneously more precise and more sensitive than either model alone โ€” achieving the best F1 score of 97.5%.

Mission Accomplished

We built a complete end-to-end predictive maintenance pipeline โ€” from raw imbalanced sensor data to a high-performance failure detection system โ€” across five systematic stages.

๐Ÿ† Best Model: Hybrid (RF + KNN) โ€” sklearn VotingClassifier
97.5%
Accuracy
97.3%
Precision
97.6%
Recall
97.5%
F1 Score
๐Ÿ“Š
Raw Data
10,000 ยท 96% imbalanced
โ†’
โš–
SMOTE
19,322 ยท 50:50 balanced
โ†’
๐Ÿ“
Normalize
Mean=0, Std=1
โ†’
๐Ÿค–
Train ร— 5 Models
80/20 stratified split
โ†’
๐Ÿ†
Hybrid Model
97.5% F1 Score
๐ŸŽฏ
High Recall (97.6%)

The system correctly identifies nearly all actual machine failures โ€” minimising undetected breakdowns, catastrophic damage, and production halts.

โšก
High Precision (97.3%)

Very few false alarms โ€” maintenance teams are not overwhelmed with unnecessary work orders, keeping operations efficient and cost-effective.

๐Ÿ“ˆ
Real-world Impact

Predictive maintenance reduces unplanned downtime by up to 50% and maintenance costs by up to 25% compared to reactive approaches in real industrial deployments.

Machine Health Tester

Analyze real-time sensor data using our proprietary **Hybrid Ensemble Model**. This diagnostic tool provides the final predictive output based on aggregated model intelligence.

hybrid_ensemble.v1
๐Ÿ“ก

Awaiting input data...