Render Free Tier is spinning up the server. This usually takes 30-60 seconds.
In industrial environments, unexpected machine failures lead to significant financial losses, production downtime, and safety risks. Traditional maintenance is either too early (wasteful) or too late (costly).
This project builds a data-driven predictive system that analyses real machine sensor readings to forecast failures before they happen โ enabling timely, targeted maintenance.
The AI4I Predictive Maintenance Dataset contains 10,000 operational snapshots from a simulated milling machine. Each row captures the machine's state at one point in time, labelled as failure or no-failure.
Before applying any ML technique, we performed a thorough EDA on the original dataset to understand the shape of each feature, detect outliers, and discover relationships between variables.
Rotational speed has a right skew โ most machines operate in the 1200โ1800 rpm range but a few run much faster. Torque shows high-end outliers beyond 65 Nm, which often correspond to overstrain failures. Tool wear is near-uniform since it increments linearly with usage time. These outliers are intentionally kept โ they carry critical failure signal and removing them would reduce model sensitivity.
Air Temp โ Process Temp (r = 0.88): Strong positive correlation. When ambient temperature
rises, internal machine temperature follows โ increasing thermal stress risk.
Rotational Speed โ Torque (r = โ0.88): Strong negative correlation. This is mechanical law
โ power = speed ร torque. Extreme combinations (very low speed + very high torque) trigger overstrain
failures.
Tool Wear โ Target (r = 0.11): Weak but real. Tool wear amplifies other failure modes
rather than causing failures on its own.
During EDA, we discovered that the dataset is severely imbalanced. This is one of the most dangerous problems in real-world ML classification tasks.
A model that always predicts "No Failure" achieves 96.6% accuracy without learning anything. This is why we must use F1-score and recall as primary metrics โ not accuracy.
In predictive maintenance, a missed failure (false negative) costs far more than a false alarm. An undetected failure can cause catastrophic equipment damage, production halt, or even worker injury. A biased model optimising for accuracy is useless โ it would miss 100% of failures. We must fix the imbalance before training.
Synthetic Minority Oversampling Technique (SMOTE) generates new synthetic
samples for the minority class by interpolating between existing examples โ preserving distribution while
balancing classes. We used random_state=42 for reproducibility.
Feature distributions remain structurally similar before and after SMOTE โ confirming that synthetic samples follow the same statistical patterns as real data. The distributions overlap closely, meaning the synthetic failure samples are realistic interpolations rather than noise. This is the key advantage of SMOTE over simple oversampling (duplication), which would cause severe overfitting.
After SMOTE, the dataset grew from 10,000 to 19,322 samples (9,661 per class). We ran a fresh EDA pass to confirm the balance and verify statistical integrity.
The correlation matrix is nearly identical to the original dataset. Air-Process temperature correlation remains at 0.86, and the Speed-Torque negative correlation stays at โ0.85. This is critical evidence that SMOTE generated statistically consistent samples โ maintaining the underlying mechanical relationships of the machine rather than introducing artificial patterns.
Features exist at vastly different numerical scales. StandardScaler transforms each feature to zero mean and unit standard deviation so no feature dominates others due to magnitude alone.
KNN computes similarity using Euclidean distance. Without normalization, rotational speed
(range ~1700) dominates over torque (range ~73) purely due to scale. After StandardScaler, all five features
contribute equally to distance calculations.
Logistic Regression uses gradient descent to fit coefficients โ unnormalized features cause
slow convergence and biased weights.
Random Forest and Decision Tree are scale-invariant by nature but normalization ensures
consistent preprocessing across the entire pipeline.
We trained four base classifiers with a stratified 80/20 train-test split (random_state=42). Stratification ensures both train and test sets maintain the balanced 50/50 class ratio.
Linear decision boundary fails to capture non-linear interactions between features like speed-torque combinations. Useful as a baseline but insufficient for complex failure patterns in sensor data.
Finds 5 nearest neighbors in normalized feature space. Achieves the highest recall (98.7%) โ misses very few failures. Lower precision (93.1%) means more false alarms. Performance entirely depends on normalization.
Learns explicit IF-THEN rules from feature thresholds. Highly interpretable. Slightly prone to overfitting which the Random Forest ensemble approach corrects. Good balance between precision and recall.
Builds 100 trees on random feature subsets, aggregates votes. Best balance of precision (96.4%) and recall (98.3%) among base models. Ensemble approach reduces variance and avoids overfitting compared to a single decision tree.
Deep-dive evaluation of every model. The confusion matrix shows exactly what each model got right and wrong. The ROC curve shows performance across all decision thresholds โ AUC closer to 1.0 is better.
The Matrix breaks down the exact number of correct and incorrect predictions. The goal is to maximize the Diagonal (Blue/Green) and minimize the others.
The red area represents the model's ability to distinguish between classes. A larger area (AUC) means the model is better at catching failures without raising too many false alarms.
The diagonal dashed line is a random classifier (AUC = 0.5). Any model above this line is better than guessing. The closer the curve hugs the top-left corner, the better. AUC = 1.0 is perfect. Our Hybrid achieves the highest AUC โ confirming it as the strongest model at every threshold.
We combined the top two models using a Hard Voting Classifier (sklearn VotingClassifier). Each model independently predicts the class and the majority vote determines the final output โ reducing the individual weaknesses of each model.
RF is strong at precision โ it avoids false alarms by being conservative. KNN is strong at recall โ it catches almost all true failures by considering local feature space proximity. When RF misclassifies a borderline failure case, KNN's proximity-based vote often corrects it, and vice versa. The combination produces a classifier that is simultaneously more precise and more sensitive than either model alone โ achieving the best F1 score of 97.5%.
We built a complete end-to-end predictive maintenance pipeline โ from raw imbalanced sensor data to a high-performance failure detection system โ across five systematic stages.
The system correctly identifies nearly all actual machine failures โ minimising undetected breakdowns, catastrophic damage, and production halts.
Very few false alarms โ maintenance teams are not overwhelmed with unnecessary work orders, keeping operations efficient and cost-effective.
Predictive maintenance reduces unplanned downtime by up to 50% and maintenance costs by up to 25% compared to reactive approaches in real industrial deployments.
Analyze real-time sensor data using our proprietary **Hybrid Ensemble Model**. This diagnostic tool provides the final predictive output based on aggregated model intelligence.
Awaiting input data...