Sampling Strategies and Metric Sensitivity in Imbalanced Predictive Maintenance: A Comparative Study with Ensemble Methods

Authors

  • Berk Mercan Tekirdağ Namık Kemal University, Çorlu Faculty of Engineering, Electrical and Electronics Engineering Department, Çorlu, Tekirdağ, Türkiye https://orcid.org/0009-0006-2440-2372
  • Eren Ağar Tekirdağ Namık Kemal University, Çorlu Faculty of Engineering, Electrical and Electronics Engineering Department, Çorlu, Tekirdağ, Türkiye https://orcid.org/0009-0000-1619-6317
  • Mert Çimentepe Tekirdağ Namık Kemal University, Çorlu Faculty of Engineering, Electrical and Electronics Engineering Department, Çorlu, Tekirdağ, Türkiye https://orcid.org/0009-0002-7643-5855
  • Meltem Apaydın Üstün Tekirdağ Namık Kemal University, Çorlu Faculty of Engineering, Electrical and Electronics Engineering Department, Çorlu, Tekirdağ, Türkiye https://orcid.org/0000-0001-9225-9455

DOI:

https://doi.org/10.58190/ijamec.2026.179

Keywords:

Machine Learning, predictive maintenance, class imbalance, feature engineering, SHAP, random forest, XGBoost

Abstract

This study presents a systematic comparative analysis of class balancing strategies for predictive maintenance of industrial equipment using the AI4I 2020 Predictive Maintenance Dataset, which comprises 10,000 samples with a severe class imbalance of 3.4% failure instances. Two tree-based ensemble classifiers, Random Forest (RF) and XGBoost, are evaluated under nine balancing configurations: Baseline, class-weight balancing, SMOTE, ADASYN, Random Undersampling (RUS), SMOTE-Tomek, SMOTEENN, SMOTE+RUS, and SVMSMOTE. Four domain-informed engineered features, Delta_T, Power_Proxy, Temp_Ratio, and Torque_per_Wear, are derived from physical process knowledge and integrated into the modeling pipeline. Models are evaluated using Precision, Recall, F1-score, and PR-AUC, with particular emphasis on PR-AUC due to the severe class imbalance. The results show that RF is robust to class imbalance without complex resampling. Tuned Baseline RF achieves the best RF performance with an F1-score of 86.4% and a PR-AUC of 0.886, while Tuned Balanced RF provides a comparable F1-score of 85.7% and a PR-AUC of 0.875. In contrast, XGBoost benefits more from sampling strategy and hyperparameter optimization: Tuned SVMSMOTE + XGB achieves the strongest tuned XGBoost performance with an F1-score of 83.0% and the highest PR-AUC of 0.914. These findings demonstrate that sampling methods do not universally improve classifier performance; rather, their effectiveness depends on the learning algorithm. In addition, SHAP-based interpretability analysis shows that Rotational Speed, Torque, Tool Wear, and engineered load-related features contribute substantially to failure prediction. Overall, PR-AUC is recommended as a primary evaluation metric for imbalanced industrial predictive maintenance datasets.

Downloads

Download data is not yet available.

References

[1] I. Hector and R. Panjanathan, “Predictive maintenance in Industry 4.0: A survey of planning models and machine learning techniques,” PeerJ Computer Science, vol. 10, Art. no. e2016, 2024. doi: 10.7717/peerj-cs.2016

[2] T. P. Carvalho et al., “A systematic literature review of machine learning methods applied to predictive maintenance,” Computers & Industrial Engineering, vol. 137, Art. no. 106024, 2019. doi: 10.1016/j.cie.2019.106024

[3] S. Matzka, “AI4I 2020 predictive maintenance dataset,” UCI Machine Learning Repository, 2020. doi: 10.24432/C5HS5C

[4] S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Proc. 31st Int. Conf. Neural Information Processing Systems (NIPS), Long Beach, CA, USA, 2017, pp. 4765–4774. arXiv:1705.07874.

[5] O. E. Hassan et al., “Induction motor broken rotor bar fault detection techniques based on fault signature analysis – a review,” IET Electric Power Applications, vol. 12, no. 7, pp. 895–907, 2018. doi: 10.1049/iet-epa.2018.0054

[6] D. Neupane et al., “Data-driven machinery fault diagnosis: A comprehensive review,” Neurocomputing, vol. 627, Art. no. 129588, 2025. doi: 10.1016/j.neucom.2025.129588

[7] A. S. Kalafatelis et al., “An effective methodology for imbalanced data handling in predictive maintenance for offset printing,” in Proc. 11th Int. Conf. Mechatronics and Control Engineering (ICMCE 2023), Lecture Notes in Mechanical Engineering, Springer, Singapore, 2024, pp. 89–98. doi: 10.1007/978-981-99-6523-6_7

[8] A. Hakami, “Strategies for overcoming data scarcity, imbalance, and feature selection challenges in machine learning models for predictive maintenance,” Scientific Reports, vol. 14, Art. no. 9645, 2024. doi: 10.1038/s41598-024-59958-9

[9] K. Patel and A. Shanbhag, “Exploring ML for predictive maintenance using imbalance correction techniques and SHAP,” in Proc. 2022 International Conference on Electrical, Computer and Energy Technologies (ICECET), Prague, Czech Republic, 2022, pp. 1–10. doi: 10.1109/ICECET55527.2022.9873073

[10] S. H. H. Zaidi et al., “A systematic review of anomaly and fault detection using machine learning for industrial machinery,” Algorithms, vol. 19, no. 2, Art. no. 108, 2026. doi: 10.3390/a19020108

[11] A. Borré et al., “Machine fault detection using a hybrid CNN-LSTM attention-based model,” Sensors, vol. 23, no. 9, Art. no. 4512, 2023. doi: 10.3390/s23094512

[12] W. Li and T. Li, “Comparison of deep learning models for predictive maintenance in industrial manufacturing systems using sensor data,” Scientific Reports, vol. 15, no. 1, Art. no. 23545, 2025. doi: 10.1038/s41598-025-08515-z

[13] K. M. A. Alghtus et al., “Short-horizon predictive maintenance of industrial pumps using time-series features and machine learning,” arXiv preprint arXiv:2508.19974, 2025. doi: 10.48550/arXiv.2508.19974

[14] L. Breiman, “Random forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001. doi: 10.1023/A:1010933404324

[15] T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, San Francisco, CA, USA, 2016, pp. 785–794. doi: 10.1145/2939672.2939785

[16] F. Velasco-Loera, M. Alcaraz-Mejia, and J. L. Chavez-Hurtado, “An interpretable hybrid fault prediction framework using XGBoost and a probabilistic graphical model for predictive maintenance: A case study in textile manufacturing,” Applied Sciences, vol. 15, no. 18, Art. no. 10164, 2025. doi: 10.3390/app151810164

[17] Y. Altintas, Manufacturing Automation: Metal Cutting Mechanics, Machine Tool Vibrations, and CNC Design, 2nd ed. Cambridge, U.K.: Cambridge Univ. Press, 2012. doi: 10.1017/CBO9780511843723

[18] N. V. Chawla et al., “SMOTE: Synthetic minority over-sampling technique,” Journal of Artificial Intelligence Research, vol. 16, pp. 321–357, 2002. doi: 10.1613/jair.953.

[19] H. He et al., “ADASYN: Adaptive synthetic sampling approach for imbalanced learning,” in Proc. IEEE Int. Joint Conf. Neural Networks (IJCNN), Hong Kong, 2008, pp. 1322–1328. doi: 10.1109/IJCNN.2008.4633969.

[20] H. M. Nguyen et al., “Borderline over-sampling for imbalanced data classification,” International Journal of Knowledge Engineering and Soft Data Paradigms, vol. 3, no. 1, pp. 4–21, 2011. doi: 10.1504/IJKESDP.2011.039875

[21] G. Lemaitre et al., “Imbalanced-learn: A Python toolbox to tackle the curse of imbalanced datasets in machine learning,” Journal of Machine Learning Research, vol. 18, no. 17, pp. 1–5, 2017. [Online]. Available: http://jmlr.org/papers/v18/16-365.html.

Downloads

Published

30-09-2026

Issue

Section

Research Articles

How to Cite

[1]
B. Mercan, E. Ağar, M. . Çimentepe, and M. Apaydın Üstün, “Sampling Strategies and Metric Sensitivity in Imbalanced Predictive Maintenance: A Comparative Study with Ensemble Methods”, J. Appl. Methods Electron. Comput., vol. 14, no. 3, pp. 120–129, Sep. 2026, doi: 10.58190/ijamec.2026.179.

Similar Articles

161-170 of 179

You may also start an advanced similarity search for this article.