Estimating Sufficient Amount of Training Data for Data-Driven Fault Diagnosis Learning Curve Extrapolation and Entropy-Based Stopping Criteria
##plugins.themes.bootstrap3.article.main##
##plugins.themes.bootstrap3.article.sidebar##
Fabian Mauthe
Peter Zeiler
Abstract
Data-driven approaches for diagnosing and predicting the health of engineering systems require sufficient training data to achieve reliable performance. However, despite its high practical relevance, the question of how much data are actually required has not yet been systematically investigated.
This paper presents and evaluates two complementary approaches to answering this question. Approach I fits parametric functions to the initial segments of learning curves and extrapolates them to estimate the amount of training data required to achieve a target classification accuracy. Approach II introduces a training-free, entropy-based stopping criterion that uses the gradient of normalized entropy (GNE). The information gain from the data collected over time is monitored to determine whether further data acquisition is necessary. Both approaches are evaluated using a dataset for fault classification in rolling bearings, employing a multi-layer perceptron classifier and randomly drawn, in-distribution test sets. The results show that learning curve extrapolation enables predictions of the required amount of data and that GNE can serve as an effective online stopping rule during data acquisition due to the strong correlation between GNE and classification accuracy. The two approaches complement each other: entropy analysis can monitor the progress of data acquisition in real time without model training, while learning curve extrapolation provides quantitative estimates of the remaining data requirements for specific accuracy targets.
##plugins.themes.bootstrap3.article.details##
entropy, fault diagnosis, learning curve, PHM, prognostics and health management, required training data, rolling bearing, training data evaluation
Bachoc, F., Gamboa, F., Halford, M., Loubes, J.-M., & Risser, L. (2023). Explaining machine learning models using entropic variable projection. Information and Inference: A Journal of the IMA, 12(3), 1686–1715. doi: 10.1093/imaiai/iaad010
Braig, M., & Zeiler, P. (2025). A study on using transfer learning to utilize information from similar systems for data-driven condition diagnosis and prognosis. IEEE ACCESS, 13, 98485–98503. doi: 10.1109/ACCESS.2025.3576435
Chen, J., Zhou, D., Guo, Z., Lin, J., Lyu, C., & Lu, C. (2019). An active learning method based on uncertainty and complexity for gearbox fault diagnosis. IEEE ACCESS, 7, 9022–9031. doi: 10.1109/ACCESS.2019.2890979
Crowther, P. S., & Cox, R. J. (2006). Accuracy of neural network classifiers as a property of the size of the data set. In B. Gabrys, R. J. Howlett, & L. C. Jain (Eds.), Knowledge-based intelligent information and engineering systems (pp. 1143–1149). Berlin, Heidelberg: Springer. doi: 10.1007/11893011 144
Estepa, R., Diaz-Verdejo, J. E., Estepa, A., & Madinabeitia, G. (2020). How much training data is enough? a case study for http anomaly-based intrusion detection. IEEE ACCESS, 8, 44410–44425. doi: 10.1109/ACCESS.2020.2977591
Figueroa, R. L., Zeng-Treitler, Q., Kandula, S., & Ngo, L. H. (2012). Predicting sample size required for classification performance. BMC Medical Informatics and Decision Making, 12(1), 8. doi: 10.1186/1472-6947-12-8
Fink, O., Nejjar, I., Sharma, V., Faghih Niresi, K., Sun, H., Dong, H., . . . Kesmen, Y. (2026). From physics to machine learning and back: Part ii - learning and observational bias in prognostics and health management (phm). Reliability Engineering & System Safety, 274, 112376. doi: 10.1016/j.ress.2026.112376
Guo, Y., Dai, J., & Zhang, J. (2025). Few-shot cross-domain fault diagnosis via adversarial meta-learning. SCIENTIFIC REPORTS, 15(1), 41876. doi: 10.1038/s41598-025-25854-z
Gwon, Y., Hwang, S., Kim, H., Ok, J., & Kwak, S. (2025). Enhancing cost efficiency in active learning with candidate set query. arXiv. doi: 10.48550/arXiv.2502.06209
Hagmeyer, S., Mauthe, F., & Zeiler, P. (2021). Creation of publicly available data sets for prognostics and diagnostics addressing data scenarios relevant to industrial applications. International Journal of Prognostics and Health Management, 12(2), 1–20. doi: 10.36001/ijphm.2021.v12i2.3087
Hogg, R. V., & Ledolter, J. (1989). Engineering statistics (Internat. ed. ed.). New York: MacMillan.
Jo, D. U., Yun, S., & Choi, J. Y. (2022). How much a model be trained by passive learning before active learning? IEEE ACCESS, 10, 34677–34689. doi: 10.1109/ACCESS.2022.3162253
Lei, Y., Yang, B., Jiang, X., Jia, F., Li, N., & Nandi, A. K. (2020). Applications of machine learning to machine fault diagnosis: A review and roadmap. Mechanical Systems and Signal Processing, 138, 106587. doi: 10.1016/j.ymssp.2019.106587
Li, C., Li, S., Feng, Y., Gryllias, K., Gu, F., & Pecht, M. (2024). Small data challenges for intelligent prognostics and health management: a review. Artificial Intelligence Review, 57(8). doi: 10.1007/s10462-024-10820-4
Mahmood, R., Lucas, J., Acuna, D., Li, D., Philion, J., Alvarez, J. M., . . . Law, M. T. (2022). How much more data do i need? estimating requirements for downstream tasks. In 2022 ieee/cvf conference on computer vision and pattern recognition (pp. 275–284). Piscataway, NJ: IEEE. doi: 10.1109/CVPR52688.2022.00037
Mahmood, R., Lucas, J., Alvarez, J. M., Fidler, S., & Law, M. T. (2022). Optimizing data collection for machine learning. In Proceedings of the 36th international conference on neural information processing systems. Red Hook, NY, USA: Curran Associates Inc.
Raj, K. K., Kumar, S., & Kumar, R. R. (2025). Systematic review of bearing component failure: Strategies for diagnosis and prognosis in rotating machinery. Arabian Journal for Science and Engineering, 50(8), 5353–5375. doi: 10.1007/s13369-024-09866-x
Wang, B., Lei, Y., Li, N., & Li, N. (2020). A hybrid prognostics approach for estimating remaining useful life of rolling element bearings. IEEE Transactions on Reliability, 69(1), 401–412. doi: 10.1109/TR.2018.2882682
Xie, Y., Ding, M., Tomizuka, M., & Zhan, W. (2023). Towards free data selection with general-purpose models. Advances in Neural Information Processing Systems, 36, 1309–1325.
Zio, E. (2022). Prognostics and health management (phm): Where are we and where do we (need to) go in theory and practice. Reliability Engineering & System Safety, 218, 108119. doi: 10.1016/j.ress.2021.108119
https://orcid.org/0000-0003-0737-8025