Planning or Learning: Reliability and Cost in Multi-Asset Maintenance

##plugins.themes.bootstrap3.article.main##

##plugins.themes.bootstrap3.article.sidebar##

Published Sep 28, 2026
Xian Yeow Lee Chandrasekar Venkatraman Ahmed Farahat

Abstract

Industrial maintenance systems increasingly involve multiple interacting assets and shared resources, making it challenging to balance reliability and operational cost using a single decision framework. While recent work has focused on reinforcement learning (RL) for maintenance scheduling, direct comparisons with planning approaches under identical settings remain limited. In this work, we empirically compare planning and RL for multi-asset bearing maintenance using run-to-failure data. We examine how these methods behave when balancing preventive maintenance against tolerable failures across a range of failure penalty scenarios. We observed a consistent behavioral difference driven by objective formulation. Planning enforces reliability as a hard constraint and produces zero-failure policies whose total cost is largely insensitive to the magnitude of failure penalties. RL agents optimize expected cost and often trade off preventive maintenance against occasional failures as penalties vary, resulting in lower costs under low-penalty regimes but persistent nonzero failures even when penalties are high. We also investigate lightweight constraint mechanisms, including reward shaping and action masking, to encourage RL’s reliability. From a practical perspective, planning may be more suitable when strict reliability is required and deployment horizons are short, whereas RL may provide cost-efficient policie when limited failures are acceptable and long-run operational efficiency is prioritized. Overall, this study clarifies the tradeoffs between reliability and cost in multi-asset maintenance, highlights the challenges of enforcing zero-failure behavior in RL, and suggests that planning and RL are complementary approaches whose applicability depends on operational objectives. Beyond these findings, the controlled benchmark protocol itself that unifies environment, cost model, and evaluation across paradigms, offers a reusable template for comparing decision-making approaches in other maintenance settings.

How to Cite

Lee, X. Y., Venkatraman, C. ., & Farahat, A. . (2026). Planning or Learning: Reliability and Cost in Multi-Asset Maintenance. Annual Conference of the PHM Society, 18(1). https://doi.org/10.36001/phmconf.2026.v18i1.4787
Abstract 0 | PDF Downloads 0

##plugins.themes.bootstrap3.article.details##

Keywords

Reinforcement Learning, Planning, Remaining Useful Life, Maintenance Scheduling

References
Alshiekh, M., Bloem, R., Ehlers, R., K¨onighofer, B., Niekum,
S., & Topcu, U. (2018). Safe reinforcement learning
via shielding. In Proceedings of the aaai conference on
artificial intelligence (Vol. 32).
Altman, E. (2021). Constrained markov decision processes.
Routledge.
Andriotis, C. P., & Papakonstantinou, K. G. (2021). Deep
reinforcement learning driven inspection and maintenance
planning under incomplete information and constraints.
Reliability Engineering & System Safety, 212,
107551. doi: 10.1016/j.ress.2021.107551
Dhungana, H., Rykkje, T., & Lundervold, A. S. (2025). Bearing
prognostics using the pronostia data: a comparative
study. IEEE Access.
Dijkstra, E.W. (1959). A note on two problems in connexion
with graphs. Numerische mathematik, 1(1), 269–271.
Feng, M., & Li, Y. (2022). Predictive maintenance decision
making based on reinforcement learning in multistage
production systems. IEEE Access, 10, 18910–18921.
Hou, Y., Liang, X., Zhang, J., Yang, Q., Yang, A., & Wang, N. (2023). Exploring the use of invalid action masking
in reinforcement learning: A comparative study of onpolicy
and off-policy algorithms in real-time strategy
games. Applied Sciences, 13(14), 8283.
Hu, Y., Wang, W., Jia, H., Wang, Y., Chen, Y., Hao, J., . . .
Fan, C. (2020). Learning to utilize shaping rewards: A
new approach of reward shaping. Advances in Neural
Information Processing Systems, 33, 15931–15941.
Mazumdar, A., Wisniewski, R., & Bujorianu, M. L. (2024).
Safe reinforcement learning for constrained markov decision
processes with stochastic stopping time. arXiv
preprint arXiv:2403.15928.
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness,
J., Bellemare, M. G., . . . others (2015). Human-level
control through deep reinforcement learning. nature,
518(7540), 529–533.
Nectoux, P., Gouriveau, R., Medjaher, K., Ramasso, E.,
Chebel-Morello, B., Zerhouni, N., & Varnier, C.
(2012). Pronostia: An experimental platform for bearings
accelerated degradation tests. In Ieee international
conference on prognostics and health management,
phm’12. (pp. 1–8).
O’Malley, C., de Mars, P., Badesa, L., & Strbac, G. (2022).
Reinforcement learning and mixed-integer programming
for power plant scheduling in low carbon systems:
Comparison and hybridisation. arXiv preprint
arXiv:2212.04824.
Ong, K. S. H.,Wang,W., Niyato, D., & Friedrichs, T. (2021).
Deep-reinforcement-learning-based predictive maintenance
model for effective resource management in industrial
iot. IEEE Internet of Things Journal, 9(7),
5173–5188.
O’Neil, R., Khatab, A., & Diallo, C. (2025). Optimizing
predictive maintenance and mission assignment to enhance
fleet readiness under uncertainty. Autonomous
Intelligent Systems, 5(1), 17.
Petchrompo, S., & Parlikad, A. K. (2019). A review of asset
management literature on multi-asset systems. Reliability
Engineering & System Safety, 181, 181–201.
Pignatelli, E., Ferret, J., Geist, M., Mesnard, T., van Hasselt,
H., & Toni, L. (n.d.). A survey of temporal credit assignment
in deep reinforcement learning. Transactions
on Machine Learning Research.
Prashanth, L. A., & Fu, M. C. (2022). Risk-sensitive reinforcement
learning via policy gradient search. arXiv
preprint arXiv:1810.09126.
Sakib, N., & Wuest, T. (2018). Challenges and opportunities
of condition-based predictive maintenance: a review.
Procedia cirp, 78, 267–272.
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., &
Klimov, O. (2017). Proximal policy optimization algorithms.
arXiv preprint arXiv:1707.06347.
Siraskar, R., Kumar, S., Patil, S., Bongale, A., & Kotecha, K.
(2023). Reinforcement learning for predictive maintenance:
A systematic technical review. Artificial Intelligence
Review, 56(11), 12885–12947.
Wachi, A., & Sui, Y. (2020). Safe reinforcement learning
in constrained markov decision processes. In Proceedings
of the 37th international conference on machine
learning (icml) (Vol. 119, pp. 9797–9806). PMLR.
Yang, K., Yang, J., & Shen, C. (2024). Average reward reinforcement
learning for wireless radio resource management.
In Proceedings of the asilomar conference
on signals, systems, and computers. (arXiv preprint
arXiv:2501.06700)
Zhang, C., Li, Y.-F., & Coit, D. W. (2022). Deep reinforcement
learning for dynamic opportunistic maintenance
of multi-component systems with load sharing. IEEE
Transactions on Reliability, 72(3), 863–877.
Zhang, Y., Xiao, W., & Bi, Y. (2025). Integrated predictivemaintenance
and mpc scheduling: Achieving high
availability in smart manufacturing. IEEE Access.
Section
Technical Research Papers

Most read articles by the same author(s)

1 2 > >>