Disentangling Prognostic Observability from Policy Stochasticity in PHM-Aware Reinforcement Learning for Semiconductor Fab Dispatch

##plugins.themes.bootstrap3.article.main##

##plugins.themes.bootstrap3.article.sidebar##

Published Sep 28, 2026
Sean Mondesire
Tori Wright Bulent Soykan

Abstract

High-mix low-volume (HMLV) semiconductor manufacturing is a demanding proving ground for prognostics and health management (PHM): frequent product changeovers, re-entrant routes, and heavy tool-qualification overhead amplify the operational cost of unexpected failures, so healthaware dispatch justify its sensing investment. This work investigates that promise in a reliability-aware HMLV fab simulator using a 2×2 ablation of two deep reinforcement learning proximal policy optimization (PPO) dispatchers (PHMaware vs. dispatch-only) and two deployment modes (deterministic argmax vs. stochastic sampling), all sharing reward, training budget, and seeds. Across 10 seeds and 10 heldout HMLV order sets, deployment mode dominated: stochastic PHM-aware PPO cut mean makespan from 48,939.9 to 24,387.2 minutes and lifted OEE from 0.707 to 0.829 (all FDR-significant), while PHM-aware and dispatch-only PPO were statistically indistinguishable under matched deployment. The result sharpens the PHM-awareness case for HMLV fabs: prognostic value is real but contingent on how learned policies turn information into action.

How to Cite

Mondesire, S., Wright, T., & Soykan, B. (2026). Disentangling Prognostic Observability from Policy Stochasticity in PHM-Aware Reinforcement Learning for Semiconductor Fab Dispatch. Annual Conference of the PHM Society, 18(1). https://doi.org/10.36001/phmconf.2026.v18i1.4907
Abstract 0 | PDF Downloads 0

##plugins.themes.bootstrap3.article.details##

Keywords

semiconductor manufacturing, reinforcement learning, policy stochasticity, prognostic health, deep learning

References
Andrychowicz, M., Raichuk, A., Sta´nczyk, P., Orsini, M.,
Girgin, S., Marinier, R., . . . Bachem, O. (2021). What
matters in on-policy reinforcement learning? a largescale
empirical study. Retrieved from https://
arxiv.org/abs/2006.05990 doi: 10.48550/
arXiv.2006.05990
Azevedo, R., Amon, M. J., Anderson, M., Mondesire, S.,
Sanz, F.-G., Sottilare, R., & Wiedbusch, M. (2024).
Human digital twins to support nurse practitioner’s
clinical decision-making using multimodal data: A theoretical,
methodological, and analytical framework. In
Digital twin (pp. 149–172). Springer Nature. Retrieved
from https://doi.org/10.1007/978
-3-031-67778-6 7 doi: 10.1007/978-3-031
-67778-6 7
Brandimarte, P. (1993). Routing and scheduling in a flexible
job shop by tabu search. Annals of Operations Research,
41(3), 157–183. Retrieved from https://
doi.org/10.1007/BF02023073 doi: 10.1007/
BF02023073
Carvalho, T. P., Soares, F. A. A. M. N., Vita, R., Francisco,
R. d. P., Basto, J. P., & Alcal´a, S. G. S.
(2019). A systematic literature review of machine
learning methods applied to predictive maintenance.
Computers & Industrial Engineering, 137, 106024.
Retrieved from https://doi.org/10.1016/j
.cie.2019.106024 doi: 10.1016/j.cie.2019
.106024
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos,
F., Rudolph, L., & Madry, A. (2020). Implementation
matters in deep policy gradients: A case study
on PPO and TRPO. In International conference on
learning representations. Retrieved from https://
openreview.net/forum?id=r1etN1rtPB
Fuller, A., Fan, Z., Day, C., & Barlow, C. (2020).
Digital twin: Enabling technologies, challenges
and open research. IEEE Access, 8, 108952–
108971. Retrieved from https://doi.org/10
.1109/ACCESS.2020.2998358 doi: 10.1109/
ACCESS.2020.2998358
Google. (2025). OR-Tools: CP-SAT solver. Retrieved
from https://developers.google.com/
optimization/cp/cp solver
Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S.
(2018). Soft actor-critic: Off-policy maximum entropy
deep reinforcement learning with a stochastic
actor. In Proceedings of the 35th international
conference on machine learning (Vol. 80, pp. 1861–
1870). Retrieved from https://proceedings
.mlr.press/v80/haarnoja18b.html
Hatfield, C., & Mondesire, S. C. (2026). Distributed proximal
policy optimization for wafer-lot dispatching in semiconductor
manufacturing digital twins with 2D3PO. In
Proceedings of the winter simulation conference (wsc).
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup,
D., & Meger, D. (2018). Deep reinforcement
learning that matters. In Proceedings of the aaai
conference on artificial intelligence (Vol. 32). Retrieved
from https://doi.org/10.1609/aaai
.v32i1.11694 doi: 10.1609/aaai.v32i1.11694
Huang, S., Dossa, R. F. J., Raffin, A., Kanervisto, A., &
Wang, W. (2022). The 37 implementation details
of proximal policy optimization. Retrieved from
https://iclr-blog-track.github.io/
2022/03/25/ppo-implementation
-details/
Immordino, A., St¨ockermann, P., Hayen, N., Altenm¨uller,
T., Susto, G. A., Gebser, M., . . . Seidel, G.
(2025). Explainable AI for reinforcement learning
based dynamic scheduling solutions in semiconductor
manufacturing. Journal of Intelligent Manufacturing.
Retrieved from https://doi.org/10
.1007/s10845-025-02631-3 doi: 10.1007/
s10845-025-02631-3
Jones, D., Snider, C., Nassehi, A., Yon, J., & Hicks,
B. (2020). Characterising the digital twin: A
systematic literature review. CIRP Journal of
Manufacturing Science and Technology, 29, 36–52.
Retrieved from https://doi.org/10.1016/j
.cirpj.2020.02.002 doi: 10.1016/j.cirpj.2020
.02.002
Kopp, D., Hassoun, M., Kalir, A., & M¨onch, L. (2020).
SMT2020: A semiconductor manufacturing testbed.
IEEE Transactions on Semiconductor Manufacturing,
33(4), 522–531. Retrieved from https://doi
.org/10.1109/TSM.2020.3001933 doi: 10
.1109/TSM.2020.3001933
Kritzinger, W., Karner, M., Traar, G., Henjes, J., &
Sihn, W. (2018). Digital twin in manufacturing:
A categorical literature review and classification.
In Ifac-papersonline (Vol. 51, pp. 1016–1022).
Retrieved from https://doi.org/10.1016/j
.ifacol.2018.08.474 doi: 10.1016/j.ifacol
.2018.08.474
Lee, J., Bagheri, B., & Kao, H.-A. (2015). A cyber-physical
systems architecture for Industry 4.0-based manufacturing
systems. Manufacturing Letters, 3, 18–23.
Retrieved from https://doi.org/10.1016/j
.mfglet.2014.12.001 doi: 10.1016/j.mfglet
.2014.12.001
Lei, Y., Li, N., Guo, L., Li, N., Yan, T., & Lin, J. (2018).
12
ANNUAL CONFERENCE OF THE PROGNOSTICS AND HEALTH MANAGEMENT SOCIETY 2026
Machinery health prognostics: A systematic review
from data acquisition to RUL prediction. Mechanical
Systems and Signal Processing, 104, 799–834.
Retrieved from https://doi.org/10.1016/j
.ymssp.2017.11.016 doi: 10.1016/j.ymssp
.2017.11.016
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness,
J., Bellemare, M. G., . . . Hassabis, D. (2015). Humanlevel
control through deep reinforcement learning. Nature,
518, 529–533. Retrieved from https://doi
.org/10.1038/nature14236 doi: 10.1038/
nature14236
Mondesire, S. C., Angelopoulou, A., Sirigampola, S., &
Goldiez, B. (2019). Combining virtualization and containerization
to support interactive games and simulations
on the cloud. Simulation Modelling Practice and
Theory, 93, 233–244. Retrieved from https://doi
.org/10.1016/j.simpat.2018.08.005 doi:
10.1016/j.simpat.2018.08.005
Mondesire, S. C., Brown, M., & Soykan, B. (2025). Fabsim:
A micro-discrete event simulator for machine
learning dynamic job shop scheduling optimizers. In
Proceedings of the winter simulation conference. Retrieved
from https://www.informs-sim.org/
wsc25papers/inv186.pdf
Mondesire, S. C., & Soykan, B. (2026a). Automated wafer
dispatching through greedy multi-heuristic scheduling
in high-mix low-volume wafer fabs. In Advanced semiconductor
manufacturing conference (asmc).
Mondesire, S. C., & Soykan, B. (2026b). Real-time diffusion
scheduling for job shop optimization. In Proceedings
of the winter simulation conference (wsc).
Mondesire, S. C., & Wiegand, R. P. (2023). Mitigating
catastrophic forgetting with complementary
layered learning. Electronics, 12(3),
706. Retrieved from https://doi.org/
10.3390/electronics12030706 doi:
10.3390/electronics12030706
Mondesire, S. C.,&Wright, T. (2025). A tool grouping-based
approach to remaining useful life prediction for predictive
maintenance in semiconductor manufacturing.
In Advanced semiconductor manufacturing conference
(asmc): Big data management and machine learning.
Nsiye, E., Wright, T., Tse, B., & Mondesire, S. C.
(2024). A micro-discrete event simulation environment
for production scheduling in manufacturing
digital twins. In Modsim world. Retrieved
from https://www.modsimworld.org/
papers/2024/MODSIM 2024 paper 27.pdf
Park, J., Chun, J., Kim, S. H., Kim, Y., & Park, J. (2021).
Learning to schedule job-shop problems: Representation
and policy learning using graph neural network
and reinforcement learning. Retrieved from
https://arxiv.org/abs/2106.01086 doi:
10.48550/arXiv.2106.01086
PHM Society. (2016). 2016 PHM data challenge: Chemical
mechanical polishing data set. Retrieved from
https://phmsociety.org/conference/
annual-conference-of-the-phm
-society/annual-conference-of-the
-prognostics-and-health-management
-society-2016/phm-data-challenge-4/
PHM Society. (2018). 2018 PHM data challenge:
Ion mill etch tool fault and remaining useful life
data. Retrieved from https://c3.ndc.nasa
.gov/dashlink/resources/1009/
Pinedo, M. L. (2016). Scheduling: Theory, algorithms,
and systems (5th ed.). Springer. Retrieved
from https://doi.org/10.1007/978
-3-319-26580-3 doi: 10.1007/978-3-319-26580
-3
Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus,
M., & Dormann, N. (2021). Stable-baselines3: Reliable
reinforcement learning implementations. Journal
of Machine Learning Research, 22(268), 1–8. Retrieved
from https://www.jmlr.org/papers/
v22/20-1364.html
Sabri, S., Aghaabbasi, M., Reay Atkinson, S., Amon, M. J.,
Hancock, P., Azevedo, R., . . . Rabadi, G. (2025).
Integrating human-machine systems and digital twin
technologies: Navigating trust, interoperability, and
ethical challenges. Retrieved from https://doi
.org/10.2139/ssrn.5189676 doi: 10.2139/
ssrn.5189676
Sargent, R. G. (2013). Verification and validation of simulation
models. Journal of Simulation, 7(1), 12–24. Retrieved
from https://doi.org/10.1057/jos
.2012.20 doi: 10.1057/jos.2012.20
Schreck, J., Matthews, G., Lin, J., Mondesire, S., Metcalf, D.,
Dickerson, K., & Grasso, J. (2025). Levels of automation
for a computer-based procedure for simulated nuclear
power plant operation: Impacts on workload and
trust. Safety, 11(1), 22. Retrieved from https://
doi.org/10.3390/safety11010022 doi: 10
.3390/safety11010022
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., &
Klimov, O. (2017). Proximal policy optimization
algorithms. Retrieved from https://arxiv
.org/abs/1707.06347 doi: 10.48550/arXiv
.1707.06347
Sharp, M., & Weiss, B. A. (2018). Hierarchical modeling
of a manufacturing work cell to promote contextualized
PHM information across multiple levels.
Manufacturing Letters, 15, 46–49. Retrieved
from https://doi.org/10.1016/j.mfglet
.2018.02.003 doi: 10.1016/j.mfglet.2018.02.003
Si, X.-S., Wang, W., Hu, C.-H., & Zhou, D.-H. (2011).
Remaining useful life estimation: A review on
13
ANNUAL CONFERENCE OF THE PROGNOSTICS AND HEALTH MANAGEMENT SOCIETY 2026
the statistical data driven approaches. European
Journal of Operational Research, 213(1), 1–14.
Retrieved from https://doi.org/10.1016/j
.ejor.2010.11.018 doi: 10.1016/j.ejor.2010.11
.018
Singh, K., Selvanathan, B., Zope, K., Nistala, S. H., &
Runkana, V. (2018). Concurrent estimation of
remaining useful life for multiple faults in an ion
etch mill: A data-driven approach. In Annual conference
of the phm society (Vol. 10). Retrieved
from https://doi.org/10.36001/phmconf
.2018.v10i1.591 doi: 10.36001/phmconf.2018
.v10i1.591
Soykan, B., Mondesire, S. C., & Rabadi, G. (2026a). Realtime
dispatching in semiconductor manufacturing: A
hybrid graph deep reinforcement learning and discreteevent
simulation approach. In Proceedings of the winter
simulation conference (wsc).
Soykan, B., Mondesire, S. C., & Rabadi, G. (2026b). A supply
chain digital twin architecture for semiconductor
industry. In Optimizing supply chains through digital
twins: Navigating the modern landscape. Springer Nature.
Soykan, B., Mondesire, S. C., Rabadi, G., & Bochenek,
G. (2025). Graph-enhanced deep reinforcement
learning for multi-objective unrelated parallel
machine scheduling. In Proceedings of the winter
simulation conference (pp. 2515–2526). Retrieved
from https://doi.org/10.1109/WSC68292
.2025.11338986 doi: 10.1109/WSC68292.2025
.11338986
St¨ockermann, P., Feudel, S., Immordino, A., Hayen,
N., Altenm¨uller, T., Gebser, M., . . . Higgins, F.
(2025). Reinforcement learning based dispatching
solutions in semiconductor manufacturing: A literature
review on validation and deployment. Production
and Manufacturing Research, 13(1). Retrieved
from https://doi.org/10.1080/21693277
.2025.2582472 doi: 10.1080/21693277.2025
.2582472
St¨ockermann, P., Immordino, A., Altenm¨uller, T., Seidel,
G., Gebser, M., Tassel, P., . . . Zhang, F.
(2023). Dispatching in real frontend fabs with industrial
grade discrete-event simulations by deep reinforcement
learning with evolution strategies. In
Proceedings of the winter simulation conference (pp.
3047–3058). Retrieved from https://doi.org/
10.1109/WSC60868.2023.10408625 doi: 10
.1109/WSC60868.2023.10408625
St¨ockermann, P., S¨udfeld, H., Immordino, A., Altenm¨uller,
T., Wegmann, M., Gebser, M., . . . Zhang, F. F.
(2025). Scalability of reinforcement learning methods
for dispatching in semiconductor frontend fabs:
A comparison of open-source models with real
industry datasets. The International Journal of
Advanced Manufacturing Technology, 139, 4395–
4415. Retrieved from https://doi.org/10
.1007/s00170-025-16117-2 doi: 10.1007/
s00170-025-16117-2
Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning:
An introduction (2nd ed.). MIT Press. Retrieved
from http://incompleteideas.net/
book/the-book-2nd.html
Tao, F., Xiao, B., Qi, Q., Cheng, J., & Ji, P. (2022). Digital
twin modeling. Journal of Manufacturing Systems,
64, 372–389. Retrieved from https://doi.org/
10.1016/j.jmsy.2022.06.015 doi: 10.1016/
j.jmsy.2022.06.015
Tassel, P., Gebser, M., & Schekotihin, K. (2021). A reinforcement
learning environment for job-shop scheduling.
Retrieved from https://arxiv.org/abs/
2104.03760 doi: 10.48550/arXiv.2104.03760
Tassel, P., Kov´acs, B., Gebser, M., Schekotihin, K., St¨ockermann,
P., & Seidel, G. (2023). Semiconductor
fab scheduling with self-supervised and reinforcement
learning. In Proceedings of the winter
simulation conference (pp. 1924–1935). Retrieved
from https://doi.org/10.1109/WSC60868
.2023.10407747 doi: 10.1109/WSC60868.2023
.10407747
Towers, M., Kwiatkowski, A., Terry, J. K., Balis, J. U.,
De Cola, G., Deleu, T., . . . Younis, O. G. (2024).
Gymnasium: A standard interface for reinforcement
learning environments. Retrieved from https://
arxiv.org/abs/2407.17032 doi: 10.48550/
arXiv.2407.17032
Wang, B., Zhou, H., Yang, G., Li, X., & Yang, H.
(2022). Human digital twin driven human-cyberphysical
systems: Key technologies and applications.
Chinese Journal of Mechanical Engineering,
35(11). Retrieved from https://doi.org/10
.1186/s10033-022-00680-w doi: 10.1186/
s10033-022-00680-w
Wright, T., Nsiye, E., & Mondesire, S. C. (2024). Generalized
vs. individualized kaplan-meier time-to-failure
predictive models for preventative maintenance in
semiconductor manufacturing. In International symposium
on semiconductor manufacturing (issm).
Wright, T., Soykan, B., & Mondesire, S. C. (2025). Real-time
remaining useful life prediction with LSTM priority resampling.
In Ieee international conference on machine
learning and applications (icmla).
Zhang, C., Song, W., Cao, Z., Zhang, J., Tan, P. S., &
Xu, C. (2020). Learning to dispatch for job shop
scheduling via deep reinforcement learning. In
Advances in neural information processing systems.
Retrieved from https://proceedings
.neurips.cc/paper/2020/hash/
Section
Technical Research Papers