Disentangling Prognostic Observability from Policy Stochasticity in PHM-Aware Reinforcement Learning for Semiconductor Fab Dispatch
##plugins.themes.bootstrap3.article.main##
##plugins.themes.bootstrap3.article.sidebar##
Tori Wright Bulent Soykan
Abstract
High-mix low-volume (HMLV) semiconductor manufacturing is a demanding proving ground for prognostics and health management (PHM): frequent product changeovers, re-entrant routes, and heavy tool-qualification overhead amplify the operational cost of unexpected failures, so healthaware dispatch justify its sensing investment. This work investigates that promise in a reliability-aware HMLV fab simulator using a 2×2 ablation of two deep reinforcement learning proximal policy optimization (PPO) dispatchers (PHMaware vs. dispatch-only) and two deployment modes (deterministic argmax vs. stochastic sampling), all sharing reward, training budget, and seeds. Across 10 seeds and 10 heldout HMLV order sets, deployment mode dominated: stochastic PHM-aware PPO cut mean makespan from 48,939.9 to 24,387.2 minutes and lifted OEE from 0.707 to 0.829 (all FDR-significant), while PHM-aware and dispatch-only PPO were statistically indistinguishable under matched deployment. The result sharpens the PHM-awareness case for HMLV fabs: prognostic value is real but contingent on how learned policies turn information into action.
How to Cite
##plugins.themes.bootstrap3.article.details##
semiconductor manufacturing, reinforcement learning, policy stochasticity, prognostic health, deep learning
Girgin, S., Marinier, R., . . . Bachem, O. (2021). What
matters in on-policy reinforcement learning? a largescale
empirical study. Retrieved from https://
arxiv.org/abs/2006.05990 doi: 10.48550/
arXiv.2006.05990
Azevedo, R., Amon, M. J., Anderson, M., Mondesire, S.,
Sanz, F.-G., Sottilare, R., & Wiedbusch, M. (2024).
Human digital twins to support nurse practitioner’s
clinical decision-making using multimodal data: A theoretical,
methodological, and analytical framework. In
Digital twin (pp. 149–172). Springer Nature. Retrieved
from https://doi.org/10.1007/978
-3-031-67778-6 7 doi: 10.1007/978-3-031
-67778-6 7
Brandimarte, P. (1993). Routing and scheduling in a flexible
job shop by tabu search. Annals of Operations Research,
41(3), 157–183. Retrieved from https://
doi.org/10.1007/BF02023073 doi: 10.1007/
BF02023073
Carvalho, T. P., Soares, F. A. A. M. N., Vita, R., Francisco,
R. d. P., Basto, J. P., & Alcal´a, S. G. S.
(2019). A systematic literature review of machine
learning methods applied to predictive maintenance.
Computers & Industrial Engineering, 137, 106024.
Retrieved from https://doi.org/10.1016/j
.cie.2019.106024 doi: 10.1016/j.cie.2019
.106024
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos,
F., Rudolph, L., & Madry, A. (2020). Implementation
matters in deep policy gradients: A case study
on PPO and TRPO. In International conference on
learning representations. Retrieved from https://
openreview.net/forum?id=r1etN1rtPB
Fuller, A., Fan, Z., Day, C., & Barlow, C. (2020).
Digital twin: Enabling technologies, challenges
and open research. IEEE Access, 8, 108952–
108971. Retrieved from https://doi.org/10
.1109/ACCESS.2020.2998358 doi: 10.1109/
ACCESS.2020.2998358
Google. (2025). OR-Tools: CP-SAT solver. Retrieved
from https://developers.google.com/
optimization/cp/cp solver
Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S.
(2018). Soft actor-critic: Off-policy maximum entropy
deep reinforcement learning with a stochastic
actor. In Proceedings of the 35th international
conference on machine learning (Vol. 80, pp. 1861–
1870). Retrieved from https://proceedings
.mlr.press/v80/haarnoja18b.html
Hatfield, C., & Mondesire, S. C. (2026). Distributed proximal
policy optimization for wafer-lot dispatching in semiconductor
manufacturing digital twins with 2D3PO. In
Proceedings of the winter simulation conference (wsc).
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup,
D., & Meger, D. (2018). Deep reinforcement
learning that matters. In Proceedings of the aaai
conference on artificial intelligence (Vol. 32). Retrieved
from https://doi.org/10.1609/aaai
.v32i1.11694 doi: 10.1609/aaai.v32i1.11694
Huang, S., Dossa, R. F. J., Raffin, A., Kanervisto, A., &
Wang, W. (2022). The 37 implementation details
of proximal policy optimization. Retrieved from
https://iclr-blog-track.github.io/
2022/03/25/ppo-implementation
-details/
Immordino, A., St¨ockermann, P., Hayen, N., Altenm¨uller,
T., Susto, G. A., Gebser, M., . . . Seidel, G.
(2025). Explainable AI for reinforcement learning
based dynamic scheduling solutions in semiconductor
manufacturing. Journal of Intelligent Manufacturing.
Retrieved from https://doi.org/10
.1007/s10845-025-02631-3 doi: 10.1007/
s10845-025-02631-3
Jones, D., Snider, C., Nassehi, A., Yon, J., & Hicks,
B. (2020). Characterising the digital twin: A
systematic literature review. CIRP Journal of
Manufacturing Science and Technology, 29, 36–52.
Retrieved from https://doi.org/10.1016/j
.cirpj.2020.02.002 doi: 10.1016/j.cirpj.2020
.02.002
Kopp, D., Hassoun, M., Kalir, A., & M¨onch, L. (2020).
SMT2020: A semiconductor manufacturing testbed.
IEEE Transactions on Semiconductor Manufacturing,
33(4), 522–531. Retrieved from https://doi
.org/10.1109/TSM.2020.3001933 doi: 10
.1109/TSM.2020.3001933
Kritzinger, W., Karner, M., Traar, G., Henjes, J., &
Sihn, W. (2018). Digital twin in manufacturing:
A categorical literature review and classification.
In Ifac-papersonline (Vol. 51, pp. 1016–1022).
Retrieved from https://doi.org/10.1016/j
.ifacol.2018.08.474 doi: 10.1016/j.ifacol
.2018.08.474
Lee, J., Bagheri, B., & Kao, H.-A. (2015). A cyber-physical
systems architecture for Industry 4.0-based manufacturing
systems. Manufacturing Letters, 3, 18–23.
Retrieved from https://doi.org/10.1016/j
.mfglet.2014.12.001 doi: 10.1016/j.mfglet
.2014.12.001
Lei, Y., Li, N., Guo, L., Li, N., Yan, T., & Lin, J. (2018).
12
ANNUAL CONFERENCE OF THE PROGNOSTICS AND HEALTH MANAGEMENT SOCIETY 2026
Machinery health prognostics: A systematic review
from data acquisition to RUL prediction. Mechanical
Systems and Signal Processing, 104, 799–834.
Retrieved from https://doi.org/10.1016/j
.ymssp.2017.11.016 doi: 10.1016/j.ymssp
.2017.11.016
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness,
J., Bellemare, M. G., . . . Hassabis, D. (2015). Humanlevel
control through deep reinforcement learning. Nature,
518, 529–533. Retrieved from https://doi
.org/10.1038/nature14236 doi: 10.1038/
nature14236
Mondesire, S. C., Angelopoulou, A., Sirigampola, S., &
Goldiez, B. (2019). Combining virtualization and containerization
to support interactive games and simulations
on the cloud. Simulation Modelling Practice and
Theory, 93, 233–244. Retrieved from https://doi
.org/10.1016/j.simpat.2018.08.005 doi:
10.1016/j.simpat.2018.08.005
Mondesire, S. C., Brown, M., & Soykan, B. (2025). Fabsim:
A micro-discrete event simulator for machine
learning dynamic job shop scheduling optimizers. In
Proceedings of the winter simulation conference. Retrieved
from https://www.informs-sim.org/
wsc25papers/inv186.pdf
Mondesire, S. C., & Soykan, B. (2026a). Automated wafer
dispatching through greedy multi-heuristic scheduling
in high-mix low-volume wafer fabs. In Advanced semiconductor
manufacturing conference (asmc).
Mondesire, S. C., & Soykan, B. (2026b). Real-time diffusion
scheduling for job shop optimization. In Proceedings
of the winter simulation conference (wsc).
Mondesire, S. C., & Wiegand, R. P. (2023). Mitigating
catastrophic forgetting with complementary
layered learning. Electronics, 12(3),
706. Retrieved from https://doi.org/
10.3390/electronics12030706 doi:
10.3390/electronics12030706
Mondesire, S. C.,&Wright, T. (2025). A tool grouping-based
approach to remaining useful life prediction for predictive
maintenance in semiconductor manufacturing.
In Advanced semiconductor manufacturing conference
(asmc): Big data management and machine learning.
Nsiye, E., Wright, T., Tse, B., & Mondesire, S. C.
(2024). A micro-discrete event simulation environment
for production scheduling in manufacturing
digital twins. In Modsim world. Retrieved
from https://www.modsimworld.org/
papers/2024/MODSIM 2024 paper 27.pdf
Park, J., Chun, J., Kim, S. H., Kim, Y., & Park, J. (2021).
Learning to schedule job-shop problems: Representation
and policy learning using graph neural network
and reinforcement learning. Retrieved from
https://arxiv.org/abs/2106.01086 doi:
10.48550/arXiv.2106.01086
PHM Society. (2016). 2016 PHM data challenge: Chemical
mechanical polishing data set. Retrieved from
https://phmsociety.org/conference/
annual-conference-of-the-phm
-society/annual-conference-of-the
-prognostics-and-health-management
-society-2016/phm-data-challenge-4/
PHM Society. (2018). 2018 PHM data challenge:
Ion mill etch tool fault and remaining useful life
data. Retrieved from https://c3.ndc.nasa
.gov/dashlink/resources/1009/
Pinedo, M. L. (2016). Scheduling: Theory, algorithms,
and systems (5th ed.). Springer. Retrieved
from https://doi.org/10.1007/978
-3-319-26580-3 doi: 10.1007/978-3-319-26580
-3
Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus,
M., & Dormann, N. (2021). Stable-baselines3: Reliable
reinforcement learning implementations. Journal
of Machine Learning Research, 22(268), 1–8. Retrieved
from https://www.jmlr.org/papers/
v22/20-1364.html
Sabri, S., Aghaabbasi, M., Reay Atkinson, S., Amon, M. J.,
Hancock, P., Azevedo, R., . . . Rabadi, G. (2025).
Integrating human-machine systems and digital twin
technologies: Navigating trust, interoperability, and
ethical challenges. Retrieved from https://doi
.org/10.2139/ssrn.5189676 doi: 10.2139/
ssrn.5189676
Sargent, R. G. (2013). Verification and validation of simulation
models. Journal of Simulation, 7(1), 12–24. Retrieved
from https://doi.org/10.1057/jos
.2012.20 doi: 10.1057/jos.2012.20
Schreck, J., Matthews, G., Lin, J., Mondesire, S., Metcalf, D.,
Dickerson, K., & Grasso, J. (2025). Levels of automation
for a computer-based procedure for simulated nuclear
power plant operation: Impacts on workload and
trust. Safety, 11(1), 22. Retrieved from https://
doi.org/10.3390/safety11010022 doi: 10
.3390/safety11010022
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., &
Klimov, O. (2017). Proximal policy optimization
algorithms. Retrieved from https://arxiv
.org/abs/1707.06347 doi: 10.48550/arXiv
.1707.06347
Sharp, M., & Weiss, B. A. (2018). Hierarchical modeling
of a manufacturing work cell to promote contextualized
PHM information across multiple levels.
Manufacturing Letters, 15, 46–49. Retrieved
from https://doi.org/10.1016/j.mfglet
.2018.02.003 doi: 10.1016/j.mfglet.2018.02.003
Si, X.-S., Wang, W., Hu, C.-H., & Zhou, D.-H. (2011).
Remaining useful life estimation: A review on
13
ANNUAL CONFERENCE OF THE PROGNOSTICS AND HEALTH MANAGEMENT SOCIETY 2026
the statistical data driven approaches. European
Journal of Operational Research, 213(1), 1–14.
Retrieved from https://doi.org/10.1016/j
.ejor.2010.11.018 doi: 10.1016/j.ejor.2010.11
.018
Singh, K., Selvanathan, B., Zope, K., Nistala, S. H., &
Runkana, V. (2018). Concurrent estimation of
remaining useful life for multiple faults in an ion
etch mill: A data-driven approach. In Annual conference
of the phm society (Vol. 10). Retrieved
from https://doi.org/10.36001/phmconf
.2018.v10i1.591 doi: 10.36001/phmconf.2018
.v10i1.591
Soykan, B., Mondesire, S. C., & Rabadi, G. (2026a). Realtime
dispatching in semiconductor manufacturing: A
hybrid graph deep reinforcement learning and discreteevent
simulation approach. In Proceedings of the winter
simulation conference (wsc).
Soykan, B., Mondesire, S. C., & Rabadi, G. (2026b). A supply
chain digital twin architecture for semiconductor
industry. In Optimizing supply chains through digital
twins: Navigating the modern landscape. Springer Nature.
Soykan, B., Mondesire, S. C., Rabadi, G., & Bochenek,
G. (2025). Graph-enhanced deep reinforcement
learning for multi-objective unrelated parallel
machine scheduling. In Proceedings of the winter
simulation conference (pp. 2515–2526). Retrieved
from https://doi.org/10.1109/WSC68292
.2025.11338986 doi: 10.1109/WSC68292.2025
.11338986
St¨ockermann, P., Feudel, S., Immordino, A., Hayen,
N., Altenm¨uller, T., Gebser, M., . . . Higgins, F.
(2025). Reinforcement learning based dispatching
solutions in semiconductor manufacturing: A literature
review on validation and deployment. Production
and Manufacturing Research, 13(1). Retrieved
from https://doi.org/10.1080/21693277
.2025.2582472 doi: 10.1080/21693277.2025
.2582472
St¨ockermann, P., Immordino, A., Altenm¨uller, T., Seidel,
G., Gebser, M., Tassel, P., . . . Zhang, F.
(2023). Dispatching in real frontend fabs with industrial
grade discrete-event simulations by deep reinforcement
learning with evolution strategies. In
Proceedings of the winter simulation conference (pp.
3047–3058). Retrieved from https://doi.org/
10.1109/WSC60868.2023.10408625 doi: 10
.1109/WSC60868.2023.10408625
St¨ockermann, P., S¨udfeld, H., Immordino, A., Altenm¨uller,
T., Wegmann, M., Gebser, M., . . . Zhang, F. F.
(2025). Scalability of reinforcement learning methods
for dispatching in semiconductor frontend fabs:
A comparison of open-source models with real
industry datasets. The International Journal of
Advanced Manufacturing Technology, 139, 4395–
4415. Retrieved from https://doi.org/10
.1007/s00170-025-16117-2 doi: 10.1007/
s00170-025-16117-2
Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning:
An introduction (2nd ed.). MIT Press. Retrieved
from http://incompleteideas.net/
book/the-book-2nd.html
Tao, F., Xiao, B., Qi, Q., Cheng, J., & Ji, P. (2022). Digital
twin modeling. Journal of Manufacturing Systems,
64, 372–389. Retrieved from https://doi.org/
10.1016/j.jmsy.2022.06.015 doi: 10.1016/
j.jmsy.2022.06.015
Tassel, P., Gebser, M., & Schekotihin, K. (2021). A reinforcement
learning environment for job-shop scheduling.
Retrieved from https://arxiv.org/abs/
2104.03760 doi: 10.48550/arXiv.2104.03760
Tassel, P., Kov´acs, B., Gebser, M., Schekotihin, K., St¨ockermann,
P., & Seidel, G. (2023). Semiconductor
fab scheduling with self-supervised and reinforcement
learning. In Proceedings of the winter
simulation conference (pp. 1924–1935). Retrieved
from https://doi.org/10.1109/WSC60868
.2023.10407747 doi: 10.1109/WSC60868.2023
.10407747
Towers, M., Kwiatkowski, A., Terry, J. K., Balis, J. U.,
De Cola, G., Deleu, T., . . . Younis, O. G. (2024).
Gymnasium: A standard interface for reinforcement
learning environments. Retrieved from https://
arxiv.org/abs/2407.17032 doi: 10.48550/
arXiv.2407.17032
Wang, B., Zhou, H., Yang, G., Li, X., & Yang, H.
(2022). Human digital twin driven human-cyberphysical
systems: Key technologies and applications.
Chinese Journal of Mechanical Engineering,
35(11). Retrieved from https://doi.org/10
.1186/s10033-022-00680-w doi: 10.1186/
s10033-022-00680-w
Wright, T., Nsiye, E., & Mondesire, S. C. (2024). Generalized
vs. individualized kaplan-meier time-to-failure
predictive models for preventative maintenance in
semiconductor manufacturing. In International symposium
on semiconductor manufacturing (issm).
Wright, T., Soykan, B., & Mondesire, S. C. (2025). Real-time
remaining useful life prediction with LSTM priority resampling.
In Ieee international conference on machine
learning and applications (icmla).
Zhang, C., Song, W., Cao, Z., Zhang, J., Tan, P. S., &
Xu, C. (2020). Learning to dispatch for job shop
scheduling via deep reinforcement learning. In
Advances in neural information processing systems.
Retrieved from https://proceedings
.neurips.cc/paper/2020/hash/

This work is licensed under a Creative Commons Attribution 3.0 Unported License.
The Prognostic and Health Management Society advocates open-access to scientific data and uses a Creative Commons license for publishing and distributing any papers. A Creative Commons license does not relinquish the author’s copyright; rather it allows them to share some of their rights with any member of the public under certain conditions whilst enjoying full legal protection. By submitting an article to the International Conference of the Prognostics and Health Management Society, the authors agree to be bound by the associated terms and conditions including the following:
As the author, you retain the copyright to your Work. By submitting your Work, you are granting anybody the right to copy, distribute and transmit your Work and to adapt your Work with proper attribution under the terms of the Creative Commons Attribution 3.0 United States license. You assign rights to the Prognostics and Health Management Society to publish and disseminate your Work through electronic and print media if it is accepted for publication. A license note citing the Creative Commons Attribution 3.0 United States License as shown below needs to be placed in the footnote on the first page of the article.
First Author et al. This is an open-access article distributed under the terms of the Creative Commons Attribution 3.0 United States License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
https://orcid.org/0000-0002-5152-1709