IndusDiff: A Latent Diffusion Framework for Industrial Audio Generation for Data-Scarce PHM
##plugins.themes.bootstrap3.article.main##
##plugins.themes.bootstrap3.article.sidebar##
Abstract
Prognostics and Health Management (PHM) systems increasingly rely on data-driven models for fault diagnosis and anomaly detection, yet their effectiveness is fundamentally constrained by the scarcity and imbalance of labeled fault condition data. This challenge is particularly acute in industrial audio monitoring, where failure events are rare, hazardous to induce, and highly variable across operating conditions. While recent advances in generative modeling offer a promising pathway for data augmentation, existing audio generation methods either incur high computational cost in waveform-domain diffusion models or fail to preserve phase-sensitive characteristics critical for capturing transient fault signatures.
To address these challenges, this paper proposes IndusDiff, a phase-aware latent diffusion framework for high-fidelity industrial audio generation tailored for PHM applications. The proposed approach combines a phase-aware convolutional autoencoder with a latent diffusion model to enable efficient and physically meaningful audio synthesis. The autoencoder is trained using a multi-resolution spectral loss that jointly enforces waveform fidelity, spectral consistency, and temporal phase coherence, thereby preserving diagnostically relevant features such as transients, harmonics, and high-frequency components in the latent representation. Diffusion is then performed in this compressed latent space, significantly reducing computational complexity while maintaining generation quality.
Unlike conventional latent diffusion methods that primarily focus on perceptual realism, the proposed framework explicitly addresses the unique characteristics of industrial audio signals, including non-stationarity, multi-scale temporal dynamics, and phase-sensitive fault signatures. By preserving both magnitude and phase information, the model generates synthetic audio that is not only perceptually realistic but also structurally consistent with real machine signals, making it suitable for downstream PHM tasks.
The framework is evaluated on two complementary datasets: a publicly available industrial motor dataset and an in-house dataset capturing multi-stage assembly operations. Generation quality is assessed using the Frechet Audio Distance (FAD) as the primary quantitative metric, along with waveform and spectrogram analyses for qualitative validation. Experimental results demonstrate that IndusDiff produces statistically consistent, high-fidelity industrial audio while achieving significant improvements in computational efficiency, generating 48 kHz audio samples lasting several seconds in just seconds on modern GPU hardware.
How to Cite
##plugins.themes.bootstrap3.article.details##
Latent Diffusion, Industrial Audio Generation, Data-Scarce PHM
Chan, W. (2021). WaveGrad: Estimating gradients for
waveform generation. In Proceedings of the international
conference on learning representations (iclr).
Choi, H.-S., Kim, J.-H., Huh, J., Kim, A., Ha, J.-W., & Lee,
K. (2019). Phase-aware speech enhancement with deep
complex U-Net. In Proceedings of the international
conference on learning representations (iclr).
Chung, Y., Lee, J., & Nam, J. (2024). T-FOLEY:
A controllable waveform-domain diffusion model for
temporal-event-guided foley sound synthesis. In Proceedings
of the ieee international conference on acoustics,
speech and signal processing (icassp). doi:
10.1109/ICASSP48485.2024.10447898
Ding, Y., Ma, L., Ma, J., Wang, M., & Lu, C. (2023). A
survey on deep learning for machinery fault diagnosis
and prognosis. IEEE Transactions on Instrumentation
and Measurement, 72, 1–20.
Evans, Z., Parker, J. D., Simon, C. J., Carr, C. J., Zukowski,
Z., & Abela, J. (2024). Stable audio open. arXiv
preprint arXiv:2407.14358.
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B.,
Warde-Farley, D., Ozair, S., . . . Bengio, Y. (2014).
Generative adversarial nets. In Advances in neural information
processing systems (neurips) (Vol. 27, pp.
2672–2680).
Griffin, D. W., & Lim, J. S. (1984). Signal estimation from
modified short-time Fourier transform. IEEE Transactions
on Acoustics, Speech, and Signal Processing,
32(2), 236–243. doi: 10.1109/TASSP.1984.1164317
Grollmisch, S., & Abeßer, J. (2021). Audio-based machine
fault diagnosis with WaveNet and data augmentation.
In Proceedings of the 9th european workshop on structural
health monitoring (ewshm).
Grollmisch, S., Abeßer, J., Liebetrau, J., & Lukashevich, H. (2019). Sounding industry: Challenges and datasets for
industrial sound analysis. In Proceedings of the 27th
european signal processing conference (eusipco) (pp.
1–5). doi: 10.23919/EUSIPCO.2019.8902941
Gui, A., Gamper, H., Braun, S., & Emmanouilidou, D.
(2024). Adapting Fr´echet audio distance for generative
music evaluation. In Proceedings of the ieee
international conference on acoustics, speech and
signal processing (icassp) (pp. 1331–1335). doi:
10.1109/ICASSP48485.2024.10446663
Hershey, S., Chaudhuri, S., Ellis, D. P. W., Gemmeke, J. F.,
Jansen, A., Moore, R. C., . . . Wilson, K. (2017).
CNN architectures for large-scale audio classification.
In Proceedings of the ieee international conference on
acoustics, speech and signal processing (icassp) (pp.
131–135). doi: 10.1109/ICASSP.2017.7952132
Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion
probabilistic models. In Advances in neural information
processing systems (neurips) (Vol. 33, pp. 6840–
6851).
Ho, J., & Salimans, T. (2022). Classifier-free diffusion guidance.
arXiv preprint arXiv:2207.12598.
Kilgour, K., Zuluaga, M., Roblek, D., & Sharifi, M. (2019).
Fr´echet audio distance: A reference-free metric for
evaluating music enhancement algorithms. In Proceedings
of interspeech 2019 (pp. 2350–2354). doi:
10.21437/Interspeech.2019-2219
Kingma, D. P., & Welling, M. (2014). Auto-encoding variational
Bayes. In Proceedings of the international conference
on learning representations (iclr).
Koizumi, Y., Kawaguchi, Y., Imoto, K., Nakamura, T., Niitsuma,
Y., Tanabe, R., . . . Harada, N. (2020). Description
and discussion on DCASE 2020 challenge task 2:
Unsupervised anomalous sound detection for machine
condition monitoring. In Proceedings of the detection
and classification of acoustic scenes and events workshop
(dcase) (pp. 1–6).
Kong, J., Kim, J., & Bae, J. (2020). HiFi-GAN: Generative
adversarial networks for efficient and high fidelity
speech synthesis. In Advances in neural information
processing systems (neurips) (Vol. 33, pp. 17022–
17033).
Kong, Q., Cao, Y., Iqbal, T., Wang, Y., Wang, W., &
Plumbley, M. D. (2020). PANNs: Large-scale pretrained
audio neural networks for audio pattern recognition.
IEEE/ACM Transactions on Audio, Speech,
and Language Processing, 28, 2880–2894. doi:
10.1109/TASLP.2020.3030497
Kong, Z., Ping, W., Huang, J., Zhao, K., & Catanzaro, B.
(2021). DiffWave: A versatile diffusion model for audio
synthesis. In Proceedings of the international conference
on learning representations (iclr).
Kumar, K., Kumar, R., de Boissiere, T., Gestin, L., Teoh,
W. Z., Sotelo, J., . . . Courville, A. (2019). MelGAN:
Generative adversarial networks for conditional waveform
synthesis. In Advances in neural information processing
systems (neurips) (Vol. 32).
Lei, Y., Li, N., Guo, L., Li, N., Yan, T., & Lin, J. (2018). Machinery
health prognostics: A systematic review from
data acquisition to RUL prediction. Mechanical Systems
and Signal Processing, 104, 799–834.
Li, X., Ding, Q., & Sun, J.-Q. (2018). Remaining useful life
estimation in prognostics using deep convolution neural
networks. Reliability Engineering & System Safety,
172, 1–11.
Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D., . . .
Plumbley, M. D. (2023). AudioLDM: Text-to-audio
generation with latent diffusion models. In Proceedings
of the international conference on machine learning
(icml) (pp. 21450–21474).
Liu, H., Yuan, Y., Liu, X., Mei, X., Kong, Q., Tian, Q., . . .
Plumbley, M. D. (2024). AudioLDM 2: Learning
holistic audio generation with self-supervised pretraining.
IEEE/ACM Transactions on Audio, Speech, and
Language Processing, 32, 2871–2883.
Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., & Zhu, J.
(2022). DPM-Solver: A fast ODE solver for diffusion
probabilistic model sampling in around 10 steps.
In Advances in neural information processing systems
(neurips) (Vol. 35, pp. 5775–5787).
McLachlan, G. J., & Peel, D. (2000). Finite mixture models.
New York: Wiley.
Mehri, S., Kumar, K., Gulrajani, I., Kumar, R., Jain, S.,
Sotelo, J., . . . Bengio, Y. (2017). SampleRNN:
An unconditional end-to-end neural audio generation
model. In Proceedings of the international conference
on learning representations (iclr).
Nichol, A. Q., & Dhariwal, P. (2021). Improved denoising
diffusion probabilistic models. In International conference
on machine learning (icml) (pp. 8162–8171).
Perez, E., Strub, F., de Vries, H., Dumoulin, V., & Courville,
A. (2018). FiLM: Visual reasoning with a general conditioning
layer. In Aaai conference on artificial intelligence.
Purohit, H., Tanabe, R., Ichige, K., Endo, T., Nikaido, Y.,
Suefusa, K., & Kawaguchi, Y. (2019). MIMII dataset:
Sound dataset for malfunctioning industrial machine
investigation and inspection. In Proceedings of the detection
and classification of acoustic scenes and events
workshop (dcase) (pp. 1–6).
Rabiner, L. R. (1989). A tutorial on hidden Markov
models and selected applications in speech recognition.
Proceedings of the IEEE, 77(2), 257–286. doi:
10.1109/5.18626
Reynolds, D. A., & Rose, R. C. (1995). Robust
text-independent speaker identification using Gaussian
mixture speaker models. IEEE Transactions on
Speech and Audio Processing, 3(1), 72–83. doi: 10.1109/89.365379
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer,
B. (2022). High-resolution image synthesis with
latent diffusion models. In Proceedings of the ieee/cvf
conference on computer vision and pattern recognition
(cvpr) (pp. 10684–10695).
Salamon, J., & Bello, J. P. (2017). Deep convolutional neural
networks and data augmentation for environmental
sound classification. IEEE Signal Processing Letters,
24(3), 279–283. doi: 10.1109/LSP.2017.2657381
Salimans, T., & Ho, J. (2022). Progressive distillation for fast
sampling of diffusion models. In International conference
on learning representations (iclr).
Song, J., Meng, C., & Ermon, S. (2021). Denoising diffusion
implicit models. In Proceedings of the international
conference on learning representations (iclr).
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon,
S., & Poole, B. (2021). Score-based generative
modeling through stochastic differential equations. In
Proceedings of the international conference on learning
representations (iclr).
Tailleur, M., Lee, J., Lagrange, M., Choi, K., Heller,
L. M., Imoto, K., & Okamoto, Y. (2024). Correlation
of Fr´echet audio distance with human perception
of environmental audio is embedding dependent.
In Proceedings of the 32nd european signal processing
conference (EUSIPCO). Lyon, France. doi:
10.48550/arXiv.2403.17508
Takaki, S., Nakashika, T., Wang, X., & Yamagishi, J. (2019).
STFT spectral loss for training a neural speech waveform
model. In Ieee international conference on acoustics,
speech and signal processing (icassp) (pp. 7065–
7069).
van den Oord, A., Dieleman, S., Zen, H., Simonyan, K.,
Vinyals, O., Graves, A., . . . Kavukcuoglu, K. (2016).
WaveNet: A generative model for raw audio. arXiv
preprint arXiv:1609.03499.
Yamamoto, R., Song, E., & Kim, J.-M. (2020). Parallel
WaveGAN: A fast waveform generation model
based on generative adversarial networks with multiresolution
spectrogram. In Proc. of icassp (pp. 6199–
6203).
Zhao, R., Yan, R., Chen, Z., Mao, K., Wang, P., & Gao,
R. X. (2019). Deep learning and its applications to
machine health monitoring. Mechanical Systems and
Signal Processing, 115, 213–237.
Ziyin, L., Hartwig, T., & Ueda, M. (2020). Neural networks
fail to learn periodic functions and how to fix it.
In Advances in neural information processing systems
(neurips) (Vol. 33, pp. 1583–1594).

This work is licensed under a Creative Commons Attribution 3.0 Unported License.
The Prognostic and Health Management Society advocates open-access to scientific data and uses a Creative Commons license for publishing and distributing any papers. A Creative Commons license does not relinquish the author’s copyright; rather it allows them to share some of their rights with any member of the public under certain conditions whilst enjoying full legal protection. By submitting an article to the International Conference of the Prognostics and Health Management Society, the authors agree to be bound by the associated terms and conditions including the following:
As the author, you retain the copyright to your Work. By submitting your Work, you are granting anybody the right to copy, distribute and transmit your Work and to adapt your Work with proper attribution under the terms of the Creative Commons Attribution 3.0 United States license. You assign rights to the Prognostics and Health Management Society to publish and disseminate your Work through electronic and print media if it is accepted for publication. A license note citing the Creative Commons Attribution 3.0 United States License as shown below needs to be placed in the footnote on the first page of the article.
First Author et al. This is an open-access article distributed under the terms of the Creative Commons Attribution 3.0 United States License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.