Resource-Aware Model Comparison for DASHlink Multi-Class Flight-Anomaly Screening

##plugins.themes.bootstrap3.article.main##

##plugins.themes.bootstrap3.article.sidebar##

Published Sep 28, 2026
Keanan Milton Ge Zhou

Abstract

Flight-data prognostics and health management models have to do more than rank anomalous windows correctly. For future edge or onboard validation, a screening model also has to be compact, fast to score, false-alarm aware, and interpretable enough that its decision statistic can be trusted. Autoencoders are attractive in this setting because they can be trained from nominal data, but a single scalar reconstruction error may be too coarse to separate operationally different non-nominal approach behaviors. We tested that tradeoff on the NASA DASHlink multi-class flight-anomaly window benchmark using a fixed split, nominal-only threshold calibration, and a held-out test set. Each sample is a 160-second, 20-variable window labeled as Nominal, Speed High, Path High, or Flaps Late Setting. We treat the primary task as nominal-versus-non-nominal screening, while preserving class-wise recall so that the rare classes are not hidden by aggregate recall. The comparison includes raw-summary classifiers, classical unsupervised baselines, scalar convolutional variational autoencoder (CVAE) reconstruction scores, CVAE residual and latent representations, and raw-plus-CVAE hybrid models. Thresholded metrics use a nominal-calibrated p95 operating point; resource proxies include total scoring time and serialized pipeline size. Scalar CVAE reconstruction error was weak as a standalone detector, with AP 0.179 and worst-class recall 0.051. Residual and latent CVAE features retained substantially more class-discriminative information, and the strongest raw-plus-CVAE hybrid reached AP 0.935, recall 0.945, and worst-class recall 0.894 at 0.158 ms/window. Full-20 raw-summary HGB remained the simpler high-performing baseline. The practical result is that compact feature-based models are strong screening baselines, while reconstruction models are most useful here as auxiliary representation generators rather than standalone scalar anomaly detectors.

How to Cite

Milton, K., & Zhou, G. . (2026). Resource-Aware Model Comparison for DASHlink Multi-Class Flight-Anomaly Screening. Annual Conference of the PHM Society, 18(1). https://doi.org/10.36001/phmconf.2026.v18i1.5060
Abstract 0 | PDF Downloads 0

##plugins.themes.bootstrap3.article.details##

Keywords

prognostics and health management; flight data monitoring; anomaly detection; DASHlink; variational autoencoder; resource-aware machine learning; onboard anomaly screening

References
Chandola, V., Banerjee, A., & Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys, 41(3), Article 15. https://doi.org/10.1145/1541880.1541882
Davis, J., & Goadrich, M. (2006). The relationship between precision-recall and ROC curves. Proceedings of the 23rd International Conference on Machine Learning, 233-240. https://doi.org/10.1145/1143844.1143874
Dudukcu, H. V., Taşkıran, M., & Kahraman, N. (2025). AnoSense: Edge computing for real-time flight anomaly detection by using embedded deep neural networks. Gümüşhane Üniversitesi Fen Bilimleri Dergisi, 15(3), 797-808. https://doi.org/10.17714/gumusfenbil.1676270
Efron, B., & Tibshirani, R. J. (1993). An introduction to the bootstrap. Chapman & Hall.
Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5), 1189-1232.
Kingma, D. P., & Welling, M. (2014). Auto-encoding variational Bayes. International Conference on Learning Representations.
Memarzadeh, M., Matthews, B., & Avrekh, I. (2020). Unsupervised anomaly detection in flight data using convolutional variational auto-encoder. Aerospace, 7(8), 115. https://doi.org/10.3390/aerospace7080115
Memarzadeh, M., Matthews, B., & Templin, T. (2022). Multiclass anomaly detection in flight data using semi-supervised explainable deep learning model. Journal of Aerospace Info rmation Systems, 19(2), 83-97. https://doi.org/10.2514/1.I010959
National Aeronautics and Space Administration. (n.d.). Sample Flight Data - DASHlink. NASA DASHlink Collaborative Sharing Network. Retrieved June 18, 2026, from https://c3.ndc.nasa.gov/dashlink/projects/85/
Reddy, K. K., Sarkar, S., Venugopalan, V., & Giering, M. (2016). Anomaly detection and fault disambiguation in large flight data: A multi-modal deep auto-encoder approach. Annual Conference of the Prognostics and Health Management Society.
Schwabacher, M., & Goebel, K. F. (2007). A survey of artificial intelligence for prognostics. Proceedings of the AAAI Fall Symposium, Arlington, VA.
Vachtsevanos, G., Lewis, F. L., Roemer, M., Hess, A., & Wu, B. (2006). Intelligent fault diagnosis and prognosis for engineering systems. John Wiley & Sons.
Section
Technical Research Papers