Demonstration-Based Visual Anomaly Detection in Manufacturing Factory Using Vision-Language Models

##plugins.themes.bootstrap3.article.main##

##plugins.themes.bootstrap3.article.sidebar##

Published Sep 28, 2026
Huimin Zhuge Song Wang Huijuan Shao Brian Chien Dipanjan Ghosh

Abstract

Flexible anomaly detection is important for identifying process deviations in manufacturing, yet conventional approaches typically depend on hand-coded inspection rules, annotated datasets, task-specific model training, or additional sensing infrastructure, all of which are costly to configure and difficult to adapt. This work investigates demonstration based visual anomaly detection using vision-language models (VLMs), in which monitoring behavior is specified by example rather than by explicit programming. Instead of enumerating normal operating rules, the model is shown images and videos of nominal operation and infers the expected visual patterns directly from these demonstrations.  During monitoring, new observations are assessed against the demonstrated patterns to detect potential abnormal states and to generate interpretable warnings and corrective suggestions. We formalize the detection problem mathematically and evaluate the approach on three representative fault types: sorting faults, where a color category is missing or misplaced; storage faults, where caps violate the expected bottom-to-top filling pattern or appear in an incorrect color region; and conveyor faults, where caps overlap, become stuck, move irregularly, or lose consistent spacing. These scenarios are realized on a small manufacturing cell equipped with pickup arms, a conveyor belt, sorting and storage units, and fixed monitoring cameras, using colored cylinder caps as representative workpieces. Results indicate that visual demonstrations combined with reasoning can substantially reduce the effort required to configure perception, anomaly interpretation, and decision support in manufacturing workflows.

How to Cite

Zhuge, H., Wang, S. ., Shao, H., Chien, B., & Ghosh, D. (2026). Demonstration-Based Visual Anomaly Detection in Manufacturing Factory Using Vision-Language Models. Annual Conference of the PHM Society, 18(1). https://doi.org/10.36001/phmconf.2026.v18i1.4811
Abstract 0 | PDF Downloads 0

##plugins.themes.bootstrap3.article.details##

Keywords

Flexible Manufacturing Systems, Industrial Automation, Multimodal Reasoning, Scene Understanding

References
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O.,
David, B., . . . others (2022). Do as i can, not as i
say: Grounding language in robotic affordances. arXiv
preprint arXiv:2204.01691.
Anthropic. (2025). Introducing claude sonnet 4.6.
Argall, B. D., Chernova, S., Veloso, M., & Browning, B.
(2009). Robot learning from demonstration. Robotics
and Autonomous Systems, 57(5), 469–483.
Bergmann, P., Fauser, M., Sattlegger, D., & Steger, C. (2019).
Mvtec ad: A comprehensive real-world dataset for unsupervised
anomaly detection. In Proceedings of the
ieee/cvf conference on computer vision and pattern
recognition (pp. 9592–9600).
Carion, N., Gustafson, L., Hu, Y.-T., Debnath, S., Hu, R.,
Suris, D., . . . others (2025). Sam 3: Segment anything
with concepts. arXiv preprint arXiv:2511.16719.
Chibani, A., Coudert, T., & Kamsu-Foguem, B. (2022). A
review of deep learning in the study of materials degradation.
Engineering Applications of Artificial Intelli-
gence, 113, 104940.
Google DeepMind. (2025). What’s new in gemini 3.5 flash.
Guan, T., Liu, F., Wu, X., Xian, R., Li, Z., Liu, X., . . . others
(2023). Hallusionbench: An advanced diagnostic
suite for entangled language hallucination and visual illusion
in large vision-language models. arXiv preprint
arXiv:2310.14566.
Kragic, D., & Christensen, H. I. (2002). A survey of visionbased
robotic manipulation. In Ieee international con-
ference on robotics and automation (pp. 1810–1815).
MMAD: A comprehensive benchmark for multimodal large
language models in industrial anomaly detection.
(2025). In International conference on learning rep-
resentations (iclr). (arXiv:2410.09453)
Newman, T. S., & Jain, A. K. (1995). Machine vision applications
in manufacturing. Computer Vision and Image
Understanding, 61(2), 231–262.
NVIDIA, :, Blakeman, A., Grattafiori, A., Basant, A.,
Gupta, A., . . . Yan, Z. (2025). Nemotron 3 nano:
Open, efficient mixture-of-experts hybrid mamba-
transformer model for agentic reasoning. Retrieved
from https://arxiv.org/abs/2512.20848
OpenAI. (2023). Gpt-4v(ision) system card. OpenAI Techni-
cal Report.
OpenAI. (2026). Introducing gpt-5.4 mini and nano. (Accessed
June 2026)
Pang, G., Shen, C., Cao, L., & Hengel, A. v. d. (2021). Deep
learning for anomaly detection: A survey. ACM Com-
puting Surveys, 54(2), 1–38.
Qwen Team. (2026, February). Qwen3.5: Towards native
multimodal agents.
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G.,
Agarwal, S., . . . others (2021). Learning transferable
visual models from natural language supervision. In In-
ternational conference on machine learning (pp. 8748–
8763).
Radke, R. J., Andra, S., Al-Kofahi, O., & Roysam, B. (2005).
Change detection techniques. IEEE Transactions on
Image Processing, 14(3), 294–307.
Xia, Z., Li, H., Li, W., & Song, B. (2024). Large language
models in manufacturing: Opportunities and
challenges. Journal of Manufacturing Systems.
Zhang, J., Huang, J., Jin, S., & Lu, S. (2024). Visionlanguage
models for vision tasks: A survey. IEEE
Transactions on Pattern Analysis and Machine Intel-
ligence.
Section
Technical Research Papers