Demonstration-Based Visual Anomaly Detection in Manufacturing Factory Using Vision-Language Models
##plugins.themes.bootstrap3.article.main##
##plugins.themes.bootstrap3.article.sidebar##
Abstract
Flexible anomaly detection is important for identifying process deviations in manufacturing, yet conventional approaches typically depend on hand-coded inspection rules, annotated datasets, task-specific model training, or additional sensing infrastructure, all of which are costly to configure and difficult to adapt. This work investigates demonstration based visual anomaly detection using vision-language models (VLMs), in which monitoring behavior is specified by example rather than by explicit programming. Instead of enumerating normal operating rules, the model is shown images and videos of nominal operation and infers the expected visual patterns directly from these demonstrations. During monitoring, new observations are assessed against the demonstrated patterns to detect potential abnormal states and to generate interpretable warnings and corrective suggestions. We formalize the detection problem mathematically and evaluate the approach on three representative fault types: sorting faults, where a color category is missing or misplaced; storage faults, where caps violate the expected bottom-to-top filling pattern or appear in an incorrect color region; and conveyor faults, where caps overlap, become stuck, move irregularly, or lose consistent spacing. These scenarios are realized on a small manufacturing cell equipped with pickup arms, a conveyor belt, sorting and storage units, and fixed monitoring cameras, using colored cylinder caps as representative workpieces. Results indicate that visual demonstrations combined with reasoning can substantially reduce the effort required to configure perception, anomaly interpretation, and decision support in manufacturing workflows.
How to Cite
##plugins.themes.bootstrap3.article.details##
Flexible Manufacturing Systems, Industrial Automation, Multimodal Reasoning, Scene Understanding
David, B., . . . others (2022). Do as i can, not as i
say: Grounding language in robotic affordances. arXiv
preprint arXiv:2204.01691.
Anthropic. (2025). Introducing claude sonnet 4.6.
Argall, B. D., Chernova, S., Veloso, M., & Browning, B.
(2009). Robot learning from demonstration. Robotics
and Autonomous Systems, 57(5), 469–483.
Bergmann, P., Fauser, M., Sattlegger, D., & Steger, C. (2019).
Mvtec ad: A comprehensive real-world dataset for unsupervised
anomaly detection. In Proceedings of the
ieee/cvf conference on computer vision and pattern
recognition (pp. 9592–9600).
Carion, N., Gustafson, L., Hu, Y.-T., Debnath, S., Hu, R.,
Suris, D., . . . others (2025). Sam 3: Segment anything
with concepts. arXiv preprint arXiv:2511.16719.
Chibani, A., Coudert, T., & Kamsu-Foguem, B. (2022). A
review of deep learning in the study of materials degradation.
Engineering Applications of Artificial Intelli-
gence, 113, 104940.
Google DeepMind. (2025). What’s new in gemini 3.5 flash.
Guan, T., Liu, F., Wu, X., Xian, R., Li, Z., Liu, X., . . . others
(2023). Hallusionbench: An advanced diagnostic
suite for entangled language hallucination and visual illusion
in large vision-language models. arXiv preprint
arXiv:2310.14566.
Kragic, D., & Christensen, H. I. (2002). A survey of visionbased
robotic manipulation. In Ieee international con-
ference on robotics and automation (pp. 1810–1815).
MMAD: A comprehensive benchmark for multimodal large
language models in industrial anomaly detection.
(2025). In International conference on learning rep-
resentations (iclr). (arXiv:2410.09453)
Newman, T. S., & Jain, A. K. (1995). Machine vision applications
in manufacturing. Computer Vision and Image
Understanding, 61(2), 231–262.
NVIDIA, :, Blakeman, A., Grattafiori, A., Basant, A.,
Gupta, A., . . . Yan, Z. (2025). Nemotron 3 nano:
Open, efficient mixture-of-experts hybrid mamba-
transformer model for agentic reasoning. Retrieved
from https://arxiv.org/abs/2512.20848
OpenAI. (2023). Gpt-4v(ision) system card. OpenAI Techni-
cal Report.
OpenAI. (2026). Introducing gpt-5.4 mini and nano. (Accessed
June 2026)
Pang, G., Shen, C., Cao, L., & Hengel, A. v. d. (2021). Deep
learning for anomaly detection: A survey. ACM Com-
puting Surveys, 54(2), 1–38.
Qwen Team. (2026, February). Qwen3.5: Towards native
multimodal agents.
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G.,
Agarwal, S., . . . others (2021). Learning transferable
visual models from natural language supervision. In In-
ternational conference on machine learning (pp. 8748–
8763).
Radke, R. J., Andra, S., Al-Kofahi, O., & Roysam, B. (2005).
Change detection techniques. IEEE Transactions on
Image Processing, 14(3), 294–307.
Xia, Z., Li, H., Li, W., & Song, B. (2024). Large language
models in manufacturing: Opportunities and
challenges. Journal of Manufacturing Systems.
Zhang, J., Huang, J., Jin, S., & Lu, S. (2024). Visionlanguage
models for vision tasks: A survey. IEEE
Transactions on Pattern Analysis and Machine Intel-
ligence.

This work is licensed under a Creative Commons Attribution 3.0 Unported License.
The Prognostic and Health Management Society advocates open-access to scientific data and uses a Creative Commons license for publishing and distributing any papers. A Creative Commons license does not relinquish the author’s copyright; rather it allows them to share some of their rights with any member of the public under certain conditions whilst enjoying full legal protection. By submitting an article to the International Conference of the Prognostics and Health Management Society, the authors agree to be bound by the associated terms and conditions including the following:
As the author, you retain the copyright to your Work. By submitting your Work, you are granting anybody the right to copy, distribute and transmit your Work and to adapt your Work with proper attribution under the terms of the Creative Commons Attribution 3.0 United States license. You assign rights to the Prognostics and Health Management Society to publish and disseminate your Work through electronic and print media if it is accepted for publication. A license note citing the Creative Commons Attribution 3.0 United States License as shown below needs to be placed in the footnote on the first page of the article.
First Author et al. This is an open-access article distributed under the terms of the Creative Commons Attribution 3.0 United States License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.