Deploying Vision Foundation Models for Scalable Object Recognition in Web-Based Production Systems: Architecture, Integration, and Optimization

Baklizi MK, Alkhazaleh M, Al Sardy L (2026)


Publication Type: Authored book

Publication year: 2026

Publisher: IGI Global Scientific Publishing

ISBN: 9798260001387

DOI: 10.4018/979-8-2600-0136-3.ch006

Abstract

Vision Foundation Models (VFMs) include CLIP, DINOv2, and SAM are exceptional in generalizing visual recognition tasks, but their use in web-based production systems is still challenging because of computational costs and architecture. The chapter is a complete guide to the integration of VFMs into scalable web applications based on the microservice-based architecture, optimization of models (quantization, pruning, knowledge distillation), and the efficient design of APIs. According to experimental analysis, INT8 quantization gives a 2.5130x speedup at only a small accuracy penalty, and knowledge distillation gives 611x faster inference. The effectiveness of the framework is also confirmed using three case studies: e-commerce visual search, quality inspection, and document processing; the latency of the framework is less than 200ms, and operations are considerably improved. The chapter is closed with the rules of optimization and new tendencies in the deployment of an edge-cloud.

Authors with CRIS profile

Involved external institutions

How to cite

APA:

Baklizi, M.K., Alkhazaleh, M., & Al Sardy, L. (2026). Deploying Vision Foundation Models for Scalable Object Recognition in Web-Based Production Systems: Architecture, Integration, and Optimization. IGI Global Scientific Publishing.

MLA:

Baklizi, Mahmoud Khalid, Mohammad Alkhazaleh, and Loui Al Sardy. Deploying Vision Foundation Models for Scalable Object Recognition in Web-Based Production Systems: Architecture, Integration, and Optimization. IGI Global Scientific Publishing, 2026.

BibTeX: Download