Design and Deployment of Custom GPT Solutions Using Azure OpenAI Service
Enterprise adoption of Large Language Models demands evaluation beyond model accuracy, encompassing cost, latency, compliance, and update adaptability. This paper introduces a Quantitative Evaluation Framework (QEF) to benchmark GPT deployment strategies within Azure OpenAI infrastructure. Five configurations baseline inference, prompt engineering, supervised fine-tuning, Retrieval-Augmented Generation (RAG), and a hybrid architecture were evaluated against 500 domain-specific queries drawn from 12,000 enterprise policy documents. Six metrics were measured: Factual Accuracy Score, Hallucination Rate, Mean Latency, Cost per 1,000 Queries, Update Flexibility Index, and a composite Enterprise Deployment Efficiency Index. Results confirm statistically significant performance differences across configurations (p < 0.001). RAG achieved the optimal balance between factual grounding and cost efficiency, reducing hallucination by over 70% relative to baseline. A governance-integrated Deployment Decision Framework is proposed, aligning technical, economic, and regulatory dimensions for enterprise AI architecture selection.
Authors
- Stefan Trajanoski
- Aleskandar Karadimce
Institutions
- University of Information Science and Technology St. Paul The Apostle (MK)
Publication Details
- Journal
- WSEAS Transactions on Computers archive
- Published
- 2026-09-22
- DOI
- https://doi.org/10.37394/23205.2026.25.14
- Primary Topic
- Artificial Intelligence in Healthcare and Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00