CognitiveCPQ: A Graph-Based Reinforcement Learning Framework for Autonomous Constraint-Aware Enterprise Deal Optimisation
Configure-Price-Quote (CPQ) systems are central to enterprise B2B sales execution, yet most deployments remain rule-based and deterministic, unable to learn from deal outcomes. This paper presents CognitiveCPQ, combining a Heterogeneous Constraint Graph (HCG) with Proximal Policy Optimisation (PPO) for autonomous, constraint-aware deal optimisation. A deterministic kill-switch guarantees zero hard constraint violations; soft constraints are navigated via a z-score-normalised penalised reward. A KernelSHAP explainability layer provides per-decision attribution. Evaluation on a large-scale synthetic simulation demonstrates 18.9% margin uplift, 94.3% constraint compliance, and a 67% reduction in configuration steps over rule-based baselines. We additionally specify two robustness evaluation protocols—response-model misspecification and state-observability noise—as immediate follow-up work; both protocols are expected to confirm graceful degradation given the kill-switch’s structural independence from the response model and state vector. All results represent proof-of-concept performance in simulation; production deployment requires shadow-mode calibration on real deal data.
Authors
- Thottempudi Pardhu (ORCID: https://orcid.org/0000-0002-9653-1951)
- Rajesh Soma (ORCID: https://orcid.org/0009-0002-5436-4777)
Institutions
- University of the Cumberlands (US)
- Koneru Lakshmaiah Education Foundation (IN)
Publication Details
- Journal
- AI
- Published
- 2026-10-05
- DOI
- https://doi.org/10.3390/ai7100407
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- article
- Field-Weighted Citation Impact
- 0.00