CognitiveCPQ: A Graph-Based Reinforcement Learning Framework for Autonomous Constraint-Aware Enterprise Deal Optimisation

Configure-Price-Quote (CPQ) systems are central to enterprise B2B sales execution, yet most deployments remain rule-based and deterministic, unable to learn from deal outcomes. This paper presents CognitiveCPQ, combining a Heterogeneous Constraint Graph (HCG) with Proximal Policy Optimisation (PPO) for autonomous, constraint-aware deal optimisation. A deterministic kill-switch guarantees zero hard constraint violations; soft constraints are navigated via a z-score-normalised penalised reward. A KernelSHAP explainability layer provides per-decision attribution. Evaluation on a large-scale synthetic simulation demonstrates 18.9% margin uplift, 94.3% constraint compliance, and a 67% reduction in configuration steps over rule-based baselines. We additionally specify two robustness evaluation protocols—response-model misspecification and state-observability noise—as immediate follow-up work; both protocols are expected to confirm graceful degradation given the kill-switch’s structural independence from the response model and state vector. All results represent proof-of-concept performance in simulation; production deployment requires shadow-mode calibration on real deal data.

Authors

Institutions

Publication Details

Journal
AI
Published
2026-10-05
DOI
https://doi.org/10.3390/ai7100407
Primary Topic
Reinforcement Learning in Robotics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

CognitiveCPQ: A Graph-Based Reinforcement Learning Framework for Autonomous Constraint-Aware Enterprise Deal Optimisation

Thottempudi Pardhu, Rajesh Soma
AI
Reinforcement Learning in Robotics
article

CognitiveCPQ: A Graph-Based Reinforcement Learning Framework for Autonomous Constraint-Aware Enterprise Deal Optimisation

Thottempudi Pardhu, Rajesh Soma
article en

Abstract

Configure-Price-Quote (CPQ) systems are central to enterprise B2B sales execution, yet most deployments remain rule-based and deterministic, unable to learn from deal outcomes. This paper presents CognitiveCPQ, combining a Heterogeneous Constraint Graph (HCG) with Proximal Policy Optimisation (PPO) for autonomous, constraint-aware deal optimisation. A deterministic kill-switch guarantees zero hard constraint violations; soft constraints are navigated via a z-score-normalised penalised reward. A KernelSHAP explainability layer provides per-decision attribution. Evaluation on a large-scale synthetic simulation demonstrates 18.9% margin uplift, 94.3% constraint compliance, and a 67% reduction in configuration steps over rule-based baselines. We additionally specify two robustness evaluation protocols—response-model misspecification and state-observability noise—as immediate follow-up work; both protocols are expected to confirm graceful degradation given the kill-switch’s structural independence from the response model and state vector. All results represent proof-of-concept performance in simulation; production deployment requires shadow-mode calibration on real deal data.

AIVol. 7(10)
University of the Cumberlands (US), Koneru Lakshmaiah Education Foundation (IN)
Openalex Percentile: Top 10%
Reinforcement Learning in Robotics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.