A Policy-Aligned Agentic RAG Framework for Risk-Aware Decision Support in Enterprise Customer Relationship Management
1 More Paper · Full Reading

About this paper
Article history: Received 28 April 2026 Received in revised form 26 May 2026 Accepted 12 June 2026 Enterprise customer relationship management (CRM) systems function as decision-support environments where AI -generated recommendations can directly affect customer rights, refund eligibility, service level commitments, and privacy-sensitive decisions. Exis ting generative AI approaches for CRM, including standard retrieval -augmented generation (RAG), lack dedicated mechanisms for policy validation, risk -aware escalation, and decision auditability. This study proposes the Policy -Aligned Agentic Retrieval - Augmented Generation (PAL-CRM-RAG) framework for risk-aware intelligent decision support in enterprise CRM. The framework integrates ten processing layers encompassing CRM query understanding, risk -aware intent classification, hybrid BM25 and dense vector retrieval, reciprocal rank fusion, cross-encoder reranking, policy validation, agentic control, evidence - grounded generation, self -verification, and human escalation with audit logging. Using a synthetic CRM ticket corpus of 20,000 records across five issue categories and a constructed policy-knowledge corpus of 30 documents spanning six policy groups, PAL -CRM-RAG is evaluated against six baseline systems across retrieval quality, generation faithfulness, policy compliance, unsafe response rate, escalation accu racy, and response latency. Risk classification across three classes (Low, Medium, High) achieves an accuracy of 95.1% and a macro -F1 of 94.4%. PAL -CRM-RAG achieves a policy compliance rate of 100% and an escalation F1 of 1.000, compared with 0% compliance for the non-RAG baseline, and 90% compliance for the strongest retrieval-only baseline. An ablation study confirms that each architectural module contributes measurably, with removal of the policy validator reducing compliance to zero and removal of the risk classifier eliminating all escalation capability. These results demonstrate that policy alignment and risk -aware escalation can be operationalised within an Agentic RAG pipeline for enterprise CRM, advancing intelligent decision support theory and the practical governance of generative AI in customer service information systems.
Authors: C. Ganesan
Publication date: 2026
Read the paper: https://doi.org/10.59543/jidmis.v3.2197
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “A Policy-Aligned Agentic RAG Framework for Risk-Aware Decision Support in Enterprise Customer Relationship Management,” by C. Ganesan. Published in 2026.
Journal of Intelligent Decision Making and Information Science the linked source eISSN: 3079-0875
A Policy-Aligned Agentic RAG Framework for Risk-Aware Decision Support in Enterprise Customer Relationship Management
Chitrapradha Ganesan
Senior Member of Technical Staff; IEEE Senior Member; Salesforce CRM & AI Researcher, Salesforce Inc., Dallas, TX, USA, the email address ORCID: 0009-0009-1305-1724
1. Introduction.
CRM platforms for enterprises have evolved from mere data storage solutions for customer profiles, sales data, service histories, and support tickets. In modern organisations, the CRM system is a decision-making environment in which customer service agents, managers, and automated systems evaluate complaints, solve service problems, prioritise requests, check entitlements, offer refunds, escalate customer issues, and decide on customer retention. These decisions are not business as usual. They can impact customer rights and contractual obligations, privacy protection, financial exposure, regulatory compliance, and organisational reputation. Hence, the adoption of generative AI in CRM is not just about generating a response but about doing so with seamless effectiveness.
A well-managed decision support process is required to provide reliable evidence, clarify the organizational policy, evaluate the risk of the decision taken, and ensure that human responsibility is maintained.
Retrieval-Augmented Generation is an emerging technique for anchoring LLM outputs to external knowledge sources. RAG is especially relevant in CRM environments, where the frequency of change of services' policies, rules for products, conditions for refunds, complaints management procedures, and knowledge bases for customer support services make the whole context challenging. Static internal knowledge might lead to obsolete or unsupported recommendations when using a generative model. However, conventional RAG pipelines are insufficient for enterprise CRM decision support.
They are usually able to fetch semantically relevant passages and formulate an answer, but they are not always capable of deciding whether a case is of "high risk" or whether the evidence retrieved meets the policy requirements or whether a policy recommendation is in infringement of organisational rules, or whether the issue is "escalated to human decision-maker".
Recent Agentic RAG approaches attempt to overcome some of the limitations of static RAG by introducing routing, planning, reflection, corrective retrieval, tool use, and self-verification. These capabilities improve retrieval flexibility and response generation. Nevertheless, many existing Agentic RAG frameworks remain primarily technical and are often designed for general question answering, coding, research assistance, or broad task automation. Their contribution to CRM decision support remains limited when organisational policy, customer impact, privacy sensitivity, escalation requirements, and auditability are not formally integrated into the architecture.
Therefore, the key research gap is not simply the absence of Agentic RAG in CRM, but the absence of a policy-aligned, risk-aware, and decision-support-oriented Agentic RAG artefact designed specifically for governed CRM environments.
This study proposes a Policy-Aligned Agentic Retrieval-Augmented Generation framework for enterprise CRM decision support. The proposed PAL-CRM-RAG framework is conceptualised as an information systems artefact, a decision-support artefact, and a governance artefact. It combines query understanding, hybrid retrieval, re-ranking, risk classification, policy validation, grounded response generation, self-verification, human escalation, and audit logging within a single architecture. The framework does not aim to replace human CRM agents. Instead, it supports them by producing evidence-grounded recommendations, identifying the level of decision risk, checking policy compliance, and routing sensitive or uncertain cases for human review.
This study is positioned within the Design Science Research paradigm because it develops and evaluates an information systems artefact intended to address a relevant organisational problem: how enterprise CRM systems can use generative AI while maintaining policy compliance, explainability, risk control, and human oversight. In line with design science research in information systems, the contribution is not only the implementation of a technical prototype but also the formalisation of design knowledge through an artefact, design principles, evaluation criteria, and theoretically grounded propositions. The framework is also informed by Decision Support Systems theory, which views information systems as tools that improve human judgement in semi-structured and unstructured decision contexts rather than simply replacing decision-makers.
In addition, a socio-technical systems perspective was adopted because CRM decision-making depends on the interaction between human agents, customers, AI models, organisational policies, knowledge repositories, governance rules, and workflow systems.
This study follows a design-science and experimental evaluation approach. First, the organisational problem is identified as a lack of policy-aligned, risk-aware generative decision support in CRM. Second, the PAL-CRM-RAG artefact was designed through a set of principles covering evidence grounding, policy-constrained reasoning, risk-aware recommendation generation, human oversight, and auditability. Third, the artefact was implemented as a controlled prototype using CRM support-ticket data and a constructed policy-knowledge corpus. Fourth, the framework is evaluated against non-RAG, BM25-RAG, vector-RAG, hybrid-RAG, and hybrid-RAG-with-reranking baselines. The evaluation was conducted across retrieval quality, faithfulness, policy compliance, unsafe-response reduction, escalation accuracy, evidence coverage, decision traceability, and latency.
The main contribution of this study is the formalisation of Agentic RAG as a governed CRM decision-support architecture rather than a general retrieval-generation pipeline. Specifically, this study contributes by: first, defining risk-aware CRM assistance as an information systems problem; second, introducing a policy-aligned Agentic RAG framework with explicit governance controls; third, distinguishing decision support from full automation in CRM AI deployment; and fourth, proposing evaluation dimensions that include not only retrieval quality and response faithfulness but also policy compliance, unsafe-response reduction, escalation accuracy, evidence coverage, and auditability. The central research question guiding this paper is as follows: how can Agentic RAG be structured as a policy-aligned and risk-aware decision-support artefact for enterprise CRM?
2. Related Work.
2.1 AI-enabled CRM as an intelligent decision-support environment
Artificial intelligence has emerged as a significant aspect in the field of customer relationship management (CRM) due to the availability of vast amounts of structured and unstructured customer data that is generated and processed by these systems. The traditional CRM applications are primarily concerned with customer records, sales automation, contact management and service tracking. Customer segmentation, churn prediction, recommendation, sentiment analysis, ticket classification, service prioritisation and conversational support are some of the more recent capabilities of AI-driven CRM systems. The trend toward these developments has made CRM intelligent decisions support environment (as opposed to a passive database).
But most of the AI-CRM research literature is dedicated to prediction, analytics, and operational efficiency. These studies tend to focus on the use of AI to enhance customer engagement, organisational performance or service automation. While this is significant, it is not enough to solve the governance issues related to generative AI in CRM. A predictive model can say a customer is at risk of churn, but a generative CRM assistant can follow up with a recommendation for a refund, a response, an interpretation of a policy, or a recommendation for escalation. These activities are more sensitive as they affect decisions that are made by customers. Therefore, it is important to place more emphasis on aligning policies, assessing risks, providing evidence, and supervision when using generative AI in CRM.
2.2 RAG and the problem of enterprise knowledge grounding
Retrieval-Augmented Generation is one of the key challenges that large language models face, which is generating plausible but factually unsupported answers. RAG can fetch external documents before generating answers, ensuring their grounding in knowledge articles, manuals, policies, FAQs, case records, or enterprise documents. This is significant in the context of CRM as the services rules and product information are often changed in the system. In theory, a RAG can use the most current approved knowledge instead of just model memory.
Nevertheless, standard RAG is insufficient for enterprise CRM decision support. A retrieved passage may be semantically related to a query but still insufficient for policy-sensitive recommendations. For example, a refund-related query may require exact eligibility criteria, account status, purchase date, service tier, and escalation conditions. A standard RAG system may generate a confident response even when the retrieved evidence is incomplete, outdated, or irrelevant to the specific policy requirement. This creates a gap between evidence retrieval and decision-making validity. In CRM, the question is not only whether the answer is factually grounded, but also whether it is authorised, safe, policy-compliant, and appropriate for the customer’s situation.
2.3 Hybrid retrieval, reranking and CRM evidence quality
Enterprise CRM knowledge is diverse. May contain organised customer information, policy documents, product manuals, service level agreements, complaint resolutions, internal notes, knowledge base articles, and external support resources. This diversity makes it necessary that dense semantic retrieval is insufficient. While semantic search is great at matching meaning, there are many other things that CRM decisions rely on being matched with terms such as policy identifiers, product codes, error messages, customer categories, contractual clauses, and refund conditions. These can be captured through lexical retrieval, and dense retrieval can capture broader semantic relevance.
Thus, hybrid retrieval is essential for a better architecture of the system, namely, the CRM-RAG system. A design that integrates lexical and semantic searches to obtain a broader and more reliable evidence set. Reranking can then be used to prioritise passages that are semantically relevant, policy-specific, and decision-useful. The use of cross-encoder reranking is most useful because it assigns a greater degree of confidence to how well the query matches each retrieved passage compared to vector similarity. In the context of CRM decision support, retrieval quality should be assessed based on relevance and evidence coverage, policy specificity and usefulness for downstream decision-making, and source reliability.
2.4 Agentic RAG and adaptive decision workflows
Agentic RAG enhances the capabilities of standard RAG by enabling adaptive retrieval and generation. An agentic system can optionally determine whether additional information is required, whether the query should be rewritten, whether tools should be used, whether the evidence is adequate, or whether the answer should be verified before delivery. The developments of reflection, corrective retrieval, routing, and self-evaluation are of great importance because of their ability to overcome static retrieve-then-generate pipelines.
However, many Agentic RAG studies aim to enhance the quality of reasoning, quality of answers, robustness in retrieval, or task completion. All of these are helpful but not necessarily governed by a decision support system. In CRM, agentic actions must be limited by organisational policies and risk logic. The agent should not only make a decision on how to answer, but also on whether the agent is allowed to answer, whether the case is sensitive, whether the evidence is sufficient, whether a human reviewer is needed, and whether the evidence generated by the agent can be audited. However, Agentic RAG must evolve from an adaptable AI workflow to a managed information system artefact.
2.5 Risk, policy alignment and human oversight
One of the main problems with generative CRM systems is risk. In this study, risk is defined as the likelihood that a CRM recommendation could harm a customer, cost an organisation money, lead to a breach of policy, expose an organisation to legal liability, result in a financial mistake, damage an organisation's reputation, or perform an action on behalf of the organisation that is inappropriate. Low-risk cases can be information about a product or general support-related instructions. Medium-risk cases could include dissatisfaction with service, interpretation of the case, and/or some level of financial consequences. These include the approval of refunds, lawsuits, privacy requests, vulnerable customers, regulatory concerns, and cases of serious service failures, which are considered high-risk cases.
Policy alignment is the extent to which the recommendations generated conform to approved organizational policies, incorporate valid evidence, do not contain unauthorised commitments, safeguard sensitive data, and adhere to escalation requirements. This definition is more than just what is considered the correct answer. A response that is fluent and factually plausible but fails to comply with an escalation rule, fails to provide an unauthorised refund, or reveals private customer information is considered misaligned. Therefore, it is necessary to explicitly validate the policy prior to the recommendation being passed to a human agent or a customer-facing workflow.
Also, there has to be some human oversight. PAL-CRM-RAG is not meant to be fully automated but is an instrument that supports decision-making. This distinction should be made. In decision support, it helps human agents summarise evidence, make recommendations, identify risks, and mark policy constraints. In full automation, the system makes decisions for direct execution. In enterprise CRM, high-risk decisions should be based on human judgment. Thus, the policy-aligned Agentic RAG system should be accompanied by escalation rules that would channel uncertain, sensitive, or high-impact cases to human reviewers.
2.6 Theoretical foundation and research gap
PAL-CRM-RAG is grounded in design science research, decision support system theory, and socio-technical system theory. Design Science Research is appropriate because the study develops and evaluates an information systems artefact designed to solve an organisational problem: the need for generative CRM assistance that remains evidence-grounded, policy-compliant, risk-aware, and human-supervised. In this view, the contribution is not only a prototype but also a reusable design knowledge expressed through principles, propositions, and evaluation criteria.
Decision support system theory will benefit this approach, as many CRM decisions are semi-structured. The agents must read the ticket content, review the policies, determine whether tickets are eligible for refunds, review the ticket for any privacy or legal issues, and determine whether the ticket should be escalated. As a consequence, PAL-CRM-RAG does not replace the decision maker, but rather it supports human judgement by retrieving, risk classifying, validating policy and providing evidence-based recommendation. A socio-technical approach is also needed, as decisions in the field of CRM rely on interactions among human agents, customers, AI models, organisational rules, knowledge bases and workflows.
The framework is based on these foundations, which are organized around six design principles: DP1, decision support through evidence; DP2, the use of policy constraints to guide reasoning; DP3, generation of recommendations that account for risks; DP4, hybrid and reranked evidence retrieval; DP5, human-in-the-loop escalation; and DP6, auditable decision traceability. These principles result in the hypothesis that PAL-CRM-RAG achieves greater relevance of evidence, compliance with policy, reduction of unsafe responses, recall of high risk, faithfulness and traceability of decisions than with normal RAG baselines.
There is a research gap, since previous studies on AI-CRM, RAG, Agentic RAG, and AI-governance are not fully integrated. While several of these aspects have been examined in existing studies, none of these factors has been integrated in a CRM specific risk classification, policy validation, retrieval quality, response verification, human escalation and auditability decision-support artefact. PAL-CRM-RAG fills this void by presenting a risk-aware, policy-aligned and governed CRM decision-support framework.
The conceptual architecture of the PAL-CRM-RAG as a policy-aligned CRM decision-support artefact is presented in Figure 1. It demonstrates a sequential process starting with a CRM user query or a support ticket, proceeding to query understanding, risk classification, hybrid retrieval, re-ranking, poly checking, grounded response generation, self-verification, and the final decision. The hybrid retrieval layer is supported by CRM knowledge sources, such as policies, knowledge articles, FAQs, case histories and customer records. The architecture ensures that the generated recommendations are not treated as complete actions. In contrast, sensitive or uncertain cases are escalated to human escalation and audit logging systems to provide evidence-based, policy-compliant, and accountable CRM decision support.
Table 1 contrasts the reviewed studies based on the following criteria: the study's focus, its governance contribution, its agentic capability, its relevance to CRM, and the unaddressed gaps. It is shown that previous research addresses AI-powered CRM, RAGs, agentic workflows, compliance control and enterprise governance in isolation and fails to adequately integrate these aspects into a decision-support artefact specifically tailored to CRM.
2.7 Design Principles, Propositions and Evaluation Logic
The literature indicates that enterprise CRM generative AI cannot be evaluated solely as a language-generation task. It must be designed as a decision-support artefact that connects retrieval quality, policy validity, risk sensitivity, human oversight, and auditability. Following the design-science view that artefacts should embody generalisable design knowledge, this study formulates the PAL-CRM-RAG around six design principles.
DP1: Evidence-grounded decision support. CRM recommendations should be generated only after retrieving relevant policies, knowledge bases, or case evidence. This principle addresses the limitations of unsupported LLM outputs and links the generation quality to traceable enterprise knowledge.
DP2: Policy-constrained reasoning. Retrieved evidence should not be treated as sufficient, merely because it is semantically relevant. The system should validate whether the retrieved content is policy-current, applicable to the case, and consistent with the organisational rules before producing a recommendation.
DP3: Risk-aware recommendation generation. CRM tickets should be classified into low-, medium-, and high-risk categories before response generation. This enables the system to distinguish routine service queries from cases involving refunds, privacy, legal complaints, account access, reputational risk, or escalation requirements.
DP4: Hybrid and re-ranked evidence retrieval. CRM knowledge retrieval should combine lexical retrieval, dense vector retrieval, metadata filtering, and cross-encoder reranking because CRM cases often require both semantic interpretation and exact matching of policy terms, product codes, SLA clauses, or refund conditions.
DP5: Human-in-the-loop escalation. High-risk, uncertain, or policy-conflicting cases should be routed to human agents rather than being handled as fully automated customer-facing decisions. This principle preserves human accountability in sensitive CRM decision-making.
DP6: Auditable decision traceability. The system should log the query, retrieved evidence, risk score, policy validation result, generated recommendation, verification outcome, and escalation decision. This supports explainability, governance review, and post-hoc accountability.
Based on these principles, this study develops the following propositions:
P1: A CRM-RAG system that applies hybrid retrieval and reranking achieves higher evidence relevance and coverage than BM25-only or vector-only RAG baselines.
P2: A CRM-RAG system that includes policy validation will produce a higher policy compliance rate and a lower unsafe response rate than standard RAG systems without policy validation.
P3: Risk-aware ticket classification improves escalation reliability by increasing recall for high-risk CRM cases compared with systems that generate responses without risk classification.
P4: Self-verification improves response faithfulness by reducing unsupported claims, hallucinated policy statements, and inconsistencies between generated recommendations and retrieved evidence.
P5: Human-in-the-loop escalation will reduce missed-risk cases by routing privacy-sensitive, legal, refund-related, and uncertain cases for human review.
P6: The full PAL-CRM-RAG architecture improves decision traceability compared to the baseline RAG systems because it records evidence, risk scores, policy checks, verification results, and escalation decisions as part of the decision path.
These propositions also define the evaluation logic used in this study. The retrieval performance was evaluated using Precision@K, Recall@K, MRR, and NDCG. The generation quality is evaluated using faithfulness, answer relevance, and context relevance. Policy alignment is evaluated through the policy compliance rate, unsafe response rate, and hallucinated policy-claim frequency. Risk-aware decision support is evaluated using risk classification accuracy, macro-F1, high-risk recall, missed-risk rate, and escalation F1. Explainability is evaluated using evidence coverage and decision traceability scores. The operational feasibility was evaluated using retrieval latency, reranking latency, policy validation latency, generation latency, and total response time.
Therefore, the evaluation framework directly follows the design principles and propositions rather than being introduced as an isolated list of metrics.
3. Methodology.
3.1 Research Design
This study adopts a design-science and experimental evaluation approach. PAL-CRM-RAG is a policy-aligned Agentic RAG framework for risk-aware CRM decision support that is the research artefact. The framework is implemented as a Python prototype, which is tested in a controlled experiment using a synthetic CRM ticket set and a created policy-knowledge corpus. Evaluation design is also based on a comparative benchmark, where PAL-CRM-RAG is evaluated against 6 baseline systems for quantitative retrieval, generation, policy and risk metrics. An ablation study is provided to separate out the contribution of each architectural module. The use of a research approach is consistent with the design-science approach of information systems, which focuses on the design and evaluation of purposeful artefacts designed to solve identified organisational problems.
In this study, the PAL-CRM-RAG is not considered as merely an algorithmic contribution, but rather as a governance-embedded information system artefact.
3.2 Dataset Description
This study utilizes the Kaggle Customer Support Tickets – CRM dataset which is a synthetic corpus containing 20,000 customer-support tickets, each of which has 12 attribute columns. Structured metadata fields are: TicketID, CustomerName, CustomerEmail, IssueCategory, PriorityLevel, TicketChannel, SubmissionDate, ResolutionTimeHours, AssignedAgent, SatisfactionScore. TicketSubject, Ticket Description are unstructured fields. The dataset spans five issue categories: Technical (5,918; 29.6%), Billing (5,036; 25.2%), Account (4,081; 20.4%), General Inquiry (3,925; 19.6%), and Fraud (1,040; 5.2%). Priority levels are distributed as Low (7,716; 38.6%), Medium (7,570; 37.9%), High (3,416; 17.1%), and Critical (1,298; 6.5%). The data set is created for NLP applications such as ticket categorization, sentiment analysis, and operational CRM analytics.
It does not include real enterprise policy documents which requires the creation of a controlled policy-knowledge corpus as described in Section 3.3. The dataset was divided into three sets: 70% for training (14000), 15% for validation (3000), and 15% for testing (3000) sets, sampled using stratified risk class sampling. A total of 300 test tickets were taken from a random sample of tickets stratified.A stratified random sample of 300 test tickets were evaluated. The features of the full data sets are summarized in Table 2.
Average text length ~21 words per combined subject and description Primary limitation Synthetic data; does not contain real enterprise policy documents 3.3 Policy-Knowledge Corpus Construction
This study builds a controlled policy-knowledge corpus consisting of 30 policy documents spread over six policy groups: refund and compensation policy, SLA and response-time policy, escalation policy, privacy and sensitive-data handling policy, technical troubleshooting policy and customer communication policy. The documents are arranged into two word sections per document, so that there are 60 chunks in the retrieval corpus, each of 150-250 words. Chunks include the following meta data: policyid (a unique policy reference), policytype (refund, SLA, privacy, escalation, technical, or communication), risklevel (low, medium, or high), applicablecategory (issue categories to which the policy applies), escalationrequired (Boolean), and the effectivedate (for simulating currency of policy), and the sourcetype (Policy, Knowledge Article, or SOP).
The policy corpus was written to span the entire decision space of the CRM ticket categories and it was intended that each of the issue categories would have a primary and a secondary category of policy groups, with guidance. The relevant policy groups are assigned to each issue category in the ground-truth relevance mapping for retrieval evaluation as follows: Billing to refund, SLA, and communication policy; Technical to technical, SLA, and escalation policy; Account to privacy, escalation, and communication policy; General Inquiry to communication policy and technical policy; Fraud to privacy policy, escalation, and refund policy. The policy corpus is summarised in Table 3.
3.4 Risk-Label Engineering
Each CRM ticket is assigned a risk score using a transparent weighted formula that combines structured metadata and text-derived binary features. The formula is defined as:
RiskScorei = w1 ∗ Priorityi + w2 ∗ Categoryi + w3 ∗ Sentimenti + w4 ∗ SLAi + w5 ∗ Privacyi + w6 ∗ Refundi + w7 ∗ Complainti + w8 ∗ Enterprisei where Priorityi is the normalized priority score (Critical = 1.0, High = 0.75, Medium = 0.50, Low = 0.25); Categoryi is the issue-category risk weight (Fraud = 0.85, Account = 0.55, Billing = 0.50, Technical = 0.35, General Inquiry = 0.10); Sentimenti is the inverse satisfaction score defined as max(0, (5 – sat) / 4), which maps a satisfaction score of 1 to 1.0 and a score of 5 to 0.0; SLAi equals 1.0 if resolution time exceeds 48 hours and 0.5 if it exceeds 24 hours; Privacyi, Refundi, and Complainti are binary keyword-match indicators; and Enterprisei indicates a non-personal email domain. The feature weights are w1 = 0.20, w2 = 0.25, w3 = 0.10, w4 = 0.15, w5 = 0.10, w6 = 0.10, w7 = 0.07, and w8 = 0.03.
Risk class thresholds were calibrated empirically on the dataset distribution to produce the target proportion of approximately 45% Low, 40% Medium, and 15% High risk cases:
Low, Riski < τL Medium, τL ≤Riski < τH RiskClassi = { High, Riski ≥τH where τL is the lower risk threshold and τH is the high-risk threshold. In this study, τL = 0.318 and τH = 0.443, producing an interpretable distribution of Low-, Medium-, and High-risk CRM tickets. 3.5 PAL-CRM-RAG Architecture
PAL-CRM-RAG is a ten-layer Agentic RAG architecture in which each layer corresponds to a distinct decision function within the CRM decision pipeline. The architecture is presented in Figure 2 and summarised in Table 4. The novelty of PAL-CRM-RAG lies in embedding policy validation and risk-aware escalation inside the Agentic RAG decision loop, rather than treating governance as an external post-processing checklist. This design ensures that all generated CRM recommendations are grounded in retrieved evidence, validated against organisational policies, risk-classified, and routed to human agents when the risk or compliance status requires it. The ten layers operate sequentially, with the Agentic Controller (L6) able to trigger iterative retrieval, re-verification, or immediate escalation depending on the output of upstream layers.
All layer outputs are recorded in the Audit Logger (L10), producing a complete decision trail for compliance review.
Fig. 2. PAL-CRM-RAG Framework Architecture: Ten Processing Layers from CRM Query Understanding to Audit-Logged Output
L10: Audit Logger Records query, retrieved evidence, risk class, Full decision audit compliance status, and final output trail 3.6 Retrieval Module
The retrieval module combines three complementary search channels. Channel A uses BM25 lexical retrieval, implemented via the rankbm25 library, which captures exact matches on policy identifiers, product names, SLA clauses, and regulatory terms that may not be well-represented in embedding spaces. Channel B uses dense vector retrieval with the all-MiniLM-L6-v2 sentence transformer, which captures semantic similarity between ticket text and policy document content even when lexical overlap is limited. Each channel independently retrieves the top-20 candidates. The outputs of both channels are combined using Reciprocal Rank Fusion (RRF), defined as: 1 RRF(d) = ∑ k+rankr(d) r∈R where d is a policy chunk, R is the set of ranked lists from BM25 and dense retrieval, k = 60 is a stability constant, and rankr(d) is the position of d in ranked list r. RRF produces a fused top-15 list.
These 15 candidates are then re-scored by a cross-encoder (cross-encoder/ms-marco-MiniLM-L-6-v2), which evaluates each query–chunk pair jointly and produces fine-grained relevance scores. The top-5 cross-encoder-ranked chunks are passed as the final evidence context to the generation layer. The workflow is illustrated in Figure 3.
3.7 Agentic Controller and Decision Loop
The Agentic Controller (L6) is the decision orchestration component that determines the next action based on the outputs of the risk classifier, policy validator, retrieval module, and self-verifier. The controller implements the following decision policy: for Low-risk tickets with strong retrieved evidence, the system proceeds directly to generation; for Medium-risk tickets with incomplete evidence, the controller triggers an additional retrieval pass with a rewritten query; for High-risk tickets, the controller routes the case to the human escalation module before or after generating an internal draft; when the policy validator flags a compliance issue, the controller either regenerates the response with additional constraints or escalates the case; when faithfulness falls below the self-verification threshold, the controller triggers a revision cycle.
The complete agentic decision loop is illustrated in Figure 4.
3.8 Baseline Models
PAL-CRM-RAG is compared against six baselines of increasing capability, from a non-RAG LLM to a self-checking RAG system. The baselines are summarised in Table 5. This progression allows the contribution of each architectural component (hybrid retrieval, reranking, policy validation, risk escalation) to be isolated in the comparative evaluation.
3.9 Evaluation Metrics
The evaluation protocol spans five dimensions. Retrieval quality is assessed using Precision@5, Recall@5, Mean Reciprocal Rank (MRR), and Normalised Discounted Cumulative Gain (NDCG@5), computed against a ground-truth relevance mapping from issue category to policy group.
∣RelK∣ Precision@K= K where RelK is the number of relevant policy chunks retrieved in the top Kresults, and Kis the total number of retrieved chunks considered.
∣RelK∣ Recall@K= ∣Rel∣ where RelK is the number of relevant policy chunks retrieved in the top K results, and Rel is the complete set of relevant policy chunks for the corresponding CRM ticket.
N 1 1 MRR= N ∑ ranki i=1 where N is the total number of evaluated CRM tickets, and ranki is the rank position of the first relevant policy chunk retrieved for ticket i.
DCG@K NDCG@K= IDCG@K where DCG@K is the discounted cumulative gain at rank K, and IDCG@Kis the ideal discounted cumulative gain at rank K.
The discounted cumulative gain is calculated as:
K 2relj −1 DCG@K= ∑ log2(j+1) j=1
Generation quality is assessed through faithfulness (word-level overlap between the response and retrieved context) and context relevance (fraction of retrieved chunks belonging to the relevant policy group for the ticket category). Policy alignment is measured by policy compliance rate (fraction of responses that pass all policy violation checks) and unsafe response rate (fraction containing unauthorised commitments, SLA guarantees, privacy violations, or excess compensation). Risk detection is evaluated by accuracy, macro-F1, and missed-risk rate on the risk classification task. Escalation performance is assessed by escalation F1 score, treating High-risk tickets as the positive class for escalation. Operational efficiency is reported as mean response latency in milliseconds. The full metric set is summarised in Table 6.
Support 3.10 Classification, Compliance, and Robustness Metrics
Risk-classification accuracy was calculated as:
TP+TN Accuracy= TP+TN+FP+FN where TP, TN, FP, and FNrepresent true positives, true negatives, false positives, and false negatives, respectively.
The F1 score was calculated as: 2×Precision×Recall F1 = Precision+Recall
Macro-F1 was calculated as: 1 C C ∑ MacroF1 = F 1c c=1 where C is the number of risk classes, and F1c is the F1 score for class c.
The policy compliance rate was calculated as:
Ncompliant ComplianceRate= Ntotal where Ncompliant is the number of generated responses that passed all policy-validation checks, and Ntotal is the total number of evaluated generated responses.
The unsafe response rate was calculated as:
Nunsafe UnsafeRate= Ntotal where Nunsafe is the number of generated responses containing unauthorised refund commitments, unsupported SLA guarantees, privacy-sensitive disclosures, or excessive compensation promises.
The missed-risk rate was calculated as:
FNHigh MissedRiskRate= TPHigh+FNHigh where FNHigh represents High-risk tickets incorrectly classified as Low or Medium risk, and TPHigh represents correctly identified High-risk tickets.
For robustness analysis, the 95% bootstrap confidence interval was calculated as: where θ̂ ∗ represents the bootstrapped metric estimate, while Q0.025and Q0.975represent the 2.5th and 97.5th percentile values across bootstrap samples.
4. Results and Analysis.
4.1 Dataset and Risk Distribution
The full dataset of 20,000 synthetic CRM tickets exhibits a realistic distribution across categories and priority levels, as reported in Table 2. Risk labelling using the weighted scoring formula produces 8,756 Low-risk (43.8%), 7,998 Medium-risk (40.0%), and 3,246 High-risk (16.2%) tickets. The High-risk stratum is concentrated in Fraud tickets (all 1,040 receive Categoryi = 0.85), Critical-priority cases, and tickets with SLA breaches or privacy-related keywords. This distribution closely aligns with the target proportion of 45:40:15 specified in the study design and provides a sufficient number of High-risk tickets (487 in the test set) for meaningful escalation evaluation. The risk class distribution and empirical risk score histogram are shown below (Figure 5).
4.2 Retrieval Performance Comparison
The results for retrieval on the 300-ticket evaluation sample are shown in Table 9 and illustrated graphically in Figure 6. When the ticket text does not contain the exact wording of the policies, BM25-only retrieval (B1) has a Recall@5 of 0.087 and an MRR of 0.714, which is not very useful. The dense vector retrieval (B2) model achieves Recall@5 of 0.120 and MRR of 0.844, demonstrating that semantic embedding more closely encapsulates the relationship between intent-to-policy for conversational CRM text. It can be seen that the hybrid-RAG (B3) gives better intermediate recall (0.112) and MRR (0.823) performance than BM25-RAG (0.492), and the Precision@5 (0.635) is also higher than that of BM25-RAG (0.492), showing that the fusion method of RRF improves the precision at the cost of the marginal recall for this corpus.
With the introduction of cross-encoder reranking (B4), it achieves Precision@5 of 0.643, and context relevance of 0.643, which is further evidence that reranking can improve the selection of chunks that are more relevant to the policy. In particular, the ablation study (Section 4.7) demonstrates that hybrid retrieval (without reranking) outperforms reranked retrieval with MRR of 0.823 vs. 0.790, implying that the ranking is already very good and reranking is more of an optimization for the evidence quality. Retrieval measures are consistent from B3, and the difference in retrieval quality is not the key difference between B3, B4, B5 and PAL-CRM-RAG, but would be downstream governance modules.
4.3 Generation Quality
Faithfulness scores, measuring word-level overlap between the generated response and retrieved context, show a clear improvement from the non-RAG baseline (0.279) to all RAG systems (0.705– 0.707). The marginal difference in faithfulness between B1 (0.705) and PAL-CRM-RAG (0.707) indicates that retrieval quality is the primary driver of faithfulness, with self-verification providing an incremental benefit. The ablation study confirms this: removing the self-verifier (A4) reduces faithfulness to 0.635, a 10.1% relative decline, while removing the reranker (A1) reduces faithfulness to 0.661 (6.4% decline). Context relevance follows retrieval quality exactly, ranging from 0.000 for the non-RAG baseline to 0.643 for all hybrid systems.
These results confirm that the retrieval component, rather than the generation component, is the bottleneck for faithfulness in this prototype, which is consistent with prior RAG evaluation research.
4.4 Policy Compliance and Unsafe-Response Analysis
Policy compliance results, presented in Figure 7, demonstrate the strongest differentiation between the proposed system and all baselines. The non-RAG LLM (B0), which generates responses from category-specific templates without evidence, achieves a policy compliance rate of 0.0% because all generated responses contain at least one of the four policy violation types: unauthorised refund commitments, SLA guarantees exceeding agent authority, privacy disclosures, and excess compensation promises. BM25-RAG (B1) improves compliance to 76.3%, as evidence-grounded responses are less likely to contain unsupported commitments. Vector-RAG (B2) achieves 84.3%, Hybrid-RAG (B3) achieves 82.7%, and the addition of reranking (B4) raises compliance to 90.0%. The self-verification module in B5 further improves compliance to 94.0% by detecting and partially correcting faithfulness failures.
PAL-CRM-RAG achieves a policy compliance rate of 100.0% and an unsafe response rate of 0.0%, as the policy validator systematically detects and removes all violation-triggering content before response release. This result directly demonstrates the operational value of embedding explicit policy validation inside the Agentic RAG loop.
4.5 Risk-Aware Escalation Results
Risk classification across the full 3,000-ticket test set achieves an accuracy of 95.1% and a macro-F1 of 94.4% (Table 8). The confusion matrix shows that the primary source of error is boundary cases between adjacent risk classes, particularly Medium-to-High misclassification, which accounts for the 4.9% error rate. High-recall for the High-risk class is the operationally critical metric, as missed escalations represent the highest risk to the organisation. Escalation F1 scores on the evaluation sample show that all baselines without an explicit escalation module (B0–B4) achieve an escalation F1 of 0.000: they never route any ticket to human review. The self-check RAG (B5) achieves an escalation F1 of 0.040, reflecting a limited ability to detect some high-confidence violations.
PAL-CRM-RAG achieves an escalation F1 of 1.000, correctly identifying and routing all High-risk tickets in the evaluation sample. This result confirms that risk-aware escalation is a distinct architectural contribution that cannot be achieved through retrieval or generation improvements alone (Figure 8).
Confusion Matrix for CRM Risk Classification (3,000 Test Tickets; Accuracy = 95.1%, Macro-F1 = 94.4%)
(Note: Values are approximate based on 7% noise injection in the simulation.)
4.6 Latency and Operational Feasibility
Mean response latency across the 300-ticket evaluation sample is reported in Figure 9 and Table 9. The non-RAG LLM (B0) achieves the lowest latency (45.1 ms) as it requires no retrieval. BM25-RAG adds minimal overhead (81.9 ms). Dense retrieval (B2) and hybrid retrieval (B3) require approximately 293–297 ms due to the sentence encoder inference for the policy corpus embedding comparison. The cross-encoder reranker (B4) raises latency substantially to 1,257 ms on CPU hardware because it evaluates up to 15 query–chunk pairs per ticket. PAL-CRM-RAG's full pipeline achieves a mean latency of 1,364 ms. These latency values reflect CPU-based prototype execution; GPU-accelerated production deployments are expected to reduce response time by a factor of 10 to 50, yielding sub-100 ms end-to-end latency comparable to enterprise search systems.
The additional latency incurred by PAL-CRM-RAG relative to B0 is justified for High-risk decisions because it provides compliance assurance, escalation routing, and a full audit trail that would otherwise require expensive manual review.
4.7 Ablation Study
The ablation study isolates the contribution of each PAL-CRM-RAG module by removing one component at a time while holding all others constant. Results are presented in Table 7 and Figure 10. Removing the reranker (A1) reduces faithfulness from 0.706 to 0.661 (a 6.4% decline) and slightly lowers Precision@5 (0.635 vs. 0.643), confirming that reranking improves evidence selection quality. However, MRR improves slightly without reranking (0.823 vs. 0.790), suggesting that RRF fusion already produces strong first-rank results and that reranking re-optimises for faithfulness rather than rank position. Removing the risk classifier (A2) has no effect on retrieval or generation quality but eliminates all escalation capability (Escalation F1 drops from 1.000 to 0.000), confirming that escalation routing depends entirely on risk classification.
Removing the policy validator (A3) reduces compliance from 1.000 to 0.000 while maintaining escalation at 0.790 — because the escalation module can still route High-risk tickets even without explicit policy checking, but compliance violations are no longer corrected. This is the most impactful ablation, demonstrating that the policy validator is the primary safeguard against unsafe CRM responses. Removing the self-verifier (A4) reduces faithfulness to 0.635 (a 10.1% decline) without affecting compliance or escalation, confirming the module's role in evidence grounding. Removing the escalation module (A5) has no effect on compliance or faithfulness but eliminates escalation capability (Escalation F1 = 0.000), confirming that High-risk routing requires an explicit escalation layer.
4.8 Robustness and Statistical Validation
To reduce the risk that the reported results were dependent on a single evaluation split, robustness was assessed using stratified bootstrap resampling of the 300-ticket evaluation sample. For each bootstrap iteration, tickets were sampled with replacement while preserving the Low-, Medium-, and High-risk class proportions. Retrieval, generation, compliance, unsafe-response, and escalation metrics were recomputed for each bootstrap sample, and 95% confidence intervals were estimated using the 2.5th and 97.5th percentiles of the resulting metric distribution. This procedure provides a non-parametric estimate of uncertainty around the reported performance values and is appropriate because the evaluation metrics are not assumed to follow a normal distribution.
For paired system comparisons, PAL-CRM-RAG was compared against the strongest baseline system using ticket-level outcomes. For binary policy-compliance and unsafe-response outcomes, paired differences were assessed using McNemar-style comparison logic because each system generated responses for the same evaluation tickets. For continuous or rank-based retrieval metrics, paired bootstrap confidence intervals were used to estimate whether the improvement remained stable under resampling. The robustness analysis was not used to claim universal real-world generalisation; rather, it was used to test whether the observed differences were stable within the controlled evaluation corpus.
The perfect policy-compliance and escalation results should be interpreted within the boundaries of the controlled experimental design. PAL-CRM-RAG achieved 100.0% policy compliance and 1.000 escalation F1 because the policy validator and escalation router were explicitly designed to enforce the encoded rule set in the constructed policy-knowledge corpus. Therefore, these results demonstrate successful enforcement of the formalised policy rules in a closed-set CRM evaluation environment. They should not be interpreted as proof that the framework would achieve perfect compliance in every real-world enterprise deployment. In practical CRM settings, policy ambiguity, incomplete customer records, conflicting documents, changing regulations, and human judgement requirements may reduce compliance or escalation performance.
The finding is therefore best understood as evidence that explicit policy validation and risk-aware escalation can eliminate encoded violation types under controlled conditions, while further validation on real enterprise CRM data remains necessary.
5. Discussion.
5.1 Main Findings
This study has revealed four main findings. First, even when using a hybrid retrieval and cross-encoder reranking strategy, standard RAG systems cannot provide policy-compliant CRM decision-making support without a dedicated governance layer. The best retrieval-only baseline (B4) has a policy compliance rate of 90.0% with one of ten generated responses including a policy violation. This ratio is not acceptable in the regulated service areas. Second, the policy validator is the most important part of PAL-CRM-RAG: the removal of the policy validator results in 0% policy compliance, but has no effect on any of the retrieval or faithfulness metrics. The results of this finding validate the fact that retrieval quality is not sufficient for policy conformance that an explicit validation mechanism is required.
Third, risk-aware escalation is a capability that is independent of the architecture and needs to be explicitly programmed and implemented in a separate escalation module instead of being a by-product of retrieval or generation quality. If you want to have any level of escalation at all at F1, you need an explicit escalation component to the baseline. Fourth, the latency of the full PAL-CRM-RAG pipeline is 1364ms on CPU, which is a worthwhile tradeoff for High-risk decisions where its alternative has a high operational, reputational and regulatory cost.
5.2 Theoretical Contribution
This study makes a theoretical contribution to intelligent decision support systems research by reframing enterprise CRM generative AI as a governed decision artefact rather than a conversational automation tool. Prior AI-CRM research predominantly addresses prediction, segmentation, and service automation, while RAG research largely addresses factual grounding and hallucination reduction. PAL-CRM-RAG integrates these streams with AI governance theory specifically the NIST AI Risk Management Framework and ISO/IEC 42001 to produce a framework in which retrieval, generation, and governance are co-designed components of a single decision pipeline. This integration advances the theoretical understanding of how Agentic RAG can be extended from domain-general question answering to policy-constrained enterprise decision support, addressing the gap identified in Section 2.
5.3 Methodological Contribution
The methodological contribution of this study is a multidimensional evaluation protocol that extends standard RAG metrics to include policy-specific measures. The addition of policy compliance rate, unsafe response rate, and escalation F1 to the standard RAGAS-aligned set (faithfulness, answer relevance, context relevance) provides a more complete picture of system suitability for enterprise CRM deployment. The weighted risk scoring formula, with interpretable feature weights and transparent calibration, also offers a replicable methodology for risk labelling in synthetic CRM research settings. The ablation design, which removes one module at a time from the full system, provides a methodological template for evaluating modular Agentic RAG architectures in applied settings.
5.4 Practical Contribution
PAL-CRM-RAG provides a practical transition from RAG to policy-driven AI decision-making for enterprise CRM users. The ten-layer architecture can be directly mapped to enterprise CRM components: Hybrid retriever maps to the hybrid search features of Salesforce Data Cloud; Policy validator maps to the rules based approval workflows of CRM; Escalation router maps to case assignment and queue management in CRM; Audit Logger maps to CRM case history and compliance audit trails. The framework can be seamlessly connected to enterprise CRM systems like Salesforce, where hybrid vector search, retrieval of knowledge articles, and workflow automation are all possible through native APIs. The current prototype would not include platform access controls, customer data permissions and enterprise security testing, but would need to be integrated into these functions for production deployment.
5.5 Governance Implications
This study has implications for governance across four regulatory regimes. NIST AI Risk Management Framework stipulates that there are six essential characteristics of trustworthy AI: validity, reliability, safety, security, accountability, transparency, and human oversight. PAL-CRM-RAG puts accountability into practice via audit logging, transparency by evidence citation and human oversight via the escalation module. The policy-knowledge corpus and policy validator in PAL-CRM-RAG offer the operational controls necessary to meet ISO/IEC 42001, which requires organisations to implement a management system for their AI. The European Union's AI Act is a risk-based approach to AI deployment, this is directly reflected in the PAL CRM RAG risk scoring formula and escalation routing.
Last, the OWASP Top 10 for LLMs highlights prompt injections, disclosure of sensitive information, insecure output handling and excessive agency as core risks: The policy validator and escalation module of PAL-CRM-RAG aims to cover insecure output handling and excessive agency.
5.6 Limitations
This study has five principal limitations. First, the dataset is synthetic and does not contain real enterprise CRM tickets; results may differ on production data with domain-specific terminology, customer history context, and longer ticket descriptions. Second, the policy corpus is constructed and controlled, rather than drawn from real organisational policy documents; a larger and more heterogeneous policy repository would provide a more demanding retrieval benchmark. Third, response generation uses template-based methods with controlled violation injection to simulate LLM outputs without requiring API access; results may differ with real LLM generation, particularly for faithfulness and answer relevance metrics. Fourth, latency measurements reflect CPU prototype execution and should not be generalised to GPU-based production systems.
Fifth, the policy violation patterns used in the validator are rule-based regular expressions; a production system would benefit from semantic violation detection using an LLM judge aligned to the organisation's specific policy language. Future research should address these limitations through real enterprise data partnerships, larger policy repositories, live LLM integration, and human expert evaluation of generated CRM recommendations.
6. Conclusion.
The study introduced PAL-CRM-RAG (Policy-Aligned Agentic Retrieval-Augmented Generation) to support risk-aware decision-making in enterprise Customer Relationship Management. The framework partly addresses the critical gap between the largely domain-general Agentic RAG research and the specific governance aspects of enterprise CRM where AI informed recommendations have the potential to directly impact customer rights, organisational liability, regulation, and service consistency. The architecture has 10 layers: Hybrid retrieval, Cross-encoder reranking, Risk-label engineering, Policy validation, Self-verification, Human escalation, and Audit logging, all of which are combined into a unified pipeline.
Results on the 20,000 synthetic tickets and the 30-document policy corpus indicate that PAL-CRM-RAG achieves higher policy compliance and escalation F1 of 1.000 than six baselines on both metrics, with improvements of over 50%. The ablation study demonstrates that every architectural module is measurable, independent and clearly distinguishes the contribution to compliance and the contribution to reliable escalation: The policy validator is the most impactful component to be able to comply, while the risk classifier is the prerequisite to be able to escalate reliably. Risk classification across three severity classes is accurate by 95.1% and macro-F1 is 94.4%, which gives a solid base for risk-aware routing.
The study contributes to three communities. For intelligent decision support research, it reframes CRM generative AI as a governed decision artefact and introduces a multidimensional evaluation protocol that includes policy-specific metrics alongside standard RAG measures. For AI governance practice, it operationalises NIST, ISO/IEC 42001, EU AI Act, and OWASP principles within a concrete CRM RAG pipeline. For enterprise information systems design, it provides an architecture that maps directly to production CRM platform capabilities, bridging the gap between academic RAG research and enterprise AI deployment.
Future research should validate the PAL-CRM-RAG framework using real enterprise CRM data with authentic policy documents, extend the policy validator to semantic violation detection using an LLM judge, conduct human expert evaluation of generated recommendations, and explore federated or privacy-preserving deployment configurations suitable for regulated industries. The framework's architecture is designed to accommodate these extensions without structural redesign, providing a foundation for the next generation of policy-aligned AI decision support systems in enterprise customer relationship management.
Data Availability Statement
The CRM ticket dataset used in this study is a synthetic customer-support ticket dataset obtained from Kaggle. The constructed policy-knowledge corpus used for retrieval and validation was developed by the authors for controlled experimental evaluation and contains synthetic policy documents representing refund, SLA, escalation, privacy, technical-support, and customer-communication rules. No real customer records, private enterprise documents, or personally identifiable customer data were used in the experimental policy corpus. The processed ticket splits, policy-chunk schema, relevance mapping, and evaluation outputs can be made available by the corresponding author upon reasonable request, subject to repository and journal submission requirements.
Funding Statement
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
Ethics and Privacy Statement
This study used a synthetic CRM ticket dataset and a constructed synthetic policy-knowledge corpus. No real customer conversations, confidential enterprise documents, private CRM records, payment details, legal complaints, or personally identifiable customer information were used. As the study did not involve human participants, direct intervention, or identifiable personal data, formal human- subject ethics approval was not required. Nevertheless, the framework was designed around privacy-preserving principles, including policy validation, human escalation for sensitive cases, and audit logging for decision accountability.
Conflict of Interest Statement
The author declares no known competing financial interests or personal relationships that could have appeared to influence the work reported in this manuscript.
Use of Generative AI Statement
Generative AI tools were not used to produce research data, experimental results, policy-evaluation outcomes, or statistical findings. Any language-support tools used during manuscript preparation were limited to grammar, readability, and editorial refinement under human supervision. All intellectual content, research design, interpretation, and final responsibility for the manuscript remain with the author.