Continuous Business Process Improvement Driven by Large Language Models
1 More Paper · Full Reading

About this paper
A full audio edition of this paper.
Authors: A.S. Araujo, H.S. Mamede, V. Santos, V. Filipe
Publication date: 2026
Read the paper: https://doi.org/10.1109/access.2026.3694469
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “Continuous Business Process Improvement Driven by Large Language Models,” by A.S. Araujo and colleagues. Published in 2026.
Abstract.
Some of the main challenges faced by organizations when applying Continuous Business Process Improvement are data fragmentation, limited explainability, weak governance, and the isolated use of Artificial Intelligence in Business Process Management. This study initially conducts a Systematic Literature Review on the topic of business process improvement enabled by Large Language Models or Artificial Intelligence in organizations, presenting a comprehensive analysis of prevailing research trends, conceptual frameworks, and persistent limitations, identifying seventeen recurring gaps that affect the effectiveness of integrating the capabilities of Large Language Models and other Artificial Intelligence technologies throughout the entire lifecycle of Continuous Business Process Improvement.
As a result, we propose a Framework and its gap-oriented reference architecture that, through modular components, facilitates data integration, reasoning, validation, execution, and monitoring within a closed loop of continuous business process improvement. The framework is operationalized through six phases: Process Understanding, Process Diagnosis, Process Redesign, Process Validation, Process Execution Support, and Continuous Monitoring. The results suggest that designing the framework and architecture directly from the identified gaps creates a coherent foundation for AI-driven process improvement, enabling more reliable, explainable, and easily governed and managed solutions. The study improves the current state of the art by creating a cohesive framework for intelligent, scalable, lifecycle-integrated, and operationally deployable process optimization systems.
Introduction.
Business processes are collections of related structured activ-ities or tasks that are used to generate a product or service. They determine how the work should be done within an orga-nization, becoming the building blocks of the organization. Consequently, business processes are vital to the operation of an organization. Due to their use, resources can be used more efficiently, products and services can be delivered on time, and quality standards are reached. There-fore, it is important for organizations to properly manage their business processes.
The associate editor coordinating the review of this manuscript and approving it for publication was Sedat Akleylek.
In this context, Business Process Management (BPM) is a systematic methodology to assess, plan, implement, mon-itor, and optimize business processes. It helps organizations streamline operations, minimize expenses, improve customer satisfaction, and ensure compliance while focusing on governance and strategy. Within the BPM lifecycle, Business Process Improvement (BPI) is a systematic approach to help an organization optimize its underlying processes for more efficient results. Currently, organizations are increasingly dependent on BPI to maintain efficiency, service quality, regulatory compliance, and adaptability in rapidly changing environments. In this context, with the advancement of Artifi-cial Intelligence (AI) technologies, organizations have started to use these technologies, such as Large Language Models
(LLMs), as a boost to promote the continuous improvement of business processes.
A Continuous Business Process Improvement (CBPI) approach seeks to maintain and improve organizational per-formance. In real settings, CBPI depends on the com-bination empirical execution evidence (e.g., event logs, KPIs, historical baselines) with contextual and normative knowl-edge (e.g., policies, procedures, regulations, risk controls) to produce diagnostic insights and actionable redesign decisions. Recent advances in AI, especially LLMs when com-bined with Retrieval-Augmented Generation (RAG), bring new opportunities for BPM and BPI, since these models have the ability to handle unstructured data, create summaries of operational evidence and textual explanations that aid in decision-making processes, particularly when related to business process redesign.. However, LLM-CBPI ini-tiatives frequently struggle with three interrelated challenges:
1) integration of heterogeneous structured and unstruc-.
tured information into a coherent end-to-end improve-ment workflow;
2) limited explainability, traceability, and governance.
of improvement recommendations—especially under compliance and risk constraints; and
3) insufficient support for closed-loop operation, where.
exante validation of redesign alternatives and ex-post monitoring evidence systematically trigger subsequent improvement cycles.
Moreover, LLMs applications related to process improve-ment are still fragmented in their focus on specific tasks. Existing studies often focus on isolated use cases, and little attention has been paid to how LLM-based capabilities can be systematically incorporated throughout the CBPI lifecycle. As a result, organizations can experiment with AI-enabled business process improvement tools without a coherent archi-tecture that considers the end-to-end improvement process.
In this context, it is necessary to better understand this sce-nario. To this end, this study initially conducts a Systematic Literature Review (SLR) on research in AI-enabled process improvement. The SLR presents the current state of the art on this topic and identifies seventeen recurring gaps that limit the effective adoption of AI and LLMs in CBPI environments.
Based on the results of the SLR, this study also presents a conceptual framework, the LLM-CBPI Framework and its gap-oriented reference architecture, as a proposal to address the gaps identified in the SLR. The Framework consists of six integrated phases that support explainable, governance-driven, and closed-loop improvement cycles. The reference architecture is based on elements widely used and empirically demonstrated in multiple contexts, such as Human-in-the-Loop, LLM/RAG Pipeline, ETL/ELT Pipeline, and Process Discovery.
Developing the solution directly from the identified gaps creates a more coherent basis for AI-driven process improve-ment, as aligning improvement needs with architectural decisions results in a more reliable, explainable, governable, and operationally manageable design. In this context, the main contributions of this article are the presentation of an SLR that shows the state of the art on AI-enabled BPI research and identifies seventeen recurring gaps in this topic; the presentation of a six-phase workflow framework that supports explainable, governance-driven, and closed-loop improve-ment cycles and its gap-oriented reference architecture to integrate LLMs into the CBPI lifecycle, addressing the gaps identified in the SLR.
II. BUSINESS PROCESS IMPROVEMENT.
Business Process Management (BPM) helps organizations streamline operations, minimize expenses, improve customer satisfaction, and ensure compliance while focusing on gov-ernance and strategy. According to Dumas et al., the key aspects of BPM are: Identifying and Documenting Existing Business Processes, Sing Technology to Automate Repetitive Tasks, Tracking Process Efficiency Using Key Performance Indicators (KPIs), and Enhancing Processes Through Methodologies such as Six Sigma, Lean, and Pro-cess Mining.
BPM relies on data-driven optimization, employing Six Sigma to minimize variability, Lean to eliminate waste, and Process Mining to examine real-time operations. To maintain improvements, enterprises should implement performance monitoring tools such as AI-driven predictive analytics, KPI tracking, and Business Activity Monitoring (BAM) dashboards to guarantee ongoing optimization and compliance.
In this context, Business Process Improvement involves analyzing current processes, identifying areas for improve-ment and implementing changes to improve performance, quality, and reducing costs. BPI is a component of BPM that aims to refine business processes to make them more efficient (doing things right) and practical (doing the right things). Several methodologies and technologies are commonly used in BPI, each offering a unique approach to process improvement. In many cases, AI models are widely employed. AI algorithms can analyze large datasets to recognize trends, forecast demand, and predict customer behavior.
This information is essential to the success of the decision-enabling process in order to anticipate market changes. More-over, by sifting through large amounts of data, AI can provide actionable insights, helping organizations optimize opera-tions and improve customer experiences.
In this sense, techniques such as Machine Learning (ML), Robotic Process Automation (RPA), Natural Language Pro-cessing (NLP), Process Mining, and LLMs have been widely used. ML techniques can discern patterns in business processes, optimize workflows, predict customer behav-ior, forecast demand, and detect anomalies, whereas RPA automates repetitive tasks such as data entry, invoicing, and document processing, excelling with structured data and rule-based workflows. RPA can be augmented with AI (Intelligent RPA) to facilitate decision making.
NLP technologies automate text-centric operations such as customer service chatbots and sentiment analysis, extract insights from unstructured data (emails, documents, and social media), and improve document classification and con-tract analysis. Process Mining approaches employ AI to examine corporate workflows, identify inefficiencies, offer insights into bottle-necks and automation opportunities, and improve BPM. LLMs, such as Generative Pre-trained Transformer 4 (GPT-4), Bidirectional Encoder Representa-tions from Transformers (BERT), and Large Language Model Meta AI (LLaMA), have transformed BPI by improving automation, decision making, and customer engagement. They excel at man-aging unstructured data, automating work-flows, and enhancing human productivity.
III. LARGE LANGUAGE MODELS.
Large Language Models are advanced AI models created to comprehend, produce and modify human language. LLMs such as GPT-3, GPT-4, and Pathways Language Model (PaLM) are based on Deep Learning (DL) architectures, par-ticularly the Transformer Model, and use large volumes of text data to produce coherent text, predict word sequences, and even perform complex reasoning tasks. Among their capabilities are Reinforcement Learning From Human Feedback (RLHF), which enhances performance and aligns responses with human preferences, and in-context learning, where models generate outputs tailored to specific prompts or stimuli.
Many aspects of daily life have changed as a result of the use of Large Language Models (LLMs). Language translation, sentiment analysis, personalized instruction, and virtual assistance are just a few of the many applications where LLMs have demonstrated utility. In the work-place, LLMs help physicians diagnose illnesses and evaluate medical literature, while also streamlining corporate procedures and automating administrative duties. In addition, they help software developers with coding and de-bugging, assist marketers with content creation and consumer trend analysis, and improve education through AI-driven tutoring and adaptive learning. Finally, LLMs support banking services, investment analysis, contract drafting, and legal research in the legal and financial domains.
Recently, organizations have started using LLMs to improve business processes. But to do so, it is necessary to have a well-structured process, understand and address the challenges and opportunities this approach presents. These organizations also face several constraints, including the need for high-quality data and data privacy compliance, integration with legacy systems, and significant initial and ongoing costs. Businesses also grapple with skill gaps that require specialized talent and training. Algorithmic bias and the need for transparency are ethical concerns, along with regulatory complexity and uncertainty regarding liability for AI-driven decisions.
Consequently, to take advantage of the opportunities offered by these approaches, it is essential to understand their characteristics and technical requirements,, as well as the challenges and restrictive aspects involved, including integration, cost, and data governance. Furthermore, it is also important to analyze their advantages and disadvantages, such as the potential for automation vs. constraints in inter-pretability, and to identify opportunities for improve-ment and refinement, especially in the integration of LLMs with robotic and intelligent automation methodologies.
IV. RESEARCH PROBLEM.
Taking into account the above, the introduction of LLMs has transformed several fields. Business processes often involve decision making under pressure, unstructured data process-ing, and repetitive duties. These fields offer a special oppor-tunity for innovation because they closely match the strengths of LLMs. The potential of LLMs is a topic of grow-ing interest among academics. Assessing LLMs is essential because it helps us understand their advantages, disadvan-tages, and effects on society and business.
Thus, we formulated the Main Research Question (MRQ) based on the ‘‘Population and its Problem, Intervention or Issue, Comparative Intervention, Results or Themes and Con-text’’ Framework proposed by Booth et al.:
MRQ: In organizational contexts of BPI, how does the integration of LLM/RAG capabilities with process ana-lytics, compared to traditional BPI approaches, ensure explainability, governance, and closed-loop effectiveness of BPI throughout its lifecycle? To answer the RQ, we first conducted a systematic literature review to investigate the current state of the art in BPI using LLMs in order to understand in which business domains and organizational processes LLMs have been used most efficiently to improve business workflows and task performance, the results that have been achieved by adopting LLMs to improve business processes workflows, and the main challenges for business sectors to adopt LLMs to improve business processes.
We also identified a set of gaps that hinder the systematic imple-mentation of LLMs in BPI, in order to enhance the explain-ability, governance, and effectiveness of a continuous closed loop in organizations. After that, we propose a conceptual framework and a reference architecture to address these gaps.
V. RELATED SURVEYS.
In order to address the MRQ subject of this study, we first applied the method described in Section VI and a set of surveys was recovered from the most pertinent digital libraries: ACM Digital Library, Elsevier, Institute of Elec-trical and Electronics Engineers (IEEE), Science Direct, Scopus, Springer Link, Web of Science, and Wiley Online Library. In addition, we defined and applied the subsequent Assessment Criteria (AC) to assess surveys.
• AC-01 Business Process Coverage: Does the study encompass a comprehensive array of BPM tasks (e.g. automation, decision making, compliance)?
• AC-02 LLM Implementation Strategies: Does the re-search address optimal implementation strategies, including fine-tuning, reinforcement learning, or hybrid AI methodologies?
• AC-03 Business-Specific Applications: Does the study examine the use of LLM across various business sectors (e.g. healthcare, banking, retail)?
• AC-04 Methodological Rigor: Does the study employ a systematic literature review or a structured research methodology?
• AC-05 Practical Recommendations: Does the study pro-vide actionable insights for organizations that integrate LLMs in BPM?
• AC-06 Ethical and Regulatory Considerations: Does the study address obstacles related to AI governance, data privacy, and ethical issues?
No single survey was found that addresses the MRQ pro-posed in this study. In this regard, we found seven surveys that touch on the topic of our investigation in part: Vidgof and Bachhofner, Grohs et al., Wang et al., Kourani et al., Zhao et al., Kampik et al., and Ziche and Apruzzese. Therefore, Tables 1 and 2 list the evaluation results.
A. OUR CONTRIBUTION
This work provides the most comprehensive cross-sectoral state-of-art of LLM-enabled BPI to date, addressing chal-lenges, governance constraints, across 49 qualified studies, in contrast to previous reviews that focus on isolated tasks, domains, or conceptual discussions. Unlike previous studies that focus predominantly on isolated activities, the solution proposed in this study adopts an end-to-end lifecycle per-spective, integrating the phases of understanding, diagnosis, redesign, validation, execution support and continuous mon-itoring of the business process.
The study delineates seventeen gaps that hinder the system-atic implementation of LLMs in BPM and BPI and proposes two novel scientific contributions to address these gaps.
A gap-oriented approach was used both in the develop-ment of the framework and in its reference architecture. This approach provides strong traceability between evidence from the literature, identified needs, and design decisions, unlike previous studies that generally introduce AI features without a systematic design justification.
Another striking characteristic of this project is the inte-gration of reasoning, governance, explainability, simulation, operational deployment, and monitoring components into a single modular architecture structure. This differs from most similar projects that present structures focused solely on analytical support. Similarly, opting for a closed-loop logical improvement structure, in which evidence monitoring con-tinuously informs new cycles of diagnosis and redesign, the proposed solution allows for adaptive and iterative process improvement, rather than one-off interventions as identified in most of the analyzed studies.
Finally, this study contributes a practical and research-oriented basis for future empirical implementations of AI-enabled CBPI systems that are reliable, explainable, and scalable.
VI. DATA AND METHODS.
An Systematic Literature Review (SLR) is a review process of existing research on a particular topic using a rigorous and comprehensive approach. It involves identifying, evaluating, and synthesizing all relevant studies to answer a specific research question or set of questions. Unlike a traditional literature review, which can rely on the knowledge and biases of the researcher, an SLR follows a predefined protocol that ensures a thorough and unbiased literature search. Con-ducting an SLR involves several steps, including defining the research question, searching for relevant studies using a predetermined set of criteria, assessing the quality of the studies and synthesizing the results.
A. A SYSTEMATIC LITERATURE REVIEW METHODOLOGY
The protocol for conducting systematic reviews proposed by Kitchenham was followed to carry out this study. The first step is to define the research question or topic that guides the entire review process. Next, a comprehensive search strat-egy is developed to identify all relevant studies. This includes searching multiple databases and using specific keywords and inclusion/exclusion criteria to narrow the results. After identifying all relevant studies, they undergo screening based on their abstracts and titles to determine if they meet the inclusion criteria. The full-text articles are then retrieved for the remaining studies and the quality and relevance are assessed. Finally, data are extracted from the selected stud-ies and synthesized to answer the research question. Figure 1 depicts the SLR protocol proposed by Kitchenham and followed to perform this review.
In order to do that, a search process was carried out by using Harzing’s Publish or Perish (Windows GUI Edition) 8.12.4612.8838 software, to search the main avail-able data sources such as ACM Digital Library, Elsevier, IEEE, Science Direct, Scopus, Springer Link, Web of Sci-ence, and Wiley Online Library, as well as the platforms of specific conference proceedings and journals. Figure 2 depicts the workflow of the search process.
In order to identify the publications to be analyzed, a search string was defined and a single search string algorithm was created using conventional Boolean operators to make the following query: (‘‘Business Process Improvement’’ OR ‘‘Business Process Management’’ OR ‘‘Business Pro-cess Reengineering’’) AND (‘‘Large Language Models’’ OR ‘‘Generative AI’’ OR ‘‘Generative Artificial Intelli-gence’’). When running the algorithm, we identified that the documents retrieved that directly reference LLMs date back only to 2020. Consequently, we define a search period from 2020 to 2025.
B. SCREENING AND ELIGIBILITY
Continuing to execute the workflows, the inclusion/exclusion criteria were used to filter documents, followed by the exe-cution of the search string algorithm. The objective is to ensure that the selection process is systematic, unbiased, and transparent. Table 3 lists the Inclusion / Exclusion Criteria (IEC). Then each document was reviewed using these criteria in the order in which they are listed in Table 3. The objective is to ensure relevance and quality.
As shown in Table 3, studies that met criteria IEC-01 to IEC-06 undergo a quality assessment (IEC-07) and the results are compiled into a search report. Studies that scored less than 6 out of 12 in the Quality Assessment Criteria (QAC) listed in Table 4 were rejected.
In order to proceed with the analysis, a three-step approach was used to read the research articles based on the method proposed by Keshav.
As a result, 534 documents were retrieved. After removing 34 duplicates, 28 studies did not meet IEC-01 (corrections and retractions). Consequently, 472 records were screened. After that, 13 studies were identified to not meet IEC-02 (non-English publications), 57 studies did not meet IEC-03 (non-academic sources), 87 studies did not meet IEC-04 (studies published by potential predatory journals or publish-ers, ) and 45 studies did not meet IEC-05 (studies not accessible by public or academic account). Finally, 128 stud-ies did not meet IEC-06 (studies that do not correspond to the topic under study) and 93 studies did not meet IEC-07 (studies that did not meet qualitative assessment).
It important to note that due to the rapid evolution of re-search in the field of LLMs, the returned results include pre-publication sources, such as arXiv. Peer-reviewed sources were prioritized throughout the manuscript to ensure aca-demic rigor and relevance. However, pre-publication sources were used selectively and only in emerging and specific topics where peer-reviewed evidence is still limited or where recent technical developments have not yet been fully reflected in indexed journal publications. It is noteworthy that whenever suitable peer-reviewed alternatives were avail-able, references were re-evaluated and updated accordingly. Figure 4 depicts the workflow of the systematic review process.
C. INCLUDED STUDIES
Since we want to understand how the integration of LLM/RAG capabilities with process analytics affects the explainability, governance, and closed-loop effectiveness of BPI, the set of Research Questions (RQ) listed in Table 5 was formulated. The purpose of this study is to be able to answer these research questions.
Consequently, according to the research problem statement described in Section IV, the 49 selected studies underwent an assessment to respond to these RQ. Some selected studies address adjacent or peripheral themes related to AI-enabled process improvement. It is important to note that these studies were identified by the search strategy and prede-fined selection criteria, and not through manual reference selection. Although not all focus directly on the topic of AI-enabled continuous process improvement or LLM, they help to illustrate the current scenario more broadly where these AI-enabled continuous process improvement or LLM approaches are being adopted, which does not alter the main analytical focus of this study.
To obtain reliable and reproducible findings from the final set of 49 included studies, we used a dual analytical approach, integrating qualitative content analysis with bibliometric and scientific mapping methodologies. This hybrid methodol-ogy facilitated both conceptual integration and quantitative assessment of the literature environment.
D. QUALITATIVE CONTENT ANALYSIS
In order to analyze the text of the sources in the final sample to obtain the results, we compiled a systematic content anal-ysis methodology from Mayring, Bengtsson, and Kuckartz to rigorously analyze the complete text of each study.
This procedure encompassed three fundamental phases:
1) Data Extraction: Each study was assessed using a stan-.
dardized extraction template corresponding to the four research questions (RQ01 - RQ04). Essential informa-tion was recorded, encompassing business domains, addressed processes, AI approaches (such as rapid engineering and RLHF), deployment strategies, empirical results, and recognized obstacles.
2) Open Coding and Thematic Classification: A manual.
open coding procedure was performed on pertinent content. The codes were developed iteratively and classified into subject areas such as ‘Prompt Engi-neering for Semantic Enhancement’’, ‘‘Hybrid AI Integration’’, and ‘‘RLHF’’. This inductive-deductive methodology guaranteed coherence with both emerg-ing discoveries and established analytical frameworks.
3) Cross-Case Synthesis: Coded data were amalgamated.
to juxtapose findings across studies, discern prevailing trends, and differentiate outcomes within sectors or methodologies. Frequency counts were used, if appli-cable, to measure patterns (e.g., the number of studies utilizing LLMs in healthcare or implementing process mining).
This qualitative process facilitated a comprehensive anal-ysis of the implementation of LLMs to enhance business processes while maintaining transparency and consistency.
VII. BIBLIOMETRIC ANALYSIS.
Bibliometric analysis evaluates research output using quantitative methods to measure the impact of research in a particular field or discipline. It can evaluate the performance of individual researchers, departments, and institutions. The goal of bibliometric analysis is to provide information on trends and patterns of research output and to identify influen-tial authors and publications.
To assess and analyze the scientific domains of the included studies, the following indicators provided by SCImago were used: Quartile, H-Index, and Scimago Journal Rank (SJR). These indicators aim to provide a more accurate and meaningful measure of the scientific prestige of the journal by considering various factors beyond simple citation counts. Table 6 lists the overall dataset with the respective quartiles and H-index scores retrieved from SCImago and Google Scholar.
Donthu et al. state that the main techniques for a bibliometric analysis are Performance Analysis and Science Mapping. According to the authors, while Performance Anal-ysis considers the contributions of each research constituent, Scientific Mapping concentrates on the connections between research constituents. It involves visualizing and understand-ing the structure and dynamics of scientific knowledge. In order to do that, the VOSviewer tool,, Google Scholar,, and Google Looker Studio were used to create visual analytics dashboards based on the data recovered by the query algorithm presented in Section VI-A. Accordingly, we present below a series of indicators sug-gested by Donthu et al., Guerrero-Bote and Moya-
Anegón, and SCImago, in order to perform the bibliometric analysis with the included studies.
A. PERFORMANCE ANALYSIS
Performance analysis aims to quantify the contributions and impact of these various entities in the research landscape. This helps identify the leading researchers, institutions and journals, as well as trends in research productivity and influence.
In order to do that, we analyzed the number of studies pub-lished yearly. Initially with little research activity in 2020 and 2022, when only one article was published each year, the data show a notable rise trend over time. Following this, there was a notable increase in 2023, with a peak of 21 studies, suggest-ing that the value of LLMs in improving business processes in organizations was becoming increasingly acknowledged. In 2024 there was consistent and increasing research activity in this area, as evidenced by the number of publications, which increased to 24 in 2024. This ongoing interest shows how relevant this topic is.
Continuing the performance analysis, the values of the indicators listed in Table 6 were analyzed. This led us to deliver two analyzes: Journal Data Analysis and Author H-Index. These analyzes and indexes are important for under-standing the landscape of academic publishing and for guid-ing research-related decisions.
1) JOURNAL DATA ANALYSIS.
Journal Data Analysis refers to the examination and interpre-tation of data collected from journals. The quartile metric is used to categorize journals according to their relative position within a specific field. It means raking journals by impact using quartiles (Q1–Q4). The Q1 journals are the most pres-tigious, highly cited, and influential. The Q2 journals have strong impact and above-average citations. The Q3 journals have moderate influence, receiving fewer citations than Q1 and Q2. The Q4 journals are newer or highly specialized and have the lowest impact. Quartile rankings use metrics such as Impact Factor and SJR.
As we can see in Table 6, most of the articles (35 articles - 71%) were published in a journal that belongs to a quartile and most of them (26 articles - 53%) were published in Q1 journals. 6 articles (12,2%) were published in Q2 journals and 3 articles were published in Q3 journals. It is impor-tant to note that the set of 14 studies that were published in unavailable quartile journals refers to articles presented at renowned international conferences whose proceedings were also published by renowned publishers such as Springer Nature, IEEE, Elsevier, ACM Digital Library, and arXiv.org, as stated in Section VI-B
2) AUTHOR H-INDEX.
The H-index is a metric used to measure productivity and impact of citation. Hirsch states that the H-index evaluates the weight, significance, and overall influence of the total research contributions of a scientist. According to Hirsch, an H-index of 20 characterizes a successful scientist. An H-index of 40 characterizes out-standing scientists, and an H-index of 60 characterizes truly unique individuals.
As we can see in Table 6, in most of the analyzed studies (30 studies – 61.22%) at least one of the authors has an H-index greater than 20. In 23 studies (47%) at least one of the authors has an H-index greater than 40, and in 16 studies (32,7%), at least one of the authors has an H index greater than 60.
3) ANALYSIS OF AUTHOR AFFILIATION DATA.
The analysis of the author affiliation involves examining information related to the affiliations of authors who have contributed to academic publications. This type of analysis can provide information on various aspects of research and academic collaboration.
Consequently, analyzing the set of 49 articles that are the subject of this study, we can identify that the affiliations of the majority of authors are from Germany, the United States of America, Romania, Slovakia, Australia and the Netherlands.
4) CITATION ANALYSIS.
A citation analysis examines how often other works have cited the article to determine its impact and influence. A higher citation count indicates a more significant influence and recognition within the field. Data was recovered from the reliable and widely used Google Scholar search engine. The included articles have been referenced, ranging from 1930 to 0 times. Table 7 lists the ten most referenced articles.
5) CONCLUSION.
Performance analysis reveals that the research landscape is marked by rapid expansion and significant scholarly cred-ibility. The significant increase in publication volume after 2022 indicates that LLM-driven process improvement has evolved from an emerging topic to an established research focus. The prevalence of publications in Q1 journals and the significant proportion of contributions by authors with high H-index values offer strong evidence of intellectual maturity and academic credibility in the discipline.
The geographical distribution of institutional affiliations indicates a concentration of research effort in technologically sophisticated countries, suggesting that the field is now influ-enced by a limited number of innovation ecosystems. Cita-tion patterns further underscore the field’s prominence, with numerous works exhibiting remarkably high effect. Nonethe-less, the broad range of citations indicates uneven dissemina-tion of knowledge and disparate levels of consolidation within subtopics.
In summary, these identified robust performance metrics provide substantial evidence that the investigated research is credible and of high academic quality.
B. SCIENCE MAPPING
According to Donthu et al., scientific mapping visualizes and analyzes the relationships and interactions between these entities to understand the structure and evolution of scientific knowledge. This helps in understanding the intellectual struc-ture of a field and identifying core topics, emerging trends, and potential areas for future research. Therefore, we present the analysis of the relationships and interactions between the studies included below: Keyword Co-occurrence Analysis.
1) KEYWORDS CO-OCCURRENCE ANALYSIS.
In research and bibliometric analysis, a keyword is a term or phrase that captures the essential topics, concepts, or themes of a research paper, article, or document. Using VOSviewer,, a visual and analytical perspective was created on the interrelationships of various concepts and Keywords addressed in the included studies. Figure 5 depicts the visualization of the co-keyword network based on occurrences, and Figure 6 depicts the visualization of the co-keyword network based on occurrences and average publication per year scores between 2022 and 2025.
As depicted in Figure 5, there are six keyword clusters based on their patterns of co-occurrence. Taking into account this image, it is evident that the centrality of Generative AI and LLMs (Green & Red Clusters) suggests that these subjects predominate in the dataset. Strong ties between AI and digitalization (Yellow Cluster) and BPI (Purple Cluster) illustrate the use of AI in automation and BPI. The prominent positions of ‘‘Generative AI’’ and ‘‘LLMs’’ demonstrate how these subjects are propelling innovation and research. The links to ‘‘decision making’’, ‘‘BPM’’, and ‘‘service individu-alization’’ suggest that AI’s potential effects on digital trans-formation and corporate efficiency are being investigated.
In contrast to Figure 5, Figure 6 shows the average year of keyword occurrence with a color-gradient bar in the bottom right corner. Green Nodes correspond to terms that were popular throughout 2023, Blue or Purple Nodes reflect themes that were more popular in early 2023 but have since declined in prominence, and Yellow Nodes represent the most recent topics, mostly occurring in 2024. ‘‘Generative AI’’, ‘‘LLMs’’, and ‘‘ChatGPT’’ continue to be the most related and important subjects. Nevertheless, more recent conversations on ‘‘abstraction’’, ‘‘hypothesis’’, ‘‘service indi-vidualization’’, and ‘‘application’’ shown by Yellow Nodes are beginning to emerge, pointing to a move toward more theoretical and applied research.
Meanwhile, expressions such as ‘‘RPA’’, ‘‘BPM’’, and ‘‘decision-making’’ appear in Green and Yellow Nodes, suggesting that more people are paying attention to the role of AI in business process automa-tion and BPI. Furthermore, the associations found between the expressions ‘‘digitalization’’, ‘‘API’’, and ‘‘automation’’ imply that the incorporation of AI into software systems and business procedures is becoming increasingly significant.
2) CONCLUSION.
The co-keyword analysis reveals that Generative AI and LLM serve as the conceptual core of the research landscape, organizing subject clusters and inter-topic connections. Their prominence indicates a domain propelled more by technolog-ical advancement than by completely established theoretical underpinnings. The temporal overlay indicates a shift from an initial experimental topic to more contemporary applications and abstraction-focused subjects, reflecting a progressive development of the field. The continued relevance of fun-damental concepts like BPM, RPA and decision-making in recent developments indicates that the incorporation of AI into established process frameworks is still an unresolved research challenge rather than a definitive paradigm change.
The evidence indicates a rapidly consolidating, yet struc-turally developing, knowledge space, marked by robust tech-nological influences, growing multidisciplinary convergence, and a nascent focus on operationalization and empirical validation.
C. OVERALL CONCLUSIONS
The information synthesized from the performance analysis and science mapping indicates that the research domain has advanced beyond initial emergence and is now transition-ing into a phase of systematic scientific consolidation. The observed increase in publication output, along with the preva-lence of contributions in high-impact venues authored by highly cited scholars from prestigious institutions, suggests that the field has attained significant academic legitimacy and intellectual authority. Simultaneously, the conceptual frameworks revealed by keyword co-occurrence analysis indicate that Generative AI and Large Language Models have emerged as crucial organizational constructs inside the knowledge network, influencing thematic clustering and inter-topic connectedness.
This pattern indicates a shift from exploratory technological discussion to a more established and conceptually robust research environment. Its is evident that this is a rapidly expanding field, however, with limited structural consolidation. The rapid diffusion of generative AI tools from 2023 onward spurred a significant increase in pub-lications on the subject in that same year. However, although much research exists on the topic, the predominant research focuses on task-specific applications, but the development of integrative frameworks is not evolving at the same pace.
The alignment of quantitative maturity indicators with con-ceptual stabilization is a significant hallmark of disciplinary creation. The findings reveal a cohesive intellectual trajectory marked by the growing integration of AI capabilities with BPI frameworks, rather than only showcasing isolated research endeavors. Residual asymmetries, including geographic con-centration and unequal citation dispersion, are evident. How-ever, these traits align with the rapid evolution of scientific fields and indicate structural growth dynamics rather than conceptual fragmentation.
The fact that we identify a concentration of studies in the areas of finance, health, and supply chains indicates that ini-tial adoption is stronger in highly regulated and data-intensive sectors. This is because the speed of decision-making and the complexity of the process create opportunities for immedi-ate value. On the other hand, the recurring gaps identified related to explainability, governance, and closed-loop mon-itoring suggest that organizational control mechanisms have not evolved at the same pace as AI capabilities. Finally, a very limited amount of evidence is identified that accompanies a solution developed over time in real-world environments, indicating that the field remains largely exploratory, with a great need for real-world validation.
VIII. RESULTS.
This study examined 49 pertinent publications chosen from an initial set of 534 based on their rigor, relevance, and contribution. The report provides a detailed examination of the impact of LLMs on BPI across various sectors and methodological frameworks. Although most of the studies demonstrate a rigorous methodology, some reveal constraints on sample size, generalizability, or the absence of longitudi-nal data.
Numerous sectors have made extensive use of traditional BPM models, including Business Process Model and Nota-tion (BPMN), Six Sigma, and Process Reengineering. On the other hand, the advent of LLMs presents new approaches to business process automation, analysis, and improvement.
This study aims to understand these new approaches and understand how the integration of LLM/RAG capa-bilities with process analytics improves the explainability, governance, and closed-loop effectiveness of BPI across its lifecycle. Furthermore, we aim to understand the character-istics of these approaches used and discover their unresolved issues. Consequently, taking this into account, we can now answer the RQs that are the subject of this study.
RQ-01: In what business domains and organizational processes have LLMs been used most efficiently to improve business workflows and task performance? LLMs have improved corporate processes and productiv-ity in several business domains. The findings indicate that LLMs applications are most prevalent in financial ser-vices (17% of total citations), particularly in the organi-zational processes of risk assessment, fraud detection, and compliance automation. The healthcare sector ranks sec-ond (15% of total citations), using LLMs for workflow automation, medical diagnostics, and patient care optimiza-tion. Supply chain management (13% of total citations) employs LLMs for demand fore-casting and inventory management, while Customer Service (13% of total citations) improves service automation, inquiry management, and user interactions.
Other business domains include manufacturing (11% of total citations), Retail & E-commerce (11% of total citations), public sector (9% of total citations) and education (6% of total citations) using LLMs for document processing, regulatory automation, and service optimization.
RQ-02: What results have been achieved by adopting LLMs to improve business processes workflows? Con-sidering the set of studies analyzed, the findings indicate that LLMs markedly improve efficiency, reduce costs, and facilitate decision making in BPI. Businesses indicate a 30–50% improvement in automation efficiency, leading to a reduction in process cycle durations and bottlenecks, coupled with a 40% reduction in human processing time, particu-larly in document-intensive workflows. AI-driven automation reduces operating expenses by 20–30%, decreasing manual labor and improving resource efficiency.
The precision of the forecast has improved by 30–40%, improving supply chain management, market analysis, and financial planning. AI-assisted compliance improves regu-latory conformance, while chatbot-driven customer service has improved resolution accuracy by 25%, hence increasing customer satisfaction and decreasing dependence on human agents.
In addition to efficiency, LLMs improve agility and adapt-ability, facilitating dynamic AI-driven workflows that adapt to market fluctuations. Enhanced user-friendliness and work-force participation further facilitate uptake.
RQ-03: What are the main challenges for business sec-tors to adopt LLMs to improve business processes?
Considering the set of studies analyzed, the findings indicate that LLMs adoption is complicated by data protection restrictions (e.g., General Data Protection Regu-lation - GDPR), high implementation costs, and information technology integration issues, especially with outdated sys-tems. As employees fear job loss, AI adoption is hindered by work-force opposition and upskilling. Bias in training data raises ethical and legal problems about AI-driven conclu-sions.
In the same way, the calculation of the Return on In-vestment (ROI) is complicated and varies by business and use case. High-performance computers raise operational expenses, and LLMs management requires AI competence. In regulated industries such as healthcare and banking, clear AI choices are essential. The adoption of LLMs sometimes re-quires redesigning workflows, which complicates operations. Although LLMs improve corporate operations, its application requires combining efficiency with technical, financial, and ethical constraints.
These findings repeatedly underscore the substantial ad-vantages that LLMs confer on organizational opera-tions, such as increased automation efficiency, cost sav-ings, higher customer satisfaction, compliance, and improved decision making. The advantages are especially evident in business sectors such as banking, healthcare, and supply chain management, where activities such as risk assess-ment, fraud detection, medical diagnostics, and customer service are data-intensive and depend on rapid information processing.
Notwithstanding its benefits, the implementation of LLMs frequently exposes enduring problems. Challenges include compliance with data privacy, model interpretability, integra-tion with legacy systems, elevated computing expenses, and labor opposition are frequently noted in the literature. These limits highlight the need to create governance structures and technical techniques to mitigate risks while optimizing benefits.
RQ-04: What are the gaps in the effective integration of LLM/RAG capabilities with Business Process Improve-ment, in order to enhance the explainability, governance, and effectiveness of the closed loop in organizations?
Taking into account the findings described above aligned with existing specialized literature, we extracted 17 gaps that organizations must address to successfully integrate LLM/RAG capabilities with process analytics, in order to enhance the explainability, governance, and closed-loop effectiveness of BPI across its lifecycle.
To identify gaps, reports and information on capability deficiencies, limitations, and contextual problems related to LLM-enabled BPM approaches were investigated in the 49 selected studies. After extracting this information, recur-ring patterns were identified in the studies analyzed. Finally, these patterns were consolidated into the 17 gaps presented below. This ensures the use of multiple sources of empirical or conceptual evidence to provide adequate support for archi-tectural decisions.
It is important to note that this mapping process is based much more on representative evidence than exhaus-tive evidence. This means that the gaps are not identi-cal to those presented in the 49 studies analyzed, nor do they have the same probative force. While some gaps are directly supported by a subset of the ana-lyzed studies, others are only partially derived from those studies and result from refinements of design abstractions informed by the literature rather than directly specified by it.
Gap-01: End-to-end BPI Lifecycle Coverage
Process improvement constitutes a cyclical and dynamic lifecycle encompassing discovery, diagnosis, redesign, implementation, and monitoring. The BPM literature regu-larly highlights lifecycle integration as essential for endur-ing organizational learning and performance improvements. Disjoint or phase-isolated enhancements typically yield localized optimizations instead of a comprehensive advancement. Addressing the entire BPI lifecycle ensures methodological consistency and facilitates closed-loop learn-ing, a fundamental characteristic of adaptive companies.
Gap-02: Heterogeneous Data Ingestion and Access
Digital processes produce data from various sources, such as event logs, structured records, papers, and policy repos-itories. Studies in analytics-driven businesses indicate that the integration of diverse data enhances decision-making quality and contextual understanding. Process mining illustrates that the integration of operational logs with con-textual knowledge yields enhanced diagnostic insights. Consequently, extensive data input is essential for evidence-based BPI.
Gap-03: Separation of Concerns and Modular Decom-position
Complex socio-technical systems are enhanced by modu-lar architectures that allow analytics, reasoning, governance, and execution components to develop independently. Fun-damental software engineering theory posits that modular breakdown enhances maintainability, scalability, and robust-ness. A distinct separation and modular decomposition diminish the connection, enhance maintainability, and facil-itate the independent development of data pipelines and AI components.
Gap-04: Event-log Readiness and Empirical Process Evidence
Process mining research identifies event logs as objective, empirical depictions of process activity, facilitating impar-tial diagnostics and compliance analysis. Evidence-based enhancement rooted in actual operational data diminishes dependence on subjective perceptions and increases the pre-cision of intervention. It depends on the event logs as the primary source of behavioral evidence for observed execu-tions.
Gap-05: Process Model Management and Traceability Process models serve as the formal foundation for BPM. Robust model governance ensures version control, evidence traceability, and compliance with execution environments. Traceable models allow organizations to understand the rela-tionship between decisions, previous baselines, and prospec-tive alterations. BPI needs clear representations of the process and the regulated evolution of the model over time..
Gap-06: Diagnostic Insight Generation and Inter-pretability
Diagnostic analytics must be comprehensible to facili-tate actionable insights. Studies on explainable analytics and decision-support indicate that stakeholders are more inclined to believe and act on findings they comprehend. Interpretability connects technical evidence with manage-ment logic, making AI-supported diagnostics operationally effective. This facilitates the prioritization of initiatives.
Gap-07: Improvement Reasoning and Alternative Gen- eration
Decision theory posits that superior decisions result from the examination of several alternatives. In BPM, redesign is fundamentally innovative and advantages from methodical examination of enhancement scenarios. AI-assisted rea-soning increases this exploratory capability. An effective CBPI needs prescriptive reasoning instead of purely descrip-tive reporting.
Gap-08: Human-in-the-Loop Accountability and Ap-proval Gates
Socio-technical research emphasizes that human supervision mitigates automation bias and ensures ethical and organizational accountability. Research on Responsible AI emphasizes the need for explicit human accountability in automated decision-making systems.
Gap-09: Ex-ante Simulation and Validation of Re- designs
Simulation is a long-established method in BPM for assessing the potential impact of process changes prior to deployment. Ex-ante evaluation mitigates imple-mentation risk and facilitates evidence-based selection of redesign alternatives.
Gap-10: Operational Implementation and Integration with Enterprise Systems
Operational implementation of process improvement is essential for value creation. Research on BPM and infor-mation systems integration indicates that the alignment of models and enterprise systems is essential for enduring digital trans-formation. Improvement is only effective when it can be implemented and assimilated into operational systems.
Gap-11: Continuous Monitoring and Closed-loop Feed-back
The theory of continuous improvement, specifically the PDCA cycle, identifies feedback as the fundamental mech-anism for ongoing learning and improvement of performance. BPM monitoring strategies apply this notion to digital contexts. CBPI needs a closed loop that juxtaposes observed outcomes with anticipated behavior and baselines.
Gap-12: Historical Baselines and Comparative Evalu-ation
Performance assessment needs historical benchmarks to differentiate genuine improvement from inherent variability.
Research in evidence-based management underscores the importance of baseline comparison in analytics. In this sense, baselines are essential to illustrate improvement effects and mitigate confounding variables.
Gap-13: Explainability, Traceability, and Auditable
Reasoning (LLM-grounded)
Research on explainable AI indicates that transparent and traceable reasoning enhances accountability and trust in auto-mated systems. Auditable reasoning is crucial in organizational and regulatory frameworks..
Gap-14: Knowledge Freshness and Adaptability With-out Retraining
Organizational knowledge evolves rapidly, and a retrieval-based grounding facilitates quick updates while main-taining constant model parameters. Retrieval-augmented reasoning enables AI systems to integrate current knowl-edge dynamically without the need for retraining. Adaptability enhances relevance in dynamic organizational contexts..
Gap-15: Compliance and Risk Constraints Enforce-ment
The literature on risk governance emphasizes that inno-vation should function within established limitations to pre-vent unmanaged uncertainty. Incorporating compliance checks ensures adherence to institutional regulations. CBPI in regulated settings must maintain compliance and risk awareness, ensuring that improvemenets do not cre-ate intolerable exposures.
Gap-16: Security, Privacy, and Access Control
AI-driven CBPI systems handle critical operational data, becoming essential security features. Data processing fre-quently involves personal and confidential information; governance mandates the principle of least privilege and regulated exposure,. Research in information security underscores that data protection and regulated access are essential for reliable information systems.
Gap-17: Reproducibility and Operational Reliability
Scientific integrity is contingent on reproducibility. Sys-tems research posits that reproducible pipelines enhance transparency and institutional confidence. Dependable operation is crucial for the long-term sustainability of CBPI. Operational discipline is essential for managing technical debt and maintaining quality in production ML systems over time.
So, in order to ensure the traceability and interpretability of the information presented here, Table 8 presents the mapping between the identified gaps and their respective literature support.
IX. DISCUSSION.
The incorporation of LLMs signifies a transformative shift from conventional BPM models, providing AI-driven flexi-bility and scalability that exceed the static or manual char-acteristics of existing methods. Once we are interested in comparing traditional BPI approaches with approaches using
LLMs, it is important to understand the scope, methods, ad-vantages, and disadvantages of using traditional BPM models with the approaches presented above using LLMs.
A. COMPARISON OF TRADITIONAL APPROACHES TO USING BPM WITH APPROACHES USING LLMS
Starting with BPMN, it is a standardized notation that visu-ally represents workflows, facilitating process standardiza-tion and effective communication between stakeholders. Nevertheless, it requires manual updates and is not adaptable in real-time. Conversely, LLMs systematically evaluate and improve workflows, pinpointing inefficiencies within unstructured data, such as emails and documents. Although LLMs automate and improve processes, they lack BPMN’s structured visual representation, making them less suitable for standardized documentation. Six Sigma employs a structured, data-driven DMAIC (De-fine, Measure, Analyze, Improve, and Control) frame-work to reduce errors and enhance efficiency, with a focus on statistical analysis and specialized teams. How-ever, it involves in-depth study, iterative improvements, and challenges with unstructured data.
In contrast, LLMs assess massive datasets instantaneously, adjust dynamically, and anticipate optimal process parameters without requiring human interaction. Whereas Six Sigma ensures precision and control, LLMs provide automation, speed, and flexibility, making them more suitable for dynamic business environments. Consequently, highly regulated sectors benefit greatly from the greater structure and rigor of Six Sigma. However, LLMs provide more automation, speed, and flexibility, especially in dynamic commercial set-tings.
Finally, Business Process Reengineering (BPR) funda-mentally restructures workflows to achieve significant per-formance improvements, enhance customer experience, and match operations to strategic objectives. Nevertheless, it faces prolonged adoption, substantial expenses, and institu-tional resistance. Conversely, LLMs improve processes incrementally, modifying workflows in real time without significant reorganization. BPR facilitates strate-gic transformation, whereas LLMs automates and improves current processes. In this sense, LLMs offer continuous AI-driven optimization, while BPR is strategic and disruptive. Businesses can combine the two strategies, using LLMs to support and maintain BPR improvements. Table 9 presents a comparative summary of LLMs vs. traditional BPM models.
The findings indicate that LLMs can significantly enhance operational efficiency, especially in business sectors that are dependent on document-intensive procedures, consumer engagement, or data-driven compliance.
Nonetheless, real application requires addressing many challenges:
• Technical Readiness: organizations must allocate re-sources towards high-performance computing, secure data pipelines, and AI governance structures.
• Workforce Transformation: successful implementa-tion relies on reskilling personnel, promoting AI knowl-edge, and alleviating concerns about job displacement.
• AI Governance: addressing bias, ensuring explainabil-ity, transparency, and adhering to standards such as GDPR are imperative prerequisites for sustainable AI integration, as well as for ensuring the ethics and com-pliance of BPI.
• Integration: achieving an adequate level of compati-bility with current BPM tools and enterprise systems presents a significant challenge that requires both tech-nological solutions and change management strategies.
In summary, LLMs do not replace standard BPM models. Instead, they enhance them with flexibility, adaptability, and automation. A hybrid methodology that uses the advantages of BPMN for visualization, Six Sigma for quality assur-ance, BPR for strategic transformation, and LLMs for AI-enhanced optimization presents the most promising path forward. As organizations increasingly pursue agility in unpredictable marketplaces, the astute integration of various techniques will characterize the forthcoming evolution of process management.
B. EVIDENCE STRENGTH AND LIMITATIONS OF THE REVIEWED LITERATURE
Although the analyzed studies report promising results regarding the use of LLMs and AI techniques to improve business processes, the strength of the available evidence is heterogeneous and should be interpreted with caution. The analysis combines conceptual articles, prototype demonstra-tions, specific case implementations, research, simulation-based studies, and empirical evaluations with varying levels of methodological rigor. As a result, not all reported benefits have the same degree of empirical support.
While some studies present controlled experiments or use cases, others propose conceptual frameworks, proof-of-concept systems, or exploratory discussions. Furthermore, the analyzed studies are distributed across multiple sec-toral contexts, which demonstrates broad applicability but also limits generalization, since the results observed in one do-main may therefore not be directly transferable to another.
It is also important to note that some inconsistencies were observed regarding the sample size and scope of the evalua-tion. While several studies rely on restricted datasets, limited process instances, pilot implementations or short-term eval-uations; long-term evaluations, large-scale deployments, and comparative studies are still scarce. This restricts the ability to infer sustained performance effects over time.
Finally, the level of empirical validation is uneven. While some studies provide measurable operational results, others infer expected benefits from technical feasibility or expert judgment.
In summary, the literature can be considered promising but is still in the process of maturing.
C. ENSURING EXPLAINABILITY, GOVERNANCE, AND EFFECTIVENESS IN A CLOSED-LOOP BPI THROUGHOUT ITS ENTIRE LIFECYCLE
Considering the above, we can now answer the Main Re-search Question that is the object of this study: In orga-nizational contexts of BPI, how does the integration of LLM/RAG capabilities with process analytics, compared to traditional BPI approaches, ensure explainability, gov-ernance, and closed-loop effectiveness of BPI throughout its lifecycle?.
Workflows can be effectively automated, optimized, and improved with the help of LLMs, which increase the intelligence, adaptability, and data-drivenness of business pro-cesses. Therefore, in order to ensure explainability, gov-ernance, and effectiveness in a closed-loop BPI throughout its entire lifecycle, organizations must address the 17 gaps described in Section VIII. In order to do that, we propose a modular and extendable six-phase conceptual LLM-driven framework for Continuous Business Process Improvement (LLM-CBPI Framework) for application in organizations. Figure 7 depicts the LLM-CBPI Framework.
1) LLM-CBPI FRAMEWORK.
As can be seen in Figure 6, this framework has six phases:
1) Process Understanding;.
2) Process Diagnosis;.
3) Process Redesign;.
4) Process Validation;.
5) Process Execution Support; and.
6) Continuous Monitoring.
1. Process Understanding.
This phase emphasizes the development of a detailed and organized depiction of the existing (‘‘As-Is’’) business pro-cess. LLMs facilitate the extraction of procedural information from various sources, including papers, procedures, commu- nications, narratives, guidelines, and event logs. They consol-idate this information into coherent descriptions, de-lineate tasks, roles, dependencies, limitations, and are capable of producing preliminary BPMN models or textual process maps. This allows companies to quickly cultivate a collective understanding of the real functioning of the process, encom-passing informal practices and undocumented deviations.
2. Process Diagnosis.
During this phase, the LLM evaluates the acquired process model and its related data to detect inefficiencies, bottle-necks, deviations, non-conformities, hazards, and resource misalignments. The LLM employs contextual reasoning to elucidate the core causes, identify discrepancies between the intended and actual implementation, and propose avenues for improvement. When logs or measurements are accessible, the model can identify trends that indicate performance deterio-ration, compliance breaches, or cycles of rework. The result is a systematic diagnostic evaluation that guides the redesign initiative.
3. Process Redesign.
This phase emphasizes the collaborative development of enhanced process alternatives (‘‘To-Be’’ models). LLMs sug-gest redesigning strategies that meet organizational goals, such as cost reduction, efficiency, quality, or compliance. The model can assess trade-offs, compare various redesign alternatives, and produce BPMN diagrams or textual pro-cess representations that integrate automation, simplification, standardization, task elimination, or role reallocation. A HITL method-ology ensures that domain specialists enhance, verify, and supplement the options produced to archive feasi-ble and contextually relevant solutions.
4. Process Simulation and Validation.
This phase involves evaluating the modified (‘‘To-Be’’) process models to ascertain their anticipated performance, practicality, and alignment with corporate objectives prior to deployment. LLMs facilitate the analysis of the ramifi-cations of proposed modifications, identify potential haz-ards, and assess if the restructured processes adhere to compliance, capacity, and resource limitations. Human specialists evaluate and contrast various scenarios pro-duced by the model, enhancing the redesign until the To-Be version exhibits an acceptable equilibrium of effi-ciency, reliability, and operational realism. The phase cul-minates with a validated process design prepared for implementation.
5. Process Execution Support.
In this phase, LLMs operate as an intelligent assistant inte-grated into routine operational processes. The model can offer task advice, interpret policies or Standard Operating Proce-dures (SOPs) as needed, assist decision-making in extraor-dinary or unclear circumstances, and recommend corrective measures when discrepancies arise. When combined with BPM or RPA systems, LLM can oversee execution traces, identify anomalies in real time, and suggest alternate paths or human interventions. The purpose of this phase is not to automate execution, but to enhance the workforce and digital systems with contextual, adaptive, and cognitively sophisti-cated process knowledge.
6. Continuous Monitoring.
In this phase, the LLM methodically collects and evaluates feedback from event logs, performance metrics, user reports, key performance indicators, and compliance assessments. It recognizes indicators of deterioration, such as increased delays, novel bottlenecks, developing dangers, role overload, or departures from the sanctioned To-Be model. By contrast-ing the actual execution with the planned design, the model can identify the drift of the process, elucidate root causes, and propose incremental modifications or novel improvement opportunities.
2) LLM-CBPI FRAMEWORK REFERENCE ARCHITECTURE.
In order to deploy the 6 phases of the LLM-CBPI Frame-work, we designed a reference architecture that incorporates diverse data sources, process mining functionalities, LLM-driven reasoning, human-centered redesign processes, and ongoing monitoring systems.
This is a modular structure that incorporates process ana-lytics, LLMs, human governance, and workflow execution functionalities to support the 6 phases of the LLM-CBPI Framework. The proposed architecture was designed using a notation inspired by the Unified Modeling Language (UML)
with a set of visual elements that represent a distinct set of functions in the architecture. Below is a description of the visual elements that make up the proposed architecture.
Component. A component is a self-contained functional element that incorporates a series of duties in the LLM-CBPI Framework. Components contain well-defined interfaces and interact with other elements of the architecture via a de-pendency relationship that allows that component to make requests to other elements of the architecture. They are more system capabilities than software implementations. These components may contain subcomponents or modules that, when combined, address a set of related issues. Within the LLM-CBPI Framework reference architecture, we have the following components:
• Human-in-the-Loop Redesign Module
• Improvement Reasoning Engine
• LLM-Based Process Intelligent Layer
• Monitoring & Analytics Layer
• Process Discovery
• Simulation / Validation Module
• Workflow Execution Engine
Packages. A package is a logical grouping of data from the same issue domain. They are not executable entities, but feed these entities with data or information, acting as a facilitating mechanism for communication between the other elements of the architecture. Within the LLM-CBPI Frame-work reference architecture, we have the following packages:
• AS-IS
• Compliance & Risk Constraints Layer
• Diagnostic Insights
• Input Data Layer
• LLM Use Cases
• Model Layer
• Process Models Repository
• Simulation / Validation Layer
• Structured Data Sources
• TO-BE
Data Source. A data source is a persistent storage or an origin of a data set used by the elements of the architecture. These are structured and semi-structured data sources, exter-nal or passive entities that are not executable but provide or retain information used or produced by other elements of the architecture. Within the LLM-CBPI Framework reference architecture, we have the following data sources:
• Event Records
• Vector Database
Pipeline. A pipeline is a sequential data-processing system that converts inputs into outputs via organized steps. They are flow-oriented semantic elements.Within the LLM-CBPI Framework reference architecture, we have the following pipelines:
• ETL/ELT Pipeline
• LLM/RAG Pipeline
External System. An external system is an application existing in the organization beyond the LLM-CBPI Frame-work reference architecture. These systems interact with elements of the proposed architecture through well-defined interfaces. They are not part of the CBPI itself, but produce or use data relevant to the CBPI process. Within the LLM-CBPI Framework reference architecture, we have the following external system:
• Business Application Layer
Dependency Relationships. For a better understanding of the proposed architecture, only one type of relationship between elements was drawn in the diagram: dependency. This is a unidirectional relationship in which one element (the client) depends on another element (the provider) for its accurate description, specification, or functionality. This relationship is conventionally represented by a dashed arrow that points from the dependent element to the element on which it is based. In the dependency relationships of the proposed architecture, the flow of requests occurs in the direction of the arrow, whereas the flow of information returned from the requests occurs in the opposite direction.
The design maintains a loose link between components by using dependency relationships instead of more rigid struc-tural ties, which is very relevant to ensure the scalability, modularity and autonomous development of individual archi-tectural pieces.
Figure 8 depicts this reference architecture.
As we can see in Figure 8, this architecture is based on elements widely used and empirically demonstrated in multiple contexts, such as Human-in-the-Loop, LLM/RAG Pipeline, ETL/ELT Pipeline, and Process Discovery. It is designed to facilitate all six phases of the CBPI lifecycle by coordinating a suite of interoperable components that collectively consti-tute the cognitive and analytical foundation of LLM-assisted continuous improvement. Every subsystem inside the archi-tecture provides distinct functionalities to the CBPI lifecycle, creating a cohesive ecosystem that connects data, models, processes, and human expertise. In the following, we present how each element of the reference architecture contributes to the workflow execution of each phase of the LLM-CBPI Framework.
Phase 1 (Process Understanding)
This phase combines human expertise with the LLMs capa-bilities to analyze information about organizational processes and transform it into workflows that reflect how processes actually unfold within the organization. The generated rep-resentations need to be validated by stakeholders to serve as a diagnostic basis for subsequent business process redesign activities.
This phase begins with the definition of scope, stakehold-ers, and information sources by human actors. For this, the elements of the Input Data Layer, Structured Data Sources, and the Process Models Repository are used to access orga-nizational records, legacy workflows, governance documents, and operational systems. Here, too, the Compliance & Risk Constraints Layer provides information on access, owner-ship, or regulatory boundaries.
After these activities are completed, the LLM comes into action when the LLM/RAG Pipeline is fed with this infor-mation, while the ETL/ELT Pipeline prepares the struc-tured operational data. In this context, the Vector Database stores embeddings that enable semantic retrieval of con-textual evidence during later interactions, and the LLM-Based Process Intelligence Layer extracts activities, roles, rules, exceptions, and dependencies from these data sources, converting this extracted information into AS-IS process descriptions.
In this way, the architecture supports the generation of preliminary process models, such as a workflow description or even a draft of a BPMN workflow, if there is support for integration with a BPMN tool, including task sequences, gateways, handoffs, and control-flow assumptions. These artifacts are stored in the Process Models Repository, which performs version control in the refinement cycles.
It is important to note that not all input data mentioned in the phase descriptions will be available within the orga-nization. Even so, LLMs should be able to generate the expected results. However, the accuracy of these output data will be greater if all suggested input data are available, which will also depend on the quality of that data. If the organization does not possess the set of input data suggested in the description of each phase or if the quality of the avail-able data is poor, the process of analyzing and evaluating the results of each phase will require greater effort from stakeholders.
These artifacts are subjected to validation by human actors through interviews, walkthrough sessions, and stakeholder reviews. These tasks are coordinated by the Human-in-the-
Loop Redesign Module. At this stage, experts analyze the generated information and may suggest further refinements. If inconsistencies persist, the workflow iterates until an acceptable representation is achieved. Finally, the approved AS-IS baseline is formally stored in the repository and released as input to the next phase.
Phase 2 (Process Diagnosis)
This phase combines managerial judgment, analytical capabilities, and LLM-enabled reasoning in order to trans-form raw execution evidence into structured information. This information is essential to guide future redesign decisions.
In this context, the workflow begins with human actors, who need to define the scope of the diagnosis and develop a set of questions that are within that scope and are properly aligned with business objectives, compliance requirements, and target KPIs. In the architecture, these activities are sup-ported by the Process Models Repository and the Compli-ance & Risk Constraints Layer. The Process Models Repos-itory provides the query base for AS-IS processes, and the Compliance & Risk Constraints Layer contains the policies, regulations, standards, and tolerance thresholds that should be considered in this process.
Next, the ETL/ELT Pipeline accesses the execution data of the AS-IS processes in the Structured Data Sources and performs cleansing, standardization, and event-trace structur-ing. These results are then processed by the Monitoring & Analytics Layer, where the previously defined target KPIs are compared with the current values stored in the Historical Baselines. The results are stored in Diagnostic Insights as empirical evidence.
This information is then interpreted by the LLM/RAG Pipeline and the LLM-Based Process Intelligence Layer to generate an interpretive diagnosis using process evidence, business rules, constraints, and contextual documentation retrieved from the Vector Database. This provides support for identifying root causes, operational risks, and cause-effect relationships that may not be immediately visible through quantitative metrics alone. The outputs of this process include narrative explanations, hypothesis generation, and consoli-dated diagnostic findings.
Subsequently, this information undergoes a review process by stakeholders, managers, and domain experts. If inconsis-tencies or areas for improvement are identified, the work-flow iterates through further analysis and clarification cycles. Once validated, it may be necessary to prioritize the issues, considering aspects such as impact, feasibility, organizational constraints, and governance thresholds. This is done with the help of the Improvement Reasoning Engine, which helps to rank opportunities based on performance gains, risk expo-sure, and strategic relevance.
In the next step, human actors must produce and approve the diagnostic report that includes confirmed bottlenecks, validated root causes, prioritized issues, and evidence-based conclusions. This artifact is stored for traceability and for-mally transferred as input to the next phase.
Phase 3 (Process Redesign)
This phase combines the judgment capabilities of human actors with LLM-enabled generative reasoning. The objective is to create, compare, refine, and formally approve re-design alternatives for later implementation.
Here, the workflow begins by defining the objectives and establishing the constraints that will guide the process re-design strategy. To do this, human actors must perform these activities considering strategic priorities, policies, legal requirements, budget limitations, and technological condi-tions. These activities are supported by the Compliance & Risk Constraints Layer.
After performing these activities, human actors use the AS-IS process model stored in the Process Models Repository and the prioritized issues identified during diagnosis to define redesign principles such as automation, control strengthen-ing, cycle-time reduction, workload balancing, etc..
Once these steps are completed, the LLM/RAG Pipeline and the LLM-Based Process Intelligence Layer sup-port the proposal of process changes and the creation of multiple TO-BE alternatives. To achieve this, these components utilize contextual knowledge retrieved from the Vector Database, prior process models, business rules, and redesign principles. Expected outputs include recommended modifications such as task elimination, activity consolidation, decision-rule simplification, role reas-signment, exception handling improvements, sequencing changes, or automation opportunities. These results are stored in the TO-BE repository for version control and traceability.
Another important point is the generation of a summary with the expected changes and trade-offs associated with each alternative, for example, expected efficiency gains, imple-mentation complexity, compliance implications, operational risks, required resources, or organizational disruption, etc. The Improvement Reasoning Engine supports these activ-ities.
The next step is the evaluation of these proposed alterna-tives and the selection of the preferred redesign option. This can be done through workshops and stakeholder feedback, supported by the Human-in-the-Loop Redesign Module.
Finally, the selected TO-BE process model is submitted for governance approval. Once approved, the redesigned process is formally stored in the repository and released as an autho-rized input to the next phase. It is important to note that, as stated in Phase 1 description, these models (AS-IS and TO-BE) can be stored as a workflow description or even a draft of a BPMN workflow if there is support for integration with a BPMN tool.
Phase 4 (Process Validation).
Before the operational deployment of the business process, it is necessary to validate the new workflow model proposed in the previous phase. In this context, this phase aims to verify whether the redesigned process is feasible, compliant, and capable of achieving the desired performance improve-ments under realistic conditions, thus reducing the risk of implementation and strengthening evidence-based decision-making.
The workflow begins with human actors defining the val-idation objectives and acceptance criteria for the approved redesign proposal. To do this, they must specify target KPIs, expected service levels, stakeholder expectations, and accept-able risk thresholds. These activities can be carried out with the support of the Compliance & Risk Constraints Layer, in order to ensure the inclusion of regulatory obliga-tions, governance criteria, and tolerance limits in the valida-tion process.
After these activities are carried out, there must be a prepa-ration stage for the validation data and simulation scenarios. This is done using historical records, prior execution evi-dence, workload assumptions, and alternative operating con-ditions obtained from the Historical Baselines, Structured Data Sources, and Process Models Repository. At this stage, the ETL/ELT Pipeline supports data cleaning and scenario preparation.
Therefore, the approved TO-BE model is then processed by the Simulation / Validation Module. The objective is to simulate the redesigned workflow under different condi-tions, such as different demand profiles, resource constraints, or policy assumptions. The expected outputs of this stage include projected cycle times, throughput levels, queue for-mation, resource utilization, Service Level Agreement (SLA) compliance, and expected cost implications. Subsequently, these outputs are assessed by the Monitoring & Analytics Layer, which compares the simulated results with acceptance criteria and historical benchmarks.
At the same time, the LLM/RAG Pipeline and the LLM-Based Process Intelligence Layer support a higher-level interpretation of outcomes validation. Using simulation evi-dence, organizational rules, and contextual documentation retrieved from the Vector Database. The goal is to identify potential implementation risks, hidden bottlenecks, policy conflicts, unrealistic assumptions, or operational depen-dencies that may not be evident through numerical metrics alone. At this stage, it is also possible to recommend refine-ments such as control adjustments, resource reallocations, sequencing changes, or exception-handling improvements.
All these findings must be submitted to review by human actors for verification, by managers and domain specialists, whether acceptance criteria have been met and whether iden-tified issues can be resolved through incremental refinements. If necessary, it is possible to return for adjustment through iterative redesign cycles until the validation requirements and objectives are met and formally proven. This approved artifact is stored for traceability and is released as authorized input to the next phase.
Phase 5 (Process Execution Support) The objective of this phase is to apply the approved re-design business process to the execution of the organization’s daily operations. To do this, it is necessary to develop execution guidance and obtain real-time evidence, in order to prepare the basis for continuous improvement of business processes and for the next monitoring phase.
This phase begins with the development of the deployment plan by human actors. For this, they used the validated TO-BE process approved in Phase 4. In this context, stakeholders must define rollout priorities, resource allocation, transition schedules, and organizational readiness actions. These activ-ities are supported by the Compliance & Risk Constraints Layer. This layer ensures that deployment decisions are aligned with governance requirements, segregation-of-duty rules, security controls, and operational risk thresholds.
After this step, human actors must configure (integration parameters, user permissions, interfaces, system dependen-cies, etc.) the execution environment of the proposed work-flow through the Workflow Execution Engine (BPMS/RPA) and connected enterprise systems represented in the Business Apps Layer, such as Enterprise Resource Planning (ERP), Customer Relationship Management (CRM), and automation platforms. The goal is to make the proposed workflow operate under realistic conditions that reflect the daily operations of the organization.
Once the execution environment is configured, the re-designed process is deployed as an executable workflow instance. At this stage, the Workflow Execution Engine is responsible for orchestrating tasks, decision logic, handoffs, escalations, and automated activities. All of this is done ac-cording to the approved TO-BE model stored in the Process Models Repository.
In this context, the LLM/RAG Pipeline and the LLM-Based Process Intelligence Layer provide contextual execu-tion support to operational users. For this, policies, SOPs, business rules, role definitions, and organizational knowl-edge retrieved from the Vector Database are used. These procedures reduce ambiguities and improve consistency in workflow execution.
Here, we also have the role of the Monitoring & Analytics Layer through the continuous collection of evidence related to the execution of the business process, such as executed tasks, decision logs, timestamps, resource utilization records, event streams and system audit trails. This involves creating an empirical basis for future evaluations of performance eval-uation and process learning.
In this step, it is also important to collect structured feed-back from stakeholders, such as execution difficul-ties, unexpected constraints, usability concerns, or sugges-tions regarding AI-generated guidance. Human-in-the-Loop Redesign Module can support these activities in order to complement the system-generated evidence captured by the Monitoring & Analytics Layer.
Finally, the consolidated data from this phase serves as input for the next phase, allowing longitudinal performance tracking and future optimization cycles.
Phase 6 (Continuous Monitoring)
This is the phase that completes a cycle of the CBPI closed-loop. In this phase, the objectives are to measure sustained performance, verify compliance, detect emerging risks or process drift, and trigger new improvement cycles when-ever operational conditions or strategic objectives change.
The workflow begins with the computation of the performance metrics (predefined KPIs, SLA thresholds, through-put indicators, etc.) by the Monitoring & Analytics Layer. These values provide an up-to-date view of the workflow’s operational performance over time.
At the same time, the Compliance & Risk Constraints Layer monitors regulatory and constraint indicators such as business rules, standards, policies, and governance require-ments. The expected values are compared with the values obtained from the current execution of the workflow. The goal is to identify deviations, control violations, segregation-of-duty conflicts, or growing risk exposure. These two layers ensure efficiency and conformity.
This information is compared with historical references stored in the Historical Baselines. In this way, it is possi-ble to detect anomalies, gradual performance deterioration, demand shifts, resource imbalances, or behavioral drift in the deployed process.
Next, the LLM/RAG Pipeline and the LLM-Based Process Intelligence Layer interpret the monitoring signals at a higher level. To do this, these components use contextual process knowledge from the Vector Database, previous di-agnostic findings, and organizational objectives. In this way, it is possible to translate the raw indicators into understand-able managerial insights. Outputs from this stage include narra-tive explanations of anomalies, likely causes of de-creased performance, identification of recurring bottlenecks, and contextualized interpretations of compliance alerts. Poten-tial recommendations are also generated, such as process recalibration, workload redistribution, control reinforcement, parameter adjustment, etc.
Finally, human actors must analyze the monitoring in-sights. This can be done through governance forums, dash-boards, reports, or data visualization, provided there is sup-port for integration with Business Intelligence and data visualization tools. Based on these analyzes, new improve-ment cycles are initiated, and the monitoring baselines are updated.
3) DESIGN RATIONALE.
The LLM-CBPI reference architecture is designed to address three sometimes competing demands in modern process improvement efforts: analytical rigor, explainable intel-ligence, and organizational governance. This subsection elucidates the reasoning behind the principal architectural selections and delineates how each selection aligns with the CBPI objectives.
A key design choice is the clear distinction between deterministic analytics and generative AI reasoning. Quan-titative duties, including KPI computation, conformance ver-ification, anomaly detection, and simulation, are assigned to the Monitoring & Analytics Layer, while interpretative and explanatory tasks are managed by the LLM-Based Process Intelligent Layer.
This distinction prevents LLMs from being erroneously regarded as numerical or statistical engines and mitigates hazards related to fabricated measurements or unverifiable outcomes. The design maintains methodological rigor by anchoring all reasoning in verifiable analytical outputs while also using the capacity of LLMs to synthesize, contextual-ize, and elucidate intricate discoveries. This design decision directly facilitates Phases 2, 4, and 6 of the LLM-CBPI Frame-work, where analytical evidence must precede inter-pretation and decision-making.
Another essential justification is the establishment of human oversight within the framework. The architecture clearly incorporates governance checkpoints through the Human-in-the-Loop Redesign Module and decision gate-ways at each level of the CBPI, rather than considering human involvement as an exception. Humans maintain dominion over:
• Determination of scope and diagnostic framework (Phases 1–2);
• Selection and approved redesign options (Phase 3);
• Decisions about acceptance and deployment (Phases 4–5); and
• Initiation of fresh improvement cycles (Phase 6).
This design ensures accountability, regulatory adherence, and institutional confidence, particularly in scenarios where automated process modifications may entail legal, ethical, or operational consequences.
The architecture is deliberately constructed as a closed-loop system, rather than a linear BPM pipeline. The results of execution and monitoring (Phases 5–6) are directly sent back to the preceding phases, especially Phase 1 - Process Understanding and Phase 2 - Process Diagnosis.
This design, informed by feedback, facilitates the ongoing acquisition of knowledge from operational data, adaptive baseline modifications, and premature identification of deviations and nascent threats. The Monitoring & Analytics Layer is important in operationalizing this loop, converting runtime evidence into organized improvement signals instead of static reports.
Another striking feature of the proposed architecture is modularization for scalability and adaptation. Every archi-tectural capability—data input, reasoning, validation, exe-cution, and monitoring—is executed as a loosely connected module. This modularity enables the architecture to progress over time without necessitating a complete system redesign. For instance: LLM models may be substituted or refined autonomously within the Model Layer. New analytical engines may be integrated into the Monitoring & Analytics Layer. Furthermore, additional governance constraints can be integrated into the Compliance & Risk Constraints Layer. This architecture facilitates long-term maintainability and scalability within actual organizational contexts.
In the same way, the suggested solution, unlike conven-tional BPM designs that emphasize either discovery or exe-cution, retains explicit representations of both As-Is and To-Be process models in the Process Models Repository. This dual representation facilitates traceability of enhancement decisions, contrast between the present and desired situations, simulation, and validation preceding deployment. Addition-ally, it facilitates LLM-Based Process Intelligent Layer and Improvement Reasoning Engine by providing structured artifacts that ground semantic explanations in formal process models.
Furthermore, explainability is regarded as a fundamen-tal architectural consideration rather than a secondary mat-ter. The architecture ensures that each AI-assisted out-put— diagnostic insight, redesign recommendation, or monitoring alert—can be traced to fundamental data sources, analytical findings, and defined restrictions or objectives. By designat-ing LLMs as elucidators and synthesizers instead of indepen-dent decision-makers, the framework improves transparency and fosters appropriate AI utilization in BPM scenarios.
Finally, the architecture is designed to seamlessly interact with existing business applications, including ERP, CRM, and RPA platforms. This prevents the formation of inde-pendent AI systems and ensures that CBPI functions within established organizational frameworks. The Compliance & Risk Constraints Layer ensures that all process modifications adhere to organizational policies, regulatory mandates, and auditability standards, thus improving the appropriate-ness of the architecture for practical implementation.
In summary, proposed LLM-CBPI Framework reference architecture is deliberately structured as a modular, regulated, and elucidative socio-technical system that harmonizes ana-lytical precision, AI-enhanced reasoning, and human over-sight. Every architectural decision immediately facilitates the ongoing, evidence-driven, and accountable enhancement of business processes.
4) FRAMEWORK AND ARCHITECTURAL RESPONSE TO.
IDENTIFIED GAPS
In this section, we summarize the correlation between the gaps identified in the SLR and presented in Section VIII and the elements of the LLM-CBPI Framework that address these gaps. This approach underscores the function of LLMs as facilitators of CBPI, focusing on automation, optimization, and decision support. By applying these elements, organizations can optimize the advantages of LLM-driven continuous process improvement.
Gap-01: End-to-end CBPI Lifecycle Coverage Through its closed-loop approach, the LLM-CBPI frame-work pro-vides clear assistance for the entire CBPI lifecycle, encom-passing its six phases: process comprehension, diagnosis, redesign, simulation/validation, execution support, and ongo-ing monitoring. In the same way, its architecture incorporates elements for analysis, reasoning, validation, enactment, and monitoring, while delineating the data and information flows between them. Moreover, each phase is characterized by specific inputs and outputs, at least one architectural com-ponent operationally supports each phase, and feedback from monitoring can initiate a renewed diagnosis and redesign.
Gap-02: Heterogeneous Data Ingestion and Access The proposed framework assimilates and offers regulated access to diverse data sources, encompassing structured (e.g., relational records, KPIs, event logs), semi-structured (e.g., JSON/XML), and unstructured data (e.g., policies, manuals, tickets, emails). The Input Data Layer provides standard-ized inputs for structured and semi-structured data. Struc-tured Data Sources and the ETL/ELT Pipeline perform the extraction, cleansing, transformation, and integration of structured operational data. Simultaneously, the LLM/RAG Pipeline and Vector Database provide semantic ingestion, indexing, and retrieval of unstructured knowledge sources. Furthermore, Event Logs offer dedicated support for event-driven evidence.
In this way, this strategy provides a suit-able approach to separating responsibilities between inges-tion, storage, retrieval, higher-level reasoning, and analytical functions.
Gap-03: Separation of Concerns and Modular Decom-position As evidenced in the previous paragraph, the LLM-CBPI reference architecture establishes a modular divi-sion of duties, segregating data access, transformation, modeling, retrieval, reasoning, validation, execution, and monitor-ing. It also features a layered structure with explicit interfaces and defined responsibilities for each layer. Each element serves a singular core function, cross-cutting issues (such as security and governance) are uniformly enforced, and compo-nents can be substituted without compromising the integrity of the entire architecture.
Gap-04: Event-log Readiness and Empirical Process Evidence In order to support event-log readiness and ensure an empirical basis for process discovery, conformance ver- ification, and performance evaluation, ETL/ELT Pipeline extracts, cleanses, transforms, and enriches execution data from enterprise systems and operational repositories. Event Logs was designed as an architectural element dedicated to storing and maintaining the resulting logs that serve as input to the Process Discovery and Monitoring & Analyt-ics Layer for downstream analysis. This approach provides quality checks, completeness and consistency assessments, standardized trace generation, and continuous availability of empirical evidence for diagnosis and monitoring activities.
Gap-05: Process Model Management and Traceability The LLM-CBPI framework reference architecture pro-vides support for centralized management, versioning, comparison, and controlled evolution across successive redesign iterations through the Process Models Repository, AS-IS, and TO-BE storage repositories. In addition, these architectural elements also provide traceability links between process models and supporting evidence, including event logs, KPIs, compliance constraints, simulation results, and redesign decisions. Thus, it is possible to track changes, compare alternative versions, diagnoses, findings, and governance and compliance require-ments. These elements also offer support for integration with redesign, validation, and execution components, ensuring continuity between modeling activities and downstream oper-ational stages.
Gap-06: Diagnostic Insight Generation and Inter-pretability The proposed solution supports the generation of diagnostic insights and interpretability. To this end, the reference architecture combines the use of the Monitoring & Analytics Layer, the Model Layer, and the LLM-Based Process Intelligence Layer. These architectural elements are responsible for transforming the execution evidence into structured di-agnostic outputs. These outputs are subse-quently consumed by downstream decision-support compo-nents, including the Improvement Reasoning Engine and the Human-in-the-Loop Redesign Module. In this way, they provide redesign prioritization and stakeholder evaluation.
In this sense, di-agnostic findings have an adequate level of explainability about what happened, where issues occur, and why they are likely emerging, while preserving traceability to underlying evidence sources (event logs, KPIs, compliance constraints, etc.).
Gap-07: Improvement Reasoning and Alternative Gen-eration The framework produces improvement hypotheses and re-design alternatives, encompassing logic, assumptions, anticipated impacts, and clear trade-offs (cost–time–quality– risk). An Improvement Reasoning Engine analyzes diag-nostics, restrictions, and contextual knowledge to generate structured redesign alternatives. In this way, alternatives are generated as organized change sets, each alternative encom-passes justification, and trade-offs and limits are clearly delineated. This architectural element works in conjunction with the LLM-based Process Intelligence Layer and the Compliance and Risk Constraints Layer. In this way, it is possible to ensure that the proposed alternatives remain contextually relevant and aligned with governance.
This approach ensures that the redesign options are always accom-panied by a clear justification, expected benefits, implementation limitations, and comparable compensation profiles for subsequent human evaluation.
Gap-08: Human-in-the-Loop Accountability and Approval Gates
The LLM-CBPI framework incorporates human-in-the-loop methods for the review, refinement, and approval of diagnostic interpretations and redesign ideas, to ensure accountability, contextual judgment, and integration of tacit knowledge. Human-in-the-Loop Redesign Module offers interaction points, explanatory views, and decision documen-tation. As a result, deployment necessitates prior approval, decisions are documented with rationale, and expert comment informs revisions of design artifacts. Furthermore, in coop-eration with the Improvement Reasoning Engine and the Process Models Repository, these three elements of the pro-posed architecture allow stakeholders to evaluate proposed changes in a structured way, request adjustments, compare alternatives, and formally approve the selected design arti-facts.
In this context, critical redesign decisions are subject to explicit approval gates, including managerial rationale and expert feedback that can iteratively refine the results before validation or downstream deployment.
Gap-09: Ex-ante Simulation and Validation of Re-designs
The proposed framework provides ex-ante simulation and validation of redesigned process alternatives before operational deployment. The objective is to mitigate the implementation risk and improve evidence-based decision-making. In this context, the Simulation / Validation Layer evaluates the TO-BE models under alternative demand con-ditions, resource allocations, policy constraints, and opera-tional assumptions. The Simulation / Validation Layer works together with the Process Models Repository, Historical Baselines, and the Monitoring & Analytics Layer, allowing redesigned processes to be compared against current per-formance levels, target KPIs, and prior execution patterns. As a result, it is possible to check projected cycle times, throughput, queue levels, resource utilization, SLA achieve-ment, and expected cost implications.
In parallel, the LLM-Based Process Intelligence Layer interprets the results of the validation process by identifying hidden bottlenecks, unreal-istic assumptions, policy conflicts, or additional improvement opportunities. Thus, redesign alternatives can be screened before implementation, high-risk proposals can be rejected or refined, and only validated process models are promoted to the downstream execution stages.
Gap-10: Operational Implementation and Integration with Enterprise Systems
In the LLM-CBPI Framework reference architecture, the Workflow Execution Engine orchestrates task routing, deci-sion logic, transfers, escalations, and automated activities according to the redesigned process approved in the TO-BE models. This work is done in conjunction with the Business Applications Layer, as this layer connects the elements of the designed architecture with business systems such as ERP, CRM, databases, legacy platforms, and RPA. In this way, it is possible to deploy workflows, exchange data, process trans-actions, and promote synchronization with existing opera-tional infrastructures. Simultaneously, the evidence generated during execution is captured and forwarded to the Monitoring & Analytics Layer. This ensures real-time traceability of operations and continuous visibility of process performance. Gap-11: Continuous Monitoring and Closed-loop Feed-back
The LLM-CBPI Framework provides a phase of continu-ous monitoring of implemented processes. The objective is to obtain feedback that allows for continuous performance evaluation as well as to detect deviations and promote iter-ative improvement. In this context, the architectural element Continuous Monitoring & Feedback Layer is responsible for obtaining real-time evidence from operational environments and continuously evaluating process behavior in relation to predefined KPIs, SLA targets, and governance boundaries. The Monitoring & Analysis Layer and Historical Baselines allow current performance to be compared with previous execution standards, expected service levels, and historical trends.
Furthermore, relevant signals can be forwarded to the elements of the architecture Diagnostic Insights and Improvement Reasoning Engine for reanalysis and gener-ation of corrective recommendations.
Gap-12: Historical Baselines and Comparative Evalu-ation
The proposed reference architecture maintains Histori-cal Baselines that contain previous operational behavior, process performance records, and previously observed KPI levels. This approach allows for before-and-after compar-isons, boundary setting, and evidence-based evaluation of improvement initiatives. These baselines provide a reference point for determining whether observed changes represent genuine improvement or normal operational variation. Also noteworthy here are the architectural elements: the Simula-tion/Validation Module, the Simulation/Validation Layer, the Monitoring & Analytics Layer, and the Continuous Monitoring Layer. These elements allow redesigned pro-cesses and real-time executions to be evaluated against his-torical performance standards.
Thus, the proposed solution allows for the calibration of simulation parameters, sup-ports post-implementation performance evaluation, enables anomaly detection, and strengthens the longitudinal monitor-ing of continuous improvement results.
Gap-13: Explainability, Traceability, and Auditable Reasoning (LLM-grounded)
To support explainability, traceability, and auditable rea-soning, the proposed solution links diagnostic and pre-scriptive outputs to the underlying evidence used during the analysis and generation of recommendations. In the reference architecture, the LLM/RAG Pipeline works with organizational context information and evidence to support the answers, pre-serving references related to supporting artifacts. To this end, the LLM/RAG Pipeline works in conjunction with the Vector Database, Process Model Repos-itory, Diagnostic Insights, and the Compliance & Risk Con-straints Layer in order to provide recommendations linked to process models, KPI evidence, business rules, historical records, and governance constraints.
This approach provides greater transparency to stakeholders as it allows the Improve-ment Reasoning Engine to present the assumptions, trade-offs, and decision criteria used to propose process redesign alternatives.
Gap-14: Knowledge Freshness and Adaptability With-out Retraining
In order to support knowledge freshness and adaptabil-ity without retraining, the LLM-CBPI Framework reference architecture provides continuous updates of organizational knowledge sources such as policies, manuals, constraints, and process documentation. To do that, the Input Data Layer, in conjunction with the Structured Data Sources, the LLM/RAG Pipeline, and the Vector Database, allow newly available content to be ingested, indexed, and retrieved incre-mentally at the time of inference. In this way, updated artifacts are immediately available for subsequent reasoning tasks, while obsolete or superseded documents can be deactivated through repository maintenance procedures. Furthermore, the Compliance & Risk Restrictions Layer offers metadata fil-ters, document provenance attributes, and access control poli-cies to provide greater control over retrieval operations.
Gap-15: Compliance and Risk Constraints Enforce-ment
The proposed Framework incorporates and applies com-pliance requirements, control obligations, and risk limits throughout the entire improvement lifecycle, including diag-nosis, redesign, validation, execution, and continuous moni-toring. This capability is primarily implemented through the dedicated Compliance & Risk Constraints Layer, which cen-tralizes business rules, regulatory requirements, security con-trols, approval policies, and organizational risk tolerances. This element of the LLM-CBPI Framework reference archi-tecture works in a coordinated manner with the
Improvement Reasoning Engine, the Simulation/Validation Layer, the Workflow Execution Engine, and the Monitoring & Analytics Layer. The goal is to ensure that analytical out-put, redesign alternatives, executable workflows, and runtime operations remain aligned with governance.
Gap-16: Security, Privacy, and Access Control
The proposed LLM-CBPI framework incorporates secu-rity, privacy, and access control mechanisms across all lay-ers of its reference architecture. The objective is to ensure that sensitive process information and organizational data are handled in accordance with internal policies and regulatory requirements. In this context, the Input Data Layer, Struc-tured Data Sources, the LLM/RAG Pipeline, and the Vec-tor Database apply authentication, authorization, metadata-based filtering, and least-privilege retrieval policies that restrict the exposure of protected content during the ingestion, indexing, retrieval, and generation stages. Data minimization principles can also be applied so that only relevant informa-tion is processed for each task.
In parallel, the Compliance & Risk Constraints Layer supports the application of privacy and security rules, while the logging mechanisms and Event Logs provide auditable records of data access, model interac-tions, and operational actions.
Gap-17: Reproducibility and Operational Reliability
In the LLM-CBPI Framework reference architecture, reproducibility is primarily enabled by the Model Layer, the ETL/ELT Pipeline, the LLM/RAG Pipeline, and the Process Model Repository, where versions of models, data trans-formations, recovery configurations, prompts, and generated artifacts can be tracked and historically recorded. This allows previous runs to be executed again with equivalent inputs and configurations whenever auditing, validation, or trou-bleshooting activities are required. At the same time, the Monitoring & Analytics Layer supports operational relia-bility by detecting model deviations, data deviations, output degradation, abnormal behavior, and service instability. This approach promotes the traceability, repeatability, and sus-tained performance of AI-enabled continuous improvement operations.
X. CONCLUSION.
The results of this study demonstrate how the incorporation of LLMs has revolutionized BPI and BPM. In order to increase productivity, save costs, and ensure compliance, traditional BPM approaches use structured rule-based process optimiza-tion techniques such as Six Sigma, Lean, and Process Mining. The use of LLMs, however, brings with it fresh perspectives that contradict and enhance these conventional BPM models. In the following, we highlight key points that show how LLMs challenge traditional BPM models and some applica-tions where the use of traditional BPM models is still relevant.
Process Flexibility and Adaptability: conventional BPM models are less flexible in dynamic business environments because they place an emphasis on strict procedures and pre-determined decision trees. The processing powers of LLMs facilitate real-time adaptability and decision mak-ing, giving organizations greater flexibility in responding to changing circumstances.
Handling Unstructured Data: emails, customer reviews, and legal documents are examples of unstructured data sources that are difficult for traditional BPM systems to inte-grate. LLMs can efficiently process and analyze such data, offering useful insights and automation possibilities that are not possible with conventional BPM tools.
Enhancing Human-Centric Workflows: BPM frame-works typically call for human participation in data processing and decision making. By automating intricate pro-cesses such as document classification, sentiment analysis, and compliance checks, LLMs reduce manual workload and change the human role in BPM from execution to strategic decision-making and monitoring.
Predictive and Prescriptive Capabilities: traditional BPM models lack strong predictive analytics, even though they optimize processes using past data. When combined with AI-driven techniques like RLHF, LLMs provide predictive insights that assist organizations in foreseeing bottlenecks and recommending changes before inefficiencies occur.
Confirming Traditional BPM Models: the research vali-dates that standardized BPM frameworks are very important for maintaining process integrity and compliance, with LLMs augmenting rather than supplanting governance structures. LLMs improve BPM automation methodologies such as RPA and Process Mining, increasing precision, decreasing vari-ability, and facilitating automated decision-making.
In AI-driven BPM, Lean and Six Sigma continue to be rel-evant, and LLMs facilitate continuous improvement through automated data analysis and real-time insights. Although LLMs improve workflows, human supervision is essential for fine-tuning, mitigating bias, and ensuring accountability.
Ultimately, LLMs enhance rather than replace traditional BPM by automating processes, improving decision making, and managing unstructured data, while still adhering to essen-tial principles of BPM such as compliance, standardization, and continuous improvement. To optimize advantages, orga-nizations must integrate LLMs with current BPM systems, harmonizing AI-driven automation with human over-sight.
In this context, this study enhances the field of AI-driven process improvement by introducing the LLM-CBPI Frame-work and its reference architecture as a systematic solution to ongoing structural deficiencies identified in current method-ologies. Although previous studies have shown the capability of AI to improve business processes, the majority of cur-rent solutions are disjointed, focusing on singular analytical or automation functions rather than offering a cohesive, lifecycle-oriented and governable framework. The findings of this study suggest that to overcome these restrictions one needs a systematic architectural approach based on clearly specified requirements that together tackle reasoning, empirical evidence, governance, and operational implementation, rather than mere incremental methodological improvements.
The proposed framework addresses these needs by delin-eating seventeen essential gaps that must be addressed to effectively operationalize key dimensions for reliable AI-enabled CBPI, encompassing explainability, traceabil-ity, compliance alignment, human accountability, lifecycle coverage, and validation mechanisms. By addressing these gaps as essential architectural design requirements instead of optional system features, the framework creates a stringent basis for the development of AI-driven process intelligence solutions that are reliable, auditable, and applicable in actual organizational contexts. The provided reference architecture enhances this contribution by converting conceptual needs into modular components and interaction frameworks, thus closing the enduring gap between theoretical models and implementable systems.
This research establishes a cohesive architectural and requirement-driven foundation that converts AI-enabled pro-cess improvement from a disparate array of promising tech-niques into a unified, governable, and operationally feasible paradigm, thereby enhancing both the theoretical and practi-cal dimensions of the field.
XI. LIMITATIONS OF THIS STUDY.
Regarding the limitations of this study, we can state that because LLMs are evolving so rapidly, some of the findings can become out of date very quickly, requiring ongoing updates and further research to take into account new devel-opments in technology. Another drawback is that, even if the research finds patterns in several business sectors, the precise influence of LLMs may differ depending on the sector, organization size, and state of the digital infrastruc-ture. Moreover, a lot of research focused on case studies and theoretical advantages rather than thorough empirical evaluation of LLM-driven organization improvements.
XII. RECOMMENDATIONS FOR FUTURE RESEARCH.
The findings indicate that the future development of AI-driven process enhancement will rely more on the avail-ability of integrative architectures that can coordinate di-verse data sources, reasoning modules, validation processes, and monitoring systems within defined governance and risk parameters, rather than solely on improvements in model performance. The LLM-CBPI Framework not only enhances current methodologies, but also creates a design-focused paradigm that redefines the systematic integration of AI inside process improvement lifecycles.
In this context, although the proposed framework and architecture are conceptually grounded in the literature and their elements have already been tested and validated in previ-ous research, the current arrangement proposed in the LLM-CBPI Framework and its reference architecture represents an innovation that needs empirical validation in real organi-zational environments. In this sense, future research should focus on confirming its practical effectiveness through case studies, pilot implementations, action research, and longitu-dinal evaluations in different sectors and process con-texts.
Another possibility for future research is the refinement and operational validation of the workflows for each of the phases of the LLM-CBPI Framework described in Sec-tion IX-B-2. The proposed workflows for Phases 1 to 6 should be tested with professionals in the field to assess usability, completeness, sequencing logic, decision points, and feasibility of integration with corporate systems. An interesting approach is to subject each workflow of each phase of the LLM-CBPI Framework to the same continu-ous improvement logic advocated in the framework itself (Figure 7). This involves iteratively evolving the LLM-CBPI Framework over time, transforming it into a self-improvement reference model, based on both theoretical advances and accumulated implementation evidence.
Finally, controlled comparisons with traditional BPM approaches and long-term studies are also needed to eval-uate the dynamics of closed-loop learning and the lasting effects of LLMs-enabled improvements on organizational performance, as well as interdisciplinary collaboration that incorporates insights from information systems, software design, governance, and ethics in AI, in order to enhance the theoretical foundations and ensure responsible implementa-tion.