1 More Paper.
Full Reading01:37:14

Conversational Enterprise Management Through Agentic AI: A Five-Pillar Framework for LLM-Driven ERP Systems

1 More Paper · Full Reading

Full Reading podcast cover
Listen to the Full Reading

About this paper

A full audio edition of this paper.

Authors: M.H. Assidiqi, D. Alghazzawi, S. Alarifi, L. Cheng

Publication date: 2026

Read the paper: https://doi.org/10.1080/08839514.2026.2700868

Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/

The authors and publisher do not sponsor or endorse this recording.

Brief episode

Transcript

You’re listening to “Conversational Enterprise Management Through Agentic AI: A Five-Pillar Framework for LLM-Driven ERP Systems,” by M.H. Assidiqi and colleagues. Published in 2026.

Applied Artificial Intelligence An International Journal

ISSN: 0883-9514 (Print) 1087-6545 (Online) Journal homepage: the linked source

Conversational Enterprise Management Through Agentic AI: A Five-Pillar Framework for LLM-Driven ERP Systems

Mohammad Hasbi Assidiqi, Daniyal Alghazzawi, Suaad Alarifi & Li Cheng

To cite this article: Mohammad Hasbi Assidiqi, Daniyal Alghazzawi, Suaad Alarifi & Li Cheng (2026) Conversational Enterprise Management Through Agentic AI: A Five-Pillar Framework for LLM-Driven ERP Systems, Applied Artificial Intelligence, 40:1, 2700868, DOI: 10.1080/08839514.2026.2700868

© 2026 The Author(s). Published with license by Taylor & Francis Group, LLC.

View supplementary material

Published online: 18 Jul 2026.

Submit your article to this journal

Article views: 614

View related articles View Crossmark data

Conversational Enterprise Management Through Agentic AI: A Five-Pillar Framework for LLM-Driven ERP Systems a, Daniyal Alghazzawi a, Suaad Alarifi a, and Li Cheng b Mohammad Hasbi Assidiqi aKing Abdulaziz University, Jeddah, Saudi Arabia; bChinese Academy of Sciences, Urumqi, China

ABSTRACT.

Traditional Enterprise Resource Planning systems impose complex menu-driven workflows requiring extensive training, limiting accessibility for non-technical users. This paper presents a five-pillar framework for agentic AI-driven ERP systems that transforms enterprise management from complexity to conversation. Integrating Large Language Models, Multi-Agent Systems, Natural Language Processing, Intelligent Process Automation, and Security Architecture, the framework enables natural language interfaces for intuitive enterprise interaction. Preliminary validation across three enterprise scenarios (finance, HR, supply chain) achieves 28 of 31 criteria passed (preliminary proof-of-concept, n = 3 scenarios). System dynamics analysis identifies four feedback loops governing adoption dynamics. Comprehensive empirical validation through enterprise deployments remains essential future work.

Introduction.

Enterprise Resource Planning (ERP) systems serve as the backbone of modern business operations, providing integrated platforms to streamline processes across finance, human resources, supply chain, and customer relationship management. In practice, ERP systems often require complex integration and impose high interaction costs, as users must navigate multiple screens and rigid, multi-step workflows to accomplish routine tasks. This complexity increases task completion time, places a substantial cognitive load on users, and elevates the risk of human error due to extensive manual data entry and form-based interactions. Moreover, effective use of traditional ERP systems typically depends on extensive user training and, in many cases, complex customization through coding.

As a result, these systems are primarily accessible to professionally trained personnel, significantly limiting their usability for non-technical users. In addition, ERP systems inherent security challenges, including risks related to unauthorized access, SQL injection attacks, and potential data leakage. These challenges become even more significant for small and medium enterprises (SMEs) face particularly acute challenges due to limited resources, making failed implementations potentially business-threatening. For clarity and consistency, a complete list of abbreviations used throughout this manuscript is provided in Appendix A.

While artificial intelligence integration into ERP systems is not a recent development, with established implementations spanning predictive analytics, supply chain optimization, and intelligent automation, these conventional AI-ERP systems exhibit persistent architectural limitations. Current implementations remain fundamentally AI-enhanced rather than AI-native, incorporating AI capabilities as supplementary layers atop traditional menu-driven architectures rather than reimagining enterprise systems around AI-first principles.

From recent literature, according to Rozenes and Cohen (2022); Munkácsi, Orosz, and Alexy (2024), five critical architectural gaps constrain these systems – inadequate natural language understanding, fragmented distributed intelligence, ineffective language-data translation mechanisms, limited operational execution autonomy, and ongoing enterprise reliability challenges – each analyzed in the Literature Review (Related Work and Gap Analysis). These deficiencies reveal that conventional AI augmentation, while improving specific functions, fails to address the fundamental complexity inherent in traditional ERP architectures; addressing this complexity requires moving beyond incremental AI augmentation toward reconsideration at a more fundamental architectural level.

According to Chawla et al. (2024); Gangapatnam (2025), the emergence of agentic AI presents a transformative paradigm shift that addresses these limitations. Characterized by autonomous decision-making and adaptive learning, agentic AI systems leverage Large Language Models (LLMs), multi-agent coordination, and natural language processing to reduce manual intervention requirements. Current implementations demonstrate practical applications across various business functions, with development roadmaps projecting extensive expansion. This technological evolution represents a fundamental shift from traditional ERP architectures to conversational, AI-driven enterprise management systems.

For example, agentic AI-based financial management applications demonstrate significant improvements in processing speed and error reduction, while supply chain functions benefits from enhanced demand forecasting capabilities.

This paper develops a conceptual framework for agentic AI-driven ERP systems embodying the vision of moving “from complexity to conversation” in enterprise management. By “complexity,” we refer to the traditional paradigm of menu-driven interfaces, technical workflow configurations, and structured form-based interactions that require extensive user training and technical expertise. The proposed shift toward “conversation” envisions natural language interfaces that enable intuitive, dialogue-based enterprise management accessible to users regardless of technical background.

Our research objectives include: designing a unified architectural framework integrating LLMs, multi-agent systems, and natural language interfaces; analyzing how this framework addresses identified gaps in current AI-ERP implementations; developing design principles for enterprise AI adoption; and establishing a foundation for future empirical validation and assessment.

This study proposes a five-pillar architectural framework for agentic AI-driven ERP systems that transforms traditional menu-driven interfaces into conversational, AI-driven enterprise management and proceeds through four phases. First, literature analysis identifies five critical architectural gaps in current AI-ERP implementations – natural language understanding, distributed intelligence, language-data translation, operational execution, and enterprise reliability (Related Work and Gap Analysis). Second, requirements-driven decomposition derives five necessary architectural components – Large Language Models, Multi-Agent Systems, Natural Language Processing, Intelligent Process Automation, and Security Architecture – proven sufficient through necessity-sufficiency-minimality analysis (Results).

Third, system dynamics methodology reveals four causal feedback loops (two reinforcing, two balancing) governing framework adoption, mathematically formalized and mapped to Technology-Organization-Environment (TOE) dimensions (System Dynamics Analysis: Causal Loop Diagrams). Fourth, literature-derived benchmarks establish evaluation targets for future empirical validation through controlled enterprise deployments (Results). The framework’s primary contribution lies in systematic architectural integration of agentic AI technologies to enable conversational enterprise management, with system dynamics and TOE analyses serving as analytical lenses characterizing adoption dynamics rather than independent contributions. Table 1 provides a visual roadmap of this four-phase methodology.

Note: The five-pillar framework represents this paper’s primary theoretical contribution. Causal loop diagrams (CLDs) and TOE analysis serve as analytical tools characterizing adoption dynamics rather than independent contributions.

The following research questions guide this study:

Research Questions

RQ1: Architectural Components: What architectural components are necessary and sufficient for conversational enterprise management, and how do they satisfy completeness requirements?

RQ2: Integration Mechanisms: How should identified components be integrated to enable unified agentic ERP systems while managing complexity and ensuring enterprise-grade reliability?

RQ3: Design Principles: What design principles govern the development of conversational ERP systems that balance automation capability with human oversight and organizational constraints?

RQ4: System Dynamics: What feedback structures govern the interaction between framework components and organizational adoption dynamics, and how do reinforcing and balancing mechanisms influence implementation success?

Having established the research objectives and questions, we now turn to existing literature to identify specific architectural deficiencies in current AI-ERP implementations that motivate the proposed framework.

Literature Review.

Related Work and Gap Analysis

Recent AI-ERP research can be categorized into three approaches: traditional AI augmentation focusing on predictive analytics within specific business functions, specialized domain applications targeting individual enterprise functions, and emerging agentic implementations with autonomous decision-making and multi-agent coordination.

Critical Success Factors (CSFs) represent organizational, technical, and environmental conditions that research has identified as determinants of implementation outcomes. Extensive ERP research has documented CSFs including top management support, user training and involvement, clear project objectives, change management, data quality, and vendor-client partnership. However, a persistent paradox exists: despite decades of CSF research establishing what enables success, ERP implementations continue to experience persistent challenges and suboptimal outcomes. Warren’s comparative analysis reveals that researchers have achieved substantial consensus on CSFs across project types, yet this knowledge has not translated into improved implementation success rates.

This knowledge-practice gap motivates exploration of alternative implementation approaches that fundamentally reduce complexity rather than requiring extensive organizational interventions to manage it.

Gap Identification Approach

The five architectural gaps emerged through thematic analysis of recent AI-ERP literature examining current implementation challenges and capability deficiencies. Each gap represents a recurring theme identified across multiple studies that document specific limitations preventing integrated conversational enterprise management. The gap identification process focused on identifying architectural-level deficiencies (missing or poorly integrated system components) rather than organizational or implementation process issues, aligning with the framework’s design science objectives of proposing technical solutions to persistent challenges.

Analysis reveals five critical architectural deficiencies preventing conversational enterprise management:

Gap 1 - Natural Language Understanding: Conversational solutions demonstrate value but lack integration with traditional enterprise systems.

Gap 2 - Distributed Intelligence: Shadow systems and distributed solutions provide functionality but hinder enterprise-wide integration.

Gap 3 - Language-Data Translation: Systems suffer from communication and information gaps between user intent and operational execution.

Gap 4 - Operational Execution: Traditional systems lack intelligent decision support for automated process execution.

Gap 5 - Enterprise Reliability: Despite extensive CSF research establishing organizational and technical conditions for successful implementation, ERP systems continue to face persistent deployment failures and reliability challenges, indicating a fundamental knowledge-practice gap.

Additionally, current implementations suffer from integration complexity with legacy systems and absence of unified frameworks integrating LLMs, multi-agent systems, and natural language interfaces. These gaps inform the architectural derivation process described in the methodology (Materials and Methods).

Table 2 demonstrates how architectural gaps inform framework component derivation. Each pillar represents a necessary architectural layer addressing a specific gap, ensuring both framework completeness (all gaps covered) and necessity (each component justified by identified need).

This gap-informed design ensures the framework’s five pillars emerge from empirical analysis of current limitations and theoretical foundations rather than arbitrary technology selection.

Research Scope and Contribution Hierarchy

This paper’s primary contribution is the systematic derivation and specification of a five-pillar architectural framework for conversational enterprise management, addressing identified gaps in current AI-ERP implementations. System dynamics analysis (causal loop diagrams) and Technology-Organization-Environment (TOE) framework serve as analytical tools that characterize adoption dynamics and contextualize requirements, respectively, rather than independent theoretical contributions. This focused approach enables architectural depth while acknowledging that comprehensive empirical validation through controlled enterprise deployments remains essential future work.

The framework provides: necessary and sufficient components for conversational ERP capability, integration specifications for pillar coordination, design principles for implementation, and adoption dynamics characterization for deployment planning.

Synthesizing the literature reveals that for agentic AI-driven ERP systems, feedback dynamics become critical as multiple technological pillars must coordinate seamlessly, creating interdependencies where improvements in one component cascade through others. The identified gaps establish what needs to be addressed, and the analytical lenses that follow (Analytical Lenses) examine these adoption dynamics systematically.

Analytical Lenses

This section presents two complementary analytical lenses for examining adoption dynamics and organizational requirements. System dynamics methodology reveals feedback mechanisms governing framework deployment, while the Technology-Organization-Environment (TOE) framework contextualizes adoption factors across technological, organizational, and environmental dimensions. The subsequent Materials and Methods section (Materials and Methods) details the systematic design science process by which the five-pillar framework is constructed.

System Dynamics Methodology

This study employs system dynamics methodology to model the complex feedback behaviors governing agentic AI-ERP adoption, where ERP systems function as socio-technical systems exhibiting dynamic interactions between technological capabilities, organizational adoption, and environmental pressures.

This study applies causal loop diagram (CLD) analysis as the primary analytical tool to reveal feedback structures governing framework adoption. CLDs visually represent how system variables influence each other through positive feedback (reinforcing loops driving exponential growth) and negative feedback (balancing loops maintaining equilibrium). This methodological approach enables examination of how the architectural framework generates emergent adoption dynamics through feedback mechanisms between technological capabilities and organizational factors, revealing non-linear patterns invisible to traditional linear adoption models.

The system dynamics approach employed in this research follows established methodology: identification of key system variables through framework component analysis, mapping of causal relationships based on AI-ERP literature, classification of feedback loops into reinforcing and balancing mechanisms, and mathematical formalization using difference equations to represent dynamic relationships. This structured methodology ensures systematic analysis of complex system behaviors prior to empirical validation.

Technology-Organization-Environment Framework Application

This study applies the Technology-Organization-Environment (TOE) framework to provide comprehensive analysis of adoption factors across three dimensions. The TOE framework enables systematic examination of how technological capabilities, organizational capacity, and environmental pressures interact to influence enterprise AI adoption.

Technological Dimension

This study examines the five architectural pillars (LLMs, MAS, NLP, IPA, Security) and their integration complexity as the technological dimension. These pillars were systematically derived through gap analysis of current AI-ERP implementations (Related Work and Gap Analysis), with each addressing a specific architectural deficiency: LLMs enable natural language understanding, MAS provides distributed intelligence coordination, NLP bridges language-data translation, IPA enables operational execution, and Security ensures enterprise reliability.

Organizational Dimension

This research analyzes organizational factors including adoption readiness, complexity management capacity, and learning capability. These factors moderate how effectively organizations can deploy and operationalize agentic AI technologies regardless of technical sophistication.

Environmental Dimension

This study considers external factors including market forces, regulatory requirements, and competitive pressures that shape adoption decisions and implementation approaches.

TOE-System Dynamics Integration

This study integrates TOE dimensions with causal loop dynamics to enable analysis of how technological pillar improvements cascade through organizational readiness and environmental constraints, generating context-specific adoption patterns. The TOE framework provides three analytical contributions: First, dimensional mapping organizes causal loop variables into interpretable categories (Technology = pillar capabilities, Organization = readiness/capacity, Environment = external pressures), enabling systematic identification of intervention points. Second, contextual differentiation reveals how the framework performs differently across organizational contexts – for example, SMEs face higher organizational constraints while enterprises face higher technology costs.

Third, adoption pathway analysis enables prediction of which feedback loops dominate under different TOE configurations, informing implementation sequencing decisions. Table 3 maps TOE dimensions to specific causal loops identified in the results.

Materials and Methods.

Design science research methodology enhanced with system dynamics analysis structures the development and characterization of the conceptual framework for agentic AI-driven ERP systems. The methodology integrates requirements-driven architectural design, causal loop diagram analysis, and TOE framework application to produce theoretically grounded framework specifications.

Phase 1: Requirements and Gap Analysis

Requirements derivation synthesizes multiple sources to ensure comprehensive framework specifications. The five architectural gaps documented in Related Work and Gap Analysis (conversational AI interfaces, distributed intelligence, language-data translation, operational execution, and reliability standards) provide critical insights into current system deficiencies. These gaps inform requirements alongside three additional sources: TOE framework dimensions (Technology-Organization-Environment Framework Application) establishing technological, organizational, and environmental considerations, system dynamics theory (System Dynamics Methodology) revealing feedback mechanisms requiring architectural support, and enterprise systems design principles including human oversight, security-by-design, and legacy compatibility.

This multi-source analysis translates identified needs into functional and nonfunctional requirements through systematic decomposition, ensuring the framework addresses actual deficiencies while satisfying broader enterprise constraints.

Requirements-Driven Decomposition

Identified gaps translate into functional and nonfunctional requirements through systematic decomposition. Requirements analysis examines necessary capabilities to address each gap, producing candidate architectural components through necessity-sufficiency analysis. This process ensures architectural completeness while maintaining minimality (no redundant components).

Phase 2: Architectural Design and Decomposition

Component Derivation Process

Architectural components emerge through requirements-driven decomposition that maps each capability gap to enabling technologies. Necessity-sufficiency analysis ensures every component is essential while preventing redundant layers. Requirements analysis maps identified gaps to architectural layers: natural language understanding! LLM layer, distributed intelligence! multi-agent system, language-data translation! NLP interface, operational execution! intelligent process automation, enterprise reliability! security framework.

The mapping of Gap 5 (Enterprise Reliability) to the Security Framework pillar warrants explicit clarification: this pillar addresses the technical reliability dimension of the gap – specifically, the architectural mechanisms required to ensure trustworthy, secure, and consistent AI-driven enterprise operations, including access control, audit logging, and hallucination safeguards. The broader organizational dimension of Gap 5, concerning management processes, change governance, and the knowledge-practice divide documented in CSF research, lies outside the architectural scope of this framework and is acknowledged as a deployment consideration rather than a design requirement. This mapping reflects gap-informed requirements synthesis rather than direct gap-to-component translation.

This phase yields the five-pillar architecture as the minimal design outcome satisfying the requirements derived in Phase 1.

Integration Design Principles

Framework integration follows eight design principles derived from enterprise systems best practices: Human-over-the-loop control, Explainable AI decision boundaries, Data-residency-aware processing, Fail-safe agent coordination, Conversational context preservation, Incremental intelligence deployment, Security-by-design integration, Enterprise legacy compatibility. These principles guide component integration specifications.

Architectural Completeness Validation

Completeness validation employs formal necessity-sufficiency proof. Necessity established through elimination analysis: removing any pillar breaks conversational ERP capability chain. Sufficiency demonstrated through coverage analysis: five pillars provide complete coverage of identified requirements. Minimality confirmed through redundancy analysis: each pillar addresses distinct capability gap.

Phase 3: System Dynamics Modeling

The framework employs causal loop diagram (CLD) analysis following established system dynamics methodology to reveal feedback structures governing agentic ERP adoption and operation. System dynamics modeling complements the architectural design by examining how identified components interact across organizational contexts and over time.

CLD Construction Approach

CLD development follows a four-stage methodology: Variable identification through framework component analysis, Causal relationship mapping based on established AI-ERP literature, Loop classification into reinforcing (R) and balancing (B) mechanisms, and Mathematical formalization using difference equations to represent dynamic relationships. This structured approach ensures conceptual completeness prior to quantitative simulation.

Coefficient Estimation

Loop influence coefficients (α, β, γ, δ, ε) are estimated from existing literature documenting individual AI-ERP component performance. For example, the α coefficient (LLM influence on MAS performance) is estimated at 0.75–0.85 based on multi-agent coordination studies, while conversational interface quality coefficients derive from empirical adoption research. All coefficients are derived from reported performance metrics in cited literature; none are directly stated as model parameters in their source publications. These literature-derived estimates provide baseline parameters for future empirical calibration through controlled enterprise deployments. Table 4 provides an explicit mapping of each coefficient to its source and derivation basis.

TOE Framework Integration

System dynamics analysis integrates the Technology-Organization-Environment (TOE) framework to map causal loops onto adoption dimensions. The

Note: All coefficients are literature-derived estimates, not empirically measured parameters of the integrated framework. Individual values for coefficients listed as shared ranges require calibration through controlled enterprise deployments.

technological dimension captures R1 (Technological Loop), the organizational dimension characterizes R2 (Adoption) and B1 (Complexity Management), while the environmental dimension moderates B2 (Training Reduction) through regulatory and market pressures. This mapping connects dynamic feedback structures to established adoption constructs and informs evaluation criteria in Phase 4.

Phase 4: Literature-Based Benchmark Analysis

Performance Metric Extraction

Literature analysis extracts quantitative performance metrics from existing AI-ERP implementations. Systematic data extraction records: metric type, reported improvement, implementation context, and source citation. Metrics categorize into: speed/processing time, accuracy/error reduction, cost savings/ ROI, user experience, and scalability.

Benchmark Qualification and Attribution

All extracted benchmarks receive explicit qualification indicating: source implementation (specific system documented in literature), component-level versus integrated system measurement, enterprise context (industry, organization size, function). Benchmarks are clearly attributed as literature-derived evaluation targets rather than validated results from the proposed framework.

Evaluation Framework Design

Benchmarks inform comprehensive evaluation framework design for future empirical validation. Evaluation dimensions include: functional capability assessment, performance comparison against traditional ERP, user experience studies, security and compliance validation, and total cost of ownership analysis. This framework remains unexecuted, representing essential future work.

Research Contribution Classification

The methodology produces a conceptual architectural framework integrating system dynamics and TOE analysis with three distinct contributions:

Theoretical Contribution

Systematic architectural derivation from literature-identified gaps, revealing necessary and sufficient components for conversational enterprise management.

Methodological Contribution

Integration of system dynamics causal loop analysis with TOE framework for analyzing enterprise AI adoption dynamics.

Design Science Artifact

Five-pillar architectural framework with design principles, integration specifications, and adoption dynamics characterization.

What This Methodology Does NOT Provide

Empirical validation, quantitative simulation results, enterprise deployment case studies, or validated performance metrics. These represent essential future research directions requiring controlled enterprise implementations.

Applying this four-phase methodology yields three interconnected results: the validated five-pillar architecture addressing identified gaps, its system dynamics characterization revealing adoption feedback mechanisms, and literature-derived performance benchmarks establishing future validation targets.

Results.

Five-Pillar Architectural Framework: Design Derivation

Addressing RQ1: Necessary and Sufficient Components

Requirements-driven decomposition and gap analysis (Phase 1–2, Materials and Methods) yield a five-pillar architecture as the minimal complete solution for conversational enterprise management.

Gap Analysis Summary. Literature synthesis reveals five critical deficiencies in existing AI-ERP implementations:

● Conversational solutions demonstrate value but lack integration with traditional enterprise systems

● Shadow systems and distributed solutions provide functionality but hinder enterprise-wide integration

● Systems suffer from communication and information gaps between user intent and operational execution

● Traditional systems lack intelligent decision support for automated process execution

● ERP implementations continue facing the deployment challenges and knowledge-practice gap documented in established CSF research (Related Work and Gap Analysis)

Requirement-Driven Design. Architectural decomposition through gap analysis yields five necessary capabilities (Table 5), forming the minimal complete solution for conversational enterprise management. The architectural completeness is formally validated through necessity-sufficiency-minimality analysis

(Table 6), ensuring the framework addresses all identified requirements without redundancy. This five-pillar decomposition provides the architectural foundation for subsequent integration design (System Dynamics Analysis: Causal Loop Diagrams) and system dynamics analysis (System Dynamics Analysis: Causal Loop Diagrams).

Architectural Specification and Integration

Theoretical Foundation

The derived architecture builds on systems theory and socio-technical perspectives that position ERP suites as adaptive enterprise ecosystems requiring tight alignment between technology, processes, and governance. ERP success research emphasizes organizational readiness, user training, and change management, informing the framework’s focus on conversational accessibility and guided adoption. Adoption constructs from UTAUT and the Technology-Organization-Environment framework ground the need to balance technological capability with organizational capacity and environmental pressures. These theoretical foundations justify the transition from monolithic ERP deployments to the distributed agentic platform derived through the requirements analysis in Materials and Methods.

Derived Conceptual Architecture

Figure 1 summarizes the five-pillar architecture produced by the design science process. The diagram highlights how LLM cognition, multi-agent coordination, NLP interfaces, intelligent process automation, and the security layer interlock to deliver conversational enterprise management. Each pillar instantiates the requirements traced in Table 2, providing explicit interfaces for data exchange, policy enforcement, and human oversight across the enterprise.

Pillar Capabilities

Large Language Models (LLMs). LLMs provide the cognitive substrate that translates enterprise intent into structured actions. The framework combines general-purpose and domain-specialized models (GPT-4, DeepSeek, Qwen, Claude-V3) with retrieval-augmented memory to preserve business context across workflows. Fine-tuning pipelines and policy-aligned guardrails ensure model outputs respect organizational governance constraints.

Multi-Agent Systems (MAS). The MAS layer orchestrates specialized agents for finance, supply chain, HR, and analytics functions using Chain-of-Actions coordination patterns. Figure 2 illustrates agent routing, arbitration, and escalation channels that maintain human-over-the-loop supervision while enabling autonomous collaboration at scale.

Natural Language Interfaces. The NLP layer exposes conversational access to enterprise data and processes while enforcing strict safety measures. Secure NL2SQL generation restricts queries to authorized schemas, validates syntactic and business rules, and logs confidence scores for auditing. Multilingual dialogue management with context retention supports cross-functional collaboration without reverting to traditional ERP screens.

Intelligent Process Automation (IPA). IPA capabilities extend robotic process automation with perception and reasoning modules that execute unstructured tasks, predict process exceptions, and trigger corrective actions. Agents employ feedback from operational data to continuously refine automation scripts, aligning with the balancing loops described in the subsequent system dynamics analysis.

Write-access operations – where agents create, update, or delete ERP records – require a layered transaction safety protocol to prevent hallucination-driven data errors. The protocol enforces four sequential controls. First, pre-execution confidence gating: the LLM layer produces a confidence score for each generated action; operations below a configurable threshold (e.g., < 0.85) are automatically escalated to human review rather than executed autonomously. Second, Human-in-the-Loop (HITL) approval: operations classified as high-risk by a rule-based risk classifier – including financial postings above a monetary threshold, bulk record modifications, and irreversible deletions – are routed to a human approver via an in-context confirmation dialogue before execution.

Third, staged execution: all write operations execute first against a transaction sandbox, generating a diff preview that is presented to the user for confirmation; only upon explicit approval is the transaction committed to the production ERP database. Fourth, error recovery and rollback: each committed transaction is wrapped in a compensating transaction log, enabling single-step rollback if downstream validation identifies inconsistencies. This four-stage protocol ensures that autonomous ERP write access is bounded by explicit human oversight at risk-proportionate checkpoints, directly addressing the reliability dimension of Gap 5 (Related Work and Gap Analysis).

Security and Compliance. The security pillar delivers zero-trust access control, AI-driven threat detection, and automated compliance attestation across the distributed agent fabric. Figure 3 depicts layered controls implementing STRIDE countermeasures, immutable audit logging, and policy enforcement for sensitive interactions. The security pillar complements the IPA transaction safety protocol described above: immutable audit logs capture the full decision chain – confidence scores, HITL approval records, staged diffs, and rollback events – providing a complete forensic trail for regulatory compliance and post-incident review.

Integrated Functional Modules

Enterprise capabilities emerge from orchestrating the five pillars across functional domains. Financial management agents achieve documented 40% process acceleration and 94% error reduction through cooperative reasoning and audit-ready reporting. Supply chain and procurement modules couple predictive analytics with conversational decision support to enhance demand forecasting and risk mitigation. Human resources assistants deliver empathetic, policy-aware interactions and secure NL2SQL access to sensitive employee records. Document and knowledge management services automate classification, compliance checks, and multimodal retrieval to sustain enterprise situational awareness. These modules demonstrate how the derived architecture operationalizes the five pillars while retaining human oversight and auditability.

Having specified the architectural framework and its functional capabilities, we now demonstrate its practical implementation through a proof-of-concept system.

Proof-Of-Concept Demonstration

To demonstrate the framework’s practical applicability, we developed a proof-of-concept implementation integrating all five architectural pillars within a representative enterprise environment. This demonstration illustrates how the theoretical framework translates into operational capability for conversational enterprise management using benchmark-adapted natural language queries.

Implementation and Methodology

Implementation Architecture. The Proof-Of-Concept Instantiates the Five-Pillar Architecture Through Contemporary Enterprise AI Technologies:

● LLM Layer: GPU-accelerated inference platform hosting transformer-based language models for enterprise query understanding

● Multi-Agent System: Distributed agent coordination framework managing specialized agents for financial, human resources, and supply chain domains

● NLP Interface: Natural language to SQL translation pipeline employing semantic parsing techniques adapted from established benchmarks

● IPA Layer: Automated query execution engine with validation and result verification capabilities

● Security Framework: Role-based access control with audit logging and query sanitization

● The system operates against AdventureWorks, a standard enterprise database schema containing representative ERP data spanning financial transactions, human resources operations, and supply chain management. This schema provides realistic complexity including multi-table relationships, organizational hierarchies, temporal data, and domain-specific business rules commonly encountered in production ERP environments.

Reproducibility Specification. the Validation Experiments Employed the Following Configurations and Resources to Enable Independent Replication. Model Configuration: The LLM layer employed NVIDIA NIM openai/gpt-oss-120b, an open-source large language model (temperature = 0.1, maxtokens = 4096) for low-variance query generation. Agent coordination implemented using Agno v2.3.8+ multi-agent framework with Python 3.11 +. Agent behaviors were configured through specification of model selection parameters, domain-specific system prompts, and routing logic for specialist coordination. NLP pipeline integrated Spider v1 benchmark queries adapted to AdventureWorks schema through systematic entity and relation mappings documented in structured scenario metadata specifying benchmark attributions, schema transformations, and validation criteria for each test case.

Infrastructure: Validation experiments executed in Docker-based environment with PostgreSQL 15 database, AdventureWorks schema, and Infisical secret management. Average end-to-end latency of approximately 6 seconds (range: 4.2s to 9.2s across scenarios) reflects this infrastructure; production deployments may exhibit different performance characteristics.

Implementation Availability: AdventureWorks database schema is publicly available from Microsoft at the linked source SQL Server Samples. Spider v1 dataset is accessible at the linked source. The proof-of-concept implementation includes complete source code, schema mapping specifications, automated validation procedures, and comprehensive scenario outcomes. Implementation materials and detailed validation documentation are available as described in the Data Availability statement.

Benchmark Adaptation Methodology. To ensure reproducible validation against established baselines, we adapted queries from the Spider v1 semantic parsing dataset. Spider v1 provides 1,034 human-labeled natural language to SQL pairs across diverse domains, enabling systematic evaluation of cross-domain semantic understanding and query generation capabilities.

Three representative queries were selected from Spider’s validation split and adapted to the AdventureWorks schema through systematic schema mapping:

Financial Domain Adaptation: Original Spider query requesting aggregated sales by entity (database: “singer,” query ID: 1024) was mapped to customer revenue aggregation. Schema mapping: “singer”! “sales.customer,” “song.sales”! “salesorderheader.totaldue.” Preserved semantic pattern: entity-level aggregation with JOIN operation.

HR Domain Adaptation: Original Spider query requesting employee counts by location (database: “employeehireevaluation,” query ID: 263) was mapped to departmental headcount analysis. Schema mapping: geographic grouping! organizational grouping via “hr.department.” Preserved semantic pattern: COUNT aggregation with categorical grouping.

Supply Chain Adaptation: Original Spider query requesting product range statistics (database: “employeehireevaluation,” query ID: 271) was mapped to inventory quantity analysis. Schema mapping: “shop. numberproducts”! “production.product.quantity.” Preserved semantic pattern: MIN/MAX aggregation functions.

This adaptation methodology preserves query complexity characteristics (joins, aggregations, grouping) while contextualizing queries to enterprise operations, enabling validation of both benchmark pattern fidelity and practical ERP applicability.

Experimental Design. Each adapted scenario validates the complete conversational workflow: natural language input reception, intent classification and entity extraction, domain-specific agent selection, structured query generation, validation and security verification, database execution, and natural language response synthesis. This end-to-end process confirms the integration completeness of all five architectural pillars.

Validation Results

Assessment Totals: Scenario 1: 10 checks (9 passed, 1 warning), Scenario 2: 10 checks (8 passed, 2 warnings), Scenario 3: 12 checks (12 passed, 0 warnings). Total: 31 assessments, 28 passed (90.3%).

Each scenario underwent systematic validation using the methodology defined in Table 7, comprising five primary validation criteria plus automated structural checks. Table 8 summarizes validation outcomes.

Aggregate Performance: 28/31 validation checks passed (90.3% success rate), 3 warnings, zero critical failures. All scenarios achieved correct agent routing (3/3) and successful query execution (3/3). Average end-to-end latency: 6.0 seconds (Scenario 1: 9.2s, Scenario 2: 4.5s, Scenario 3: 4.2s).

Warning Analysis: Scenario 1 warning involved date format display preference (YYYY-MM-DD vs. MM/DD/YYYY). Scenario 2 warnings involved non-semantic column ordering and aggregation label formatting. All warnings represent presentation refinements rather than functional deficiencies – queries returned semantically correct data with minor display inconsistencies. No failures in core functional validation (agent routing, SQL generation, execution success, result validity, benchmark alignment).

Methodology Note: Validation employed 5 primary criteria (agent routing, SQL quality, execution success, result validity, benchmark alignment) plus automated structural checks. Check counts varied by scenario complexity: Scenarios 1–2 applied 10 checks each, Scenario 3 applied 12 checks. Total: 31 assessments across 3 scenarios. See Table 7 for detailed criteria definitions.

Statistical Context: With n = 31 total assessments across 3 scenarios, the observed 90.3% success rate (28/ 31) has a 95% confidence interval of [74%, 97%] using the Wilson score method. This sample size provides adequate power (1 β = 0.80, α = 0.05) to detect performance below 65% but insufficient statistical power for fine discrimination between 85% and 95% performance levels. The n = 3 scenario sample enables cross-domain feasibility demonstration (Finance, HR, Supply Chain) but not precise performance benchmarking; production validation would require substantially larger scenario sets (n  30) to enable statistically powered performance characterization with narrow confidence intervals.

Representative Execution Analysis: The supply chain scenario illustrates complete five-pillar integration:

Natural Language Input: User submits conversational query: “What are the minimum and maximum inventory quantities?”

LLM Processing (Pillar 1): Language model classifies intent as inventory-analysis operation, identifies supply-chain domain context, and extracts semantic entities: inventory measurement, aggregation type (min/max), and scope (all products)

Agent Coordination (Pillar 2): Multi-agent orchestrator routes request to specialized supply chain agent based on keyword matching and domain classification (confidence: 0.97)

Query Translation (Pillar 3): NLP pipeline generates structured query implementing MIN/MAX aggregation functions over product inventory table, preserving the semantic pattern of the original Spider benchmark query

Execution and Validation (Pillars 4–5): Intelligent process automation executes query with concurrent validation: result completeness verification, data type consistency checking, value range validation, and security constraint enforcement

Response Synthesis: Natural language generation produces conversational summary: “Inventory levels range from a minimum of 2 units to a maximum of 1,013 units across the product catalog”

This execution sequence demonstrates how the five pillars coordinate to transform conversational input into structured enterprise operations, validating the framework’s core capability for natural language enterprise management.

Having demonstrated the framework’s practical implementation through a proof-of-concept system, we now examine the feedback mechanisms that govern its adoption dynamics and organizational deployment.

System Dynamics Analysis: Causal Loop Diagrams

Causal loop diagram analysis reveals four dynamic feedback structures governing the interaction between the five framework pillars: two reinforcing loops (R1, R2) driving exponential growth and two balancing loops (B1, B2) maintaining system equilibrium.

Technological Reinforcing Loop (R1)

The Technological Reinforcing Loop demonstrates how the five pillars create self-amplifying innovation cycles. Enhanced LLM capabilities enable more sophisticated multi-agent reasoning (Ct! MASt), which improves NLP accuracy through better coordinated language understanding (MASt! NLPt). Improved NLP enables more effective intelligent process automation (NLPt! IPAt), while automated processes enable consistent security controls (IPAt! St). Finally, secure environments enable safer LLM experimentation, completing the reinforcing cycle (St! Ctþ1).

This feedback loop can be expressed mathematically to quantify the cascade effect:

Practical interpretation: Each equation represents how strongly one pillar’s improvement influences the next. For example, α 1⁄4 0:75 means a 10% improvement in LLM capability yields 7.5% improvement in multi-agent performance. Coefficients α-ε (estimated 0.60–0.95 from literature quantify these cascade effects, enabling prediction of how investments in one pillar propagate through the system. The loop’s reinforcing nature means improvements compound exponentially rather than linearly – each cycle amplifies the previous cycle’s gains, as illustrated in Figure 4.

Having characterized how technological pillar improvements create self-amplifying innovation cycles, we now examine how conversational interface quality drives user adoption through network effects.

User Adoption Reinforcing Loop (R2)

The User Adoption Reinforcing Loop demonstrates how conversational interface quality drives exponential user adoption. High-quality conversational interfaces increase user adoption rates (Qt! Ut), more users generate increased usage data (Ut! Dt), greater data volumes enable enhanced system learning (Dt! Lt), and improved learning enhances interface quality, completing the reinforcing cycle (Lt! Qtþ1).

The adoption dynamics can be quantified as:

Practical interpretation: These equations model network effects in AI systems. Higher interface quality (Q) attracts more users (U), who generate more usage data (D), enabling better system learning (L), which improves quality (Qtþ1) – creating a virtuous cycle. Coefficient a (0.80–0.90) indicates that 10% quality improvement yields 8–9% adoption increase. High coefficients b (0.85–0.95) and D (0.75–0.85) from literature suggest strong network effects: each new user substantially improves the system for all subsequent users (Figure 5).

While reinforcing loops R1 and R2 drive exponential growth in technological capability and user adoption, balancing mechanisms prevent runaway complexity and maintain system usability.

Complexity Balancing Loop (B1)

The Complexity Balancing Loop illustrates how the framework manages system complexity through conversational abstraction. Increased system complexity raises user cognitive load (SCt! CLt), high cognitive load drives demand for abstraction (CLt! ANt), abstraction needs motivate conversational interface development (ANt! CSt), and conversational simplification reduces cognitive load (CSt! CLt, negative polarity), creating balancing dynamics.

The complexity management mechanism is formalized as:

Practical interpretation: Unlike reinforcing loops, this balancing loop maintains system stability. The minus sign in Equation 13 is crucial – it creates negative feedback that reduces cognitive load rather than amplifying it. When system complexity (SC) increases cognitive load (CL), organizations respond by investing in conversational simplification (CS), which reduces the load. This self-correcting mechanism prevents runaway complexity that would make the system unusable, automatically driving toward equilibrium between functionality and usability (Figure 6).

Training Cost Balancing Loop (B2)

The Training Cost Balancing Loop shows how conversational interfaces reduce training requirements. High training requirements increase training costs (TRt! TCt), elevated costs drive investment in conversational automation (TCt! CIt), conversational investment enhances natural language capabilities (CIt! NLCt), and improved NL capabilities reduce training requirements (NLCt! TRt, negative polarity), creating balancing dynamics that minimize training burden.

The training cost dynamics are expressed as:

Practical interpretation: This balancing loop quantifies ROI on conversational interface investment. High training requirements (TR) create cost pressure (TC), motivating investment in natural language capabilities (NLC), which reduce future training needs (Equation 17). The coefficient δ 1⁄4 0:64 is derived from the 64% training requirements reduction reported by Gangapatnam (2025)(see Table 9); it is not directly stated as a model parameter in that source but is adopted here as a normalized influence coefficient consistent with the reported empirical figure. It has concrete business meaning: every $100K invested in conversational interfaces reduces training costs by $64K annually. This creates a compelling business case for conversational ERP adoption, particularly for organizations with high employee turnover (Figure 7).

Integrated Multi-Loop System Dynamics

The four loops interact to create complex system behaviors beyond individual loop dynamics:

R1 $ R2 Interaction: Technological improvements (R1) enhance user experience, driving adoption (R2). Increased adoption generates data enabling further technological improvement, creating compounding exponential growth.

R2 $ B1 Interaction: User adoption (R2) increases system usage and feature demands, raising complexity. This triggers abstraction mechanisms (B1) that maintain usability, enabling sustained adoption growth.

B1 $ B2 Interaction: Complexity management (B1) and training reduction (B2) both contribute to lowering total cost of ownership, creating synergistic cost optimization.

R1 $ B2 Interaction: Technological advancement (R1) enables better conversational capabilities, accelerating training cost reduction (B2). Lower training costs reduce adoption barriers, enabling more extensive technological deployment.

The system tends toward equilibrium characterized by: High technological capability through R1, Broad user adoption through R2, Manageable complexity maintained by B1, Low training requirements maintained by B2. Figure 8 presents the integrated causal loop diagram illustrating these multi-loop interactions.

System Dynamics and TOE Framework Integration

The causal loops map directly to TOE framework dimensions:

Technology Dimension. R1 (Technological Loop) represents technological maturity and integration complexity. Loop coefficients α- ε reflect technological readiness and component compatibility.

Organization Dimension. R2 (Adoption Loop) reflects organizational readiness for AI transformation. B1 (Complexity Loop) represents organizational capacity to manage technological complexity. B2 (Training Loop) reflects organizational learning capability and change management effectiveness.

Environment Dimension. External market forces moderate loop dynamics. Competitive pressure influences adoption rate (parameter a in R2). Regulatory requirements affect security investments (parameter δ in R1). Economic conditions impact conversational interface investment (parameter β in B2).

This integration provides comprehensive framework for analyzing enterprise AI adoption across technological capabilities, organizational factors, and environmental pressures.

To contextualize these theoretical dynamics within empirical performance expectations, we present literature-derived benchmarks that establish evaluation targets for comprehensive enterprise validation.

Literature-Derived Performance Benchmarks

Table 9 presents performance benchmarks derived from existing literature documenting individual agentic AI-ERP component implementations. These metrics represent potential capabilities based on isolated component studies rather than validated results from the integrated framework. They serve as evaluation targets for future empirical validation.

Important Clarification: These benchmarks document performance achieved by individual AI-ERP implementations in cited literature. The proposed integrated framework has not been empirically validated. Actual performance requires comprehensive evaluation through controlled enterprise deployments as outlined in the methodology section.

Discussion.

Framework Limitations and Validation Requirements

We acknowledge several fundamental limitations that constrain the current contribution and establish requirements for future validation:

Conceptual Framework without Empirical Validation

This work presents a theoretical framework without empirical validation through enterprise deployments. The proposed architecture represents a conceptual model requiring comprehensive implementation and testing to validate practical viability. All assertions regarding operational benefits represent design intentions rather than validated outcomes.

Literature-Derived Evidence

Quantitative performance metrics cited – including 40% processing time reduction, 97% accuracy in document processing, and 64% training reduction – derive from literature on individual AI-ERP components rather than original validation of the integrated framework. While components demonstrate proven effectiveness individually, comprehensive empirical validation is required to substantiate synergistic benefits of the unified five-pillar architecture.

Causal Loop Coefficients Limitation

The influence coefficients (α, β, γ, δ, ε) in the four causal loops are estimated from heterogeneous literature sources rather than empirical measurements from integrated system deployments. Coefficient ranges (e.g., α = 0.75–0.85 for LLM influence on MAS) represent educated estimates requiring calibration through controlled experiments. Actual coefficient values may vary significantly across enterprise contexts, organizational maturity levels, and implementation approaches.

System Dynamics Model Limitations

The causal loop diagrams represent qualitative system dynamics models without stock-flow formalization or simulation validation. While CLDs reveal feedback structures and dynamic hypotheses, they do not provide quantitative predictions of system behavior over time. Future research should extend CLDs to full system dynamics stock-flow models with numerical simulation and sensitivity analysis.

Implementation Complexity Limitation

The framework’s integration of five advanced AI technologies (LLMs, multi-agent systems, NLP, IPA, security) presents extraordinary implementation challenges that may prove prohibitive for many organizations. Beyond technical complexity, this represents organizational transformation requiring fundamental changes to business processes, governance structures, and employee roles. Many enterprises may lack technical expertise, financial resources, or organizational resilience to successfully implement such comprehensive AI transformation.

Enterprise Context Variability

The framework is designed generically without industry-specific customization or regional regulatory consideration. Actual effectiveness across different organizational sizes, industry sectors, regulatory environments, and cultural contexts requires future empirical validation. TOE framework dimensions may exhibit different relative importance across contexts, requiring context-specific adaptation.

Technology Maturity Limitation

The framework relies on rapidly evolving AI technologies whose capabilities, limitations, and enterprise readiness continue to develop. Current LLM capabilities, multi-agent coordination mechanisms, and enterprise security standards may evolve significantly, potentially affecting framework design requirements and implementation strategies. The framework may require substantial revision as underlying technologies mature.

Validation Methodology Limitation

While the research proposes comprehensive validation methodologies, these remain unexecuted. Actual enterprise validation faces challenges including: Difficulty securing organizations willing to deploy experimental AI systems in production environments, Extended evaluation timelines required for longitudinal adoption assessment, Confounding factors in real-world deployments that complicate causal attribution, Ethical considerations in comparing AI-driven against traditional systems for critical business functions.

Pillar-Specific Validation Findings

Despite the conceptual nature of this framework, systematic analysis of the proof-of-concept reveals preliminary evidence for each architectural pillar’s empirical contribution. Pillar 1 (Large Language Models) achieved perfect intent classification across three scenarios spanning distinct enterprise domains, validating the pillar’s necessity for bridging conversational abstraction and enterprise operations. Pillar 2 (Multi-Agent Systems) achieved 100% domain routing accuracy with confidence scores between 0.89 and 0.97, confirming distributed intelligence coordination capability. Pillar 3 (Natural Language Processing) achieved 100% first-attempt query execution success while preserving Spider benchmark patterns, substantiating language-data translation necessity.

Pillar 4 (Intelligent Process Automation) completed all scenarios without manual intervention with 6.0-second average latency, confirming operational execution capability. Pillar 5 (Security Architecture) enforced role-based access controls with zero security violations across validation checks, confirming enterprise reliability requirements.

Framework Implications and Adoption Challenges

The proposed framework offers transformative potential through distributed scalability (500% workload capacity, conversational accessibility (64% training reduction, and comprehensive automation (99% invoice processing acceleration). However, these literature-derived metrics require validation within the integrated five-pillar context to substantiate synergistic benefits.

Implementation faces three critical challenges: Legacy integration complexity–architectural incompatibilities between traditional and agentic systems necessitate substantial redesign and investment; Explainability and trust–autonomous decision-making introduces “black box” concerns requiring transparent audit mechanisms for regulatory compliance and stakeholder confidence; Distributed security– multi-agent coordination creates novel attack surfaces beyond traditional enterprise security models, necessitating comprehensive zero-trust architectures and governance frameworks.

Scalability and Computational Efficiency

Multi-agent reasoning and NL2SQL translation introduce latency overhead that warrants explicit analysis, particularly for large enterprises operating under high-concurrency conditions. The agentic pipeline decomposes a single user request into sequential stages – intent classification, agent routing, query generation, validation, and response synthesis – each adding processing time relative to a direct database query. The proof-of-concept implementation measured an average end-to-end latency of 6.0 seconds (range: 4.2–9.2 seconds across scenarios), compared with sub-second response times typical of direct structured query execution in traditional ERP systems.

However, this comparison requires contextualization: traditional ERP task completion involves substantial navigation overhead – locating the correct screen, entering parameters across multiple form fields, and interpreting raw tabular output – that is not captured in raw query latency. For complex cross-domain tasks requiring data from multiple ERP modules, the conversational pipeline’s single-turn interaction may yield lower end-to-end task completion time than traditional multi-screen workflows.

Table 10 summarizes the key computational trade-offs between traditional ERP query execution and the proposed agentic framework across five performance dimensions, together with the architectural mitigation strategies available at each stage.

Note: PoC latency figures are from the proof-of-concept deployment described in Proof-Of-Concept Demonstration. Production deployment performance will differ based on infrastructure configuration, model selection, and caching strategy. Literature-reported cloud-native AI deployments demonstrate 500% workload capacity scaling relative to traditional ERP, suggesting that horizontal scaling can address high-concurrency requirements at production scale.

The primary scalability risk lies in LLM inference throughput under simultaneous user load. Architectural mitigations include: semantic response caching that serves repeated or structurally similar queries without re-invoking the full pipeline; asynchronous agent execution enabling parallel processing of independent sub-tasks within complex multi-domain queries; model quantization and distillation that reduce inference latency and memory footprint without proportionate accuracy loss; and a shared inference server architecture that amortizes GPU costs across all agents rather than allocating dedicated compute per agent instance. These design patterns are consistent with established cloud-native AI deployment practices and should be incorporated into production implementations of the framework.

Empirical measurement of throughput, latency distribution, and cost-per-query under realistic enterprise load profiles remains essential future validation work.

Limitations and Future Research

Proof-of-Concept Scope and Limitations: The proof-of-concept validates implementability-in-principle– demonstrating that the five derived pillars can be integrated to achieve conversational enterprise capability – rather than production-grade performance optimization. Three benchmark-adapted scenarios provide representative cross-domain validation but do not constitute comprehensive enterprise coverage. Production deployment would necessitate extensive scenario expansion encompassing diverse query types, multi-turn conversations, ambiguity resolution, and exception handling across complete ERP functional breadth. The demonstration employs Spider v1 benchmark adaptation to enterprise schema rather than custom organization-specific model training.

While this approach enables reproducible validation against established baselines, production systems would require domain-specific fine-tuning incorporating organization-specific terminology, business rules, and conversational patterns. Validation occurred within controlled development environment using standard enterprise sample data (AdventureWorks). Real-world deployment introduces substantial additional complexity: legacy system integration requirements, regulatory compliance constraints, organizational change management challenges, production-scale performance optimization, and data quality variability that extend beyond this technical demonstration scope. This demonstration provides preliminary empirical evidence supporting the framework’s core theoretical claim – that the five systematically derived pillars can be integrated to enable conversational enterprise management.

The 100% routing accuracy and execution success are consistent with the necessity-sufficiency -minimality analysis, demonstrating technical feasibility. However, establishing the framework’s sufficiency across diverse enterprise contexts and query complexities requires comprehensive production validation.

This conceptual framework advances architectural understanding but requires extensive empirical validation. Future research priorities include: Controlled enterprise deployments across diverse organizational contexts to validate performance claims and calibrate system dynamics coefficients, Extension of qualitative causal loop diagrams to quantitative stock-flow models enabling numerical simulation and sensitivity analysis, Development of explainable AI mechanisms with comprehensive audit trails for high-stakes business decisions, Investigation of industry-specific adaptations addressing sector-specific regulatory and operational requirements.

Emerging opportunities include integration with extended reality interfaces for immersive enterprise visualization, federated learning approaches enabling collective intelligence across organizational boundaries while preserving data privacy, and development of autonomous enterprise ecosystems coordinating complex multi-party business processes without human intervention.

Conclusions.

This work proposes a conceptual framework for agentic AI-driven enterprise systems that addresses identified gaps in current AI-ERP implementations through a unified five-pillar architecture integrating Large Language Models, Multi-Agent Systems, Natural Language Processing, Intelligent Process Automation, and Security mechanisms. System dynamics analysis reveals four emergent causal loops-two reinforcing (R1: Technological Innovation, R2: User Adoption) and two balancing (B1: Complexity Management, B2: Training Cost Reduction) – that govern framework adoption dynamics.

The framework’s theoretical contribution lies in its systematic architectural derivation from literature-identified gaps and TOE-based system dynamics analysis. However, a critical limitation is that quantitative benefits are derived from existing literature documenting individual AI-ERP implementations rather than original experimental validation of the integrated framework. The synergistic benefits of the unified architecture require comprehensive empirical assessment through controlled enterprise deployments.

A concrete validation roadmap is proposed to guide future empirical work. An initial pilot phase should engage two to three organizations spanning both SME and large-enterprise contexts, executing a minimum of 30 end-to-end scenarios (n30, as indicated by the statistical context note in Table 8) distributed across the finance, HR, and supply chain domains to achieve statistically powered performance characterization. Participating organizations should represent diverse sectors (manufacturing, services, and public sector) and regional regulatory environments to surface context-specific adoption dynamics.

A subsequent scaled phase should expand to ten or more organizations, enabling longitudinal assessment of reinforcing and balancing loop dynamics (R1, R2, B1, B2) under real deployment conditions and empirical calibration of the system dynamics coefficients currently estimated from literature. Both phases should employ a mixed-methods evaluation design combining quantitative performance metrics (latency, accuracy, error rate, training cost reduction) with qualitative organizational readiness assessments aligned with the TOE framework dimensions.

The paradigm shift “From Complexity to Conversation” represents a fundamental reimagining of enterprise management. Future research must prioritize controlled enterprise validation studies, comprehensive security assessments, and development of explainable AI mechanisms to establish the framework’s effectiveness and ensure responsible deployment in real-world business environments.

Acknowledgements.

The project was funded by KAU Endowment (WAQF) at King Abdulaziz University, Jeddah, Saudi Arabia. The authors, therefore, acknowledge with thanks WAQF and the Deanship of Scientific Research (DSR) for technical and financial support.

During the preparation of this manuscript, the authors used generative AI tools, including Claude Sonnet 4.5, Perplexity Sonar, and ChatGPT 5.2, as synthesis and drafting aids for literature synthesis and manuscript writing. An AI-powered academic search tool (Consensus) was used for literature discovery and reference verification. The core intellectual contributions of this work – including the gap analysis, derivation of the five architectural pillars, necessity-sufficiency-minimality reasoning, and system dynamics modeling – were conceived and developed by the authors. Generative AI tools were not used to generate original conceptual content; their use was limited to supporting literature organization, prose drafting, and iterative text refinement under continuous author oversight.

The authors have reviewed and edited all AI-assisted output and take full responsibility for the content of this publication.

CRediT: Mohammad Hasbi Assidiqi: Investigation, Methodology, Software, Validation, Writing – original draft, Writing – review & editing; Daniyal Alghazzawi: Conceptualization, Funding acquisition, Project administration,

Author contributions

Supervision; Suaad Alarifi: Project administration, Supervision, Validation, Writing – review & editing; Li Cheng: Supervision, Validation.

Disclosure statement

The authors report there are no competing interests to declare.

Funding

This project was funded by KAU Endowment (WAQF) and the Deanship of Scientific Research (DSR) at King Abdulaziz University, Jeddah, Saudi Arabia.

Data and Code Availability

Benchmark Datasets: AdventureWorks enterprise database schema is publicly available from Microsoft SQL Server Samples. The Spider v1 semantic parsing dataset is accessible athttps://yale-lily.github.io/spiderunder CC BY-SA 4.0 license. differentially expressed

Proof-of-Concept Implementation: Complete source code implementing the five-pillar architectural framework, including:

● Multi-agent system with domain-specialized agents (Financial, Human Resources, Supply Chain) and coordination logic.

● Natural language to SQL translation pipeline with Spider v1 benchmark adaptation.

● Schema mapping specifications documenting entity and relation transformations from Spider benchmark databases to AdventureWorks enterprise schema.

● Agent configuration files specifying model selection parameters, domain-specific system prompts, and routing logic.

● Automated validation framework implementing structural checks (SQL pattern detection, result set verification) and semantic checks (value ranges, logical constraints).

● Comprehensive validation report documenting methodology, scenario outcomes, and benchmark alignment analysis.

● Complete scenario results including natural language queries, generated SQL, execution outputs, and assessment outcomes.

The data that support the findings of this study are openly available in Zenodo athttps://doi.org/10.5281/zenodo. 18598339, reference number. These data were derived from the following resources available in the public domain: AdventureWorks enterprise database schema from Microsoft SQL Server Samples (the linked source:) and the Spider v1 semantic parsing dataset (the linked source).

Reproducibility: The implementation includes comprehensive documentation for environment setup (Docker containerization with docker-compose.yml), dependency management (Python 3.11+ with uv package manager), database initialization (PostgreSQL 15 with AdventureWorks schema), and validation procedure execution. Detailed reproduction instructions are provided in the repository README. Average reproduction time for all three validation scenarios: approximately 15minutes on standard development hardware (excluding initial environment setup).

Software Dependencies: Python 3.11+, Agno v2.3.8+ (multi-agent framework), PostgreSQL 15, Docker Engine, Infisical CLI (secrets management). Complete dependency specifications with pinned versions provided in pyproject. toml configuration file.

Appendix.

Abbreviations

Download transcript ↗