1 More Paper.
Full Reading00:30:18

Intelligent Agents For Enterprise Knowledge Management-Construction And Application Based On Large Language Models

1 More Paper · Full Reading

Full Reading podcast cover
Listen to the Full Reading

About this paper

A full audio edition of this paper.

Authors: S. Tian, H. Li, Y. Chen

Publication date: 2026

Read the paper: https://doi.org/10.6180/jase.202610_33.039

Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/

The authors and publisher do not sponsor or endorse this recording.

Transcript

You’re listening to “Intelligent Agents For Enterprise Knowledge Management-Construction And Application Based On Large Language Models,” by S. Tian, H. Li, and Y. Chen. Published in 2026.

Intelligent Agents For Enterprise Knowledge Management-Construction And Application Based On Large Language Models

Shuo Tian, Hongmei Li, and Yufei Chen∗

Department of Economics, Qinhuangdao Vocational and Technical College, Qinhuangdao Hebei, 066000

∗ Corresponding author. E-mail: the email address

This paper presents an intelligent agent for enterprise knowledge management, combining Large Language Models (LLMs), retrieval-augmented generation (RAG), and semantic embeddings to address challenges in unstructured communication data. The system integrates preprocessing, embedding generation, retrieval, summarization, question answering, and recommendations using the Enron Email Dataset. Evaluation results show high BERTScore (0.9275) for summarization with good semantic coherence despite low lexical similarity (low ROUGE scores). Retrieval performance includes Precision @k = 0.65, Recall @k = 0.9286, and nDCG@k = 0.9544, with strong contextual relevance. The model also demonstrates robust hallucination mitigation, with low hallucination rates (0.10) and high factual consistency. Scalability tests show a mean latency of 1.96 seconds and memory usage of 7.55 GB.

Overall, the agent proves to be an effective solution for reducing information overload, enhancing decision-making, and preserving organizational memory.

© The Author(’s). This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.

1. Introduction.

Managing the growing volume of unstructured data gener-ated from communication platforms such as emails, docu-ments and collaborative tools is an issue for organizations in the present digital world.Rule-based processes, static databases and keyword searches are the features of out-dated knowledge management systems, systems incapable of capturing the fuller context of the workplace conversa-tion. Huge filtering LLMs with advanced natural lan-guage understanding and reasoning are the GPT, LLaMA and Falcon. Their ability to process volumes of text, as-sess the semantics of associations from text and generate logical output makes them effective at developing smart agents to structure and operationalize enterprise data. The first of these is that data overload is now a significant problem as employees tend to spend large amounts of time sifting through duplicate or irrelevant information.

The second problem is referred to as decision latency. When managers lack access to relevant aspects of organizational history in a timely manner, it affects their strategic planning and organizational productivity. Thirdly, organizations must contend with knowledge storage, which captures information-rich knowledge as part of a team or depart-ment and leads to efforts that are duplicated, white also leading to a loss of institutional memory. finally, in order to deal with compliance and governance needs across sec-tors such as finance, health care and energy, it necessary to develop systems that enable an intelligent means of au-thenticating the credibility of, summarizing or synthesizing and searching for relevant documents without compromis-ing integrity, privacy and security. Fine-tuning enhances domain flexibility by adapting to organizational language.

• Combines retrieval, summarization, question answer-ing, and recommendations to enhance enterprise knowledge management.

• Utilized as a representative corpus for training, fine-tuning, and testing the agent with realistic enterprise communication.

• BERT models preserve contextual relationships in un-structured email data, enabling efficient retrieval and clustering.

• Integrates fine-tuned LLMs for factual grounding, minimizing hallucinations, and generating context-sensitive responses.

The paper is structured as follows: Introduction, outlin-ing the problem and contributions; 2. Literature survey; 3. Materials and Methods, 4. Results and Discussion, and 5. Conclusion.

2. Literature survey.

The rapid growth of large language models (LLMs) has generated new possibilities across several industries where knowledge management and sound decision-making are critical. Zhang et al. critically discussed the utilization of LLMs in the AECO sector for scheduling, risk evaluation, and contract analysis, whereas Li et al. proposed inte-grating knowledge graphs with LLMs to perform compli-ance checking of construction schemes in the construction and engineering fields. Although Kirk et al. extended this idea to "personalized alignment," its benefits, and its ethical issues, Chen et al. argued that LLMs are going to transform personalization systems from passive recommen-dation to active participation. Using their "Programmer’s Assistant," Ross et al. demonstrated how conversational LLMs can assist software developers by facilitating multi-turn, context-dependent coding assistance.

However, de-spite these developments, issues such as privacy concerns, reinforcement of bias and the absence of empirical vali-dation remain substantial hurdles to large-scale adoption in commercial environments. LLM-facilitated innovation and organizational knowledge processes are emphasized in another very important body of research. While Bousch-ery et al. employed the AI-supplemented double dia-mond approach to emphasize how LLMs can assist innova-tion teams during new product design. While Mokander et al. emphasized, however, the growing need for fact-checking and auditing to ensure exact LLM responses. Kasirzadeh and Gabriel also speak about the ethical and philosophical aspects of conversational agents, provid-ing rules for conversational purposes to coordinate LLMs with human ethics.

Their study is rather theoretical and lacks sufficient application in industry workflows in real-world context. Anisuzzaman et al. develop applied adaptation in construction with the demonstration of how to refine procedures for adapting pretrained LLMs to more specific domains such as medicine with also paying atten-tion to overfitting and the high expenditure of resources for the required work. Unlike standard RAG pipelines, the proposed system integrates dense semantic embeddings with enterprise-specific optimization, enabling improved contextual relevance in organizational communication data. This hybrid design enhances knowledge retrieval accuracy and reduces hallucinations compared to generic RAG im-plementations.

3. Materials and methods.

Today’s businesses generate large volumes of unstructured communication data, such as emails, which pose challenges in knowledge management. Rule- and keyword-based systems lack semantic understanding, leading to irrelevant or incomplete results, and employees spend excessive time sorting through redundant information. Collaboration is hindered by fragmented knowledge across departments, and remote work further complicates knowledge retention.

The methodology for developing the intelligent agent for enterprise knowledge management is outlined in Fig. 1. The Enron Email Dataset undergoes preprocessing, includ-ing entity identification, cleaning, tokenization, normaliza-tion, and thread organization. The Enron Email Dataset, available on Kaggle, is a large, well-known dataset contain-ing over half a million emails exchanged by Enron employ-ees before the company’s failure in 2001. Hallucination is reduced by grounding responses with RAG, applying con-fidence scoring, and validating entities, ensuring factual accuracy and trustworthiness in generated outputs. Raw emails contain extraneous metadata, headers, footers and duplicate messages that make it difficult to interpret the content. Cleaning removes unnecessary data so that the model works only with textual data that is useful to it.

Let E = {e1, e2..., en } be the set of all email messages. Define a cleaning function fclean such that,

Eclean = fclean (E) = {ei ∈ E | ei contains valid content }

This function filters out emails that contain headers, signatures, or duplicates, resulting in a cleaned email set Eclean suitable for analysis. Email content is tokenized to facilitate easier understanding of the LLM by dividing it into smaller sentences or words. For a cleaned email e, the tokenization function f token can be expressed as,

T = ftoken (e) = {t1, t2..., tm }

Were ti represents the ith token (word or sentence). This converts the text e into a sequence of tokens T, which can be used in NLP models for embedding or feature extraction. Text is normalized to standardize it for uniform processing. This includes, conversion of text to lowercase, removing stopwords (such as "the," "is," and," etc.) and handling punctuation or out-of-the-way characters. Let a token ti be transformed by a normalization function fnorm, t′ i = fnorm (ti)

Each token ti is converted to a normalized form t′ i, ensur-ing consistent representation for LLM training or embed-dings. Knowledge graph construction and intelligent agent retrieval and reasoning rely greatly on extracted items. For a tokenized email T = {t1, t2..., tm }, the entity recogni-tion function fNFR is,

ET = fNHR (T) = {(ti, ci) | ti ∈ T, ci ∈ C}

Were C is the set of entity categories (e.g., Person, Project, Date). Each token ti is mapped to an entity category ci if applicable, producing a set of recognized entities ET for the email.

Emails are often part of discussion threads. Context is maintained for LLM reasoning and summarization using thread reconstruction. Formally, given a pre-processed email document d ∈ D, Were D is the corpus of enterprise emails, the embedding function f embed maps text into a vector representation, vd = fembed (d), vd ∈ Rn

Were n is the embedding dimension. These embeddings preserve semantic similarity, i.e., if two emails discuss re- lated topics, their embeddings vd1 and vd2 will have a high cosine similarity, vd1 · vd2 sim (d1, d2) = vd1 vd2

Cosine similarity, a typical measure in natural language processing for the semantic similarity of two documents (or e-mails, in this case), is given by this formula. Here, vd1 and vd2 are the vector embeddings of the documents d1 and d2 generated by models such as SBERT or BERT. The numerator, vd1 · vd2 tells the direction in which the vectors are aligned by calculating the dot product of the vectors.

vd1 vd2

, To ensure that similarity is The denominator, based only on orientation and not on length, normalizes the vectors by their magnitude. The output value is between [−1, 1][−1, 1], with higher semantic similarity represented by values closer to 1. The final email embedding is obtained via mean pooling, m vd = 1 ∑ hi m i=1

The system uses LLaMA-2-7B as the base LLM, fine-tuned on the Enron Email Dataset with RAG-based em-beddings. Training is performed for 3 − 5 epochs with batch size 8 − 16, learning rate 2e − 5, using the AdamW optimizer and gradient accumulation. A PEFT approach via LoRA is applied to efficiently adapt the model while minimizing computational cost. Semantic comprehension of business correspondence is provided by the pretrained LLM. For email threads, context-aware answers are cru-cial. There is no flexibility for business applications that require domain-specific improvements. The model base is represented as,

ˆy = fLLM (x)

Were, x is the input email/query and ˆy is the created response.

A specifically selected sub-quantity of the Enron Email Dataset is used to tailor the LLM to the business email domain. This allows the agent to learn organizational con-texts, dictionaries and communication conventions. Using supervised training, supervised training can be the LLM can be trained on labeled data, such as email summaries and question-answer pairs.

Loss Function: Cross-entropy loss is employed to opti-mize the model,

N ∑ L = − yi · log (ˆyi) i=1

Were, yi is a true label (target token), ˆyi is predicted probability of token i, N number of tokens in the sequence and Hyperparameter Optimization-Training epochs and batch sizes are selected based on validation accuracy to avoid overfitting while maintaining efficiency. The dataset is divided into training (70%), validation (15%), and testing (15%) sets, with model training conducted over 3-5 epochs. Key hyperparameters include batch size, learning rate (2e5), and AdamW optimization to ensure reproducibility and stability. To improve factual ground-ing and reduce hallucinations, the agent integrates a RAG framework.

Embedding Generation: Each email/document is con-verted into a dense vector embedding using BERT or Open AI embeddings, v = fembed (x)

Were, v is the input email text and v ∈ Rd is the embed-ding in a d-dimensional space.

Context Retrieval: Given a query q, the system retrieves the top- k most relevant embeddings based on cosine simi-larity, q · v sim(q, v) = ∥q∥∥∥v∥

Were, q is query embedding and v is a stored document embedding.

Response Generation: The retrieved contexts {v1, v2..., vk } are concatenated with the query and passed into the LLM,

ˆy = fLLM (q ⊕{v1, v2..., vk })

Where ⊕ represents concatenation of query and re-trieved context. Fig. 2 depicts the design of the Enterprise Email Analysis Intelligent Agent with LLM integration. The workflow starts by submitting a user question or email to a base model for initial language comprehension. The model is fine-tuned for enterprise email communication using a cross-entropy loss function. This tiered approach reduces hallucinations, improves retrieval accuracy, and supports scalable decision-making, Q&A, and summariza-tion. The process begins with receiving enterprise emails or queries, which are then mapped to dense vector repre-sentations using pretrained models like BERT or SBERT.

Fig. 3 illustrates the smart agent for corporate knowl-edge management, where a user query is converted into vector embeddings to retrieve contextually relevant infor-mation from the enterprise email corpus. The LLM pro-cesses the context and triggers functional modules for sum-marization, question answering, and recommendations. While the Enron Email Dataset is used as a benchmark, its limitations in capturing modern enterprise complexities highlight the need for real-world validation with propri-etary data.

4. Results and discussion.

The proposed Intelligent Agent for Enterprise Knowledge Management was evaluated on fact accuracy, scalability, retrieval, and summarization. Precision@k: Measures the proportion of relevant items in the top-k retrieved results. Recall@k: Indicates the proportion of relevant items re-trieved in the top-k results out of all relevant items in the dataset. MAP@k (Mean Average Precision at k): The mean of the average precision scores at different cut-off ranks, considering the top-k results. nDCG@k (Normalized Dis-counted Cumulative Gain at k): A metric that evaluates ranking quality by considering the position of relevant doc-uments, normalized to the ideal ranking. BERTScore:

Measures semantic similarity between generated and reference text using BERT embeddings, focusing on factual consistency and contextual accuracy. Scalability analysis confirmed the system’s suitability for business environ-ments, with tolerable latency and memory usage. Hal-lucination rates were low, and fact accuracy and entity precision were high. Overall, the system demonstrated reliability, effectiveness, and strong support for business decisionmaking and knowledge management.

Fig. 4 compares the original email with its automatically generated summary, showing a significant reduction in length (from 5,000 to 300 characters). The summarizing model effectively captures key information while omitting irrelevant content, ensuring clarity and readability. This re-duces information overload, aids knowledge assimilation, and supports high-quality decision-making, highlighting the value of automated summarization in business knowl-edge management.

Fig. 5 compares Precision @k and nDCG @k for varying k values, showing that both measures improve as more documents are used. At k = 1, Precision @1 is low (0.065), indicating difficulty in retrieving a single rel-evant document, while nDCG@1 is slightly better (0.095), considering ranking quality. Precision@10 reaching 0.65 and nDCG@10 at 0.954. The low Precision@1 highlights a limitation in early-ranking, suggesting that enhancing the ranking layer (e.g., re-ranking or hybrid methods) could im-prove first-result accuracy. Despite this, the system excels at ranking relevant emails, emphasizing the importance of relevance ordering over exact precision for better user experience.

Fig. 6 compares baseline and embedding-based recom-mendation methods using Precision@k and NDCG metrics. The baseline method shows modest performance (Preci-sion = 0.40, NDCG = 0.50), indicating limited retrieval of relevant documents. In contrast, the embedding-based method significantly outperforms, with Precision at 0.65 and NDCG at 0.95, highlighting the power of semantic embeddings for accurate, personalized recommendations. The large disparity emphasizes the limitations of keyword-based retrieval and supports embedding-based approaches for improved knowledge discovery in enterprise settings, enhancing organizational effectiveness.

Fig. 7 compares five key retrieval metrics: nDCG, Pre-cision, Recall, F1, and MAP. The system performs best in nDCG (0.95) and Recall (0.93), indicating strong retrieval and ranking of highly relevant documents. Precision is 0.65, showing room for improvement in relevance con- sistency, while F1 (0.76) reflects a good balance between precision and recall. MAP is low at 0.24, suggesting that ranking consistency can be improved. Overall, the system excels in quality ranking and recall, making it well-suited for enterprise knowledge tasks, with room for precision enhancement.

Table 1 shows the performance of summary evaluation using ROUGE and BERTScore. The low ROUGE values (0.1102 for ROUGE-L, 0.1542 for ROUGE-2, and 0.1972 for ROUGE-1) indicate minimal lexical overlap between the generated and reference summaries, suggesting the model focuses on rephrasing rather than exact word match-ing. However, the high BERTScore of 0.9275 demonstrates strong semantic similarity, emphasizing the model’s fo-cus on meaning preservation. These results highlight the system’s ability to produce concise, semantically accurate summaries, making it ideal for business knowledge management and decision support.

Table 2 shows the system’s performance in controlling hallucinations and ensuring factual consistency. The faith-fulness metric (0.85) indicates that most responses adhere closely to the original context, while factual consistency is 0.80, with minor inconsistencies. Entity accuracy is high at 0.90, reflecting accurate tagging of names, dates, and organizations. The hallucination rate is very low at 0.10, showing minimal erroneous outputs. These results high-light that the retrieval-augmented technique effectively reduces hallucinations, making the framework valuable for organizational knowledge management and decision-making.

The ablation study evaluates the impact of key compo-nents on system performance in Table 3. The Full Sys-tem (with RAG, semantic embeddings, and reranking) achieved the highest performance across all metrics (Pre-cision @10 = 0.75, Recall @10 = 0.85, nDCG@10 = 0.92, BERTScore = 0.91). Removing RAG resulted in a sig-nificant drop, reducing Precision@10 to 0.50, showing its importance for context-driven retrieval. Excluding seman-tic embeddings lowered performance further, with Preci-sion@10 at 0.45, highlighting the need for semantic under-standing in document retrieval.

5. Conclusion.

This article introduces a smart agent that improves retrieval capabilities and LLM power, overcoming traditional bar-riers in enterprise knowledge management. It effectively handles information retrieval, summarization, question answering, and recommendations using deep learning, semantic embeddings, and preprocessing. The system achieves high performance, with summarization coherence (BERTScore=0.93), retrieval recall (0.93), ranking (nDCG @ k = 0.95), and low hallucination rates (0.10) along with high entity precision (0.90). Scalability tests show the system efficiently manages large email volumes, enhancing organizational knowledge continuity and decision-making while reducing information overload.

Future research directions include domain adaptation (health, banking, law), multimodal fusion for knowledge graphs, adaptive learning for continuous improvement, the integration of explainable AI for transparency, and efficiency improvements through lean embeddings and distributed architectures for scaling deployment. These advancements will make the system more scalable, flexible, and applicable to industry needs in enterprise knowledge management.

6. Acknowledgement.

7. Declarations.

8. Data availability: not applicable.

Conflicts of Interest: The authors declare that they have no conflicts of interest regarding the publication of this paper.

Funding Statement: This research received no external funding.

9. Author contribution.

Shuo Tian: Conceptualization, methodology, data curation, writing - original draft.

Hongmei Li: Investigation, validation, writing - review and editing.

Yufei Chen: Supervision, project administration, for-mal analysis, writing - review and editing, corresponding author.

Ethical Approval: This study does not involve human participants or animals. Ethical approval was not required. Consent to Participate: Not applicable.

Consent to Publication: Not applicable.

Competing Interests: The authors declare that they have no competing interests.

Download transcript