Integrating generative AI and machine learning classifiers for solving heterogenous MCGDM: a case of employee churn prediction
1 More Paper · Full Reading

About this paper
A full audio edition of this paper.
Authors: H.G. Abu-Faty, A. Kafafy, M.M. Hadhoud, O. Abdel-Raouf
Publication date: 2025
Read the paper: https://doi.org/10.1038/s41598-025-99119-0
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “Integrating generative AI and machine learning classifiers for solving heterogenous MCGDM: a case of employee churn prediction,” by H.G. Abu-Faty and colleagues. Published in 2025.
Integrating generative AI and OPEN machine learning classifiers for solving heterogenous MCGDM: a case of employee churn prediction
Hagar G. Abu-Faty1, Ahmed Kafafy2, Mohiy M. Hadhoud3 & Osama Abdel-Raouf2
Employee churn is a critical issue for companies and organizations, as it directly impacts productivity, efficiency, and overall operational success. High turnover rates increase recruitment and training costs, and disrupt workflows, making it a top concern for institutions aiming to maintain stability, growth and continuity. This study presents a methodology to address the employee churn prediction problem in heterogeneous environments by framing it as a Multiple Criteria Group Decision Making (MCGDM) problem. The proposed methodology integrates generative AI, Traditional MCGDM techniques, and machine learning classifiers to handle this problem type. The proposed methodology is structured into four main stages: data collection, generative AI for creating expert profiles, MCGDM for employee ranking, and machine learning for predictive modeling.
ChatGPT-4 is used as the generative AI model to simulate expert profiles from diverse fields related to churn prediction. The Analytical Hierarchy Process (AHP) is employed to calculate criteria weights, while the Technique for Order Preference by Similarity to Ideal Solution (TOPSIS) is concerned with alternatives’ ranking and employee’s classification into churn likelihood categories. These rankings data are then used for different nine machine learning classifiers, reducing the computational complexity for future predictions. The results reveal that Neural Networks, Gradient Boosting, and Random Forest outperform other used models in predicting employee churn in terms of accuracy.
The proposed methodology offers a scalable, data-driven solution for addressing MCGDM problems, particularly employee churn prediction, by integrating advanced AI techniques with traditional decision-making frameworks.
Group decision making involves multiple individuals coming together to evaluate options and make a collective decision. In such cases, the challenge lies in integrating experts’ varied opinions into a unified decision1. When the decision involves assessing alternatives based on several criteria, it evolves into MCGDM problem. In MCGDM, decision-makers must evaluate options using multiple, often conflicting criteria, and each participant may prioritize these criteria differently. This adds another layer of complexity, as the process requires balancing conflicting inputs, assigning weights to the criteria, and combining the group’s preferences into a final decision. MCGDM techniques are essential for addressing such problems, ensuring that diverse viewpoints and criteria are considered in the decision-making process2.
A wide range of traditional methods is available for solving MCGDM problems. These methods include popular approaches like AHP3, TOPSIS4, VIKOR5, ELECTRE6, and PROMETHEE7, each offering a unique framework for evaluating alternatives based on multiple criteria. One of the most widely used traditional techniques in MCGDM is the AHP technique8. AHP helps decision-makers break down complex problems into a hierarchy of simpler sub-problems. By using pairwise comparisons, it assigns weights to different criteria, which can then be aggregated to evaluate the alternatives. Another popular method is the TOPSIS, which ranks alternatives based on their distance from the positive ideal and the negative-ideal solution. The goal of TOPSIS is to find the alternative closest to the ideal solution and furthest from the least preferred solution9.
The integration of AHP and TOPSIS provides a structured, transparent, and scalable approach to solving MCDM problems. AHP ensures logical and consistent weight assignment, while 2Machine Intelligence Department, Faculty of Artificial Intelligence, Minufiya University, Shebeen El-Kom, Egypt. 3Information Technology Department, Faculty of Computers and Information, Minufiya University, Shebeen El-Kom, Egypt. email: the email address
TOPSIS offers a fast and efficient ranking process. This hybrid approach is highly effective for making data-driven, interpretable, and objective decisions across various industries10,11. While traditional decision-making techniques are effective, they often face challenges in complex, heterogeneous environments. Heterogeneous Group Decision Making (HGDM) involves decision-makers with differing expertise, preferences, opinions, or strategies, each bringing unique perspectives and using varied criteria to assess alternatives12. Key features of HGDM include diverse expertise, where decision-makers contribute based on their distinct knowledge and skills, and varied criteria, where each participant may prioritize different factors. Moreover, inconsistent preferences can emerge as participants assign different weights to the same criteria, complicating the aggregation of opinions.
These complexities necessitate the use of MCGDM methods to combine diverse inputs and guide the group reach a consensus13. To overcome the challenges posed by heterogeneous environments, newer approaches and methods, often leveraging advanced algorithms and artificial intelligence, are being explored to enhance decision-making accuracy and adaptability in complex group decision-making scenarios14.
Churn prediction is a critical issue for organizations and companies, as high employee turnover can lead to increased costs, operational disruptions, and reduced productivity. The churn prediction problem involves identifying patterns and factors that influence an employee’s decision to leave, using historical data and analytical techniques15. Churn prediction can be effectively addressed as a MCGDM problem, as it requires evaluating employees based on various factors. This process requires input from multiple decision-makers each contributing unique perspectives and priorities which enhance the overall decision quality but it comes with several limitations16,17. Gathering a group of experts for decision-making is time-consuming and costly, requiring coordination, compensation, and resources.
This makes traditional group decision-making impractical in situations where quick and efficient decisions, like churn prediction, are needed18. Many research works have been targeted toward solving Churn prediction problem as a MCDM problem by integrating Machine Learning (ML) techniques to address the limitations of traditional methods as presented in19–28.
The authors in19 explores the use of different predictive models for customer churn prediction problems in the telecom industry using six phases and highlights the model’s high accuracy through k-fold cross-validation20. proposes a novel framework for employee churn prediction, implemented using the WEKA software. It compares the efficiency and performance of Decision Tree and Logistic Regression models, demonstrating their respective strengths and predictive capabilities in forecasting employee turnover. The use of various machine learning models, such as Naive Bayes, Decision Tree, and Random Forest, is examined to predict employee churn and concludes that Random Forest is the best performing model for employee turnover prediction in21–23 and 24. In25 a machine learning model for predicting customer churn in the telecom sector using big data is discussed.
It explores feature engineering and social network analysis (SNA) to enhance prediction accuracy, with the XGBoost algorithm achieving the highest performance. A deep learning model based on convolutional neural networks (CNN) for predicting turnover in the workforce, proving its superior accuracy compared to traditional statistical methods is explored in26. The authors in27 introduce a method to predict employee churn by integrating MCDM techniques like CRITIC and MARCOS, and machine learning algorithms, showing that categorizing employees based on their importance improves churn prediction and retention strategies, with the CatBoost algorithm demonstrating the best performance. The authors of28 present a novel approach using MCDM techniques, the De-Pareto principle, and machine learning algorithms for categorizing employees and predicting churn.
Generative AI, particularly models like GPT-4, are highly effective tools for creating virtual experts by utilizing their large language model (LLM) capabilities. These AI-driven experts can help address many of the challenges in traditional group decision-making, such as those found in churn prediction problems. By simulating expert decision-making processes, AI can integrate knowledge from various domains and provide a consensus-driven approach without the delays typical of human group dynamics. AI-generated experts are able to process vast amounts of data quickly, utilizing machine learning models, and offer unbiased, data-backed recommendations. Additionally, they can help reduce conflicts stemming from human biases, desires, and backgrounds, as AI systems prioritize objective, data-driven analysis.
However, while generative AI enhances decision-making with consistent, scalable outputs, it lacks the intuitive judgment and contextual understanding that human experts provide. As a result, continuous validation is necessary to ensure the AI’s recommendations align with real-world complexities and conditions. The authors of29 evaluate GPT-4’s ability to provide medical advice by comparing its responses to those from human medical experts. A thorough evaluation of GPT-4’s translation capabilities compared to human translators is conducted in30. The study examines performance across various languages, subject domains, and expertise levels. The authors of31 propose a framework integrating GPT-4 with AHP to enhance automated multi-criteria decision support, enabling faster and more consistent decision-making while reducing the need for human intervention.
Generative AI, like GPT-4, acts as a virtual expert in employee churn prediction by analyzing complex data sources to provide real-time, data-driven insights and retention strategies, reducing reliance on human experts and enabling faster decision-making.
Predicting churn helps organizations take early steps to keep valuable employees by identifying and addressing the key reasons that might lead them to leave. The challenge is further heightened in heterogeneous environments, as each factor may carry different weights depending on the employee’s position, department, or personal circumstances, making it difficult to arrive at a single, clear decision. This research addresses these complexities by using generative AI, machine learning, and MCGDM techniques. Generative AI emulates experts to generate employees’ evaluation datasets, which are then analyzed using traditional MCGDM methods to rank employees based on churn risk. Machine learning models are subsequently trained to streamline future predictions, reducing computational complexity.
This integrated approach helps organizations gain deeper insights into the drivers of turnover and enhances the effectiveness of retention strategies.
The structure of the paper is organized as follows: Sect. 2 provides a detailed explanation of the proposed methodology, Sect. 3 presents the results and discussions related to the proposed methodology, and finally, Sect. 4 concludes the paper and outlines potential directions for future research.
Methodology.
A combination of real and imitative experts is utilized to address the MCGDM problem and rank the alternatives through a series of stages in heterogenous environment. The proposed technique comprises four main stages: data collection, generative AI, Multiple Criteria Group Decision Making (MCGDM), and machine learning for predicting employee churn as illustrated in Fig. 1. In the first stage, Data Collection; relevant employee data is gathered from the actual dataset to be used as a real expert and important features needed for the employee churn prediction process are selected. The second stage, Generative AI utilizes ChatGPT-4’s capabilities to create multiple virtual expert profiles with diverse fields and experiences to simulate real expert decision making process.
Key steps in this stage include the selection of relevant criteria for evaluating employee churn, performing pairwise comparisons to determine the relative importance of these criteria, and evaluating employees based on the selected criteria for each virtual expert. These evaluations are then used to generate datasets for subsequent analysis. In the third stage, MCGDM Techniques are employed to solve the churn prediction problem. AHP is used to determine the weight of each selected criterion, followed by TOPSIS to rank alternatives based on the weighted criteria. Employees are then classified into categories (A, B, or C) reflecting their likelihood of churn. The output obtained from this stage is utilized to gather data for training and testing machine learning algorithms in the subsequent stage.
Finally, Stage 4 focuses on ML Classifiers, where various machine learning models are trained and tested on the datasets to function as experts for predicting employee turnover. The performance of these models is thoroughly evaluated to ensure prediction accuracy.
Data collection and preprocessing
The proposal is based on the IBM Watson dataset, published in 2017 by the Smarter Workforce Institute at IBM, which contains approximately 19,479 employee records and 31 data attributes with no missing values. This structured, tabular dataset is specifically designed for employee churn prediction, helping to determine whether an employee is likely to leave the company based on various factors. The dataset includes a range of work-related attributes, such as job performance, role within the company, job involvement, monthly income, overtime status, and years at the company, as well as personal factors, including age, education level, marital status, and work-life balance. Additionally, it incorporates employee feedback on job satisfaction, workplace environment, and relationships within the company, making it a comprehensive resource for churn analysis.
The dataset’s features are categorized into numerical, ordinal, and categorical types, each playing a crucial role in churn prediction. Numerical features include Age, Number of Companies Worked, Hourly Rate, Monthly Income and other attributes. Ordinal features include Education Level (ranging from 1 to 5), Job Satisfaction (ranging from 1 to 4) and Work-Life Balance (from 1 to 4) and etc. Categorical features can include Gender (Male/ Female), MaritalStatus (Single, Married, Divorced), and Department (HR, Sales, or R&D) and so on. Given its depth and reliability, the dataset is treated as a real expert in the proposed technique, serving as a foundational source for model training and evaluation.
Data preprocessing is a vital step in preparing the dataset to be clean, structured, and well-prepared for machine learning models. The process begins with checking for missing values to confirm that the dataset is complete. Next, duplicate records are detected and removed to eliminate bias and redundancy, improving the model’s learning efficiency. Ordinal features are then validated and mapped into appropriate numerical values for accurate representation. Following this, all data is standardized using the Saaty scale, ensuring consistency across features, as machine learning models perform optimally when numerical attributes are on the same scale. This transformation enhances the dataset’s suitability for ranking alternatives in the MCGDM problem.
Finally, the dataset is split into two parts (80:20 ratio), with 80% allocated for training and 20% for testing machine learning classifiers, facilitating robust model evaluation and improving predictive performance.
Feature selection
The proposed technique employs a multi-stage feature selection approach that incorporates the capabilities of Generative AI (ChatGPT-4) and MCDM techniques to identify the most significant features for employee churn prediction. Initially, ChatGPT-4 was utilized to generate multiple virtual experts including HR specialists and data analysts, who contributed domain-specific insights to analyze the IBM Waston dataset and identify the key attributes influencing employee churn. Each expert contributed insights based on their domain knowledge ensuring that feature selection was data-driven rather than assumption-based, enhancing objectivity and eliminating human bias.
By analyzing the IBM Watson dataset, the AI model identified 15 key features with the highest impact on employee churn, including Job Satisfaction, Monthly Income, OverTime, Years at Company, Work-Life Balance, and Environment Satisfaction. As the employee evaluation is a heterogeneous MCGDM problem, different experts with varied expertise and backgrounds employ distinct evaluation criteria and methods to assess alternatives. Each expert determines the most suitable criteria based on their experience, ensuring a comprehensive and balanced evaluation. Once the key features were selected for each expert, AHP method was applied to prioritize these features by performing a pairwise comparison, assigning relative importance scores, and ensuring logical consistency through Consistency Ratio (CR) calculations. The primary steps of the feature selection process are illustrated in Fig.
1, while Sect. 2.3 provides a detailed explanation of the feature selection methodology based on the Generative AI, and Sect. 2.4 elaborates on the criteria weighting method.
Generative AI
In this research, a newly released feature from OpenAI, which enables the creation of custom GPT models, is utilized. This feature allows the model to be tailored with specific instructions, foundational documentation to be provided, and defined limitations to be set, optimizing the model for particular tasks.
Generative AI represented by GPTs, is employed to generate a set of imitative experts based on the basic information about employees in the provided dataset. On the other hand, the insights and evaluations of the real expert are derived directly from the data within the dataset.
Firstly, a custom ChatGPT model named “Expert Guide” was created to act as the primary decision-maker for the key steps in the AHP model development. This includes identifying the optimal number of virtual experts to be generated and their areas of expertise for solving the problem. Also, Identifying the most relevant criteria from the dataset for the real expert. It can also be used to conduct pairwise comparisons needed by AHP to determine the weights of criteria.
“Expert Guide” GPT was consulted to determine the optimal number of experts required for solving the churn prediction problem and responded with a recommendation of seven experts from various relevant domains. This number is suggested because it balances the need for diverse perspectives with the practicality of managing inputs and synthesizing the results.
The recommended experts and their respective fields are summarized as follows: Expert 1: Behavioral Data Scientist. Expert 2: Workforce Analytics Specialist.
Expert 3: Senior HR Consultant. Expert 4: Labor Economist. Expert 5: Employee Engagement Strategist. Expert 6: Employee Relations Specialist. Expert 7: Data Scientist and Machine Learning Expert. By incorporating expert knowledge across these areas, a well-rounded and insightful AHP model will be built for solving employee churn prediction problem.
The “Expert Guide” recommendation emphasizes the importance of the Behavioral Data Scientist and Senior HR Consultant, as their expertise is crucial for understanding the reasons behind employee turnover and designing effective interventions. Significant importance is also given to the Workforce Analytics Specialist and Employee Engagement Strategist, who offer vital data-driven insights and strategies that can directly impact churn outcomes. Meanwhile, the Labor Economist, Employee Relations Specialist, and Data Scientist play valuable and supportive roles. Their contributions, while essential, are more context-specific, helping to fine-tune the solutions developed by the primary experts.
After considering the number of virtual experts needed to address the problem, custom GPT personas were created based on the virtual profiles provided by the “Expert Guide.” After determining the number of virtual experts. Each expert focuses on a unique area related to churn prediction and employee retention. Expert 1 specializes in data science to analyze employee behavior patterns, identifying early predictors of churn and proactively addressing potential risk factors. Expert 2 develops and applies performance and engagement metrics specifically for churn prediction, offering actionable insights to reduce turnover. Expert 3 advises on HR practices that mitigate turnover risk, using churn analysis insights to enhance retention strategies. Expert 4 examines economic trends and labor market dynamics affecting turnover rates, helping to study churn within broader labor market conditions.
Expert 5 identifies and addresses factors related to engagement and job satisfaction that directly impact employee retention. Expert 6 intervenes in potential churn cases by managing the relationship between employees and the organization, handling conflicts, and ensuring workplace regularity through effective communication and policies. Expert 7 utilizes machine learning and predictive models to analyze historical data, identifying patterns and factors that signal turnover risk. Each expert contributes a distinct set of insights and approaches to build a comprehensive strategy for predicting and mitigating employee churn. As a result of generating the seven GPTs, now the problem has eight experts, one is a real expert based on the information given in the dataset and seven imitative by ChatGPT capabilities.
After that, The “Expert Guide” is tasked with identifying the most effective criteria considering information about all employees in the dataset to be utilized by the real expert. A set of criteria is selected that has the greatest impact on the employee turnover process according to the dataset. The selected criteria are EmpEnvironmentSatisfaction, EmpJobInvolvement, EmpJobSatisfaction, RelationshipSatisfaction, and EmpWorkLifeBalance. For simplicity, these criteria are listed as Work Environment, Engagement Levels, Job Satisfaction, Peer Interactions, and Work-life Balance.
Once the seven experts are generated, each of them is asked to propose a number of evaluation criteria for employee churn prediction according to their specific field of expertise. Our imitative experts each suggested an extensive list of 5 primary criteria, ultimately resulting in a total of 15 unique top-level criteria without duplication. The suggested criteria for all experts are summarized in Table 1. The distribution of decision-makers’ evaluations across 15 criteria (C1 to C15) is shown in Fig. 2. Each slice of the pie chart represents the percentage of votes attributed to each criterion. C7, C8, and C9 have the highest percentage of votes, indicating that these criteria are the most significant in the decision-making process. On the other hand, C3, C4, C6, C11, C12, C13, and C14 have each received 2.5% of the total votes, making them less significant compared to others.
Following this, each expert is asked to conduct pairwise comparisons of their selected criteria to calculate the criteria weights using the AHP technique.
As a final step for generative AI, the experts are asked to evaluate the employees in the provided dataset using the Saaty scale, ranging from 1 to 9. Each expert assesses the criteria as if they were human, drawing on their background and experience to provide a comprehensive evaluation.
MCGDM techniques
AHP and TOPSIS are two widely used techniques for solving MCGDM problems. In this stage, AHP is applied to calculate the criteria weights for each decision maker. These weights are then passed to TOPSIS, which ranks the employees in the dataset based on their likelihood of turnover.
The integration of AHP and TOPSIS provides a structured, transparent, and efficient framework for ranking employees based on their likelihood of churn. AHP determines the importance of different churn factors through a hierarchical structure and pairwise comparisons, ensuring that qualitative factors are systematically evaluated. Additionally, the Consistency Ratio test enhances the reliability of feature selection by maintaining logical consistency in expert judgments. Once AHP assigns criteria weights, TOPSIS ranks employees based on their proximity to an ideal retention scenario and distance from a worst-case churn scenario, offering clear, data-driven predictions using Euclidean distance calculations. TOPSIS considers all criteria simultaneously, providing a balanced ranking and ensuring computational efficiency, making it ideal for large-scale HR analytics.
By combining AHP’s structured weight assignment with TOPSIS’s efficient ranking system, this approach merges subjective decision-making with objective ranking, enabling interpretable and data-driven employee retention strategies. Furthermore, integrating Generative AI (GAI) with AHP-TOPSIS automates expert evaluations, enhances ranking accuracy, and enables real-time churn predictions. This combination minimizes human bias, improves decision-making, and ensures a scalable, data-driven workforce management strategy, allowing organizations to proactively identify churn risks and implement effective retention strategies.
AHP technique
After generating decision matrices for each expert using generative AI, the problem now consists of eight distinct decision matrices; seven obtained from the virtual experts and one from the original dataset, which serves as the real expert. The complete AHP hierarchy is displayed in Fig. 3. Each expert GPT is tasked with conducting pairwise comparisons of their selected criteria to calculate the criteria weights using the AHP technique. Additionally, the “Expert Guide” is asked to perform a pairwise comparison of the criteria selected from the original dataset, which will be assigned to the real expert. The AHP technique is applied to calculate the criteria weights based on the pairwise comparisons provided by the experts. Considering the overlaps between the criteria, the final weight for each criterion will be determined using Eqs. and.
Consider a multi-criteria group decision-making problem with m alternatives, n criteria, and p decision makers. Let DM = { } be the set of p decision makers. A = {A1, A2..., Am} DM 1, DM 2..., DM p be a set of m alternatives and C = {C1, C2..., C n} be the set of n criteria. Wij = {W1, W2..., W n} be the set of criteria weights assigned by decision maker i to criterion j
1. Calculate the weighted sum for each criterion by Eq.
∑ p Sj = Wij ∗ Nij i=1
Where Sj is the weighted sum for criterion Cj. Wij is the weight assigned by decision maker i to criterion j. Nij is the number of decision makers evaluating criterion j.
2. Calculate the final weight for each criterion by Eq.
Where Wj is the final weight for criterion Cj.
Sj Wj = ∑n j=1 Sj
The weights determined by each expert along with final criteria weights are presented in Table 2. The criteria weights represent the significance of the features used in conjunction with the machine learning models. The importance of these criteria or features is displayed in Fig. 4.
TOPSIS
The problem is now formulated as a MCGDM problem involving 8 experts, 15 criteria, and a set of different alternatives or employees. The weights for all criteria have been determined using the AHP technique. The primary role of TOPSIS is to rank employees in order to identify those most likely to leave the organization. TOPSIS seeks to find the alternative that is as close as possible to the best ideal solution while being as far as possible from the worst ideal solution. This makes it a powerful and intuitive method for solving MCGDM problems. The main steps of TOPSIS are summarized as follows:
1. Construction of the decision matrices based on decision makers’ assessments.
Each decision maker (DM) creates their own decision matrix based on their expertise and perspective, using the generative AI framework as explained earlier. Each matrix includes the evaluation of all alternatives (employees) against the selected criteria, with values representing the performance or suitability of each alternative relative to each criterion.
2. Construction of the aggregated decision matrix Using the Geometric Mean Operator:.
The aggregated decision matrix D = [dij ] is calculated using Eq.
)1/nij (nij ∏ D = [dij ] = Xijk k=1
Fig. 4. Criteria Importance.
Where [dij] is the aggregated value for alternative i under criterion j. Xijk is the evaluation provided by expert k for alternative i under criterion j. nij is the number of experts who evaluated alternative i under criterion j.
3. Normalization of the Aggregated Decision Matrix:.
The normalized matrix [dij N ] is computed as Eq.
4. Construction of the weighted normalized decision matrix.
The weighted normalized decision matrix [dij ∗] is calculated as in Eq.
[ N ] dij = √∑m dij i=1 dij 2
5. Identifying of the positive-ideal (PIS) and the negative-ideal (NIS) solutions using Eq.
PIS (d+) is the best possible solution, where NIS (d−) is the worst possible solution. They can be obtained by: d+ = max (dij ∗) for benefit criteria, min (dij ∗) for cost criteria d− = min (dij ∗) for benefit criteria, max (dij ∗) for cost criteria
6. Calculation of the distance measure:.
For each alternative, the distance measure to the PIS is measured by Eq., and to the NIS S + by S − i i Eq. based on the Euclidean distance measure.
n
+ = (dij ∗ − d+)2
Si j=1 n
(dij ∗ − d−)2 − =
Si j=1
7. Calculation of the relative closeness coefficient to the ideal solution by Eq.
− Si CCi = Si+ + Si− where, 0 ⩽ CCi ⩽ 1
8. Ranking the alternatives.
Alternatives are ranked based on their closeness coefficient value, with employees ranked higher being more likely to leave the organization.
After ranking the employees, they are classified into three categories based on data derived from a box plot shown in Fig. 5. The first category includes those with the greatest likelihood of turnover. Table 3 presents the classification of employee categories. The first category, Class A, with the highest rankings, identified by a closeness coefficient greater than the third quartile value (75th Percentile) (Q3 = 0.64732) obtained from the box plot. This group represents employees at high risk of turnover based on churn prediction. Class C represents the lowest risk group for leaving the organization, with closeness coefficients below the first quartile (25th Percentile) (Q1 = 0.45300). Lastly, Class B, comprises those at moderate risk, with closeness coefficients between the first and third quartiles (Q1 and Q3).
Machine learning classification
After obtaining the employee rankings and classification categories for churn prediction using TOPSIS, the next step is to apply various machine learning algorithms to the datasets to validate the effectiveness of using ML for solving this problem.
ML classification is based on four key concepts: classes, features, training, and testing. The classes represent the target categories the model aims to predict. For the churn prediction problem, there are three classes: A, B, and C. Features are the criteria used to make the predictions about the target. In the case of study, 15 different criteria are considered for prediction. The datasets provided by decision makers are used in training process where the classification algorithm learns from labeled data by identifying patterns and relationships between the features and their corresponding classes. After training, the model is evaluated on a testing dataset to assess its accuracy in predicting the correct class.
In this approach, nine different types of supervised learning classifiers are used to categorize employees into one of three categories, similar to TOPSIS method. The training process relies on evaluations obtained from both the real dataset and the generated datasets from imitative experts. The primary goal for applying ML classifiers to the churn prediction problem is to accurately predict potential turnover and deliver actionable insights to help minimize losses.
ML classifiers’ parameters Nine different classifiers are applied to the churn prediction problem to categorize employees into three groups, identifying those most likely to leave the organization.
These classifiers are AdaBoost, Gradient Boosting (CatBoost), and Random Forest, Logistic Regression, KNN, and Neural Network, SVM, Decision Tree, and Naïve Bayes.
In the proposed methodology, the working parameters of ML classifiers are stated as follows: AdaBoost utilizes 60 estimators using decision trees as the base estimator, the Samme. R algorithm for classification, and a linear loss function for regression tasks. Gradient Boosting (CatBoost) employs 170 trees, a learning rate of 0.3, a maximum tree depth of 15, with replicable training and a regularization strength of 3. Random Forest runs with 150 trees, without limits on features or tree depth. It stops splitting nodes with 5 instances, but no replicable training is supported. Logistic Regression takes a simple approach with Ridge (L2) regularization, a C value of 1, and no class weights, ensuring balanced predictions. KNN keeps things basic with 5 neighbors, utilizing the Euclidean distance metric, and uniform weighting for classification.
The Neural Network model is configured with two hidden layers of 50 neurons each. It uses the tanh activation function with the Adam solver, allowing up to 500 iterations with some replicability in training. For SVM, a radial basis function (RBF) kernel with a C-value of 1.0 and a tight numerical tolerance of 0.001is used. It’s set for 100 iterations, ensures a balance between precision and computational efficiency. The Decision Tree is pruned to keep at least 2 instances in leaves and 5 in internal nodes, with a maximum depth of 100, stopping splits when 95% of instances in a node belong to the same class. All these models have been fine-tuned to maximize precision, efficiency, and accuracy, making them highly effective for churn prediction problems.
Results and discussion
The Pareto Principle, also known as the 80:20 Rule, suggests that approximately 80% of the effects come from 20% of the causes. In line with this concept, the dataset was initially split into 80:20 ratio, with 80% used for training the classifiers and 20% reserved for testing the ML models. To further evaluate the performance of the models on training data, a 10-fold cross-validation was performed. Table 4 presents the performance measures of the ML models after splitting the dataset into an 80:20 ratio and categorizing the employees into three distinct groups.
Based on the results from the 80:20 split, Neural Networks and Gradient Boosting consistently demonstrate the best performance across all classes, with high F1-scores, precision, recall, and accuracy. Random Forest also performs exceptionally well, ranking just slightly behind Gradient Boosting and Neural Networks. Naïve Bayes, however, shows the weakest performance across all metrics, particularly struggling with precision, making it less suitable for this churn prediction problem. SVM and KNN offer moderate performance, with SVM facing challenges in Class A, while KNN provides more consistent results across all classes. Logistic Regression and AdaBoost are both provide a solid balanced performance in all categories.
In terms of accuracy:
• Neural Networks achieved the highest accuracy at 95.3%,
• Gradient Boosting followed with 92.8%,
• Random Forest achieved 90.8% accuracy.
These three models stand out as the top performers, with the choice of classifier depending on the specific needs of the churn prediction problem. The percentage of accuracy for all classifiers, ordered descending, is displayed in Fig. 6. The confusion matrices for the three top classifiers are illustrated in Figs. 7, 8 and 9, respectively. Additionally, the evaluation metrics of the ML models were assessed using 10-fold cross-validation to ensure robustness and reliability. The results of this evaluation, including key performance indicators such as accuracy, precision, recall, and F1-score, are comprehensively illustrated in Table 5.
Based on the results from 10-fold cross-validation, Neural Networks and Gradient Boosting emerge as the top-performing models, achieving the highest accuracy (95.3% and 92.2%, respectively) and consistently high F1-scores, precision, and recall across all classes. Random Forest also performs exceptionally well, with 90.2% accuracy and strong metrics overall. On the other hand, Naïve Bayes significantly underperforms, with the lowest accuracy (63.7%) and poor precision, recall, and F1-scores, making it less reliable for this specific dataset. SVM, KNN, Logistic Regression, and AdaBoost deliver moderate performance, with accuracy between 81.9% and 86.5%. While these models provide solid results, they are not as effective as Neural Networks or Gradient Boosting.
In terms of consistency, most models show balanced performance across all classes, with minor variations in precision, recall, and F1-scores for specific classes (A, B, and C). However, SVM shows some struggle in Class A, while Naïve Bayes struggles across the board, particularly in Class B. For this churn prediction problem, Neural Networks, Gradient Boosting, and Random Forest are the top choices due to their high accuracy, precision, recall, and F1-scores. Naïve Bayes is the least suitable model, and the remaining models (SVM, KNN, Logistic Regression, and AdaBoost) offer reasonable performance but are not as strong as the top performers. The final model selection should consider both the task’s accuracy requirements and computational efficiency, with Neural Networks and Gradient Boosting being the most reliable for this problem.
The differences in accuracy percentages using 80:20 split and 10-fold cross validation is shown in Fig. 10.
Results based on another dataset
To assess the generalization capability of the proposed methodology, it has been applied to an additional dataset from INX Future Inc. This dataset comprises 1,200 employee performance records with 28 attributes, reflecting key employee characteristics relevant to churn prediction. The dataset is publicly available on Kaggle. The same methodology integrating feature selection using ChatGPT-4, criteria weighting and employee ranking using AHP-TOPSIS, and classification of the employees using various ML classifiers was tested on the new dataset to evaluate its scalability and adaptability.
Table 6 presents the performance measures of the different ML classifiers when applied to the INX Future Inc. dataset after splitting the dataset into an 80:20 ratio and categorizing the employees into three distinct groups. Additionally, the evaluation metrics of the aforementioned ML models using 10-fold cross-validation are illustrated in Table 7. Figure 11 presents the accuracy comparison across the two datasets used in this study based on 80:20 split and 10-fold cross validation.
Based on the results from the 80:20 split of the second dataset, the ranking of top-performing classifiers differs from the first dataset, highlighting the dataset-dependent nature of classifier performance. Unlike the first dataset, where Neural Networks and Gradient Boosting were the dominant models, the second dataset sees Neural Networks (92.9%) and Logistic Regression (92.5%) achieving the highest accuracy. Additionally, Gradient Boosting (91.7%) and Decision Tree (91.2%) continue to perform well but do not outperform Logistic Regression as they did in the first dataset.
The same performance was obtained Based on the results from 10-fold cross-validation for the second dataset. Neural Networks remain the best model for both datasets. Conversely, Random Forest, which was a strong performer in the first dataset (90.2%), drops in ranking for the second dataset (85.8%), appears more dataset-dependent indicating that its effectiveness varies with different data distributions.
The shifts in classifier rankings between datasets highlight that model performance is influenced by dataset characteristics. While Neural Networks consistently excel, models like Logistic Regression, Random Forest, and Gradient Boosting fluctuate in ranking, reinforcing the need for statistical validation. These variations justify applying non-parametric tests like the Friedman test, which evaluates classifier performance across multiple datasets to determine whether differences in accuracy are significant or dataset-dependent. This ensures a more informed and reliable selection of the best-performing model.
Friedman test for classifier performance evaluation
To assess the statistical significance of the differences in classifier performance across multiple evaluation conditions, the Friedman test was conducted. This non-parametric statistical test evaluates whether significant differences exist among multiple classifiers over different experimental conditions32. It was applied to the accuracy results obtained from four evaluation scenarios: 80:20 split (Dataset 1), 10-fold cross-validation (Dataset 1), 80:20 split (Dataset 2), and 10-fold cross-validation (Dataset 2). The test yielded a Friedman statistic of 2.045 with a p-value of 0.563 (χ2 = 2.045, p = 0.563), indicating that the observed differences in classifier performance across evaluation conditions are not statistically significant at the 0.05 level. This suggests that the choice of evaluation method (80:20 vs.
10-fold cross-validation) and dataset variation did not significantly impact the relative ranking of classifiers. Neural Networks (NN) remains the best-performing classifier in terms of raw accuracy, making it ideal for high-precision applications. Since no statistically significant difference exists, organizations can choose a model based on practical considerations such as training time, scalability, and business requirements to determine employee churn prediction.
Time complexity
The proposed methodology follows a multi-stage process integrating Generative AI (ChatGPT-4), AHP, TOPSIS, and Machine Learning classifiers, each contributing to the overall computational complexity. Feature selection using GPT-4 operates with a complexity of O(n log n) due to ranking and sorting operations. AHP for criteria weighting is the most computationally expensive step, requiring O(n3) due to the construction of the pairwise comparison matrix, eigenvector computations, and consistency checking. TOPSIS, used for employee ranking, has a complexity of O(m × n) + O(m log m), where m is the number of employees and n is the number of features, making it more efficient compared to AHP.
In the machine learning model training, Random Forest and Gradient Boosting have a training complexity of O(k × m log m), while Neural Networks, which achieve the highest predictive accuracy, require O(i × m × n2) due to iterative weight updates. The overall complexity of the methodology is O(n3) + O(m × n) + O(m log m) + O(i × m × n2), where AHP and Neural Network training’s contribute the highest computational burden. To improve efficiency, future research should focus on optimizing AHP’s weight computation and exploring faster ranking and deep learning techniques to reduce computational costs while maintaining accuracy. A summary of the overall computational complexity time is provided in Table 8.
Conclusion.
By employing a combination of real and imitative experts, the methodology effectively captures diverse perspectives to evaluate and rank alternatives. This heterogeneity is reflected not only in the diversity of the data but also in the varied expertise of the real and imitative experts involved in the process. The methodology follows a structured approach, comprising four main stages: data collection, generative AI, MCGDM, and machine learning models for predicting employee churn. By utilizing a combination of real-world datasets and generative AI models representing experts from different fields, the methodology ensures that multiple viewpoints and criteria are considered in the decision-making process. GPT-4’s capabilities were used to generate a group of imitative experts with various areas of expertise, enriching the decision-making process.
The use of AHP and TOPSIS ensures that criteria weights and rankings are determined systematically, leading to the development of reliable predictive models for employee churn. AHP is employed to calculate the criteria weights, which are then used by TOPSIS to rank alternatives and classify employees into three categories based on their likelihood to leave work. The results from TOPSIS enable machine learning models to capture these categorization patterns, reducing the computational burden in future predictions by using trained ML models instead of recalculating with TOPSIS.
Neural Networks achieved the highest predictive accuracy for employee churn across both evaluated datasets, outperforming other machine learning classifiers. However, statistical analysis via the Friedman test confirmed no significant differences in performance among the models. Consequently, organizations may prioritize practical criteria including computational efficiency, scalability, and alignment with operational requirements when selecting an optimal model for employee churn prediction tasks.
Overall, the methodology, from data collection to final classification, provides a scalable and adaptable solution for organizations seeking to reduce employee turnover through data-driven decision-making. It stands out by integrating advanced AI techniques with traditional decision-making frameworks, enabling organizations to make informed and strategic decisions based on both expert knowledge and insightful data analysis. To summarize, the proposed methodology demonstrates practical applicability and offers a viable framework for integration within corporate environments. By employing this approach, organizations can systematically track employee turnover patterns and implement strategies to mitigate it effectively.
In conclusion, employee churn prediction is utilized as a test case for solving MCGDM problems using the proposed methodology; however, this approach can be effectively extended to a wide range of other MCGDM applications.
To enhance the effectiveness and applicability of the proposed methodology, future research should focus on several key improvements. Conducting cross-industry validation using datasets from healthcare, finance, retail, and IT will improve the model’s adaptability beyond a single corporate environment; churn prediction and ensure its applicability in diverse workforce settings. To optimize the computational efficiency of AHP and TOPSIS, alternative MCDM techniques can be explored to enhance AHP’s time complexity and the multiple steps in TOPSIS. Additionally, validating AI-generated feature selection by comparing GPT-4-selected features with traditional methods such as regression will enhance the reliability and consistency of the feature selection process.
Furthermore, integrating deep learning models and hybrid AI models to enhance the model’s ability to capture long-term behavioral trends in employee churn prediction, as traditional ML models may struggle with complex temporal patterns. By implementing these improvements, future research can significantly increase the scalability, adaptability, and predictive power of churn prediction models, making them more effective for real-world HR analytics applications.
Data availability.
IBM Watson dataset published in 2017 from the Smarter Workforce Institute at IBM is available in: https://git hub.com/shailysaigal/Job-Satisfaction-on-IBM-Watson-dataset/blob/80ea0fc85152e1047163deaaac462e95b0d a46c1/IBMHRFinalcleanedData.xlsm. INX Future Inc. dataset is publicly available at Kaggle in: https:// www.kaggle.com/datasets/abhishekgupta0695/inx-future-inc-employee-performance. In addition, The generat ed datasets from all virtual experts utilized in the current study are available from the corresponding author on reasonable request. For further inquiries, please contact Hagar G. Abu-Faty at the email address.
1. Dong, Q., Sheng, Q., Martínez, L. & Zhang, Z.
An adaptive group decision making framework: individual and local world opinion
2. Boix-Cots, D., Pardo-Bosch, F. & Pujadas, P.
A systematic review on multi-criteria group decision-making methods based on weights: analysis and classification scheme. Inform. Fusion. 96, 16–36 (2023).
3. Chaube, S. et al.
An overview of multi-criteria decision analysis and the applications of AHP and TOPSIS methods. Int. J. Math.
Eng. Manage. Sci. 9, 581 (2024).
4. Taherdoost, H. & Madanchian, M.
A Comprehensive Survey and Literature Review on TOPSIS, International Journal of Service
Science, Management, Engineering, and Technology (IJSSMET), vol. 15, pp. 1–65, (2024).
5. Mardani, A., Zavadskas, E.
K., Govindan, K., Amat Senin, A. & Jusoh, A. VIKOR technique: A systematic review of the state of the art literature on methodologies and applications, Sustainability, vol. 8, p. 37, (2016).
6. Figueira, J.
R., Mousseau, V. & Roy, B. ELECTRE methods, Multiple criteria decision analysis: State of the art surveys, pp. 155–185,
(2016).
7. Brans, J.
P. & De Smet, Y. PROMETHEE methods. Multiple Criteria Decis. Analysis: State Art Surv., pp. 187–219, (2016).
8. Canco, I., Kruja, D. & Iancu, T.
AHP, a reliable method for quality decision making: A case study in business, Sustainability, vol. 13, p. 13932, (2021).
9. Çelikbilek, Y. & Tüysüz, F.
An in-depth Review of Theory of the TOPSIS Method: an Experimental Analysisvol. 7p. 281–300 (Taylor
& Francis, 2020).
10. Supraja, S. & Kousalya, P.
A comparative study by AHP and TOPSIS for the selection of all round excellence award, in International
Conference on Electrical, Electronics, and Optimization Techniques (ICEEOT), 2016. (2016).
11. Sivalingam, C. & Subramaniam, S.
K. Cobot selection using hybrid AHP-TOPSIS based multi-criteria decision making technique for fuel filter assembly process, Heliyon, vol. 10, (2024).
12. Zhang, L. & Su, W.
A Double Heterogeneous Multi-Criteria Group Decision-Making Method Based on Decision Makers, Available at SSRN 4930172.
13. Morente-Molinera, J.
A., Wu, X., Morfeq, A., Al-Hmouz, R. & Herrera-Viedma, E. A novel multi-criteria group decision-making method for heterogeneous and dynamic contexts using multi-granular fuzzy linguistic modelling and consensus measures. Inform. Fusion. 53, 240–250 (2020).
14. Ali, R., Hussain, A., Nazir, S., Khan, S. & Khan, H.
U. Intelligent decision support Systems—An analysis of machine learning and multicriteria decision-Making methods, Applied Sciences, 13, p. 12426, (2023).
15. Yadav, S., Jain, A. & Singh, D.
Early prediction of employee attrition using data mining techniques, in IEEE 8th international advance computing conference (IACC), 2018. (2018).
16. Li, Y., Kou, G., Li, G. & Peng, Y.
Consensus reaching process in large-scale group decision making based on bounded confidence
17. Badrul, M., Marlinda, L. & Dalis, S.
B. Santoso and others, The employee promotion base on specification job’s performance using:
Mcdm, ahp, and electre method, in 6th International Conference on Cyber and IT Service Management (CITSM), 2018. (2018).
18. Li, X. & Liao, H.
A large-scale group decision making method based on Spatial information aggregation and empathetic
19. Lalwani, P., Mishra, M.
K., Chadha, J. S. & Sethi, P. Customer churn prediction system: a machine learning approach, Computing,
20. Dahiya, K. & Bhatia, S.
Customer churn analysis in telecom industry, in 4th International Conference on Reliability, Infocom
Technologies and Optimization (ICRITO)(Trends and Future Directions), 2015. (2015).
21. Alamsyah, A. & Salma, N.
A comparative study of employee churn prediction model, in 2018 4th international conference on science and technology (ICST), (2018).
22. Musanga, V. & Chibaya, C.
A Predictive Model to Forecast Employee Churn for HR Analytics, in Proceedings of NEMISA Digital
Skills Conference, (2023).
23. Chowdhury, S.
J., Mahi, M. I., Saimon, S. A., Urme, A. N. & Nabil, R. H. An Integrated Approach of MCDM Methods and Machine
Learning Algorithms for Employees’ Churn Prediction, in 2023 3rd International Conference on Robotics, Electrical and Signal Processing Techniques (ICREST), (2023).
24. Alcala, G.
E. S. & Murcia, J. V. B. Machine Learning Techniques in Employee Churn Prediction, TWIST, vol. 19, pp. 382–387,
(2024).
25. Ahmad, A.
K., Jafar, A. & Aljoumaa, K. Customer churn prediction in Telecom using machine learning in big data platform. J. Big
Data. 6, 1–24 (2019).
26. Pekel Ozmen, E. & Ozcan, T.
A novel deep learning model based on convolutional neural networks for employee churn prediction.
27. Chaudhary, M., Gaur, L., Jhanjhi, N.
Z., Masud, M. & Aljahdali, S. Envisaging employee churn using MCDM and machine learning.
Intell. Autom. Soft Comput., 33, (2022).
28. Al Abid, F.
B., Bakri, A. B., Alam, M. G. R., Uddin, J. & Chowdhury, S. J. Simplified novel approach for accurate employee churn categorization using MCDM, De-Pareto principle approach, and machine learning. Baghdad Sci. J. 21, 0706–0706 (2024).
29. Jo, E. et al.
Joo and others, assessing GPT-4’s performance in delivering medical advice: comparative analysis with human experts.
JMIR Med. Educ. 10, e51282 (2024).
30. Yan, J. et al.
GPT-4 vs. Human Translators: A Comprehensive Evaluation of Translation Quality Across Languages, Domains, and 31. Svoboda, I. & Lande, D
Enhancing Multi-Criteria Decision Analysis with AI: Integrating Analytic Hierarchy Process and GPT-4 for Automated Decision Support, arXiv preprint arXiv:2402.07404, (2024).
32. Tapia, C
E. F. & Cevallos, K. L. F. Kruskal-Wallis, Friedman and mood nonparametric tests applied to business decision making. Espirales Revista Multidisciplinaria De Investigación, 6, (2022).
H.G.A., and O.A. conceived the idea for the paper.H.G.A. wrote the main manuscript text and handled the machine learning and AI programming. O.A., and A.K. handled data collection and preparation process.O.A., M.M.H. and A.K. reviewed the manuscript. All authors reviewed and approved the final manuscript.
Funding
Open access funding provided by The Science, Technology & Innovation Funding Authority (STDF) in cooperation with The Egyptian Knowledge Bank (EKB).
Declarations
Competing interests
Additional information
Reprints and permissions information is available at the linked source.
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
© The Author(s) 2025
Author contributions
The authors declare no competing interests.