1 More Paper.
Full Reading01:32:51

Exploring Large Language Model AI tools in Construction Project Risk Assessment: Chat GPT Limitations in Risk Identification, Mitigation Strategies, and User Experience

1 More Paper · Full Reading

Full Reading podcast cover
Listen to the Full Reading

About this paper

A full audio edition of this paper.

Authors: H. Martin, J. James, A. Chadee

Publication date: 2025

Read the paper: https://doi.org/10.1061/jcemd4.coeng-16658

Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/

The authors and publisher do not sponsor or endorse this recording.

Brief episode

Transcript

You’re listening to “Exploring Large Language Model AI tools in Construction Project Risk Assessment: Chat GPT Limitations in Risk Identification, Mitigation Strategies, and User Experience,” by H. Martin, J. James, and A. Chadee. Published in 2025.

Abstract.

The last 3 years have witnessed an increasing awareness and consensus on using artificial intelligence (AI) to enhance decision-making in the construction sector. This study explores the integration of ChatGPT (Generative Pre-trained Transformer) into traditional risk management frameworks within the construction industry, contributing to the ongoing discourse on AI’s role in enhancing risk identification, analysis, and mitigation. Using a mixed-method approach comparing ChatGPT-assisted to human evaluations, interviews, and a case study, the research develops a better understanding of construction risk analysis processes and discusses decision-making errors of omission, over-and underestimation of probabilities and impacts, and treatment of the residual risk after proposed mitigation strategies.

Results suggest that users’ experience of ChatGPT is primarily favorable, characterized by quick responses and an intuitive interface that enhances decision-making efficiency. Findings indicate that GPT may be especially beneficial for less experienced practitioners since it provides comprehensive risk awareness. However, experienced professionals contend that the software lacks contextual depth. The study contributes a ChatGPT-4 prompt to evaluate infrastructure risk for a given project scope. An evidenced case study on a road upgrade project in Ireland demonstrates a lessened dependence on the quality of user prompting skills and emphasizes the quality of project scope data input.

A collaborative approach, including Chat-GPT early involvement and human refinement, promises to enhance conventional risk management speed and efficiency and reduce bias and inflexibility while maintaining the adaptability and ethical rigour required in the industry’s evolving risk landscape. DOI: 10.1061/JCEMD4.COENG-16658. This work is made available under the terms of the Creative Commons Attribution 4.0 International license, the linked source.

Practical Applications: The abilities to accurately identify risk early in project development and develop effective contingencies to mitigate potential impacts are highly desired skills among project managers and leaders. Careful implementation of risk measures can minimize their immediate and long-term impact on project performance and viability. Established risk evaluation methods and protocols are tenuous and require expertise beyond an individual’s knowledge capacity, leading to unidentified risks and implementation inefficiencies. This research draws on the unique experiences of AI (ChatGPT) and practitioners to compare their evaluation of project risk. Identified gaps and limitations were analyzed with the aim of developing a GPT prompt as an immediate practical recommendation that all project personnel could use.

The findings can be applied to expedite the identification of hidden risks, analyze their effects on projects and society, and design control strategies to limit their immediate and long-term impact. The developed prompt aids in better understanding the underlying safety, cost, time, quality, and environmental risks across construction projects and their supply and logistics chains. It also reduces reliance on individual competencies and aligns more with the quality of project context information provided by risk evaluators. Moreover, while the research adds to the growing body of knowledge of AI in construction, a pathway is presented for construction organizations seeking to integrate AI into their risk management workflows.

Author keywords: Artificial intelligence (AI); ChatGPT; Risk management; Natural language processing; Project management; Construction; Prompt; Large language model.

Introduction.

Despite extensive efforts to establish comprehensive risk management frameworks, the construction sector continues to face 1Lecturer, School of Natural and Built Environment, Queens Univ. Belfast, Elmwood Bldg., Belfast BT7 1NN, UK (corresponding author). ORCID: the linked source. Email: the email address 2Dept. of Civil Engineering, Queens Univ., Elmwood Bldg., Belfast BT7 1NN, UK. Email: the email address 3Lecturer, Dept. of Civil and Environmental Engineering, Univ. of the West Indies, St Augustine, St Augustine Circular Rd., Trinidad and Tobago. Email: the email address

Note. This manuscript was submitted on November 19, 2024; approved on March 19, 2025; published online on June 24, 2025. Discussion period open until November 24, 2025; separate discussions must be submitted for individual papers. This paper is part of the Journal of Construction En-gineering and Management, © ASCE, ISSN 0733-9364.

significant challenges. A popular lens for investigating project fail-ure is the identification and analysis of individual risk factors in this sector and associated nuisances within context-specific indus-tries and organizations. There is an extant positivist theme that although great strides have been made in construction risk management processes and practices, the outcomes of most construc-tion projects are frequent project delays, cost overruns, compromised quality, and, sometimes, project failure. Deliberate or not, even with established methodologies and improved processes, recent data suggest persistent and significant risk exposure. Alhammadi et al. (2024) reported that an astonishing 93% of construction projects encounter substantial risk events, leading to budget overruns of 5% to 10% and considerable schedule delays.

These reported metrics are not favorable for any client or investor, as overruns and delays erode profit margins. Financially, ineffective risk management costs the industry an estimated $122 million for every $1 billion invested

. Beyond the direct and quantifiable financial losses, these risks also generate extensive consequences for stakeholders, end-users, and communities, creating social and environmental repercussions. This pressing context sounds a clarion call for inno-vative approaches to risk mitigation to enhance infrastructure proj-ect outcomes.

Risk management is fundamental to construction project management and aims to minimize uncertainty’s impact on project out-comes, including on-time delivery, adherence to budget, and quality standards. However, modern infrastructure projects’ increasing complexity—featuring numerous intercon-nected activities, a diverse set of stakeholders, and an evolving risk landscape—has exposed the limitations of traditional approaches. Established frameworks, such as Project Management Body of Knowledge (PMBOK) and ISO 31000, rely on qualitative and quantitative methods to identify, assess, mitigate, and monitor risks. Yet these frameworks are con-strained by static assessments, subjective risk evaluations, labor-intensive processes, and limited integration of advanced technology.

This static nature often results in outdated as-sessments, necessitating a shift toward more dynamic, data-driven, and technology-enhanced risk management approaches.

The emergence of artificial intelligence (AI) has introduced promising advancements, with large language models such as ChatGPT, Claude.ai, and Perplexity representing a transformative development in the field. These models handle complex linguistic tasks, generate human-like responses, and maintain contextual aware-ness over extended interactions. Such capabilities enable AI to contribute to data analysis, risk identification, resource optimization, and communication—functions that are invaluable in high-stakes environments like construction, where decisions influence multiple stakeholders. With sparse-sample learning, AI demon-strates flexibility, adapting to new tasks with minimal training data, which is advantageous in construction’s dynamic risk environment.

However, the limitations of AI-driven language models present considerable challenges. Models like ChatGPT occasionally gener-ate convincing but inaccurate or misleading information, raising concerns about reliability in high-stakes applications. These models operate as statistical prediction engines without true understanding, making them susceptible to embedded biases in data, which can perpetuate stereotypes or yield ethically questionable outputs. Consequently, deploying AI in risk management necessi-tates careful balancing to harness efficiency without compromising precision, equity, or ethical standards.

This need for balance is ech-oed in research and philosophical perspectives, where cognitive biases (e.g., prospect theory) and ethical frameworks (e.g., Beck’s risk society theory) highlight the complexities of risk perception and the imperative of equitable decision-making in stakeholder-rich environments.

Despite AI’s potential to automate routine tasks, synthesize data, proactively identify risks, and facilitate collaboration, empirical validation of its effectiveness in construction remains sparse. While AI has demon-strated success in healthcare, finance, and agriculture, the construction sector has yet to document its substantial impact. Previous stud-ies primarily emphasized ChatGPT efficiency gains and construc-tion hazard recognition, supported safety education, training capabilities, and scheduling risk, and compared human and AI risk management in image data analysis and cost reduction constraints in project planning without thoroughly examining how AI could directly address complex risk management in infrastructure projects.

Consequently, Nyqvist et al. (2024) asked GPT three ques-tions in their qualitative design: what risks exist, how to prioritize them, and what can be done to control the identified risk. Residual risks and their treatment pose a significant threat to infrastructure projects. Prior studies omitted such considerations, limiting useful insights about potential risks that may have an undesirable impact on achieving appropriate and effective risk management. Nyqvist et al. (2024) discussed the feasibility of implementing ChatGPT without guiding practitioners on how to do this. Furthermore, questions remain about how to integrate AI with existing risk management frameworks to ensure better contextual understanding of construc-tion, as espoused by some researchers.

These gaps em-phasize the need for additional research to understand AI’s scope, benefits, and constraints within the construction domain and in-sights to guide the responsible implementation of AI-driven risk management systems.

This study addressed these critical gaps using an experimental design by rigorously evaluating ChatGPT performance compared to experts’ evaluation of infrastructure project risks. Specifically, it assessed the accuracy, relevance, and applicability of AI-generated recommendations compared to traditional approaches, as well as their usability and potential for adoption by construction professio-nals. The study advances our understanding of risk management in behavioral science, management science, and psychology and offers empirical evidence of AI’s potential to enhance risk management by providing usable guidelines for improving risk identification, analy-sis, and mitigation. These findings provide industry stakeholders with insights into integrating AI tools into established processes, supporting the development of a more resilient construction indus-try.

Ultimately, this work aims to inform construction industry prac-tices, shaping a future where AI-driven risk management is central to facilitating efficiency and innovation.

Risk Management Theory—Exploration of AI Knowledge Gap

The psychology of risk, which is fundamental to decision-making processes, reveals how cognitive biases—such as loss aversion and affect heuristics—often skew individuals’ perceptions, leading to deviations from rationality, as demonstrated by foundational theo-ries like Kahneman and Tversky’s (2013) prospect theory and Slovic’s (2016) risk perception theory. Such biases, while univer-sal, carry particularly high stakes in fields like construction, where risk miscalculations can lead to costly overruns and safety issues. Here, emerging AI models like ChatGPT present opportunities to counteract these biases by providing objective, data-driven assessments grounded in historical probabilities, thereby reducing human error and offering a cognitive scaffold.

However, this promising solution also introduces new challenges; cognitive biases extend to AI interactions, where users’ trust or skepticism toward AI-generated insights influences their interpretation, an underex-plored psychological variable that could compromise the tool’s ef-fectiveness. The philosophical dimensions of risk further complicate this landscape, framing it not only as a technical but also as a social and ethical challenge. Risk theories from Beck (2002) and Jonas (1984) argue that risk decisions have moral and societal implica-tions, demanding accountability, a perspective that holds particular weight in construction, where choices directly affect worker safety, environmental impact, and stakeholder interests.

While ChatGPT might streamline risk evaluation, its potential lack of contextual sensitivity raises ethical concerns, particularly around “algorithmic opacity” and accountability in AI-driven decisions, suggesting that a hybrid approach—AI-guided but human-validated—would better address ethical considerations and maintain transparency. Economic theories provide additional context, with von Neumann and Morgensterns’s (2007) expected utility theory positing rational, cost-optimized decision-making, whereas behavioral economics, championed by Thaler (1980) and Sunstein et al. (2022), suggests that real-world decisions often di-verge from pure rationality. ChatGPT’s probabilistic modeling sup-ports cost efficiency and real-time adaptability, a vital function in construction, where fluctuating factors like material costs can swiftly alter risk profiles.

Yet, its reliance on rationality may oversimplify the subjective behaviors that im-pact economic decisions, suggesting that integrating AI with hu-man oversight remains crucial for capturing the full scope of market variations. Management theories, including enterprise risk management (ERM), emphasize structured risk approaches align-ing with broader strategic goals, yet ChatGPT’s adaptive capabil-ities align better with Holling’s (1973) resilience theory, which values flexibility in complex, evolving environments. Incorporating ChatGPT into an ERM framework allows construction projects to anticipate disruptions, although the challenge of translating AI in-sights into organizational actions underlines the need for alignment between AI outputs and established ERM processes.

Ultimately, the interdisciplinary study of risk provides the challenge of reconciling quantitative and qualitative perspectives—a gap that ChatGPT’s data-processing power could help bridge by integrating insights from psychology, philosophy, and economics, aligning with Luhmann’s (1995) and Gell-Mann et al.’s (2010) systems and complexity theo-ries that view risk as an emergent, interconnected property. Thus, a hybrid model that combines AI’s analytical rigor with human con-textual judgment represents a powerful paradigm for high-stakes in-dustries like construction, where adaptive, ethically grounded, and psychologically informed risk management is essential for navigat-ing the sector’s multifaceted challenges. Table 1 outlines the critical theoretical and practical knowledge gaps for incorporating AI into risk assessments.

State of the Art in AI-Driven Risk Management

The integration of AI in construction project management offers transformative opportunities, such as improved accuracy in project planning, enhanced safety through real-time hazard detection, and streamlined processes that boost efficiency. However, these benefits are accompanied by significant challenges, including high initial costs, fragmented data, and a lack of standardized protocols. Small firms struggle with the financial investment required, while the need for specialized skills widens the indus-try’s skill gap. Moreover, data security concerns and ethical implications pose additional barriers. A stra-tegic and collaborative approach among stakeholders is essential to maximize AI’s impact while mitigating risks. The field of AI-driven risk management is rapidly evolving, with various experimental and methodological studies highlighting both its potential and limitations.

See Table 2 for opportunities and chal-lenges of AI in the construction sector.

While AI applications in risk management have advanced in pre-dictive capabilities, speed, and automation, critical methodological gaps remain unresolved, especially concerning these systems’ accu-racy, contextual relevance, and ethical deployment. Current research, as shown in Table 3, reveals a broad spectrum of approaches, from experimental designs to systematic reviews, that explore the potential of AI tools like ChatGPT in enhancing risk management in construc-tion and project management disciplines. However, these studies also highlight the pressing need for more comprehensive and scalable evaluation frameworks.

One such study by Prieto et al. (2023) evaluated ChatGPT’s ef-fectiveness in project scheduling through a controlled experiment involving six participants. While the study highlighted the tool’s speed and adaptability, it revealed significant inaccuracies due to a lack of domain-specific training, pointing to a crucial gap in con-textual knowledge. These findings emphasize the need for AI mod-els to be tailored more specifically to industry contexts, as generic models may fail to provide reliable and accurate risk assessments. Prieto et al. (2023) myopically emphasized the need for domain-specific AI training to overcome these shortcomings, overlooking the potential of developing a prompt to accelerate the lack of con-textual depth.

Similarly, Yazdi et al. (2024) explored the use of convolutional neural networks (CNNs) in AI-driven risk management, focusing on image processing for hazard identification. Their methodology demonstrated AI’s strength in rapid data processing and pattern rec-ognition across diverse sectors, yielding detailed insights into risk assessment. However, this study also identified significant draw-backs, such as the high costs of AI implementation and concerns over data protection. Furthermore, Yazdi et al. (2024) raised issues around AI’s limited understanding of contextual and nuanced data, which is essential for making informed and unbiased decisions in complex environments like construction. While their study con-firmed the utility of AI in predictive risk management, it underscored the ethical and operational challenges that must be addressed to fully realize AI’s potential.

Uddin et al. (2023) assessed the impact of ChatGPT on hazard recognition in an educational context, using a pre- and postinter-vention design with undergraduate students. Their findings demon-strated a significant improvement in students’ ability to identify hazards after using ChatGPT, suggesting that AI could serve as an effective educational tool. However, the study was limited by its con-trolled environment and small sample size, raising concerns about external validity and generalizability. This highlights the importance of conducting real-world studies incorporating larger and more di-verse participant pools to validate AI’s effectiveness in broader contexts. The educational value of AI in safety risk assessment is evident, but real-world validation remains a critical challenge.

Another key study by Aladağ (2023) evaluated ChatGPT’s pro-ficiency in quantitative risk management through a multiphase methodology that included expert focus groups and key performance indicators (KPIs). This approach thoroughly assessed AI’s capabilities but was constrained by potential biases inherent in ex-pert judgments, limiting its generalizability. The study pointed to the necessity of more objective and scalable metrics to evaluate AI’s effectiveness across different contexts. While valuable for depth, expert-driven evaluations introduce subjectivity and context-specific limitations that could hinder broader applicability. Thus, developing standardized evaluation frameworks is crucial for advancing AI risk management methodologies.

The most recent work by Nyqvist et al. (2024) qualitatively ex-plored ChatGPT-4’s performance in risk assessment, comparing it to human evaluations on identifying, prioritizing, and mitigating risks. Despite claiming a mixed-methods approach, their study is purely qualitative and lacks grounding in risk management theory or alignment with practical evaluation methods. Focused on generic questions, it underestimates ChatGPT-4’s capabilities and does not provide useful guidance for practitioners. This raises several critical questions that must be addressed. How can AI-based risk assess-ment tools like ChatGPT-4 be better integrated into practical con-texts, such as construction management, where risk evaluation involves quantification through probabilities, impacts, and mitiga-tion strategies? What methods can be employed to evaluate residual risks and their treatment systematically?

How can the comprehen-siveness of AI in risk management be examined from both human and AI perspectives to identify synergies and limitations? More-over, what guidance or prompts can be developed to support practi-tioners in effectively implementing AI tools for risk management? Addressing these gaps could provide a more comprehensive under-standing of AI’s potential as a practical and complementary risk management tool.

In addition to the aforementioned experimental designs, Taboada et al. (2023) conducted a systematic literature review (SLR) on AI in project management, using PMBOK7 principles as an evaluation framework. This review provided a comprehensive overview of AI applications in the field, highlighting the time-consuming nature of reviews and the dependency on the quality of existing literature. The review highlighted a methodological gap in synthesizing literature more dynamically and efficiently, suggesting that AI-driven risk management research could benefit from automated literature syn-thesis techniques. Furthermore, the researchers called for a more dynamic and adaptive literature review process to keep pace with the rapidly evolving nature of AI technologies.

Despite the various approaches, a consistent theme across these studies is the need for improved contextual awareness and bias mit-igation in AI applications. AI-driven risk management systems often struggle to incorporate the nuanced and situational aspects of risk, leading to inaccuracies and potential biases, especially when ap-plied across different sectors or regions. Ethical concerns, such as data protection and the transparency of AI decision-making proc-esses, remain pervasive challenges. Addressing these issues will re-quire the development of more holistic and robust methodologies that can account for both the technical and ethical dimensions of AI-driven risk management.

Collectively, the current body of research indicates that while AI has the potential to revolutionize risk management in construction and other sectors, significant methodological limitations need to be addressed. These include the need for domain-specific training of AI systems, scalable and objective evaluation metrics, and ethical frameworks to govern AI’s deployment in high-stakes environ-ments. Bridging these gaps will be essential for advancing AI-driven risk management tools’ practical utility and ethical deployment. Developing more comprehensive and interdisciplinary methodol-ogies will ensure that AI systems are effective, equitable, and con-textually relevant in real-world applications.

Method

This study employed a mixed-methods experimental design to evaluate the effectiveness of AI-driven language models, such as ChatGPT, in mitigating risks in infrastructure construction projects (Fig. 1). Following Creswell and Clark’s (2017) guidelines, the methodology integrates quantitative and qualitative data to compre-hensively analyze the problem. By combining structured performance metrics with rich qualitative data from participant interviews, this approach addresses measurable outcomes and contextual elements of user experience. The study was divided into two phases: literature review and data collection and analysis. Triangulation of data from these phases ensured validity, reliability, and a well-rounded evaluation of the AI tools being studied.

Research Design

A concurrent triangulation mixed-methods design was adopted for this study, aligned with a pragmatic philosophical framework described by Creswell and Clark (2017). This approach integrates objective and subjective measurements, which are essential for understanding the complex dynamics between AI technology and construction industry processes. The research was grounded in an abductive reasoning approach, which allows the iterative refine-ment of theories based on observed data. The design combines experimental elements and case study methodol-ogies, enabling a cross-sectional analysis of problem-solving tech-niques in risk management. Groups were formed to test AI-assisted (Group 1) versus traditional (Group 2) approaches, with industry experts (Group 3) evaluating the solutions.

This methodology al-lowed for both exploratory and confirmatory insights into AI’s ap-plication in construction.

Literature Review.

A comprehensive literature review provided the theoretical basis for formulating study objectives and hypotheses. It also established cri-teria for assessing ChatGPT’s effectiveness compared to traditional techniques in construction risk management. The review focused on recent peer-reviewed studies published in the last 5–10 years, utilizing databases such as Google Scholar and Scopus to explore key topics, including “ChatGPT in construction,” “AI in risk management,” “artificial intelligence construction problem solving,” and “language models in project management.” This review in-formed the experimental phase by summarizing AI capabilities and outlining current industry-standard practices.

Participant Selection and Sampling

The study involved three distinct participant groups, ensuring diverse insights. Group 1 consisted of construction and project management personnel with field experience of 1–2 years using ChatGPT, repre-senting future industry professionals likely to encounter AI tools in practice. Group 2 comprised industry professionals with 6–19 years of experience using traditional risk evaluation methods. Group 3 in-cluded senior industry experts with over 15–20 years of experience who were responsible for evaluating the solutions proposed by the other groups. Groups 1 and 2 were picked based on their educational backgrounds and industrial experience, whereas Group 3 evaluators were chosen for their field reputation and academic achievements (Table 4). Purposive sampling was employed to select participants based on their expertise and relevance to the research question.

Semi-structured interviews were conducted with a sample size of 11 par-ticipants to gather qualitative insights, while structured evaluations provided quantitative data.

Quantitative Method

Structured Evaluations of Problem-Solving Outcomes

The primary quantitative method involved a structured problem-solving exercise focused on risk management in construction proj-ects. Both Groups 1 and 2 completed identical tasks, including identifying risks, assessing their likelihood and impact, and propos-ing mitigation strategies. Group 1 had access to ChatGPT, while Group 2 used traditional methods. The tasks were evaluated based on accuracy, comprehensiveness, and problem-solving efficiency, with outcomes compared across groups. A risk matrix (Fig. 2) standard-ized the assessment process, allowing for systematic comparisons.

Performance Metrics

Performance metrics were used to compare the two groups’ time efficiency, risk identification accuracy, suitability of mitigation strat-egies, and overall problem-solving effectiveness. These metrics pro-vided a quantitative basis for evaluating the advantages of AI-assisted methods over traditional ones. Work completion times were recorded to assess efficiency, while risk identification accuracy and the ef-fectiveness of mitigation strategies were scored against predefined criteria. The amount of correctly detected hazards determined risk identification accuracy, as effective risk identification greatly minimizes project failures. The appropriateness of mitigation measures was applied to assess the suitability and effectiveness of proposed strategies, which are critical for minimizing risk impact.

Quantitative Questionnaire Development

A structured questionnaire was developed to evaluate the efficacy and usability of AI-assisted versus traditional problem-solving methods in construction risk management. The questionnaire in-corporated a 5-point Likert scale and open-ended questions to comprehensively understand participants’ perceptions and experi-ences. The survey was specifically designed to assess various di-mensions of the problem-solving process, including ease of use, tool satisfaction, and practical aspects such as interface intuitive-ness. The questionnaire comprised 12 questions addressing the pri-mary research objectives (Table 5).

The same Google Form was administered to all participants immediately after completing their respective tasks to ensure uni-form data collection. This approach facilitated consistent data col-lection and simplified comparative analysis across different groups. The Google Form enabled systematic data aggregation, enhancing the reliability and consistency of responses. A 5-point Likert scale provided a balanced range of options, enabling detailed responses without induced respondent fatigue associated with more extensive scales, and at the same time facilitated effective statistical analysis of ordinal data and the identification of trends across groups.

There-fore, the scale was integral to the questionnaire design, allowing respondents to express their choice with statements ranging from “strongly disagree” to “strongly agree.” Despite the study’s explor-atory nature and the sample size of 11 participants, the insights gained were deemed valuable for understanding preliminary trends and the effectiveness of AI tools like ChatGPT in construction risk management. A small sample size is typical in initial exploratory research phases, providing foundational insights that can guide subsequent larger-scale studies.

Qualitative Method

Qualitative Questionnaire Development

Semistructured interview questions were formulated to explore par-ticipants’ qualitative experiences. These questions focused on the usability, effectiveness, and overall user experience of ChatGPT for Group 1 and the effectiveness of traditional methods for Group 2. Group 3’s questions evaluated the comparative efficacy of both ap-proaches. See Appendix for the qualitative questions.

Semistructured Interviews

Semistructured interviews were conducted with Group 1 and Group 2 participants after a task, providing detailed insights into their problem-solving experiences. These interviews were held via video conferencing and transcribed for analysis. A thematic analysis was conducted using NVivo, identifying key themes related to usability, task challenges, and perceptions of AI versus traditional methods. The qualitative data supplemented the quantitative findings, providing a deeper understanding of user ex-periences with AI tools in risk management.

Data Analysis Techniques

Quantitative Data Analysis

Quantitative data analysis involves the application of statistical tools to evaluate and interpret data derived from performance metrics, systematic evaluations, and survey responses. This study employed objective, numerical measures to assess the efficacy of AI-assisted and traditional problem-solving methods in construction risk management. Data collected via a Google Form were analyzed using SPSS (Statistical Package for the Social Sciences), a widely used software for reliable statistical analysis. SPSS was selected due to its robust capabilities in handling small data sets and its user-friendly interface, which is critical for producing accurate and interpretable results.

Descriptive statistics were computed to summarize key charac-teristics of the data, including the mean, median, standard deviation, and frequency distributions of participants’ responses. These mea-sures provided a comprehensive overview of trends in performance metrics, such as the time taken for task completion, accuracy of risk identification, appropriateness of risk mitigation strategies, and overall problem-solving efficiency. Descriptive statistics were cru-cial in establishing a baseline comparison between the AI-assisted and traditional approaches.

For inferential statistics, a one-sample t-test was conducted to evaluate whether the performance of AI-assisted methods signifi-cantly differed from that of traditional methods. This test compared the mean performance scores from participants against a benchmark value (m0 1⁄4 3), indicating general agreement or disagreement with the efficacy of the method. The analysis was conducted at a 95% confidence interval (p < 0.05), ensuring that statistically significant differences were identified, with confidence intervals calculated to quantify the potential variability in popula-tion means.

Additionally, an independent sample t-test was used to assess differences between two groups—those using AI-assisted techniques and those relying on standard risk management strategies. This test compared mean scores related to risk identification and mitigation between the groups, providing insight into whether AI assistance of-fered measurable improvements.

Qualitative Data Analysis

The qualitative data analysis explored participants’ experiences, in-sights, and decision-making processes through semistructured in-terviews. Thematic analysis was used to systematically identify, analyze, and report patterns within the qualitative data. The coding process involved several stages:

1. Initial coding: Data were divided into discrete segments, with.

codes assigned to each based on content. This process facilitated the organization of data and helped identify recurring patterns.

2. Axial coding: Related codes were grouped into larger, more ab-.

stract categories, allowing for the identification of overarching themes and the establishment of relationships between them.

3. Selective coding: The final stage involved refining the catego-.

ries and selecting core themes that captured the essence of the participants’ experiences and perspectives.

NVivo software supported qualitative data analysis, offering ad-vanced tools for coding, categorizing, and visualizing data. NVivo’s capabilities enhanced the thoroughness and precision of the analysis by ensuring that the qualitative data were systematically managed and explored.

Validity and Reliability Measures

Rigorous validity and reliability testing of both quantitative and qualitative data ensured the robustness of the study. For quantitative data, Cronbach’s alpha was used to evaluate the internal consistency of survey instruments, with a value above 0.70 deemed acceptable for reliability. Ensuring high internal consistency was critical, given the subjective nature of certain sur-vey items.

For qualitative data, credibility was ensured through member checking, allowing participants to verify the accuracy of findings, and triangulation, which involved cross-verifying multiple data sources and perspectives. Additionally, transferability was addressed by providing detailed descriptions of the research context, enabling other researchers to assess the appli-cability of the findings in similar settings. Depend-ability was maintained through the documentation of an audit trail that recorded key decisions and changes made throughout the re-search process. Confirmability was ensured via reflexive journaling and peer debriefing, thereby mitigating po-tential researcher bias and ensuring that the findings were grounded in the participants’ experiences.

Ethical Considerations

Ethical approval was obtained from the university’s Institutional Review Board, which ensured the study complied with relevant eth-ical standards to safeguard participant rights and data confidentiality. Before participation, individuals were provided with a detailed informed consent form outlining the study’s objectives, methods, potential risks, and benefits. Participants were required to sign the form, con-firming their voluntary involvement and understanding of their rights, including the option to withdraw at any point. Measures were taken to minimize participant risks, maintain transparency, and up-hold data privacy and anonymity throughout the research process.

Results and Analysis

Quantitative Evaluation

The overall Cronbach’s alpha of 0.806 for the 12-item scale sug-gests strong reliability, surpassing the usually recognized threshold of 0.7. This threshold was also exceeded when the reliability scale was determined when each item was removed from the instrument, confirming a meaningful contribution to measuring the construct of interest. The results, presented in Table 6, generally indicate a positive reception of the AI-assisted approach, with mean scores consistently exceeding the midpoint of 3 on a 5-point scale. The evaluation of the AI-assisted risk management approach revealed high ratings for promoting AI in risk management (mean 1⁄4 4.27), integration into workflows (mean 1⁄4 4.09), overall experience (mean 1⁄4 4.09), and process time efficiency (mean 1⁄4 4.00), indi-cating that participants found the method efficient, well integrated, and recommendable.

It was also rated positively for supporting mit-igation strategies and improving project success (mean 1⁄4 3.91). However, features like capturing complexities and comprehensive risk identification received slightly lower scores (mean 1⁄4 3.55). The response variability based on standard deviation from the mean was most pronounced for time efficiency (SD 1⁄4 1.414) and risk plan clarity (SD 1⁄4 0.934), suggesting a lack of shared understand-ing among respondents—probably due to variability in how these issues affect different demographic groups. Overall, the AI-assisted approach was perceived as effective and innovative, though indi-vidual experiences varied, highlighting opportunities for further improvement.

Results from the one-sample t-test showed that 11 of the 12 fac-tors ranked according to their means were determined to be signifi-cant. Overall, these findings indicate that participants viewed the AI-assisted risk management approach as significantly beneficial in the majority of cases, particularly in terms of overall experience, impact on project success, and likelihood of recommendation, while also highlighting specific areas where the approach may need to be refined or where its benefits are less pronounced.

An independent sample t-test was conducted to investigate the differences between individuals using ChatGPT and those using traditional methods. The means were then compared across groups using the independent sample test. There were noticeable differences between those who used ChatGPT and those who used the tradi-tional method, particularly regarding threats (construction, tech-nology integration, and reliability). Levene’s test for equality of variances yielded a significant result (F 1⁄4 9.000, p 1⁄4 0.024), sug-gesting uneven variances between groups. As a result, we refer to the “equal variances not assumed.” The t-test showed a statistically significant difference between the two groups. The average differ-ence between the groups was 0.75000 (SE 1⁄4 0.25000), with a 95% confidence interval of −0.04561 to 1.54561.

These findings indicate that threats were more of a concern for AI-assisted tech-niques using ChatGPT than the traditional method.

Comparative Analysis of Mind Map Results: AI-Assisted Risk Management (Group 1) versus Traditional Risk Management (Group 2)

Risk management in the construction sector is increasingly marked by the integration of AI tools, notably ChatGPT, alongside tradi-tional methodologies. A comparative analysis of two groups—one using AI for risk evaluation (Group 1) and the other relying on con-ventional approaches (Group 2)—revealed critical insights into their respective strengths and limitations. This analysis in Figs. 3 and 4 highlights how these divergent pathways reflect broader trends in the construction industry and may inform future practices.

Both groups acknowledged the necessity of effective risk evalu-ation; however, their approaches differed significantly. Group 1 emphasized the user experience of ChatGPT-4, with participants reporting its intuitive interface as a key factor facilitating engage-ment, particularly for those unacquainted with traditional tools. The chat-based format was appreciated for lowering barriers to entry and enhancing user interaction. For instance, participant P-02 noted that the AI’s responsive design allowed smoother access to vital insights. In contrast, Group 2 emphasized the effectiveness of tradi-tional risk management methods, where participants such as P-03 and P-08 valued the systematic and structured nature of established practices. These methods are perceived as comprehensive in iden-tifying known risks and are deeply embedded in industry norms.

However, a significant limitation noted by Group 2 participants is their reliance on historical data and expertise, which can stifle adaptability to new and unforeseen challenges within dynamic project environments.

Transitioning from user experience and effectiveness, both groups identified inherent challenges within their respective frame-works. Group 1 expressed concerns about the generic nature of ChatGPT’s responses, as illustrated by participant P-07, who re-ported the frequent need to rephrase inquiries for more tailored in-sights. This suggests that while AI can serve as a valuable initial resource, its current capabilities may lack the detailed contextual understanding required for complex construction risk assessments. Additionally, P-05 emphasized the necessity of external verifica-tion of AI outputs, reinforcing the pressing challenge of ensuring accuracy in fields where precision is paramount. On the other hand, Group 2 highlighted the resource-intensive nature of traditional methods, where participants P-03 and P-08 pointed to manual documentation as a significant hurdle.

The cumbersome processes inherent in these approaches can hinder responsiveness, with par-ticipant P-06 warning of “tunnel vision,” where focusing on famil-iar risks may obscure critical threats that are less obvious.

The analysis of decision-making tools further illustrates the con-trast between the two groups. Participants from Group 2 identified traditional instruments such as risk registers and strength opportu-nities threats analysis as foundational elements for prioritizing risks and formulating mitigation strategies. However, participants P-06 and P-08 noted that these static tools often required real-time insights to effectively manage complex and evolving risks, highlighting a critical gap in modern construction risk management. In contrast, Group 1 recognized AI’s potential to enhance decision-making through predictive analytics and real-time data integration. Participants P-06 and P-08 articulated a shared vision of AI providing instanta-neous, data-driven insights, which could significantly improve the responsiveness of risk management processes.

The necessity of continuous learning emerged as a shared theme among both groups, though the implications diverged. Group 2 em-phasized the importance of ongoing professional development through workshops, training sessions, and collaborations with ex-ternal experts, as highlighted by participant P-03. This commitment underscores the need to effectively integrate new insights into risk management practices to address emerging challenges. Conversely, Group 1 participants advocated for enhancements to ChatGPT, em-phasizing the importance of industry-specific training and real-time data access. Participants P-01 and P-04 noted that refining AI’s capabilities to respond to complex scenarios would be crucial for improving its effectiveness in construction risk management.

Lastly, the discourse surrounding ethical and privacy concerns in Group 1 clarified apprehensions regarding data protection and the potential for job displacement due to automation. Participant P-04 highlighted the need for robust data privacy practices, especially in sensitive information sectors. This contrasts with Group 2’s discus-sions on potential enhancement, wherein participants recognized

AI’s capacity to enhance traditional methods without necessarily replacing them. While acknowledging that AI could enhance ef-ficiency and accuracy, participants in Group 2 pointed out that tra-ditional methods remained vital, particularly in contexts where historical insights guide decision-making.

Evaluation of Group 1 and Group 2 Risk Assessment Outputs

An evaluation of AI-supported (Group 1) and traditional (Group 2) risk assessment methodologies conducted by a panel of expert eval-uators (Group 3) revealed a critical trade-off between the efficiency of AI-assisted methods and the depth and contextual relevance achieved through traditional approaches. Group 1, which utilized ChatGPT, demonstrated a marked advantage in time efficiency, with an average task completion time of 36 min compared to 157 min for Group 2. This efficiency was accompanied by a dynamic, iterative engagement with AI, as participants used an average of four prompts to refine their inputs. However, while AI-generated matrices were produced quickly, they tended to lack the specific detail character-istic of expert-driven evaluations (Fig. 5) (participant P-05, Group 1) and Fig. 6 (participant P-06, Group 2).

Group 1’s matrices effec-tively identified broad risk categories such as natural disasters, cybersecurity, and regulatory compliance, reflecting the AI’s sys-tematic capacity to integrate diverse data inputs. Yet these outputs were often general, lacking the tailored specificity required for project-specific risk mitigation. In contrast, Group 2’s matrices, while significantly more time-intensive, benefited from extended periods of introspection and professional reflection. One participant, for instance, took an additional 1–2 days to deliberate on past ex-periences, resulting in matrices that were deeply contextual and aligned with the realities of specific project conditions.

This finding suggests that the extended timeframe in traditional methods facilitates the incorporation of professional tacit knowl-edge, producing a more comprehensive understanding of localized risks. However, this human reflection process is susceptible to re-cency bias, wherein recent experiences disproportionately influence risk perception, potentially skewing the risk profile toward more im-mediate or memorable issues rather than a balanced representation of all potential hazards. While lend-ing greater situational awareness, this bias may inadvertently cause human respondents to underweight less salient but critical risks, leading to contextually rich but selectively narrow assessments.

Moreover, the comparative analysis between the groups revealed distinct risk prioritization and categorization patterns (Table 7). Group 1 captured detailed insights into underrepresented areas, such as site-specific environmental risks (e.g., endangered species and soil contamination) and long-term operational risks (e.g., technologi-cal obsolescence and cybersecurity), which were either omitted or inadequately addressed in Group 2’s responses. Conversely, Group 2 emphasized social and political risks, including community resis-tance and public expectations, while neglecting complex financial and technological interdependencies like market fluctuations and supplier risks.

This divergence recognizes a fundamental difference in the nature of AI-driven versus human evaluations: While AI system-atically identifies a wide array of complex, data-driven risks, hu-man experts are more attuned to sociopolitical dynamics and context-specific concerns, resulting in a more holistic yet localized risk profile. These findings indicate that Group 1’s reliance on data-driven algorithms produces a comprehensive yet general as-sessment, whereas Group 2’s evaluations are shaped by experien-tial knowledge, offering depth but risking oversight of broader systemic issues.

Ultimately, Group 3’s evaluation highlights the necessity of in-tegrating AI’s efficiency and breadth with human expertise’s depth and contextual insight. Though advantageous for rapid risk iden-tification, AI-supported methods require iterative refinement by users to address context-specific gaps. On the other hand, while time-intensive, traditional methods with human-only input excel in generating highly targeted and reflective risk assessments. Integrat-ing these approaches can harness AI’s ability to capture complex, multidimensional risks and mitigate human biases—such as recency bias—in risk perception, thereby improving overall project resilience and decision-making accuracy.

Therefore, this study advocates for a synergistic risk management model that uses AI’s computational power to complement, rather than replace, human expertise, thereby achieving a balanced, efficient, and contextually robust approach.

The comparative evaluation between Group 1 (computer-aided replies) and Group 2 (human responses) was structured around five primary risk categories: construction delays and budget overruns, environmental regulatory compliance, social and community is-sues, technology risks, and health, safety, and operational risks. This grouping adheres to risk classification theories that categorize risks according to their characteristics and effects on project per-formance and results. Initial risk assess-ments for construction delays and budget overruns were similar among both groups, indicating a common foundational comprehen-sion of possible interruptions to project schedules and financial oversight.

However, Group 1 (AI-based) consistently exhibited a higher residual risk, suggesting a more conservative or cautious as-sessment of the effectiveness of mitigation measures following the intervention. This corresponds to the “optimism bias” concept, which suggests that human responders (Group 2) often underesti-mate prospective obstacles owing to cognitive biases, resulting in a systemic distortion in risk perception.

Group 1 perceived higher initial risks in environmental regula-tory compliance, suggesting that AI-based models may more rig-orously account for external regulatory variables and compliance costs. Nonetheless, both groups had comparable residual risk eval-uations, indicating that both AI and human participants aligned in recognizing effective mitigation techniques. This result can be ex-plained by Wild’s (2016) risk convergence theory, which posits that the incorporation of more structured and data-driven insights (Group 1) and heuristic human judgments (Group 2) results in com-parable mitigation outcomes, although the initial risk perceptions are different.

Regarding social and community risks, Group 1 rated both ini-tial and residual risks higher, which could point to the limitations of AI in fully grasping the nuances of social dynamics and stake-holder engagement. This inconsistency prompts inquiries about AI’s capacity to simulate “complex adaptive systems,” in which human elements, social behaviors, and emergent features are essential risk drivers. The heightened sensitivity in AI-driven risk assessment may stem from its dependence on data inputs that emphasize worst-case situations, whereas human responses could include qualitative assessments or contextual information that ad-justs perceived risks.

In contrast, Group 1 assessed technical dangers as lower initially and for the residual risk. This may be due to AI’s systematic in-clusion of redundancies and adaptive processes in technological risk evaluations, indicating a more comprehensive review frame-work that prioritizes control and predictability. The elevated tech-nical dangers recognized by Group 2 may arise from a “perceived complexity effect,” whereby human respondents may exaggerate uncertainty linked to new technology owing to insufficient knowl-edge or previous adverse experiences.

Finally, health, safety, and operational hazards exhibited compa-rable initial and residual ratings across both groups, demonstrating a mutual comprehension of risk severity in this domain. However, the marginally lower residual risk values in Group 1 suggest that AI might better capture the probabilistic reduction of incidents following mitigation, consistent with systematic risk evaluation.

In general, the comparative analysis demonstrated subtle discrep-ancies in the risk perceptions and mitigation outlooks of Group 1 (AI-assisted) and Group 2 (human) respondents. Group 1’s con-servative approach is emphasized by its propensity to retain higher residual risk estimates for complex, socio-environmental categories. It also draws attention to the possibility that AI models might sys-tematically dismiss the efficacy of qualitative mitigation methods used by human responders. Group 2’s decreased residual risk ratings in several categories may suggest “optimism bias” and “heuristic-driven decision-making,” where subjective assessments influence risk outcomes, diverging from objective risk models.

Consequently, the amalgamation of AI-driven in-sights with human experience may provide a balanced methodology, using AI’s systematic precision while considering contextual human comprehension.

Discussion.

Effectiveness of ChatGPT in Risk Identification and Analysis (O1)

This study of ChatGPT’s role in construction risk management highlights both its substantial potential and inherent limitations, shedding light on the complex relations between AI-driven analysis and the need for nuanced human expertise. ChatGPT’s systematic approach to categorizing risks demonstrates its strength in identi-fying general hazards and processing extensive data sets efficiently, aligning with theories that emphasize AI’s ability to mitigate hu-man biases and process information exhaustively. The capability for generalization served more as a checklist, to ensure the common risks are captured. This capability, however, contrasts with its struggle to recognize unique, context-specific risks, which are the risks often overlooked and of higher severity or criticality rankings on projects.

Without tailored inputs, AI’s data-driven models often lack the adaptability to ac-count for dynamic project-specific challenges in complex construc-tion environments. In practice, this tension between breadth and context leads to mixed perspectives on AI’s reliability, which leads to two observations. First, while structured methods yield more consistent outputs (P-09), other respondents noted inefficiencies from redundant risk lists, reflecting AI’s limi-tations in prioritizing interrelated risks—a challenge further empha-sized by Rampini and Cecconi (2022). Second, a recognition of risk relationships and trade-offs among dominant performance metrics should be established.

From the analysis, though ChatGPT allows rapid data processing and enhancing time efficiency, the model’s effectiveness depends heavily on user engage-ment and understanding, suggesting a learning curve for optimally leveraging AI. This observation resonates with Zou et al. (2007), who asserted that AI could not replace human expertise but rather would serve as a complement to it. Interestingly, despite widespread claims that AI struggles with project-specific factors, ChatGPT occasionally identified context-dependent risks, like environmental hazards, when given the right data. This finding challenges assumptions about AI’s rigidity in contextual risk analysis. On the other hand, this capability also stresses AI’s reli-ance on human inputs to generate meaningful, tailored insights, re-affirming the critical need for expert oversight to adapt AI outputs within fluid project conditions.

Thus, while ChatGPT enhances ef-ficiency and breadth in risk identification, its utility in construction risk management is maximized only when integrated with informed human judgment, positioning AI as an essential, albeit complemen-tary, tool in navigating the complexities of large-scale projects.

Effectiveness of ChatGPT’s Risk Mitigation Strategies (O2)

There are times when the effectiveness of AI-driven risk mitigation strategies are questioned. The findings from this study reveal the complex relationship between AI-driven risk mitigation strategies, such as those provided by ChatGPT, and the contextual expertise required in construction risk management. The argument is that AI offers a structured framework for identifying and addressing risks but often lacks the context-specific applicability that human judgment provides. Based on the preceding discussion, context-specific environmental risks were identified by ChatGPT, but other salient risks identified by the experts were not captured. When pre-sented to the expert participants, the consistency of the recommen-dations output from the AI model were critiqued.

Participants frequently described ChatGPT’s mitigation recommendations as “theoretical” or “textbook-style” (P-09, P-11), underscoring AI’s limitations in capturing real-world decision-making complexities. This critique aligns with Tixier et al.’s (2017) finding that AI’s com-prehensive data processing often fails to address unique, context-dependent project factors. For example, ChatGPT’s suggestion to manage labor issues by involving unions, noted by P-10, reflects this disconnect in construction contexts where union presence may be minimal or nonexistent. These observations suggest that AI-generated strategies offer breadth but require human interpreta-tion to achieve practical viability.

Nonetheless, ChatGPT’s ability to produce broad, creative strategies—praised by P-11 for its align-ment with “academic best practices”—shows promise in environ-ments where wide-ranging risk assessments are needed, as echoed by Chenya et al. (2022), who emphasized AI’s strength in process-ing large data sets to yield diverse strategies. The specificity and adaptability of human input, emphasized by P-09, are often crucial for managing complex projects, supporting Zou et al.’s (2007) assertion that human expertise is essential for translating broad AI-driven mitigation insights into actionable, context-specific so-lutions. This theme of integrating AI with human expertise appears in Afzal et al. (2021) and is echoed by P-11’s advocacy for a hybrid approach, where “experts work with AI” to maximize effective-ness.

Interestingly, ChatGPT also broadened perspectives, as P-01 noted AI’s ability to suggest novel risks and mitigation strategies, supporting Chenya et al.’s (2022) findings that AI uncovers nonobvious risks. How-ever, human oversight remains critical, as P-03 emphasized the im-portance of human judgment in interpreting and contextualizing AI-generated mitigation insights, particularly in complex, ethically sensitive projects, aligning with Rampini and Cecconi (2022). Furthermore, findings suggest that less experienced professionals might derive even greater benefit from AI mitigation insights, while experts found ChatGPT’s mitigation strategies lacking depth. P-01 noted that junior professionals could gain from AI’s novel, com-prehensive approaches, indicating AI’s potential as a tool for ex-panding risk awareness and training in early-career settings.

Usability, User Experience, and Ethical Considerations (O3)

This study’s findings reveal that ChatGPT’s user interface (UI) was widely praised for its intuitiveness and user-friendliness, with par-ticipants such as P-01 and P-04 appreciating its simplicity and clean design, which facilitated efficient, organized workflows in risk management; this aligns with Xia and Chen (2024), who highlighted the role of UI in driving AI adoption in professional settings. ChatGPT’s structured format streamlined processes, as noted by P-09, indicating an advantage in accessibility and usability, particularly for users with limited technical expertise.

However, the quality of ChatGPT’s out-puts largely depended on user input, especially in crafting effective prompts; P-11 observed that novice users might face a learning curve, revealing a disparity in results between inexperienced and advanced users, pointing to potential improvements through more guided query-building tools. For this reason, the results from this study were analyzed by GPT to develop a prompt that can accelerate this process (Fig. 7).

Despite its accessible interface, ChatGPT lacks functionalities critical for complex, large-scale projects; P-06’s comparison of ChatGPT to traditional risk management software highlights Chat-GPT’s limitations in offering specialized features such as real-time data integration, consistent with Smith et al. (2014), who argued that while AI could be efficient, it might not replace specialized tools in complex domains. Additionally, ethical and security con-cerns emerged prominently, with P-01 and P-11 raising concerns about data confidentiality when using a cloud-based AI system for sensitive information—a risk that Montasari (2022) argued could be mitigated with robust data protection protocols. Participants like P-08 emphasized the necessity of secure data processing and encryption, warning against prioritizing convenience over security.

ChatGPT’s potential for algorithmic bias further complicates its use, as P-11 noted instances of irrelevant risk suggestions, like natural disasters in improbable areas, raising concerns about the contextual validity of AI outputs; this reflects broader issues of algorithmic bias identified by Noble (2018). Integrating ChatGPT into risk management thus requires combining it with human over-sight, as suggested by P-03, to address biases and improve contex-tual accuracy. Addressing these challenges involves implementing strict data protection policies, adhering to general data privacy regu-lation, and ensuring that sensitive data are encrypted and securely stored.

Fig. 7. Risk assessment prompt generated by Chat-GPT AI.

Case Study Validation of Developed Prompt

A real-world case study in Ireland validated the prompt to ensure it produced the necessary output. The N2 Ardee to Castleblayney Road Upgrade Project is a vital infrastructure initiative aimed at addressing key deficiencies in road safety, operational capacity, and regional connectivity along a 32-km section of the N2/A5 Dublin-Derry Road in Ireland (Fig. 8). Serving as a critical link for North–South connectivity, the road is part of the Trans-European Transport Network (TEN-T) and is aligned with both European Union goals for improved mobility and sustainability and Ireland’s National Planning Framework (NPF) 2040. The project seeks to enhance economic viability by reducing travel times, improving safety by modernizing road infrastructure, and minimizing environ-mental impacts in line with Ireland’s Climate Action Plan.

It also aims to promote social inclusion by ensuring better access for vul-nerable populations and strengthening regional integration, particu-larly in the post-Brexit context. The project addresses significant challenges, including high collision rates, traffic congestion, outdated infrastructure, and environmental sensitivity. Through advanced GIS and traffic modeling tools, the project employs innovative, sustain-able practices, such as a dual carriageway design with dedicated in-frastructure for cyclists and pedestrians, while ensuring minimal ecological disruption. By aligning with EU environmental directives and using stakeholder feedback, the project aims to create safer, more efficient transport networks that bolster both local and transnational economic resilience.

The prompt was inputted into ChatGPT, and GPT requested in-formation, as shown in Fig. 9(a). The project “Option selection report—Volume 0” was provided to GPT. GPT then produced a summary of the project scope [Fig. 9(b)] and proceeded to conduct a risk evaluation of the road infrastructure project (Fig. 10). Upon completing the evaluation, GPT requested additional details to im-prove or refine the risk evaluation. For brevity, no further requests were made to expand on the number of risks identified and a risk expert was asked to comment on the suitability of the outcome given the project context.

The evaluator P-12 noted that it was very good to see GPT categorization, table headings, and consideration of residual risk; however, the evaluator suggests that AI-driven risk identification provides a high-level, generalized framework comparable to that of a junior professional, categorizing risks broadly (e.g., environ-mental impacts, stakeholder conflicts) but lacking specificity in context-dependent challenges such as ground conditions, land ac-quisition, or unexploded ordnance. P-12 suggested that traditional methodologies, by contrast, use expert workshops, historical data, and consensus-based approaches to capture secondary and local-ized risks more effectively.

The reviewer noted that while AI accel-erates initial risk identification, its utility remains constrained in the qualitative assessment of probabilities, validating impacts, and addressing complex risks, such as political opposition, but it ex-cludes design modifications arising from environmental impact as-sessments. Enhancing AI’s decision-support capabilities requires further domain-specific training with construction data sets, geo-graphical customization to reflect regulatory and environmental constraints, iterative refinement through targeted prompts, and in-tegration with real-time data on contractor availability and material supply chains.

P-12 expressed that in risk mitigation, AI-driven rec-ommendations should incorporate granular strategies, including early procurement of raw materials to hedge against market volatility, structured procurement models to minimize conflicts, and weather-adjusted scheduling, particularly in climates such as Ireland’s.

These comments led the team to prompt GPT to expand the first risk item as follows: “The first risk, ‘Environmental impact on sen-sitive areas,’ was noted for being very high level. Identify 10 more granular ‘Environmental impact on sensitive areas’ risks specific to the project and present the evaluation in a new table following the headings of previous outputs. The results are presented in Fig. 11.

Overall, the case study confirms the efficient use of the prompt in risk evaluation and challenges the notion of AI contextual

Fig. 10. GPT—Risk Matrix for N2 Ardee to Castleblayney Road Upgrade Project.

Fig. 11. Top 10 project-specific risks for environmental impacts on sensitive areas.

limitations in producing human-like evaluations in a two-staged process. The first, a general understanding of risk given a context, and the second involves individual AI-assisted refinement.

Conclusion and Recommendations

The results contribute to the broader discussion regarding the role of AI in professional decision-making processes. This study high-lights the promise and limitations of integrating AI tools like ChatGPT-04 into the evolving landscape of construction risk management. While traditional risk management methods are valued for their thoroughness and structure, they face growing criticism for in-efficiency, inflexibility, and an inability to quickly adapt to emerging risks. AI tools offer significant promise in addressing these short-comings, particularly in enhancing quick decision-making, estima-tion of probabilities, impact, and mitigation strategies. However, human expertise, collaboration, and continuous learning remain indispensable, as AI alone lacks the depth to fully address some context-specific risks and complexities without tailored inputs.

While AI can streamline risk identification and provide structured, comprehensive insights, it falls short in addressing some project-specific insights, which remain the domain of human expertise. As such, a hybrid approach that combines AI’s efficiency and data-processing power with the intuitive judgment and contextual understanding of professionals emerges as the most effective strat-egy. This dual approach can enhance adaptability and decision-making in the construction sector, ensuring that risk management practices are efficient, robust, and contextually sound. Moreover, the ethical considerations and data security concerns associated with AI adoption highlight the need for rigorous guidelines and safe-guards to ensure safe and responsible implementation.

The user experience of AI technologies like ChatGPT is mostly favorable, characterized by intuitive interfaces that enhance operational effi-ciency. Accordingly, discussions suggest that ChatGPT technologies may be especially beneficial for less experienced workers since they lack comprehensive risk awareness, whereas experienced professio-nals have indicated a deficiency in contextual depth. The quality of outputs, previously considered to be greatly affected by the user’s ability to craft effective prompts, is now one step closer to being resolved by the developed prompt presented in this research.

Despite its valuable insights, this study has limitations, includ-ing its small sample size and the project context explored, which may limit the generalizability of the findings. However, larger sam-ples are constrained by GPT’s current data upload limits. Risk management often involves collaboration among various stakeholders. The study’s design excluded this dynamic, limiting insights into how AI or traditional methods perform in team-based decision-making contexts. Future research should focus on longitudinal studies, more diverse samples, and more real-world applications to better under-stand the complementary roles of AI and human expertise.

Also, further investigations are needed to refine the developed prompt to improve AI’s capabilities to incorporate more project-specific data and context while exploring scalable strategies for integrating AI into traditional risk management methodologies. However, sense making and logical intuitive reasoning are complex representations of human behavior and are still challenging to replicate in AI risk assessments. Additionally, comparative evaluation of outcomes for different generative AI models could prove useful in guiding the suitability of generative AI models for risk evaluation. Also, evi-dence is lacking where ChatGPT accounts for cognitive biases and decision-making tendencies within projects, where it may suggest impractical mitigation strategies given a company’s risk appetite or financial constraints.

Understanding whether such tendencies are embedded in ChatGPT’s recommendations and how to mitigate them is crucial to comprehend AI’s strength in navigating the com-plex and evolving risk and decision-making landscape in construc-tion. Ultimately, the synergy between AI and human expertise is key to advancing risk management practices, positioning the construc-tion sector to tackle emerging risks with precision and agility.

Data Availability Statement

The data presented were extracted from a thesis. Some or all data, models, or codes that support the findings of this study are available from the corresponding author upon reasonable request.

Hector Martin: Conceptualization; Data curation; Formal analysis; Funding acquisition; Investigation; Methodology; Project adminis-tration; Resources; Software; Supervision; Validation; Visualiza-tion; Writing – original draft; Writing – review and editing. Jennifer

James: Conceptualization; Data curation; Formal analysis; Investi-gation; Methodology; Project administration; Software; Validation; Visualization; Writing – original draft. Aaron Chadee: Formal anal-ysis; Investigation; Validation; Writing – original draft; Writing – review and editing.

Download transcript ↗