Causal mapping of the risks of using generative AI in software development
1 More Paper · Full Reading

About this paper
A full audio edition of this paper.
Authors: David Kinnberg Hein, John Stouby Persson, Victor Vadmand Jensen, Anders Rysholt Bruun, Martin Gilje Jaatun
Published in: Information and Software Technology
Publication date: 2026-09
Read the paper: https://doi.org/10.1016/j.infsof.2026.108202
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “Causal mapping of the risks of using generative AI in software development,” by David Kinnberg Hein and colleagues. Published in Information and Software Technology in September 2026.
Aalborg Universitet
Causal mapping of the risks of using generative AI in software development
Hein, David Kinnberg; Persson, John Stouby; Jensen, Victor Vadmand; Bruun, Anders Rysholt; Jaatun, Martin Gilje
Published in:
Information and Software Technology
Creative Commons License CC BY 4.0
Publication date: 2026
Document Version
Publisher's PDF, also known as Version of record
Link to publication from Aalborg University
Citation for published version (APA):
Hein, D. K., Persson, J. S., Jensen, V. V., Bruun, A. R., & Jaatun, M. G. (2026). Causal mapping of the risks of using generative AI in software development. Information and Software Technology, 197, Article 108202. the linked source
General rights and it is a condition of accessing publications that users recognise and abide by the legal requirements associated with these rights.
Take down policy
If you believe that this document breaches copyright please contact us at the email address providing details, and we will remove access to the work immediately and investigate your claim.
Contents lists available at ScienceDirect
Information and Software Technology journal homepage: the linked source
Causal mapping of the risks of using generative AI in software development
David Kinnberg Hein a,, John Stouby Persson a, Victor Vadmand Jensen b, Anders Rysholt Bruun a, Martin Gilje Jaatun c a Aalborg University, Department of Computer Science, Selma Lagerl ̈ofs vej 300 9000, Aalborg, Denmark b Aarhus University, Department of Clinical Medicine, Jens Chr. Skous Vej 4 8000, Aarhus, Denmark c SINTEF Digital, Software Engineering, Safety and Security, Strindvegen 4 7034 Trondheim, Norway
A R T I C L E I N F O
Abstract.
Context: Generative AI tools can enhance the productivity and effectiveness of software development. These tools are evolving rapidly, as are the associated risks. Therefore, software organizations adopting these tools need to understand their risks, why they occur, and how they evolve over time.
Objective: Previous research has highlighted risks related to using generative AI in software development; however, the underlying causes of these risks remain largely unexplored. To effectively manage these risks, we need to understand what causes them. Therefore, we examine the causal explanations behind the risks of using generative AI in software development organizations.
Methods: In a multi-case study of three Danish software organizations during the early stages of generative AI tool adoption, we interviewed software developers and managers within these organizations to uncover the causal explanations of risks associated with generative AI. We employed a causal mapping approach to understand and visualize the differences and similarities across software organizations’ causal explanations of risks. A year later, we returned to validate the risks and causal maps with representatives from each software organization. Results: Our study shows that the software organizations exhibit circular reasoning regarding a perpetual uncertainty in the validity of AI-generated code, with the three risks: Mistrust of AI, Insufficient validation of AI output, and Insufficient/insecure AI output.
They are also concerned about similar root risks, which solely cause other risks without being caused by any risks themselves, such as Lacking AI competencies. While they differed in emphasis on tail-end risks, which are understood as not causing any additional risks. Conclusion: The causal mapping approach proved appropriate for showing similarities and differences in software organizations’ perceptions of risks and causal explanations related to the use of generative AI in software development. The same applies to visualizing how the emphasis on risks can change over time.
1. Introduction.
Recent developments in artificial intelligence (AI) have led to the proliferation of generative AI like ChatGPT or GitHub Copilot. Generative AI tools are trained on large amounts of data, like text or images, and can use training data to generate new content based on users' prompts. With prospects of increased productivity and time savings, software organizations have now adopted generative AI in software development. However, the introduction of AI to software engineering is not an entirely new phenomenon. In 2012, Harman argued that software engineers were becoming increasingly interested in AI due to the growing complexity and interconnectedness of software systems and a wish to optimize the processes and products underlying such complex software.
As such, traditional AI techniques, like machine learning (ML) and deep learning (DL), have been applied to many areas of software engineering. One tertiary study of ML for software engineering (ML4SE) highlighted that eleven knowledge areas were covered in the ML4SE literature, with the five most addressed areas being software quality, testing, processes, management, and requirements. Similarly, a systematic literature review found that ML and DL had been applied to over half of the reviewed literature for either analyzing defects or supporting the maintenance and evolution of software. However, researchers have also highlighted the possible risks of AI in software engineering argued that AI becomes riskier the closer it is
0950-5849/© 2026 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY license (the linked source).
applied to software's runtime due to costlier consequences. Further, end-users, experts, and other stakeholders may experience different negative impacts from the application of AI, emphasizing the variety of possible risks here. For example, using ML for software bug reports poses the risk of coordinators losing the ability to diagnose and manage bugs.
The adoption of generative AI in software development has similarly introduced a range of risks for software organizations, such as inaccurate AI output, a lack of trust and control, and being inappropriate for more complex tasks. Consequently, researchers have emphasized the need to raise awareness among software engineers regarding the risks of using generative AI and to provide guidance for mitigating these risks. However, to effectively address the risks posed by generative AI in software development, we first need to understand its underlying causes. Attempts have been made to understand the root causes of the risks in software development qualitatively and quantitatively. However, generally, qualitative research on machine learning practices remains underexplored. To the best of our knowledge, particularly for the risk causes associated with generative AI in software development.
Against this backdrop, we will address the following research question:
• How do software organizations explain the risks associated with using generative AI in software development?
To answer this question, we conducted an in-depth multi-case study of the causal reasoning for the risks associated with using generative AI in three Danish software organizations. A risk, in itself, can be understood as an entity comprising an explanatory causal relationship between a “risk object” that poses the threat and an “object at risk” that represents the value being threatened. Our study aims to expand on the systemic view of risks as interconnected; when one risk occurs, it can lead to the occurrence of another. To unfold how software organizations see causal connections between the risks of using generative AI, we apply the causal mapping approach. Causal mapping has been used in previous research for understanding and visualizing causal explanations of project risks, barriers to knowledge sharing, and technology adoption.
This case study extends this line of research regarding the risks associated with using generative AI in software development. Nevertheless, the fast technological evolution of generative AI also shapes software organizations’ explanations of the associated risks. Against this backdrop, we returned to the three software organizations after they had employed generative AI for a year to assess whether their explanations of risks had shifted in emphasis, to address the second research question:
• How do the explanations of risks associated with using generative AI in software development evolve after software organizations’ initial adoption?
The paper detailing our response to these two research questions is organized as follows: First, we review the current research on generative AI and risk management in software development. Second, we introduce and justify the study’s research approach, followed by an outline of our findings and causal maps. Finally, we discuss the paper's contributions to research, implications for practice, threats to validity, and future research directions.
2. Related work.
2.1. Generative AI in software development.
Generative AI developments in recent years have led to this technology being widely considered for professional software engineering. Researchers have highlighted the breadth of applications of AI for software engineering purposes across a wide range of software engineering activities. As potential AI pair programmers in software development, AI-generated code is set to significantly influence the quality of the entire software project. A survey study highlighted how software engineers perceived GitHub Copilot-generated code to be shorter and more accurate, indicating that it is of higher quality. Similarly, using GitHub Copilot may lead software engineers to realize that the suggestions from generative AI tools are unique and contain elements that might otherwise be overlooked.
In cases of quality issues in AI-generated code, it has been shown that ChatGPT can utilize static analysis outputs to correct some of the code itself. Furthermore, generating code with GitHub Copilot can result in fewer vulnerabilities than human-written code, thus limiting security issues that attackers can exploit. Researchers have also suggested that user experience (UX) designers may employ generative AI to identify usability issues in their designs, thereby making applications more straightforward to use. Researchers have underscored the potential of using generative AI to enhance the experiences within the software engineering process, typically by automating mundane tasks and allowing more time for the more creative aspects of software engineering.
Additionally, researchers contend that software engineers may find it easier to learn and improve their programming skills without frequent context switching.
However, while researchers have highlighted generative AI’s prospects of improving code and productivity, there is also research suggesting challenges in using generative AI to improve software did not find quality differences in the code generated with either ChatGPT or GitHub Copilot, and they did not decrease code quality compared to code produced by a human demonstrated that buggy solutions from GitHub Copilot need to be managed by expert software engineers rather than inexperienced engineers, which is corroborated by other findings that human supervision of GitHub Copilot increases code quality. Generative AI's potential here is also challenged by the fact that “ChatGPT-generated code is prone to various code quality issues, including compilation and runtime errors, wrong outputs, and maintainability problems'' [p. 21].
In continuation, recent studies have found that AI-generated code does not create the same vulnerabilities or types of bugs as human-written code. At the same time, various research efforts indicate that generative AI and human developers may prioritize different aspects of the prompt when addressing coding challenges. Further, software engineers may exhibit uncertainty toward the copyright of AI-generated code, rely too eagerly on generative AI without understanding its limitations, or have security concerns about generative AI.
Moreover, concerns have been raised about leveraging generative AI in the software engineering process. For instance, UX designers may be apprehensive about AI in their design tools, fearing the perpetuation of existing biases in their design work. Others have pointed out the risk that generative AI might lead to negative experiences within the software engineering process, such as primarily relying on generative AI outputs. Furthermore, software engineers may struggle to grasp AI-generated code fully. This challenge is compounded by the necessity for existing software engineering knowledge to use generative AI effectively for coding. The literature emphasizes both the opportunities presented by generative AI in software engineering and the associated risks of utilizing these tools.
However, these risks related to generative AI are usually considered in isolation described how software engineering management must critically assess the role of generative AI to prevent employees and organizations from facing the consequences of risks tied to this technology. Therefore, understanding how the risks of generative AI are interconnected is essential for effective risk management to address the appropriate risks and prevent risks from cascading and resulting in multiple coinciding issues. Especially as these generative AI tools are evolving at a tremendous rate, as recently seen with the introduction of Agentic AI.
Recent studies, such as Zhou et al., examined the causes of Copilot usage problems. They found that the most frequent causes are
“Copilot Internal Error,” “Network Connection Error,” and “IDE Compatibility Issue”. Other researchers examined why software engineers choose to use generative AI. It was discovered that the compatibility of generative AI with existing development workflows and organizational values serves as the main driver for its adoption. Although these generative AI tools may align with the organization's requirements, it remains essential for organizations to recognize the potential risks associated with the use of AI. Sergeyuk et al. found that the most frequent reasons for developers to avoid using generative AI are “Lack of need for AI assistance”, “AI-generated output is inaccurate,” “Users’ lack of trust and the desire to feel in control”.
Despite attempts to investigate the causes related to generative AI adoption, the literature on the causes of the risks and challenges associated with using generative AI in software development remains scarce. Therefore, we aim to bridge this knowledge gap.
2.2. Risk management in software development.
Risk management in software engineering has existed for many years and can be traced back to the introduction of risk management approaches by the prominent figure, Barry Boehm. Traditionally, Boehm defines a risk as “the possibility of loss” [pp. 33], while others merely view it as an undesirable future event. Simplistically, a risk can be viewed as the sum of the likelihood that the event will occur and the consequences of the event. There still seems to be some ambiguity regarding the definition of risk. In recent years, attempts have been made to unify risk management concepts, presenting terminology to address the aforementioned ambiguity and inconsistencies surrounding the definition of risk.
The argument for integrating risk management in software development has traditionally been to combat high software project failure rates and cost overruns, leading to a cascade of frameworks for identifying and addressing project risks. The issue of project failures and cost overruns is still prevalent in modern-day software projects. Kula et al. noted that factors like refining requirements, task dependencies, organizational alignment, and politics are seen as significantly influencing on-time delivery. In contrast, proxy measures including project size, the number of dependencies, past delivery performance, and team familiarity can account for a substantial portion of scheduling deviations.
Furthermore, the rise of agile methodologies has led the risk management literature to focus on risk management within agile development, examining risks specific to software projects that follow agile principles. Several frameworks have been created to integrate risk management into the iterative nature of agile development, and frameworks tailored for addressing the challenges of distributed agile development. Although research suggests that adopting risk management activities is an adequate way to reconcile traditional waterfall software development with agile development. The agile risk management frameworks face the issue of conflicting with a core value within the agile manifesto, “individuals and interactions over processes and tools,” as adopting risk management activities introduces new processes and tools for an agile team.
Agile software teams face new challenges in implementing risk management practices, alongside noticeable distinctions compared to the key risks traditionally identified by Boehm, which include "personnel shortfalls,” “unrealistic schedules and budgets,” and “developing the wrong functions and user interface." In contrast, the primary risks in agile projects are recognized as “Technical debt,” “Lack of knowledge retention,” and “Separation of development and IT operations”. Recent research has also revealed risks associated with data management in agile software development, including challenges in managing data integration, gathering various data types, automating data collection, and satisfying real-time analysis demands. To mitigate these risks, they suggest strategies such as adopting appropriate automation tools, decentralized data management practices, and ontology-based methods.
Clearly, the risk management literature is not scarce in proposing risk management frameworks. However, in the field of software engineering, little attention has been devoted to exploring the concept of risk in greater detail, particularly its causality. Apart from a few qualitative studies on root causes of risks, and the abundance of statistical studies predicting the causes of risks. This suggests an inherent objective view on risk in the software engineering field. Beyond the realm of software engineering, Boholm and Corvellec describe risk as a concept shaped by cultural and social contexts, representing a harmful possibility for an actor. This concept comprises an interrelated "risk object" and an "object at risk." The risk object signifies the potential danger, whereas the object at risk denotes the value in jeopardy.
The notion of risk does not arise by itself; instead, it is formed through a semantic connection between these objects. The identities of the risk object, the object at risk, and their interrelation are established only through their connections. A risk object is meaningful only in relation to an object at risk, which, in turn, is identified by its link to a risk object.
In 1998, Lyytinen et al. introduced an attention-shaping framework for software risks, grounded in Leavitt's socio-technical model. They propose that software risks are part of a system comprising four interconnected elements: actor, technology, task, and structure. A change in any of these components or their relationships can result in changes to the others and the software development process. Therefore, software risks in development can be traced back to any of these four elements. In continuation, the idea of systemic risk has been introduced to overcome the complexities and challenges associated with risk emergence, emphasizing that systemic risks are interconnected and that when one part fails, it can cause failures in another.
In closing, extant research has explored the causes of project risks by employing the concept of systemic risks and the causal mapping technique to illustrate the systemic network. They showed that risks in software projects within agile teams are interconnected, suggesting that causal mapping is suitable for treating these risks collectively instead of individually. This has similarly been found to be the case for barriers to knowledge sharing in software projects and visualizing the causality of perceptions of adopting a new technology. Based on the effective use of causal mapping in the risk management literature, we adopt this approach to address the gap in research on understanding risk causes in the risk management literature.
3. Research approach.
We adopted a multi-case study approach, collaborating with three Danish software companies to uncover the risks and causes of using generative AI in software development. This procedure provided both richness and depth in comprehending each software organization’s risks while allowing for comparisons among them. The software organizations in this study were chosen based on their early stage of adoption of generative AI tools to examine the evolution of the risks of using generative AI a year after our initial investigation. To allow for comparison among the case organizations, they were selected as a result of variation in their contexts on account of their industry, size, customer segment, and software development process.
This enabled us to investigate the causes of risks across various types of software organizations and to compare the causes of the identified risks based on the specific context of each organization. We employed the causal mapping technique to visualize the causes of risks and facilitate more straightforward comparison of risks and their associated causes among software organizations. All the researchers work for a university and have no proximity to the case companies we collaborated with. None of the researchers work for the case companies, and neither the researchers nor the case companies solely dictated the end goal of this study; instead, it was developed in unison. We anonymized the collected data, signed an NDA with the companies, and sent the completed article to one representative from each case.
3.1. The cases.
The Bankers is a leading Danish bank with a workforce exceeding 4000 employees, offering a variety of banking and housing investment products and services. They focus on providing mortgage credit and banking operations. The Bank is owned by a Danish association of homeowners and corporations. It employs around 400 staff members dedicated to developing software used both internally and by other banks that provide mortgage products. Among these, 100 employees are part of a multi-team effort to create a mortgage platform. The teams work together within the Scaled Agile Framework (SAFe), collaborating towards shared objectives and solutions. This is accomplished through the Agile Release Train, which enables continuous value delivery. The process adheres to a fixed and consistent schedule guided by the program increment rhythm.
Each program increment lasts 10 to 12 weeks, with teams initiating a new system increment every two weeks. All teams are synchronized within the same program increment, with established start and end dates. A key event within SAFe is the Program Increment Planning session, where teams align on the business context and vision, prioritize essential development areas, and identify risks that may impact the success of the program increment. In addition to the program planning session, the Bankers’ software teams utilize the Scrum event known as the "sprint retrospective” every second week to discuss past challenges and future risks in the software project and to find solutions.
The Automaters is a software vendor focused on logistics visibility and automation. Established in the early 1950s, the company expanded its automation solutions from Denmark to a global clientele in the 2000s, opening offices in Canada and Romania and employing over 200 staff. Approximately 100 of these are directly involved in software development. Utilizing RFID (radio-frequency identification), it tracks complex assets in both private and public sectors. Its clients range from healthcare providers to airports, airlines, postal services, and libraries. RFID technology uses radio waves to identify and track items via tags that store data, which are read by an RFID antenna. The company employs a homegrown development lifecycle that primarily follows plan-driven principles for its automation solutions.
At the Automaters, software development is divided among specialized teams, each focused on different software architecture components, such as a team dedicated to Internet of Things (IoT) devices to build expertise in the field. The sales team plays a vital role early in the development process, collecting client requirements and facilitating the approval process. Approved requirements are then forwarded to the software development teams for execution. The organization adheres to a plan-driven development model rooted in traditional waterfall methodology, emphasizing structured development stages with clear checkpoints. The development process ensures systematic progress and thorough validation throughout the entire software development process.
The Insurers is a Danish non-life insurance company recognized for providing affordable, customer-focused insurance solutions. Established over 100 years ago, it employs >500 people in total. The company has a long history of offering various insurance products, including auto, homeowners, accident, and travel insurance. The company was initially owned by Danish trade unions, with a strong emphasis on ensuring fair and accessible insurance for union members. They have recently been acquired by one of the largest insurance providers in Scandinavia, yet they continue to operate under their own brand. They are noted for their competitive pricing, efficient digital services, and high customer satisfaction ratings in the Danish insurance market.
The Insurers’ software development department follows the agile methodology "Scrum," enabling iterative development, flexibility, and continuous feedback loops to enhance internal systems and digital services. By implementing Scrum, the Insurers facilitate collaboration across development teams, accelerate software delivery, and align their digital solutions with evolving customer and business needs. The Insurers’ software teams follow the typical Scrum activities outlined in the Scrum Guide, including daily scrums, sprint retrospectives, and sprint reviews.
3.2. Data collection and analysis.
As shown in Table 1, we initiated our data collection process by conducting interviews with a wide range of software practitioners from our three case companies to identify the risks associated with using generative AI in software development. We conducted 15 initial semi-structured interviews, each lasting between 32 and 67 min. The interviews were all recorded and transcribed with Whisper. The participants were selected based on their experience and interest in generative AI tools. We also aimed to interview individuals with diverse backgrounds and roles in the cases to capture the richness of various perspectives and enhance data triangulation. During the analysis, we employed cross-case coding, where the first author and second author produced codes that were mutually reviewed by both to enhance data triangulation.
We conducted a cross-case analysis using NVivo to ensure consistency in the risks and causes across the three cases and to identify emerging differences among them.
The data analysis process was initiated by reviewing the transcriptions to identify risks that adhere to the definition of “the possibility of loss.” Throughout the data analysis, we noticed a pattern in which participants explained risks that caused other risks; therefore, we adopted the causal mapping method to illustrate the systemic implications of these risks. Causal mapping allows comparative analysis of different organizational actors and reveals both differences and similarities in their perspectives. Additionally, we selected this method because it has proven to be an appropriate tool for visualizing the causes of risks in similar research beyond software engineering, particularly regarding perceptions of technology adoption.
Causal mapping enables practitioners to articulate their causal reasoning regarding specific concepts within a real-world context. This method visually illustrates patterns of concepts and causal beliefs, capturing the explicit statements of various stakeholders. The process of eliciting causal relationships emphasizes identifying instances where one concept (A) is seen as leading to another concept (B) or where B is understood as a result of A. These causal connections are depicted through models utilizing nodes and arrows, with nodes signifying concepts within a specific context and arrows depicting the actors' perceived causal relationships among these concepts. Using the causal mapping approach, we constructed three causal maps in total, one for each of the case companies.
We linked risks in the causal maps when two statements, each describing a risk, suggested a cause-and-effect relationship between them, indicated by phrases such as “because,” “due to, ” or “as a result of”. The first author initially inferred the causal links, which the second and third authors then reviewed. As a specific example, the following quote: “I have no idea what it bases its knowledge on or how it stores the data you sent it to. It’s somewhat hidden in terms of ChatGPT” alludes to the risk of inadequate data management by AI, as the statement shows ambiguity in how generative AI tools handle their users’ data. This quote further explains the risk posed by Inadequate data management of AI: “The use of large language models introduces a security risk because we are not entirely in control of what it sends elsewhere,” thereby causing the risk of Lack of control over AI.
The word “because” acts as the link connecting these two risks. No additional preset criteria, frameworks, or measurable metrics were employed for developing the causal links. Our focus was on ensuring the trustworthiness and reliability of our findings through validation with our participants.
Following the creation of the maps, we validated these with representatives from the case companies to increase their trustworthiness and reliability. We explained the rationale of the risks and their causal relationships to the representatives, continuously probing them for feedback. In the validation interviews, we were particularly interested in validating the risks and causal relationships that exhibited
Overview of our participants, the three cases, and data collection procedures. The presence of a time duration indicates the participant’s presence at a specific data collection procedure.
ambiguity during analysis. For instance, the Automaters sought to establish a relationship between the two risks: Lack of AI competencies and Insufficient validation of AI output. Additionally, we conducted these validation interviews a year after the initial interviews to enable insights into the evolution of risks and causal relationships. Each of the companies provided feedback on risks that had become less relevant or increased in relevance, while also identifying a few new risks. Finally, after validating and finalizing the causal maps, we conducted a workshop with representatives from the case companies. This workshop aimed to ensure the maps' cohesiveness and to facilitate engagement among the case companies, exploring whether their causal maps would support discussion or reflection on their own.
The workshop enabled us to conduct a combined causal map capturing all the risks and causes from the case companies. The risk nodes and relationships were not determined by frequency, salience, or co-occurrence. Instead, we included all the risks and causal relationships identified in the initial causal maps. We used participants’ feedback during validation interviews and workshops to assess the salience of these risks and relationships. This process helped decide whether risks or causal links should be added or removed to accurately reflect the risks associated with using generative AI in software development. As a result of the feedback, the causal maps changed slightly (e.g., risks were dropped, or new causal relationships were added) for each case. Examples of these changes will be provided throughout the findings section.
4. Findings.
Risks represented in blue indicate that they were noted to have decreased in emphasis during the validation interview or workshop one year after initial adoption. In contrast, risks represented in red denote that they have increased in emphasis. We distinguish between root risks, tail-end risks, chaining risks, and isolated risks in organizations’ causal reasoning regarding the use of generative AI in software development. Root risks are defined as risks that solely cause other risks without being caused by any risks themselves, serving as the root of a causal reasoning chain in software organizations’ perceptions of risk. Tail-end risks, conversely, do not cause any additional risks, representing the tail-end of a causal reasoning chain. Chaining risks are part of a reasoning chain involving risks that are neither root nor tail-end.
In other words, chaining risks are those risks that cause them and also cause other risks. A chain of risks within a causal map can form a circular reasoning loop, without a tail-end risk; in such cases, there is no definitive conclusion in the causal rationale, creating an explanation that may appear to be perpetually ongoing. Finally, isolated risks are neither caused by other risks nor do they cause additional risks; they represent standalone concerns in the organization’s perception of risk.
4.1. The bankers.
At the start of our investigation, the Bankers had just begun using generative AI tools, with ChatGPT and GitHub Copilot as the primary tools. During the validation interview, they decided to remove an initial concern about bypassing usage protocols, as the risk was deemed not relevant to their software development.
Fig. 1 shows that Inadequate data management of AI is a root risk, with no other risks identified as its causes. This risk pertains to the uncertainty surrounding how AI tools store and manage data, and it is considered a contributing factor to four other risks. E.g., the tail-end risk, Lack of control over AI, which has seen a decreasing emphasis after the Bankers started hosting their generative AI locally. Nevertheless, the IT-operations manager noted: “The use of large language models introduces a security risk because we are not entirely in control of what it sends elsewhere” (P2). The Bankers, also see Inadequate data management of AI as the cause for the tail-end risk, Lack of transparency: “I have no idea what it bases its knowledge on or how it stores the data you sent to it. It’s somewhat hidden in terms of ChatGPT” (P4)... “it remains a black box how ChatGPT truly works” (P1).
Mistrust of AI is a central chaining risk in the Bankers’ causal explanations of their concerns regarding generative AI, especially for the technology’s ability to manage users’ data. One of the developers explained: “We own our software, and since I do not know how or where the data I sent to it is stored, I wouldn’t use it to transfer code to it in that manner. It’s a matter of trust” (P5). During the validation interview, one year after our initial interviews, two developers explained that the risks of Lack of control over AI and Inadequate data management of AI have diminished in emphasis. Since the Bankers are now hosting their version of generative AI internally, they mentioned that the risk of Inadequate data management of AI has taken on new emphasis. Managing the data internally has proven to be challenging, as they have never attempted anything like this before.
Therefore, the Bankers’ causal map only indicated that Lack of control over AI has decreased in emphasis, not Inadequate data management of AI.
The Mistrust of AI is further seen as caused by other risks, such as the root risk, Lack of policies: “It clearly must be used with caution, no matter what. I think people are a bit nervous about using it right now because we do not have any policies for using it internally yet” (P5). A year after implementing generative AI tools at the Bankers, they now have internal policies for using generative AI. However, the risk was not seen as having decreased emphasis in the validation interview. Moreover, the root risk, AI replacing developers, is also cited as causing Mistrust of AI: “I have considered potential unemployment and being replaced by AI. I do not believe that is where we are currently, but considering how rapidly these tools are evolving, I am quite concerned about how my future will look" (P4).
During the validation interview, the Bankers noted that they are not nearly as worried about being replaced by generative AI. After using generative AI tools for a year, they have repeatedly observed that they are unsuitable for more complex tasks and, therefore, cannot always assist them. Thus, the tail-end risk, Inadequate for complex tasks/ industry-specific tasks, is marked with red to indicate an increased emphasis in the Bankers’ causal map.
Consequently, this Mistrust of AI underscores the emphasis on validating the output of generative AI. It thus causes the chaining risk Insufficient validation of AI output: “It may be beneficial to use, but it also presents a threat because who made this (code)? What is it we do not see at first glance in the code, and that’s why we must validate the code very thoroughly before we use it” (P2). One of the developers presents an example where code generated by generative AI can lead to insecure code being integrated into their codebase if it is not adequately validated: “I asked ChatGPT for help with some regex code, and I committed the code, but our testing software flagged several warnings for potential issues introduced by the AI.
Then, I asked it to revise the code to eliminate those security threats, and the code successfully passed through the build process without any warnings” (P4). The developer and the Bankers are fortunate to have capable testing tools. If the output is improperly validated, it results in Insufficient or insecure AI output, causing Mistrust of AI and necessitating revalidation. The three chaining risks: Mistrust of AI, Insufficient validation of AI output, and Insufficient/ insecure AI output form a circular chain of reasoning, as it is difficult to determine when the output from the AI is sufficiently validated.
Moreover, Insufficient/insecure AI output was explained to be caused by the root risk, Lacking AI competencies: “I am not very experienced in using ChatGPT or similar tools. It’s unclear to me how I use it properly and generate the output that I want” (P6). The representative from the Bankers at the workshop stated that the risk continues to be an issue and likely will persist in the future. In an effort to minimize this risk, they have invested heavily in addressing the risk: “We are inviting all of our developers to an event soon, where we have also invited representatives from IBM to teach us more about using the tools and how to properly generate the output you want” (P7). The risk is viewed as something that needs to be addressed, as they are utilizing numerous resources to bring together their developers and IBM in a collaborative effort to reduce it.
Consequently, we have marked the risk in red to indicate its increased emphasis. In addition, the risk of Inadequate management of AI data was noted to cause hallucinations, causing the chaining risk Insufficient/ insecure AI output. They explained, "With these AI tools, you need to be aware that they often misrepresent information due to how they handle the data on which they base their output" (P3). This highlights the need to validate the output from generative AI to avoid integrating insecure code or solutions.
Lastly, the Bankers identified the isolated risk Falling behind without AI, which did not have any adjacent causes. They insisted on its inclusion in their causal map, as one manager explained in the initial interview: “I can’t imagine anything other than if generative AI tools become a market standard and we do not use them, we will fall behind our competitors” (P2). When generative AI tools were first introduced, the Bankers were quite concerned about missing potential improvements and being outclassed by their competitors if they did not adopt these tools as fast as possible. However, during the validation interview, they explained that this risk is not as concerning as it once was because: “What do we really miss out on if we do not use generative AI?” (P5).
A year after implementing generative AI tools at the Bankers, the hype surrounding productivity increases and promises of replacing developers has diminished. As the closing remark from one developer at the validation interview eloquently states: “The fear of the unknown has significantly decreased since we started using it” (P5). Generative AI is no longer perceived as a magical black box technology; like all technologies and people, for that matter, they have flaws that need to be considered and addressed.
In summary, the Bankers’ causal map revealed a circular chain of reasoning among Mistrust of AI, Insufficient validation of AI output, and Insufficient/insecure AI output. Meanwhile, the tail-end risks of AI replacing developers, Lack of control over, and the isolated risk of Falling behind without AI have diminished in emphasis since the initial adoption of generative AI tools. The latter risk has been lessened in emphasis as they opted to host generative AI internally, thereby gaining more control over the AI. In addition, the former two risks have decreased in emphasis as they now possess a greater understanding of the limitations of these tools.
4.2. The insurers.
They experimented with various tools, primarily using ChatGPT and GitHub Copilot. The Insurers discarded the risk of extensive AI maintenance, noting during the validation interview that it was more related to machine learning than generative AI. Additionally, they discarded insufficient AI performance, which was considered part of the Insufficient/insecure AI output risk.
The chaining risk of Insufficient/insecure AI output was explained to be caused by a few of the other risks in the causal map, such as the root risk,
Lacking data novelty. One participant noted, "Currently, you cannot use it for everything because it relies on older data”. (P8). However, Lacking data novelty was later discussed during the validation interview a year later, where it was reported to have decreased in emphasis: “ChatGPT and other similar tools are now using the latest data available, so it’s really not as important as before” (P9). The participants in the validation interview were hesitant to rule out the risk entirely, as they were uncertain whether this applied to all the tools they were using and if it was also applicable to topics they had not yet explored.
Additionally, the root risk Lacking AI competencies was similarly highlighted as a cause for Insufficient/insecure AI output: “I do not know enough about it to use it properly” (P10) and: “I believe that a lack of AI competencies can lead to Insufficient/insecure AI output because you need experience to use the tools optimally” (P14). Moreover, the chaining risk Insufficient validation of AI output was most frequently mentioned by the Insurers, explaining that it causes Insufficient/insecure AI output: “Some of this AI-generated code can be really good, but it can also be of very poor quality, so we have to ensure that we review the code or use our integrated testing tools” (P9).
Unreliable generative AI outputs breed mistrust, highlighting the need for validation, as shown in Fig. 2. This quote captures the link between Insufficient/insecure AI output and Mistrust of AI: “That would make it more trustworthy if you were guaranteed to get the right answer. It’s something that undermines my trust in it” (P12). The same is true for the chaining risk, Inadequate data management of AI, causing mistrust: “It’s a black box how it works in the machine room... Then, it becomes a matter of trust in those maintaining the AI that they don’t look at all our code. That’s where the problem arises if you send the AI all your code to optimize it or otherwise transfer sensitive data” (P9). The uncertainty of generative AI functionality and insecure technology use fosters mistrust.
Developers also fear being replaced by AI: “Because it's just AI that should handle those junior tasks like generating classes... In general, I think our work could become quite irrelevant... I'll be out of a job” (P8). During validation, it was noted that the risk of AI replacing developers had decreased as they recognized the limitations of generative AI for complex, industry-specific tasks: “As the concern for AI replacing developers has decreased in emphasis, the risk of being inappropriate for complex and industry-specific tasks has increased. They are not only opposites but also contradictory, which seems clear to me” (P14). Thus, the two risks are marked in red, indicating a shift in emphasis since the initial experimentation.
Moreover, the risk of Insufficient/insecure output of AI was highlighted as causing the tail-end risk, Inappropriate for complex/industry-specific tasks from various angles. For instance, a developer explained: “I was using Copilot and wanted to work on an insurance policy for a specific data object. If I wanted a method that removes, adds, adjusts, or handles errors for something that is suddenly domain-specific and highly tailored to our solution, it wasn’t much help” (P8). The AI also proved inappropriate for complex design tasks: “I would not use it to generate an entire design prototype. I might risk that it’s based on a different major insurance company’s design” (P11) or tasks specific to an insurance company: “I would be hesitant to use it for something related to insurance as it does not understand our business or customers” (P10).
Thus, the risk, Inappropriate for complex/industry-specific tasks, was a persistent concern for the Insurers both during and a year after their initial experimentation with generative AI tools.
Furthermore, the risks of Insufficient/insecure AI output and Inadequate data management of AI were raised as concerns that could cause the chaining risk Organizational sanctions should the generative AI wrongfully receive sensitive data or produce insecure code: “A directive or a reprimand from the Financial and Data authorities, which often leads to negative press coverage” (P13). They are concerned about the potential for sanctions, which could have a significant financial impact on them. Still, they are especially concerned about the poor publicity: “It can result in some severe fines, but the negative coverage really damages our image” (P13). Hence, we connected Organizational sanctions with the tail-end risk, Reputational repercussions.
Finally, the Insurers were also concerned about the isolated risk, Falling behind without AI: “If you look at the traction that AI-generated code has gained, we need to keep up. If we don’t, we will fall behind on many fronts. We’ll see what kind of impact it has, but if that impact is significant, we’ll be far behind our competitors” (P9). We were unable to identify any causes related to the risk, nor could our interviewees at the validation interview. However, the Insurers remarked at the validation interview that they are not as worried about this risk as they once were since the impact of AI has not been as substantial as they initially believed.
However, unlike the Bankers, they do not believe the risk has shifted in emphasis simply because generative AI tools have become very popular among developers: “We want to attract employees, and it is important to stay on top of the current trends. We are not concerned about the customer side, but it is crucial to appear attractive by demonstrating that we utilize AI to draw in new employees” (P9). Even though they are no longer worried about being outperformed by competitors or their ability to attract customers, it is vital not to fall behind in attracting new developers who want to engage with the latest technologies.
The Insurers’ causal maps revealed that the risks of AI replacing developers and Lacking data novelty have decreased in emphasis. The latter has become less significant since generative AI tools now have access to the newest data, and the former has diminished in relevance due to the awareness of the tools’ limitations, particularly regarding how Inappropriate for complex/industry-specific tasks these tools can be. The Insurer’s causal map also repeats the same chain of reasoning between the risks: Mistrusts of AI, Insufficient validation of AI output, and Insufficient/ insecure AI output, as the Bankers.
4.3. The automaters.
Throughout our investigation, the Automaters adopted ChatGPT and GitHub Copilot for their developers, maintaining a trial-and-error approach to utilizing these tools.
Several risks were identified, causing the chaining risk Insufficient/ insecure AI output. For example, the root risk, Lacking AI competencies, was cited: “If you don’t understand it, then you might get incorrect results that lead you to take actions that are not ideal” (P16). During the validation interview, the Automaters noted that this risk had gained emphasis, as generative AI tools are rapidly evolving, making it challenging to keep up with the tools’ development. They stated, “There are currently no reliable guidelines for using generative AI” (P16). The fast-paced evolution and introduction of new tools not only make it hard to stay current but also to find trustworthy guidelines for using generative AI in software development. As seen in Fig. 3, we colored the risk red accordingly.
Additionally, during the validation interview, they identified a causal relationship between Lacking AI competencies and the chaining risk, Insufficient validation of AI output, noting that a lack of competencies can similarly cause an undesirable AI output and inadequate validation of the output. Similarly, the risk of Insufficient validation of AI output could cause Insufficient/insecure AI output: “I believe the challenge in using AI is to avoid engaging blindly; we must validate both the input and the output. After all, garbage in means garbage out” (P15).
Considering the root risk Inadequate data management of AI, the Automaters expressed concerns about the ambiguity regarding how the generative AI tools store and manage data, as well as how this can cause untrustworthy responses from the AI. On the other end of the spectrum, the risk of Insufficient/insecure output of AI can cause the tail-end risk Inappropriate for complex/industry-specific tasks: “It makes sense to use AI for the simple tasks, in contrast to the complex tasks... where one must determine how things are supposed to fit together, and other aspects that are even more complex” (P17). Their generative AI tools are particularly unsuitable for operating systems development for Linux: “Currently, there are no AI tools appropriate for operating systems development... especially regarding the hardware platform and managing the Wi-Fi component” (P17).
As part of the validation interview, they noted that they have become increasingly aware of the inefficiencies of their generative tools when applied to complex tasks. They are now more hesitant to use them for complex and industry-specific tasks, such as operating systems development, as they frequently find themselves wasting time on generative AI for these responsibilities. Therefore, we found it fitting to color-code the risk to reflect this increased emphasis.
Like the other two cases, the Automaters displayed mistrust stemming from the risks of Inadequate data management of AI and Insufficient/ insecure AI output. They highlight the data management aspect: “The most critical aspect is data security and avoiding leaks of our source code, because if you have that, you can also break into it” (P16). Another participant elaborates: “We do not know what happens as soon as you install Copilot. What do they have access to? What data are they stealing?” (P15). They also highlight concerns about AI’s output: “You might have a simple while loop that you aren't allowed to use because someone has a license on it. We then need to ensure that we don’t just take it” (P15). This mistrust of AI leads them to validate the output before integrating the code.
In line with the risk of AI replacing developers, they express: “The fact that I can provide it with functional requirements, and it can generate something useful means that if it becomes more advanced, we risk it replacing developers” (P16). Amid the validation interview, they noted that the risk had significantly decreased since they began using the tools: “It is no longer much of a concern for us” (P18). The risk became less significant as they grew more familiar with the tool’s limitations, particularly for complex and industry-specific tasks.
Moreover, their concern about the root risk, Lack of AI policies, indicates an underlying Mistrust of AI, as explained: “We currently lack policies internally that dictate how we can use it. It’s quite a barrier for us that we do not have them” (P16). In continuation of the former statement, the participant explains how their Lack of policies and Lacking AI competencies can cause the risk tail-end risk, Costly implementation of AI: “I am skeptical of whether we as an organization are equipped to use it efficiently, so that it creates value, and we do not waste money needlessly” (P16). They highlighted the risk of Costly implementation of AI during the validation interview and reiterated at the workshop that its emphasis has substantially increased.
In the validation interview, one participant remarked, “It’s important for us that it meets our expectations regarding the economic aspects of using it” (P15). This was also restated at the workshop, where a critical aspect for them is ensuring their investment in the technology aligns with their economic expectations for its implementation. There was considerable hype surrounding generative AI when it first launched, and it appears to have fallen short of those early expectations. During our initial interviews, they expressed a concern for the isolated risk Falling behind without AI: “I’m afraid we’ll be left behind if we don’t start implementing these tools. Denmark has high salary costs. We must compete on efficiency and job satisfaction as well" (P15).
However, a year later, in the validation interview, they indicated that the risk had diminished in emphasis: “The hype around generative AI has significantly lessened. We still need to monitor developments in the AI field and acknowledge that our customers are urging us to do so” (P15). The risk of Falling behind without AI was a concern for them early in our examination, as they worried about their ability to compete with others in terms of efficiency. However, as the hype surrounding generative AI has stagnated and they have become more aware of its limitations, they are now more focused on aligning their economic expectations with the implementation of generative AI tools to justify their investment.
The Automaters’ causal map disclosed that the root risk Lacking AI competencies has gained emphasis, as they struggle to keep up with the rapidly changing generative AI landscape. After a year of experimentation, they lack reliable guidelines for using these tools. Additionally, the tail-end risk, Costly implementation of AI, is now more significant as they assess the balance between efficiency and cost to justify their investment. Further, the risk of Inappropriate for complex/industry-specific tasks is increasingly recognized due to their limitations, especially in areas like operating systems development. The Automaters’ causal map reflects the same concerns as the Insurers and Bankers regarding risks such as Mistrust of AI, Insufficient validation, and Insufficient/insecure AI output.
5. Combined causal map of the three cases.
In Fig. 4, we have aggregated all the identified risks and causes into a comprehensive causal map encompassing all cases. This aggregated map presents comparisons among the three company cases and should not be seen as an attempt at generalization. The thick arrows indicate causal explanations common to all cases, while the thin arrows represent causal explanations derived from only one or two cases. The map illustrates the root cause explanations used in all three software organizations. These explanations include a consistent emphasis on inadequate data management of AI, Lack of AI policies, and a decreasing emphasis (marked in blue) on AI replacing developers. Meanwhile, the Bankers and the Automaters recognized an emerging emphasis (marked in red) on Lacking AI competencies, while the Insurers maintained their emphasis.
The root risk, Lacking data novelty, with a decreasing emphasis, was uniquely identified by the Insurers.
The map illustrates a circular reasoning loop regarding the risks associated with AI, which all three software organizations repeat. This reasoning relates to the Mistrust of AI. The three organizations infer this Mistrust of AI from the risk of Inadequate data management of AI, Lack of AI policies, Insufficient/insecure AI output, and AI replacing developers. This mistrust compels software organizations to validate the quality and credibility of AI outputs, preventing the integration of low-quality or insecure code into their codebase. If validation is not conducted or is inadequate, it causes Insufficient/insecure AI output due to the inherent risk that the AI's output may be flawed or insecure. The Insufficient/insecure AI output, in turn, causes Mistrust of AI, thereby perpetuating the circular causal chain of reasoning until the AI output has been adequately validated.
However, they cannot be certain when AI-generated code is sufficiently validated, creating this circular reasoning loop.
Tail-end risks that do not serve as the reason for other risks (costly implementation of AI, lack of control over AI, lack of transparency of AI, reputational repercussions, and inappropriate for complex/industry-specific tasks) were revealed by the causal maps. All three organizations identified two key risks: Inappropriate for complex/industry-specific tasks and Lack of control over AI. While all organizations had increased their emphasis on Inappropriate for complex/industry-specific tasks, only the Bankers decreased their emphasis on Lack of control over AI. Only the Automaters recognized the risk, Costly implementation of AI, and did so with an emerging emphasis on the risk. The Insurers were the only case to acknowledge the risk Reputational repercussions, and the risk had a consistent emphasis.
The tail-end risks reveal distinctive differences between the cases as the Bankers are uniquely preoccupied with controlling AI, the Insurers are uniquely concerned with their image, and the Automaters are succinctly worried about the cost-efficiency of adopting generative AI in software development.
Finally, Fig. 4 highlights one isolated risk: Falling behind without AI, which is not involved with any other risks in the map. All three software organizations recognized the risk, but the Bankers and the Automaters share a declining emphasis, while the Insurers maintained a consistent emphasis on it.
6. Discussion.
6.1. Contributions to research.
We contribute to the research on generative AI in software development by showing how causal mapping can reveal explanations of risks that three software organizations associate with generative AI. Causal mapping has been shown to be useful for causally explaining project risks, barriers to knowledge sharing, and perceptions of technology adoption. Our study extends this body of research by demonstrating that causal mapping is also suitable for illustrating the causality of risks associated with generative AI and the perceived evolution of these risks. Moreover, we contribute to the literature on the concept of causality of risks by introducing the terminology of root risks, chaining risks, and tail-end risks to categorize the causality of risks. Specifically, we expand Boholm and Corvellec’s relational theory on risk by stating that a risk is interpreted within a specific context.
Additionally, we extend their concept of a risk consisting of an object at risk and a risk object to also apply to the relationships between risks. Our causal mapping with the three software organizations reveals substantial similarities in the risks they identified and their causal explanations. These similarities are common among software organizations, with variations in context based on industry, size, customer segment, and software development process. This similarity was particularly evident in the circular reasoning involving Mistrust of AI, Insufficient validation of AI output, and Insufficient AI output. All three case organizations’ causal maps (c.f. Figs. 1, 2, 3) show this circular explanation with these three chaining risks.
As generative AI continues to evolve rapidly, we also investigated the risk explanations of three software organizations a year after adopting generative AI tools. The three software organizations then emphasized the identified tail-end risks differently. The Bankers reduced their emphasis on the Lack of control over AI but increased their focus on Lacking AI competencies to accommodate their Scaled Agile Framework. The Automaters uniquely demonstrated an emerging concern for the risk of Costly AI implementation, which reflects their concern for cost efficiency as a smaller software organization with limited opportunities for economy-of-scale AI tool use. Finally, the Insurers distinctly recognized and consistently emphasized the risk of Reputational repercussions, which reflect the organization’s need for a trustworthy public image as a retail insurance provider.
As such, these tail-end risks highlight how the different organizational contexts of the three cases influence their explanations of the risks associated with using generative AI for software development.
6.2. Implications for practice.
Our findings present some implications for practice. First, causal mapping may be valuable not only as an analytical tool for illustrating the causes of the risks associated with generative AI in software development. Similar research highlights that causal mapping can serve as a valuable aid in decision-making. By visualizing the causes and relationships of AI-related risks, causal maps can promote internal alignment, guide prioritization, and assist managers in communicating risks to upper management while deciding on appropriate mitigation strategies. The case organizations, all at the early stages of adoption, expressed uncertainty about managing the emerging risks, underscoring the need for practical frameworks to help them proactively oversee their AI usage.
For software organizations looking to create their own causal maps, we offer these practical step-by-step guidelines:
1. Define the decision focus.
Frame a concrete question the map should inform (e.g., “Which gen-AI risks most threaten our delivery pipeline in the next two quarters, and where should we intervene first?”). Set scope (product (s), teams, lifecycle phases) and timebox the exercise.
2. Assemble a cross-functional group.
Include developers, architects, operations/SRE, product, UX, data/ ML engineers, and IT managers. Assign a facilitator (neutral), a scribe (captures statements verbatim), and a timekeeper.
3. Identify risks.
Pull from interviews, incident/postmortem logs, threat models, pen-test findings, and developer feedback. Use a uniform sentence pattern to keep inputs crisp:
“Because [cause], the risk of [risk object] may harm [object at risk] by [impact].”
4. Align on mapping syntax.
Nodes = risks.
Arrows = perceived causation (“A leads to B”).
Tags/Colors = Leavitt’s components (Actors, Tasks, Structure, Technology) to maintain a sociotechnical map. Optionally mark the risk object (source of risk) vs the object at risk (what holds value and could be harmed).
5. Run a fast-mapping workshop.
Start with silent sticky-note brainstorming of risks, clustering by theme and Leavitt component, and then draw arrows to capture “what causes what.” Press for mechanisms (“how, exactly?”) rather than correlations.
6. Classify risk types to guide attention.
Identify root risks (only cause others), tail-end risks (only outcomes), chaining risks (both cause and are caused), and isolated risks (neither cause nor are caused—still important if high impact).
7. Detect loops and propagation paths.
Circle feedback loops (e.g., “model misuse → incident load → rushed patches → weaker reviews → more misuse”). These often signal compounding dynamics that merit priority.
8. Revisit the causal map.
We recommend that software organizations revisit the first causal map after three months in a workshop to accommodate any changes since its creation; if this map is stable, it should be revisited yearly instead.
Second, while the nature of the risks remained consistent, differences arose in how organizations perceived and prioritized them. Given the rapid pace of evolution in generative AI technologies, software organizations need to adopt a mindset of continuous reassessment. Risks will likely evolve swiftly, and what is considered low-impact today may become high-priority tomorrow. Regarding our three cases, the risks of AI replacing developers, Falling behind without AI, Lacking data novelty, and Lack of control over AI have decreased in emphasis. On the other hand, the three risks, Lacking AI competencies, Inappropriate for complex/ industry-specific tasks, and Costly implementation of AI, have increased significantly since adoption.
Ongoing monitoring and reflection will be necessary to maintain a current and practical risk management approach for managing the risks associated with generative AI in software development. Software organizations planning on adopting generative AI tools for software development should be attentive to the root risks identified by all our cases, such as Inadequate data management of AI data, as this risk can cause a variety of other risks. The Bankers attempted to address this risk by hosting generative AI internally in their organization, gaining more control over the data when interacting with AI.
Third, the circular reasoning loop between Inadequate data management, Mistrust of AI, and Inadequate validation of AI output signifies a mutual exacerbation. The introduction of an external tool for code generation could worsen the practice of validating code in software organizations, making it even more challenging to ensure that the code has been adequately validated. To break this loop, the Bankers have automated static analysis tools, such as SonarQube, at key CI/CD integration points to detect insecure AI-generated code before it is merged. The Automaters developed AI-specific code reviews, such as verifying prompt handling, ensuring no sensitive information is exposed, and flagging insecure library usage.
6.3. Threats to validity.
Drawing on Runeson et al.'s recommendations on case study methodology, we evaluate the study’s validity across four dimensions: construct validity, internal validity, external validity, and reliability.
6.4. Construct validity.
Construct validity concerns whether we accurately captured the concepts we studied. In this case, the risks associated with generative AI in software development. The concept of risk may have been interpreted differently by participants, as it was not explicitly defined in strict terms. This could lead to inconsistencies in what respondents understood as a risk, affecting the comparability of their responses. Since data collection relied heavily on semi-structured interviews, there is a risk that researchers might have misinterpreted participants' intentions or inadvertently assigned unintended meanings to their statements.
Additionally, we validated the causal maps through follow-up interviews with company representatives and a cross-case workshop, which enhanced the alignment between the collected data and our interpretations. Nonetheless, construct validity could still be threatened by participants' own evolving understanding of generative AI over time, as well as our interpretation of what constitutes a risk. Lastly, the causal maps derived from the initial interviews were presented to participants approximately one year after the original data collection. While this enabled us to examine longitudinal developments, it may have compromised the accuracy of feedback on the original representations due to memory decay or evolving organizational contexts.
Presenting the maps earlier, closer to the initial interviews, along with a year after the first examination, could have strengthened triangulation and ensured a more immediate confirmation of the concepts and causal relationships as understood by participants at that time.
6.5. Internal validity.
Internal validity pertains to the accuracy of identified causal relationships and is especially pertinent given our reliance on causal mapping. Studying multiple companies allowed for cross-case comparison, but it may have limited the depth of insight into the contextual nuances of each case. A single, in-depth case could have offered more details on the causal chains of specific risks. Although causal mapping helps to visualize risk causality, it is subject to interpretation by the researchers, who may have imposed structure on relationships that were less explicit in the raw data. While the causal maps were constructed from participants’ statements and refined through validation interviews, these maps are ultimately interpretations rather than objective truths.
Lastly, causal mapping presented interpretive challenges, especially concerning the directionality of relationships between concepts. In several cases, determining whether a specific risk caused another or was instead a result of it proved ambiguous, resembling a “chicken or the egg” dilemma. Although these decisions were based on participant narratives and researcher consensus, the risk of misinterpreting causal direction remains a limitation, particularly considering the controversies of perceived cause–and–effect in organizational contexts.
6.6. External validity.
External validity relates to the generalizability of the findings. Although the case companies were selected for their size, industry, and variation in customer base, all three are Danish software organizations in the early stages of adopting generative AI. Therefore, the findings may not apply to companies in different national, regulatory, or technological contexts or those with more advanced AI integration. Nevertheless, the relative consistency in the types of risks identified across the three organizations indicates that some findings might have relevance beyond the immediate context of the cases. Our exclusive use of qualitative methods means that replication may be challenging, as another researcher might not elicit the same responses or interpret the data in the same manner. The absence of quantitative measures limits the ability to statistically validate the findings.
We prioritized in-depth detail and richness in our three cases over generalizability.
6.7. Reliability.
Reliability refers to the extent to which the study can be replicated with similar results. The coding, causal mapping, and cross-case analysis depended on the research team's interpretations, potentially threatening reliability. Despite discussions to reach a consensus, researcher bias could still influence the results. We adopted specific strategies to enhance reliability: all interviews were recorded and transcribed, coding was performed and collaboratively reviewed by two authors, and the analysis was thoroughly documented. Furthermore, member checking was utilized through map validation and a follow-up workshop. However, the qualitative nature of the research and the interpretive aspects of causal mapping still pose challenges to replicability, particularly considering the evolving landscape of generative AI.
Therefore, we chose to revisit the cases one year after adopting generative AI tools to determine if anything had changed since the cases’ early adoption.
6.8. Future research.
This study focused on understanding how software organizations explain the risks associated with generative AI through a multi-case qualitative approach. Our results revealed that the rapid evolution of generative AI has implications for transforming these risks in software development. Therefore, future research could build on our insights and explore how these risks evolve over time by conducting longitudinal studies that trace organizational responses as generative AI becomes more integrated into development workflows. This should be viewed in consideration of the future evolution of AI that software companies are exploring. The expectation for the subsequent significant development in AI is that software companies will adopt Agentic AI.
Future research must consider this new technology, as it may introduce additional implications for practice, particularly regarding the lack of coding control and the actions taken by agentic AI.
Additionally, while this study identified risks qualitatively, follow-up research could develop and apply quantitative instruments, such as structured surveys or risk assessment tools, to validate and generalize the findings across a wider sample of organizations. By grouping the risks identified in our study into categories based on similar characteristics, we encourage further research to employ statistical methods, including Bayesian networks or structural equation modeling, to explore predictions or correlations between the risks associated with using generative AI in software development. Finally, little is known about potentially breaking out of the causal reasoning loop, and the literature on risk mitigation strategies for addressing risks of using generative AI is scarce.
Thus, we call for more actionable research on risk mitigation strategies of generative AI risks in software development, for instance, through action research to test mitigation strategies in practice.
7. Conclusion.
We employed a multi-case study to examine software organizations’ explanations of the risks of using generative AI in software development and how these evolve. The case study was a collaborative effort involving three Danish software organizations, and we identified and assessed the risks and their underlying causes through interviews, validation sessions, and workshops. We employed the causal mapping technique to illustrate the causal reasoning behind these risks. Our study contributes to software engineering research by:
• Demonstrating that causal mapping effectively reveals the causation of risks linked to generative AI and the perceived evolution of these risks.
• Presenting a novel terminology of root risks, chaining risks, and tail-end risks to categorize the causality of risks. • Showing substantial similarities in the risks they identified and their causal explanations. This similarity was particularly evident in the circular reasoning loop involving Mistrust of AI, Insufficient validation of AI output, and Insufficient AI output.
• Clarifying how generative risks can change in importance over time and thus need to be revisited.
• We introduce guidelines for software teams and organizations to develop their own causal maps.
• Highlighting software organizations need for actionable frameworks and methods to address generative AI risks. This is also evident in the lack of research in this area.
CRediT authorship contribution statement
David Kinnberg Hein: Writing – review & editing, Writing – original draft, Visualization, Validation, Software, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. John Stouby Persson: Writing – review & editing, Validation, Supervision, Resources, Project administration, Methodology, Funding acquisition, Conceptualization. Victor Vadmand Jensen: Writing – review & editing, Validation, Formal analysis, Data curation. Anders Rysholt Bruun: Writing – review & editing, Validation, Resources, Funding acquisition, Conceptualization. Martin Gilje Jaatun: Writing – review & editing, Validation.
Declaration of competing interest
John Stoby Persson reports financial support was provided by Digital Lead, Denmark’s national cluster for digital technologies. John Stouby Persson reports a relationship with Digital Lead that includes: funding grants. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Acknowledgements.
The authors of this paper have used ChatGPT-4 solely as a sparring partner to generate ideas and provide feedback, thereby enhancing the article's linguistic quality while ensuring overall readability. We reviewed and edited the content as needed, and we take full responsibility for the content of the submitted article.
This article is the result of collaboration with the three software development organizations and funding from DigitalLead, Denmark's national cluster for digital technologies.
Appendix A. Sample interview guide
● Could you describe your role in the organization?
● How long have you been working with software development in this organization?
● Are you currently using any generative AI tools (e.g., ChatGPT, GitHub Copilot)? If so, which ones, and in what capacity?
● From your perspective, what are the main benefits of using generative AI in software development?
● What concerns or risks do you associate with using generative AI tools?
◦ Why do you see that as a risk?
● Are there particular situations where using generative AI might cause problems?
The authors declare the following financial interests/personal relationships which may be considered as potential competing interests:
◦ Why?
● Have you encountered any unintended outcomes when using generative AI (e.g., incorrect output, security concerns, legal issues)? ◦ Do you know why it happened?
● Are there risks that impact certain parts of your workflow or organization more than others (e.g., security, productivity, collaboration)?
◦ Why?
Appendix B. Validation interview guide
• Do you have any questions about the causal map?
• Do the risks in the map explain your concerns?
• Are there any risks missing in the causal map?
• Are there any risks that should not be in the map?
• Which are the most important risks for you?
• Do you think the arrows between the risks explain the cohesion among the
• risks in the map?
• Are there any arrows missing among any of the risks?
• Are there any arrows that you think should be added?
Data availability.
Data will be made available on request.