Between Policy and Practice: GenAI Adoption in Agile Software Development Teams
1 More Paper · Full Reading

About this paper
A full audio edition of this paper.
Authors: M. Neumann, L. Bischof, N.E. Hinz, A. Altun, L. Stockmann, D. Schrader, A.C. Ahaus, E.C. Demirci, B. Gabel, M. Rauschenberger, P. Diebold, H. Fritzemeier, A. Przybyłek
Publication date: 2026
Read the paper: https://doi.org/10.1007/978-3-032-22375-3_18
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “Between Policy and Practice: GenAI Adoption in Agile Software Development Teams,” by M. Neumann and colleagues. Published in 2026.
Michael Neumann1, Lasse Bischof2, Nic Elias Hinz1, Abdullah Altun1, Luca Stockmann1, Dennis Schrader1, Ana Carolina Ahaus1, Erim Can Demirci1, Benjamin Gabel1, Maria Rauschenberger2, Philipp Diebold3,4(B), Henning Fritzemeier5 lek6 1 University of Applied Sciences and Arts Hannover, Hannover, Germany the email address 2 University of Applied Sciences Emden/Leer, Emden/Leer, Germany the email address 3 Bagilstein GmbH, Mainz, Germany the email address 4 IU International University, Erfurt, Germany the email address 5 Volkswagen AG, Wolfsburg, Germany the email address 6 University of Galway, Galw ay, Ireland
Abstract. Context: The rapid emergence of generative AI (GenAI) tools has begun to reshape software engineering practice. Yet, their adoption within agile environments remains underexplored. Objective: This study investigates how agile practitioners adopt GenAI tools in real-world organizational contexts, focusing on regulatory conditions, role-specific use cases, benefits, and barriers. Method: An exploratory multiple case study was conducted in three German organizations, involving 17 semi-structured interviews and document analysis. A cross-case thematic analysis was applied to identify GenAI adoption patterns. Results: Find-ings reveal that GenAI is primarily used for creative tasks, documenta-tion, and code assistance. Benefits include efficiency gains and enhanced creativity, while barriers relate to data privacy, validation effort, and lack of governance.
Using the Technology-Organization-Environment (TOE) framework, we find that these b arriers stem from misalignments across the three dimensions. Regulatory pressures are often translated into poli-cies without accounting for actual usage patterns or organizational con-straints, leading to systematic policy-practice gaps and shadow IT behav-ior. Conclusion: GenAI offers significant potential to augment agile roles but requires alignment across TOE dimensions, including pragmatic poli-cies, data protection measures, and user training to ensure responsible and effective integration.
The rapid advancement of generative artificial intelligence (GenAI) has catalyzed profound transformation across diverse economic and technological domains. While impacts are evident in institutional contexts such as higher education and public administration, the proliferation of GenAI since the release of ChatGPT in November 2022 has significantly reshaped daily work across indus-tries.
Currently, GenAI tools such as ChatGPT or GitHub Copilot are perv asive in software engineering, fundamentally altering how software developmen t teams operate. Professionals now routinely apply GenAI for tasks ranging from code generation and test case creation to requirements elic-itation and analysis. While the literature highlights benefits such as increased team performance, it simultaneously warns of risks, e.g., on source code quality. Consequently, scholars emphasize the importance of human-integration, exemplified by the ethical guidelines proposed in the Copenhagen Manifesto.
Within agile software development (ASD), GenAI tools promise to auto-mate routine tasks, foster creativity, and enhance the productivity o f software developers and project stakeholders. Agile methods, characterized by itera-tive development cycles, cross-functional collaboration, and continuous feedback, provide a conducive environment for the p iloting and integration of GenAI tools into established workflows. However, deploying these tools also introduces a range of new challenges. According to Bahi et al., practitioners are confronted with critical questions such as: “Is the use of such tools permissible?” and “Can sensitive data be safely disclosed to GenAI services?”.
As agile methods focus strongly on social interaction and communication, one may assume that various agile related practices with a strong emphasis on social interaction such as pair programming are affected when specific tasks are performed with GenAI support instead of humans. Despite this potential disruption, empirical research on how agile teams specifically adopt GenAI to ols in their practice remains underrepresented in the literature.
Thus, in this study, we address the follo wing research questions:
– RQ 1: What organizational and regulatory conditions shape GenAI adoption in practice?
– RQ 2: Which use cases of GenAI tools are adopted in agile software devel-opment teams?
– RQ 3: What benefits and barriers do agile team members associate with GenAI adoption?
The paper is structured as follows: Sect. 2 reviews related work and outlines our contributions. In Sect. 3, we explain how we applied the multiple case study. Next, we present our results and answer the research questions in Sect. 4. The paper concludes in Sect. 6.
2 Literature Review
The study most closely related to ours is by Kemell et al., who conducted a multiple case study across seven companies to examine GenAI adoption, identify-ing relevant use cases and challenges. They identified six organizational barriers, including data privacy and legacy system concerns, as well as issues related to the black-box nature of LLMs. From a user perspective, they reported four further challenges, notably the complexity of learning to apply GenAI tools, particu-larly regarding prompt formulation in daily work. Our work builds on this by shifting the analytical lens from their broad categorization of “individual users”, primarily focused on coding tasks, to a granular analysis of agile-specific roles. This reveals distinct adoption patterns where GenAI extends beyond technical assistance to support agile ceremonies and soft skills.
Furthermore, while Kemell et al. document data privacy and legislative concerns as adoption barriers, we extend this by examining their behavioral consequences, showing how rigid poli-cies drive practitioners to circumvent compliance to maintain agile velocity. Con-sequently, we reframe validation not merely as a user capability challenge, but as a systematic process bottleneck.
Karlovs-Karlovskis corroborates the technical emphasis in current research: among 117 studies reviewed, about 65% focus on GenAI-based code generation, revealing major gaps in empirical research on real-world integration and a lack of studies in other domains with potential for cost reduction, innova-tion, and optimization.
A recent research agenda by Nguyen-Duc et al. provides broader theoret-ical grounding, identifying 78 open research questions across 11 knowledge areas. While implementation and testing dominate current discourse, significant poten-tial exists in requirements engineering, software design, and engineering manage-ment. The authors note persistent challenges including reliability of GenAI out-puts, data privacy and accessibility, transparency in AI-assisted decision-making, and s ustainability. Critically, they observe that managerial, organizational, and socio-technical dimensions of GenAI adoption remain significantly underrepre-sented compared to purely technical considerations, a gap our study addresses within agile contexts.
Overall, existing research on GenAI in software engineering especially within ASD remains limited. While some studies (e.g., Ulfsnes et al.; Coutinho et al. ) highlight productivity gains and tool versatility, they also point to persistent issues in knowledge sharing, reliability, and security. Bahi et al. similarly show that GenAI can support Scrum practices but emphasize ongoing challenges across the development lifecycle, underscoring the paucity of empirical evidence from real-world agile environments.
Recent XP conference editions also hosted two workshops on GenAI and Agile (2024; 2025), producing exploratory papers on GenAI use in specific agile practices, such as Test-Driven Development, and broader topics like resp onsi-bility and teamwork effects. While informative, their exploratory nature and workshop format limit their applicability to our research questions.
To address the identified gaps, our study examines GenAI adoption within agile software teams. It explores how GenAI is applied in practice, the organizational implications of its adoption, and relevant use cases along with associated benefits and barriers. By grounding the investigation in a real-world agile set-ting, this study provides empirical insights into how GenAI is reshaping software development dynamics in mid-sized enterprises.
3 Research Design
Guided by Runeson and H ̈ost’s guidelines, this study adopts an exploratory qualitative multiple-case design to investigate how GenAI is embedded, applied, and experienced within the various roles operating in agile software-development teams. Case studies are well-suited for examining modern software engineering practices in real-world contexts, especially when the phenomenon cannot be easily separated from its context. For this study, the units of analysis are the employees operating in the agile software development teams.
3.1 Case Selection and Context
Our study was conducted across three German organizations, selected to capture diversity in roles, agile methods, experience levels, and expertise. All identifying information—including company, site, and team names has been anonymized for confidentiality. The three cases are described below:
Case 1: Dinoco is one of the world’s largest automobile manufacturers, operat-ing more than nine vehicle brands and producing automobiles, motorcycles, and trucks worldwide. Our study took place at one of its German software develop-ment sites, the Dinoco Software Development Expertise Site (DSDES), which spans multiple locations and employs a round 450 people. As part of Dinoco’s digitization strategy, DSDES develops strategic software products such as cock-pit software and over-the-air update processes across more than ten agile teams. These teams follow eXtreme Programming principles within the Scaled Agile Framework (SAFe).
Case 2: Gray Matter Technologies (GMT) is a medium-sized company in Germany and Austria with around 150 employees across three sites. GMT devel-ops software products for external clients, including an ERP system and web shop applications. Its software development division consists of four agile teams, each dedicated to specific customer projects and working with an adapted Scrum approach. GMT’s clients range from large retail and manufacturing companies to medium-sized craft businesses.
Case 3: Insight Inc. is a small, fully remote consulting company in Germany with eight employees. It provides agile coaching and training and offers interim roles in which staff embed within client teams for six months to two years, typically as Scrum Masters and occasionally as Product Owners. Insight Inc.’s clients span various industries; for the employees involved in this study, the relevant clients were in the banking sector.
3.2 Data Collection
Runeson & H ̈ost emphasize that case studies must combine direct, indirect, and independent data sources so that weaknesses inherent in a single method can b e counterbalanced through data source triangulation. Our data collection comprised two primary data sources:
Semi-structured Interviews: We conducted 17 semi-structured interviews across all three cases; interviewee profiles are summarized in Table 1. The inter-views were conducted using an interview guideline which we created based on the Goal-Question-Metric (GQM) approach. The interview guide consists of four core clusters considering the personal background of the interviewee and the three main themes among the research questions: GenAI-related contextual information, use cases of GenAI adoption, and barriers/benefits of the daily use of GenAI. The interview guide is available in our research protocol. All inter-views were held online using Microsoft Teams between March and May 2025 and have been recorded. The interviews lasted around 30 min for Dinoco, between 30 and 45 min for Insight Inc. and Gray Matter Technologies.
To increase the qual-ity, we performed a pre-test to verify the structure of the interview guide. The interviews were conducted by two researchers, while one researcher performed the interview based on the interview guide and the other researcher took an observing role during the interview and made notes.
Document Analysis: We supplemented the interviews with analysis of inter-nal documents, particularly, GenAI usage guidelines for software development processes. Notably, only GMT maintains a formalized internal GenAI guideline; the other organizations provide informal guidance, functioning as living artifacts (Fig. 1).
3.3 Data Extraction and Analysis
Thematic Analysis and Coding: Recorded interviews were transcribed using Microsoft Teams and Whisper AI. The resulting transcripts were subsequently reviewed and edited by the second and third authors to ensure a ccuracy and completeness. These validated transcripts formed the basis for thematic anal-ysis following Braun and Clarke to identify adoption patterns. The process involved six phases: familiarization, initial coding, theme i dentification, theme review, theme definition, and reporting.
In order to establish a reliable and consistent coding process, multiple authors (authors 24 for the Dinoco case; 57 for GMT; and author 8 for Insight Inc.) independently coded a subset of the material. Their coding outputs were then compared and discussed with the first author. Any divergences were resolved through collaborative deliberation, fostering a shared understanding of the data and enhancing the stability of the thematic structure. This iterative procedure ensured that the resulting themes were both empirically grounded and aligned with the overall aims of the study.
Cross-Case Analysis: To strengthen the results of the multiple-case study, we conducted a cross-case analysis following Yin. For each case, the data were transferred into Microsoft Excel and examined for thematic overlaps, using separate sheets for contextual information, role-specific use cases, and identi-fied barriers and benefits. We first compared context-related information (e.g., licensed GenAI tools or usage guidelines) across all three cases. Second, we ana-lyzed the use cases, all of which appeared in at least two cases, meaning no borderline cases occurred. Third, we compared categories and codes from the interviews to identify overlaps in r eported barriers and benefits. The analysis was carried out by the first author and validated by one additional researcher (the second author), with no discrepancies identified.
3.4 Threats to Validity
As with every study, ours has specific limitations that are inherent to the chosen research approach. Although we conducted our study in accordance with the relevant guidelines, we acknowledge the limitations and explain the actions we took to mitigate their impact. To do so, we used the Threats to Validity schema according to Wohlin et al. and Runeson and H ̈ost. Due to space limi-tations, we focus on the most important aspects here. A detailed explanation of the limitations of this study can be found in our research protocol.
Construct Validity: Our interview participants may have interpreted specific terms related to GenAI (LLM or GPT) differently or assigned varying meanings to key concepts such as prompting strategies and “guidelines”. To address this, we provided clarification and used “such as Github CoPilot” or asked partici-pants for specific examples to clarify their understanding and ensure consistency across responses.
Internal Validity: We took specific actions to follow the same approach in conducting the interviews to mitigate potential bias. First, the interview guide-line was composed of neutral, non-leading questions. Additionally, the interviews were semi-structured, allowing us to explore topics in depth based on the inter-viewee’s responses. E ach interview was conducted by two researchers, ensuring rigor. Furthermore, we used different data sources to strengthen the internal validity of our findings.
External Validity: Our findings are constrained by the context of three Ger-man organizations operating under similar regulatory conditions. While the cases differ in size and structure, they do not represent the full range of software development environments. To support transferability, we provided detailed case descriptions and contextual information, enabling readers to assess the applica-bility of findings to their own settings.
4 Results 4.1 Contextual Information on GenAI Adoption
This section addresses RQ1: What organizational and regulatory conditions shap e GenAI adoption in practice?
Organizational conditions cover both licensing of GenAI tools and train-ing or coaching for professional use of GenAI in practice. Notably, none of the three cases offers training or coaching for GenAI usage in their software devel-opment teams. Interestingly, at GMT, while company policy mandates at least two annual training sessions on AI applications and legal frameworks, intervie-wees were either unaware of these offerings or reported that the training f ocused exclusively on explaining the AI guideline itself rather than developing practi-cal skills. This reveals a substantial gap between formal policy and practitioner needs.
Regarding licensing, an overview is provided in Table 2. The governance mechanisms for tool selection differ markedly across cases. At Dinoco, licensing decisions are led by compliance and top management, following a risk-averse strategy that favors established vendors; for instance, Co-Pilot was licensed through existing Office 365 agreements. GMT, by contrast, follows a formal company guideline that classifies tools according to the EU AI Act’s risk pyra-mid (Minimal, Limited, High Risk), maintains an official whitelist, requires management approval, and involves an internal expert group for validating newly requested tools. Key criteria include customer data protection, legal compliance, data-leakage prevention, and control o ver tools transmitting usage data exter-nally. At Insight Inc., no comparable governance structure exists.
A gap exists between sanctioned and actual tool usage. Tools such as DeepL Write, Gemini, and NotebookLM are not licensed in any of the three cases, yet unsanctioned tool usage is common. At both Dinoco and GMT, ChatGPT func-tions as a form of “shadow IT,” despite lacking approval. At Dinoco, an informal agreement permits developers to use ChatGPT during “private time.” At GMT, survey and interview data confirm ChatGPT as the most frequently used tool. In contrast, Insight Inc. licenses ChatGPT, though employees occasionally rely on private accounts for sensitive topics to ensure data separation.
Internal guidelines for GenAI use constitute a critical regulatory dimension. At Dinoco, the use of ChatGPT is not explicitly regulated, whereas clear policies exist for GitHub Co-Pilot and Microsoft Co-Pilot. Nevertheless, several intervie-wees reported using the Pro version of ChatGPT for work-related tasks. The specific use of these tools varies considerably across roles: while Agile Coaches and Product Owners use ChatGPT frequently, developers rely on it less often. Due to the integration of Co-Pilot tools into the IDE, their use is more seamless and compliance can be more easily monitored. Furthermore, regulatory aspects also include data protection, which is explicitly governed through agreements between Dinoco and the respective tool providers. For instance, P03 stated: “Dinoco [...] has made an agreement with Copilot that the data may not be used to evaluate or train the model.”
At GMT, several interviewees highlighted concerns that certain GenAI tools or models may not ensure sufficient data sovereignty, as user inputs and the information they contain can be used for model training. Given the opaque nature of AI model operations, it cannot always be determined how such data are further processed. Consequently, GenAI tools cannot be applied to code seg-ments involving sensitive information, such as encryption keys or personal data. Without formal guidelines, users “[...] often rely on personal judgement, experi-ence or informal principles ” (P11). While this approach allows for flexibility, it also entails risks related to consistency, legal certainty, and accountability. P14 emphasized that formal frameworks are frequently missing within organizations, resulting in uncertainty, inconsistent implementation, and highly individualized approaches.
At the same time, there is a strong demand for clear, pragmatic and legally sound guidance. Such guidelines are particularly essential for fostering acceptance of the technology among less experienced users.
Concluding our findings for RQ1, we observe that while ChatGPT emerges as the most frequently used tool across all three cases, the organizational conditions governing GenAI adoption differ considerably. Licensing arrangements vary, with Insight Inc. providing company-licensed ChatGPT accounts while Dinoco and GMT license Co-Pilot tools instead. Governance mechanisms range from formal-ized approval processes with internal expert groups (GMT) to compliance-driven top-lev el decisions (Dinoco) to largely absent formal frameworks (Insight Inc.). Crucially, even where policies exist, adherence is inconsistent; some employees circumvent regulations by using private accounts for work-related tasks.
4.2 Applied GenAI Use-Cases in Agile Software Development
This section addresses RQ2: Which use cases of GenAI tools are adopted in agile software development teams?
The cross-case analysis revealed three clusters of use cases and three roles applying in total 21 use cases. We categorized these use cases thematically, iden-tifying 6 as creative tasks, 7 as documentation-related tasks, and 8 as coding-related tasks.
Use case adoption varies substantially by role and experience level. Prod-uct Owners and Agile Coaches reported higher-intensity usage, particularly for text-based creative tasks, while developers’ adoption ranged from minimal to moderate depending on codebase characteristics. At GMT, one senior developer with 15 years’ experience self-rated usage intensity at 1 out of 4, attributing low adoption to established workflows and extensive domain knowledge that reduced perceived need for AI assistance. Conversely, newer team members and those in product-focused roles reported more extensive integration of GenAI into daily work.
Figure 2 depicts the identified use cases by role. The complete cross-case analysis including a detailed overview per role and case, as well as systematic use case descriptions following a structured schema is available in our research protocol.
Across all three organizations, GenAI is primarily leveraged as a personal pro-ductivity assistant, augmenting individual tasks rather than transforming team collaboration. As shown in Fig. 2, we identified a clear pattern of role-specific adoption, with Product Owners, Agile Coaches/Scrum Masters, and Develop-ers applying the tools to distinct aspects of their work. We characterize these patterns through three archetypes: the Research Assistant for creative a nd con-ceptual tasks, the Virtual Tutor for documentation and communication, and the Pair Software Developer for coding and technical activities.
Research Assistant: At Dinoco, GenAI is used specifically to support creative processes. P02 uses GenAI tools for brainstorming and idea validation: “A lot for brainstorming, to have a practical communication partner and to validate a few ideas.” P1 also reports typical use cases in which GenAI helps to structure thoughts: “But then you often have these thought blocks: How do I get this into sensible words now?”. In these use cases, GenAI serves less as a pure source of information, but is seen as a discursive partner that provides impulses and removes creative blocks. A particularly prevalent use case across all roles and organizations is GenAI as an enhanced search mechanism, effectively replacing traditional web searches and Stack Overflow queries.
Multiple developers charac-terized ChatGPT as their “first point of contact” for technical questions, valuing consolidated answers over navigating multiple search results.
Virtual Tutor: Product Owners use GenAI to formulate requirements faster. P01 describes this succinctly: “[...] writing features and stories ”. This example shows that GenAI is primarily used here to elaborate the content of user stories and thus simplifies writing work. From an Agile Coach/Scrum Master perspec-tive, tools like ChatGPT were frequently mentioned as primary assistants for content generation, including drafting texts, brainstorming ideas, and refining formulations. P12 described GenAI as a “very helpful sparring partner ” when initiating tasks or formulating complex content. P11 noted that while the tool does not necessarily complete tasks for them, it significantly accelerates their work by “kicking off ideas faster ”.
Furthermore, several Agile Coaches/Scrum Masters (P01, P12, P14) mentioned that they use GenAI to prepare for r et-rospectives by generating creative formats or by refining formulations in team surveys. Others highlighted that it helps them organize thoughts when faced with a blank page, particularly for internal presentations or training materials (e.g., P13). These scenarios suggest that GenAI is mainly used to support the early phases of cognitive work (ideation, structuring, and polishing) rather than full automation.
Beyond agile artifacts, Product Owners and coaches employ GenAI exten-sively for stakeholder communication. At GMT, one Product Owner (P8) reported using Microsoft Copilot daily to formulate customer emails, particularly for business English communication where uncertainty about professional phras-ing exists. The tool serves as both translator and style consultant. At Dinoco, similar patterns emerge for refining internal communications and ensuring pro-fessional tone in external correspondence. This communication support function operates largely invisibly improving message quality without transforming the underlying work process.
Pair Software Developer: For developers, GenAI primarily simplifies stan-dard tasks. Its most valued contributions include generating “boilerplate code ” (P03), writing repetitive unit tests (P03), and providing “auto-completion within the IDE ” (P04). Beyond code generation, a common use case is understand-ing existing code; developers frequently ask the AI to explain complex snippets, especially when working with large, legacy codebases. Most interviewees reported using GenAI as a “pair programmer” to improve, refactor, or get a “second opin-ion ” on code segments (P04). This interaction sometimes evolves into a collabo-rative problem-solving dialogue. Documentation, a notoriously time-consuming task is another key application, used for generating both inline comments and commit messages. One GMT developer noted documenting more because AI makes it faster.
These findings confirm that GenAI’s value in development lies in automating routine work and augmenting the developer’s understanding, allow-ing them to focus on more complex problem-solving.
Interestingly, none of the interviewees reported using GenAI for decision-making or for directly facilitating meetings, citing concerns about quality, ethics, and team acceptance. Instead, this technology is viewed as a background tool that empowers but does not replace the human element of agile ways of working. Overall, GenAI tools appear to be most effective when they are seamlessly inte-grated into existing processes and support specific tasks. Perceived value largely depends on the user’s role, the specific use case, and i ndividual approaches to the technology.
4.3 Benefits & Barriers of GenAI Adoption
This section addresses RQ3: What benefits and barriers do agile team members associate with GenAI adoption?
Benefits can be grouped into three categories: 1) efficiency gains in coding, 2) overall task efficiency through automation, and 3) enhanced creativity in early project phases.
First, GenAI enables efficiency gains in coding by simplifying or optimizing code segments through access to a broader knowledge base than individual devel-opers possess. Developers reported reduced time for boilerplate code generation, faster prototyping for customer requests, and accelerated resolution of unfamiliar technical challenges. GMT developers also reported learning new language fea-tures, design patterns, and optimization techniques from AI suggestions. Second, GenAI enhances overall task efficiency by automating routine activities and gen-erating code skeletons, thereby accelerating development processes. It can also handle tedious tasks such as documentation, allowing developers to focus more on value-adding activities. Beyond efficiency gains, GenAI to ols also foster cre-ativity, particularly in early project phases such as requirements definition and concept development.
They provide inspiration, structure, and support out-of-the-box thinking. Some interviewees also mentioned the use of prompt templates to streamline their work. These insights suggest that GenAI contributes not only to implementation but also to ideation and conceptual design.
Perceived benefits, however, vary by role. While Product Owners and Requirements Engineers described GenAI as a valuable support tool, developers expressed more reserved attitudes, indicating that the impact of GenAI depends strongly on professional role and context of use. This divergence is particu-larly pronounced at Dinoco, where developers characterized productivity gains as “modest or even neutral ”, with one senior developer (P04) asserting that “pro-ductivity cannot meaningfully be measured ” calling into question the empirical basis for claimed efficiency improvements.
Alongside these benefits, three clusters of barriers emerged: 1) validation and quality assurance effort, 2) effort required to effectively use and integrate GenAI tools, and 3) bureaucratic and governance-related barriers.
The first concerns the effort required to validate GenAI outputs and the risk of uncritically adopting generated results. At GMT, this validation burden is formalized through the AI guideline, which mandates that AI-generated con-tent may only be used after two independent, authorized persons have reviewed and approved it via email. However, this rule is unworkable in practice, as the overhead of a two-person review for every AI-generated snippet would create unsustainable bottlenecks. Several interviewees perceived this as problematic, noting that hidden code dependencies demand a thorough understanding of func-tionality for correct operation and debugging. As GenAI occasionally produces inaccurate outputs, critical evaluation remains essential.
The second cluster relates to the effort involved in effectively using GenAI. Interviewees emphasized that meaningful results often require substantial input and context, yet the outcome may still lack value. Developers at Dinoco stressed that supplying sufficient context is time-consuming, and incomplete prompts frequently lead to generic or unusable responses. Enhancing usability, by sim-plifying interaction and reducing reliance on prompt engineering, was seen as crucial. Better integration into existing tools, such as Copilot in Visual Studio Code, was cited as an effective solution, as it allows suggestions directly within the IDE.
The third cluster involves bureaucratic challenges, primarily stemming from GMT’s AI guidelines. The mandatory evaluation of GenAI models prior to use was perceived as a barrier to experimentation. Additionally, employees at Dinoco and GMT must manage their own accounts without centralized support, requir-ing them to independently identify suitable models and verify approval status. Some interviewees also expressed concern that GenAI use might prompt uncom-fortable questions or require justification to supervisors, particularly when no tangible results are achieved. For larger organizations, a central obstacle lies in the organizational restrictive attitude towards GenAI tools. This was addressed by several respondents.
P01 emphasized: “My experience with [Dinoco] - or with other companies - is that everything is always so restrictive and hostile to inno-vation.” This points to structural hurdles that can slow innovation due to safety concerns or bureaucratic processes. At Dinoco, the compliance department’s approval requirement acts as a b ottleneck; ChatGPT’s non-approved status leads developers to use it anyway, reframed as “private time” usage. Although official regulations exist, the actual implementation is often left to the employees them-selves. Responsibility for compliant behavior thus becomes highly individual-ized, while organizational support remains limited. The barrier, therefore, is not merely bureaucratic friction, but the lack of organizational pathways to reconcile security requirements with practical needs forcing employees to choose between compliance and productivity.
In summary, interviewees reported time savings, enhanced creativity, and support for routine tasks as major benefits, but also concerns about data privacy, dependency on AI-generated output, and lack o f transparency. These mixed perceptions underscore the importance of context-sensitive implementation and support mechanisms.
5 Discussion
Our findings contribute to the growing body of empirical research on GenAI adoption in software development by offering a practice-centered perspective from within agile teams. To s ystematically interpret and extend these insights, we employ the Technology-Organization-Environment (TOE) framework as a theoretical lens. The TOE framework posits that technology adoption is shaped by three interrelated contextual dimensions: the technological con-text, including characteristics of available and internal technologies; the orga-nizational context, encompassing structural attributes, resources, and human capital; and the environmental context, referring to external pressures such as industry conditions and regulatory requirements.
By applying this lens, we structure our discussion around how each dimension influences GenAI adoption in agile settings and how tensions across these dimensions give rise to the compliance gaps observed in our study.
Technology Context: The technological dimension centers on tool capabilities and limitations. In line with Kemell et al., we observe that GenAI is predom-inantly used as a personal productivity assistant, augmenting individual performance rather than fundamentally altering team-level processes. Recent studies by Russo and Banh et al. highlight the importance of seamless work-flow integration as a prerequisite for sustained GenAI adoption. Our findings corroborate this, demonstrating that tools embedded into existing development environments, such as Copilot in VS Code, facilitate coding tasks with mini-mal workflow disruption.
However, whereas Kemell et al. limit their scope predominantly to GenAI as a programming assistant, we find that the techno-logical context extends to conversational tools leveraged for communicative and facilitation tasks including retrospective preparation and team alignment.
The technological context also reveals critical limitations regarding valida-tion. The validation burden associated with AI-generated outputs represents a fundamental technological constraint, current GenAI to ols cannot guarantee accuracy, necessitating human review that partially offsets efficiency gains. While Kemell et al. frame this primarily as a user capability issue (evaluation diffi-culty), our findings suggest validation becomes a process bottleneck when orga-nizations impose rigid review procedures that conflict with agile velocity. This constraint is thus not merely technological or individual, but emerges from the interaction between tool limitations a nd organizational policy responses.
Fur-thermore, interviewees emphasized that meaningful results require substantial context, yet supplying it is time-consuming, and incomplete prompts frequently lead to generic or unusable responses.
Our results further indicate that perceived benefits vary substantially by experience level and organizational context. Senior developers working with legacy systems reported minimal productivity gains, challenging universal effi-ciency narratives. This suggests that GenAI’s value is c ontingent on factors including domain expertise, codebase maturity, and established workflow effi-ciency rather than being universally transformative.
Organization Context: The organizational dimension encompasses role struc-tures, internal capabilities, and the cultural readiness of teams to adopt GenAI. While Kemell et al. broadly distinguish between”individual users” and “management” and emphasize governance mechanisms, our explicit focus on agile roles (Product Owners, Agile Coaches/Scrum Masters, and Developers) highlights how role-specific affordances and agile values shape GenAI usage.
Based on our analysis, we identify three archetypes that characterize how GenAI is appropriated across agile teams: the Research Assistant, where GenAI sup-ports creative and conceptual tasks such as brainstorming, idea validation, and strategic exploration; the Virtual Tutor, where GenAI aids documentation, communication refinement, user story formulation, and stakeholder correspon-dence; and the Pair Software Developer, where GenAI assists with code generation, code explanation, debugging, and technical problem-solving. These archetypes reveal that GenAI adoption in agile contexts varies systematically by role, with Product Owners and Agile Coaches leveraging GenAI primarily as Research Assistants and Virtual Tutors, while Developers engage it predom-inantly as a Pair Software Developer.
The alignment between these archetypes and role-specific task demands determines adoption intensity and perceived value. Crucially, teams appropri-ate GenAI to align with agile values. Rather than automating core processes, the technology supports—but does not substitute, activities like requirement elicitation and team facilitation. This subtly reshapes work to amplify efficiency while preserving underlying soc ial dynamics. However, organizational readiness remains limited; the absence of adequate training across all cases forces practi-tioners to develop competencies independently, creating a significant capability gap. This finding resonates with Kucharski’s observation that agile mindset leadership is critical for developing dynamic capabilities, particularly “sensing” (identifying technological opportunities) and “reconfiguring” (adapting processes to integrate innovations).
Our cases suggest that organizations are failing to sense GenAI’s potential systematically or to reconfigure learning pathways accordingly, leaving adoption dependent on individual initiative rather than organizational strategy.
Environment Context: The environmental dimension captures external reg-ulatory pressures, industry norms, and competitive dynamics that shape GenAI adoption. Our findings reveal that data protection requirements constitute the most significant environmental constraints on adoption. Across all roles, GenAI remains subordinate to human judgment, reflecting both epistemic caution and adherence to human accountability. This is not merely a practitioner preference but, in some cases, an institutional mandate, at GMT, internal policies explic-itly frame AI outputs as “recommendations” requiring human review, i nstitu-tionalizing a human-in-the-loop practice. These institutional requirements reflect broader environmental pressures stemming from the EU AI Act’s risk classifica-tion framework, data protection regulations (GDPR), and sector-specific com-pliance mandates.
While this formalized response creates structured pathways for adoption, it simultaneously introduces bureaucratic friction.
A central con-Cross-Dimensional Tensions and the Compliance Gap: tribution of our study is identifying systematic tensions across TOE dimensions that manifest as compliance gaps. While Kemell et al. document data privacy and regulatory concerns as adoption barriers, their analysis focuses on describ-ing these challenges rather than examining their behavioral consequences. We extend this by revealing how practitioners systematically respond to and circum-vent restrictive policies when compliance conflicts with productivity demands. Viewed through the TOE lens, this gap arises when environmental pressures (reg-ulatory compliance) are translated into organizational policies without adequate consideration of technological characteristics and organizational constraints.
Our analysis reveals that a misalignment between top-down governance and bottom-up work practices creates systemic friction and leads to specific forms of Shadow IT. When organizational policies, designed to mitigate risk, fail to account f or the day-to-day productivity needs of practitioners, they are often perceived as impractical and are informally bypassed. This policy-practice gap manifests in distinct ways across our cases, but consistently leads to sub-versive adoption practices. For instance, GMT’s mandatory two-person review for all AI outputs illustrates how strict compliance can conflict with the need for agile velocity. At Dinoco, developers reframe unapproved ChatGPT usage as “private time”, an adaptive response that bypasses official regulations to resolve the conflict between p olicy and productivity.
Furthermore, even where organizations provide licensed tools, we found that data protection constraints prohibiting realistic work content can render them impractical, driving practi-tioners toward unlicensed alternatives for convenience.
Ultimately, when tools offer substantial productivity benefits, prohibition without practical alternatives merely shifts usage into informal arrangements that may increase organizational risk. This challenges the assumption that restrictive policies alone ensure compliance, suggesting that effective governance requires creating frameworks practitioners can and will actually follow.
6 Conclusion and Future Work
This multiple-case study investigates how agile practitioners employ GenAI tools in organizational settings. The findings show that GenAI is already integrated across a wide range of agile practices from repetitive tasks such as code imple-mentation and user-story writing to more strategic activities like team facili-tation and reflection. Yet its adoption also presents challenges, including legal ambiguities, the need for governance frameworks, and concerns about data secu-rity and trust. Overall, while GenAI offers considerable potential to strengthen agile roles, its integration must be carefully managed to align with organizational structures and policies.
Given that GenAI is already used by software development teams across vari-ous organizations, several further questions arise: Do other organizations employ GenAI in similar ways, and does access to adequate training improve team effec-tiveness? Addressing these questions requires broader empirical research that compares current practices across the wider industry. Future studies could, for example, examine additional organizations across different countries and regu-latory regimes to contrast their approaches and outcomes.
Acknowledgements % Final Remark. A. Przybylek’s contribution was funded by Taighde ́Eireann—Research Ireland (Grant No.13/RC/2094 2). Co-funded by the Euro-pean Union (SyMeCo, Grant No.101081459). The views expressed are solely those of the author(s). Neither the European Union nor the European Research Executive Agency can be held responsible for them.
As a final note, we acknowledge that the text of this paper was refined with the assistance of LLMs. The authors reviewed and edited all content generated using this service and take full responsibility for the published article.
1. Alami, A., Ernst, N.: Human and machine: how software engineers perceive and.
engage with ai-assisted code reviews compared to their peers. In: Proceedings of the 18th Int ernational Conference on Cooperative and Human Aspects of Software Engineering, pp. 63–74 (2025)
2. Arora, C., Herda, T., Homm, V.: Generating test scenarios from nl requirements.
using retrieval-augmented LLMs: an industrial study. In: Proceedings of the 32nd International Requirements Engineering Conference, pp. 240–251 (2024)
3. Bahi, A., Gharib, J., Gahi, Y.: Integrating generative AI for advancing agile soft-.
ware development and mitigating project management c hallenges. Int. J. Adv. Comput. Sci. Appl. 15 (2024)
4. Banh, L., Holldack, F., Strobel, G.: Copiloting the future: how generative ai trans-.
5. Basili, V., Caldiera, G., Rombach, D.: The goal question metric approach, pp.
528–532 (1994)
6. Braun, V., Clarke, V.: Using thematic analysis in psychology.
Qual. Res. Psychol.
3, 77–101 (2006)
7. Coutinho, M., Marques, L., Santos, A., Dahia, M., Fran ̧ca, C., Souza Santos, R.:.
The role of generative ai in software development productivity: A pilot case study. In: Pro ceedings of the 1st ACM International Conference on AI-Powered Software, pp. 131–138 (2024)
8. Depietro, R., Wiarda, E., Fleischer, M.: The context for change: organization.
technology and environment. Process. Technol. Innov. 199, 151–175 (1990)
9. Diebold, P.: From backlogs to bots: generative ai’s impact on agile role ev olution.
J. Softw. Evol. Process 37, e2740 (2025)
10. Dong, C., Jiang, Y., Zhang, Y., Zhang, Y., Liu, H.: Chatgpt-based test generation.
for refactoring engines enhanced by feature analysis on examples. I n: Proceedings of the 47th International Conference on Software Engineering, pp. 2714–2725 (2025)
11. Ebert, C., Louridas, P.: Generative ai for software practitioners.
IEEE Softw. 40, 30–38 (2023)
12. European Union: Regulation (EU) 2024/1689 on harmonised rules on artificial.
the linked source: intelligence. 32024R1689 (2024), Accessed 29 N ov 2025
13. Jarzebowicz, A., ́Slesi ́nski, W.: Assessing effectiveness of recommendations to.
requirements-related problems through interviews with experts. In: 2018 Feder-ated Conference on Computer Science and Information Systems (FedCSIS), pp. 959–968. IEEE (2018)
14. Karlovs-Karlovskis, U.: Generative artificial intelligence use in optimising software.
engineering process: a s ystematic literature review. Appl. Comput. Syst. 29, 68–77 (2024)
15. Kemell, K.K., Saarikallio, M., Nguyen-Duc, A., Abrahamsson, P.: Still just per-.
sonal assistants? – a multiple case study of generative ai adoption in software organizations. Inf. Softw. Technol. 186, 107805 (2025)
16. Kucharski, M., Kucharska, W., Jussila, J.: Impact of agile software development.
team leaders’ mindset on dynamic capabilities for achieving organizational agility. In: Lukovi ́c, I., et al. (eds.) Empowering the Interdisciplinary Role of ISD in Addressing Contemporary Issues in Digital Transformation (ISD2025 Proceedings) (2025)
17. Kwok, Y.T.C., Adil, M.: Ai and teamwork in agile software development: a sys-.
tematic mapping study. In: Agile Processes in Soft ware Engineering and Extreme Programming – Workshops, pp. 32–40. Springer Nature Switzerland, Cham (2026)
18. Mendes, W., Souza, S., De Souza, C.: You’re on a bicycle with a little motor:.
benefits and challenges of using ai code assistants. In: Proceedings of the 17th International Conference on Cooperative and Human Aspects of Software Engi-neering, pp. 144–152 (2024)
19. Mock, M., Melegati, J., Russo, B.: Generative ai for test driven development: pre-.
liminary results. In: Agile Processes in S oftware Engineering and Extreme Pro-gramming – Workshops, pp. 24–32. Springer, Cham (2025) 20. Neumann, M., et al.: Research protocol (2025). the linked source. 18520760
21. Neumann, M., Rauschenberger, M., Sch ̈on, E.M.: We need to talk about chatgpt’:.
the future of ai and higher education. In: Proceedings of the 5th International Workshop on Software Engineering Education for the Next Generation, pp. 29–32 (2023)
22. Nguyen-Duc, A., et al.: Generative artificial intelligence for software engineering—a.
23. Oertel, J., Kl ̈under, J., Hebig, R.: Don’t settle for the first! how many github copilot.
24. Rauf, I., Sharp, H., Lopez, T., Wermelinger, M.: Human-machine teaming and.
team effectiveness in ai tools for software engineering. In: Proceedings of the 18th International Conference o n Cooperative and Human Aspects of Software Engi-neering, pp. 75–80 (2025)
25. Runeson, P., H ̈ost, M.: Guidelines for conducting and reporting case study research.
26. Russo, D.: Navigating the complexity of generative ai adoption in software engi-.
neering. A CM Trans. Softw. Eng. Methodol. 33 (2024)
27. Russo, D., et al.: Generative ai in software engineering must be human-centered:.
the copenhagen manifesto. J. Syst. Softw. 216 (2024)
28. Salah, M., Abdelfattah, F., Halbusi, H.A.: Generative artificial intelligence (chatgpt.
& bard) in public administration research: a double-edged sword for street-level bureaucracy studies. Int. J. Public Adm. 0, 1–7 (2023)
29. Sami, M.A., et al.: A multi-agent llm system for automated requirements analysis:.
A study on user story generation and prioritization. In: Proceedings of the 51st Euromicro Conference on Software Engineering and Advanced Applications, pp. 178–187 (2025)
30. Santos, I., Felizardo, K.R., Steinmacher, I., Gerosa, M.A.: Great power brings.
great responsibility: Personalizing conversational ai for diverse problem-solvers. In: Proceedings of the 18th International Conference on Cooperative and Human Aspects of Software Engineering, pp. 93–95 (2025) 31. Sauvola, J., Tarkoma, S., Klemettinen, M., Riekki, J., Doermann, D.: Future of software development with generative ai. Autom. Softw. Eng. 31 (2024) 32. Stray, V., Moe, N., Ganeshan, N., Kobbenes, S.: Generative ai and developer work-flows: how github copilot and chatgpt influence solo and pair programming. In: Pro-ceedings of the 58th Hawaii International Conference on System Sciences (2025) 33. Tanzil, M.H., Khan, J.Y., Uddin, G.: Chatgpt incorrectness detection in software reviews
In: Proceedings of the 46th International Conference on Software Engi-neering (2024) 34. Ulfsnes, R., et al: Responsible ai in agile software engineering - an industry per-spective
In: Marchesi, L., et al. (eds.) Agile Processes i n Software Engineering and Extreme Programming – Workshops. pp, 33–41. Springer, Cham (2025). the linked source 4 35. Ulfsnes, R., Moe, N.B., Stray, V., Skarpen, M.: Transforming Software Develop-ment with Generative AI: Empirical Insights on Collaboration and Workflow, pp. 219–234 (2024) 36. Wohlin, C., et al.: The success factors powering industry-academia c ollaboration 37. Yin, R.K.: Case study research: Design and methods, applied social research meth-ods series, vol. 5. Sage, Los Angeles, 4. ed. edn. (2009) 38. Zhang, S., et al.: Empowering agile-based generative software development through human-ai teamwork
A CM Trans. Softw. Eng. Methodol. 34 (2025) 39. Zhang, Z., Rayhan, M., Herda, T., Goisauf, M., Abrahamsson, P.: LLm-based agents for automating the enhancement of user story quality: an early report. In: Proceedings of the 25th International Conference on Agile Software Development, pp. 117–126. Springer, Cham (2024)
Open Access This chapter is licensed under the terms of the Creative C ommons Attribution 4.0 International License (the linked source), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license and indicate if changes were made.
The images or other third party material in this chapter are included in the chapter’s Creative Commons license, unless indicated otherwise in a credit line t o the material. If material is not included in the chapter’s Creative Commons license and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder.