Responsible Engineering of Information Systems Based on Generative Artificial Intelligence: An Action Design Research Study at a German Premium Car Manufacturer
1 More Paper · Full Reading

About this paper
Abstract Generative artificial intelligence (GenAI) is a rapidly evolving field that enables novel and valuable information systems. However, developing GenAI-based information systems poses unique challenges that need to be addressed by researchers and practitioners. This paper proposes a method that supports organizations in the responsible engineering of GenAI-based information systems aimed at improving internal operations and value-creation processes. To develop the method enabling organizations to exploit GenAI’s enormous potential while mitigating associated risks, elaborated action design research was adopted. One diagnosis cycle, two design cycles, and one implementation cycle were conducted, involving the evaluation of the method’s design through 48 interviews with practitioners and based on the realization of three real-world GenAI use cases. The theoretical contribution includes prescriptive design knowledge for the responsible engineering of GenAI-based information systems. From a practical standpoint, the paper extends organizations’ methodological toolbox with a holistic, prescriptive guideline for engineering GenAI-based information systems.
Authors: Dominik Fetzer, Henner Gimpel, Oliver Meindl, Jan Strickmann
Published in: Business & Information Systems Engineering
Publication date: 2025-09-04
Read the paper: https://doi.org/10.1007/s12599-025-00950-6
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “Responsible Engineering of Information Systems Based on Generative Artificial Intelligence: An Action Design Research Study at a German Premium Car Manufacturer,” by Dominik Fetzer and colleagues. Published in Business & Information Systems Engineering on September 4, 2025.
• Henner Gimpel • Oliver Meindl • Jan Strickmann Dominik Fetzer
Received: 19 July 2024 / Accepted: 14 April 2025 / Published online: 4 September 2025 The Author(s) 2025, corrected publication 2025
Abstract Generative artificial intelligence (GenAI) is a rapidly evolving field that enables novel and valuable information systems. However, developing GenAI-based information systems poses unique challenges that need to be addressed by researchers and practitioners. This paper proposes a method that supports organizations in the responsible engineering of GenAI-based information sys-tems aimed at improving internal operations and value-creation processes. To develop the method enabling orga-nizations to exploit GenAI’s enormous potential while mitigating associated risks, elaborated action design research was adopted. One diagnosis cycle, two design cycles, and one implementation cycle were conducted, involving the evaluation of the method’s design through 48 interviews with practitioners and based on the realization of three real-world GenAI use cases.
The theoretical con-tribution includes prescriptive design knowledge for the responsible engineering of GenAI-based information
Accepted after two revisions by O ́ scar Pastor.
D. Fetzer (&) H. Gimpel O. Meindl Branch Business & Information Systems Engineering of the Fraunhofer FIT, Augsburg, Germany e-mail: the email address
D. Fetzer H. Gimpel O. Meindl FIM Research Center for Information Management, Stuttgart, Germany
D. Fetzer H. Gimpel O. Meindl Chair of Digital Management, University of Hohenheim, Stuttgart, Germany
J. Strickmann Dr. Ing. h.c. F. Porsche Aktiengesellschaft, Stuttgart, Germany systems. From a practical standpoint, the paper extends organizations’ methodological toolbox with a holistic, prescriptive guideline for engineering GenAI-based infor-mation systems.
Recent advancements in generative artificial intelligence (GenAI) are astonishing, with GenAI demonstrating capa-bilities and output quality many experts did not (yet) consider possible. GenAI models can generate novel, meaningful content across various modalities, including texts, images, videos, audio files, and process models almost indistin-guishable from those of human origin. GenAI, therefore, offers significant economic potential, which many organizations have already recog-nized. Text-generating GenAI models can, for example, power organizational knowledge management. They can also automate the time-con-suming manual extraction and consolidation of essential information from documents such as contracts or filed certificates. This highlights that GenAI is not only technologically groundbreaking but also holds a wealth of novel entrepreneurial potential (cf.
Section 2.1), which is why profound transformations in numerous fields are predicted.
To leverage this vast potential, particularly medium- and large-sized organizations are confronted with GenAI’s two unique technical characteristics: its generative variability, and the integration of large pre-trained models from external providers. Furthermore, organizations must adopt and adapt to this new technology without compromising their competitive advantage. That involves integrating GenAI into existing products, processes, and IT landscapes quickly and efficiently, as an inadequate or slow intro-duction pace poses a genuine risk of falling behind. The aforementioned complex task has two central challenges. On the one hand, organizations must utilize GenAI responsibly to avoid unintended, neg-ative consequences.
This requires organizations to address the new risks of GenAI compared to traditional software or other forms of artificial intelligence (AI), some of which may still be unknown to many employees. Organizations must adopt appropriate measures such as checking for possible copyright infringements or implementing an output restriction and control to mitigate the new risks and ensure the responsible utilization of GenAI. Errors could have profound and irreversible consequences, including reputa-tional damage or legal liability. On the other hand, organizations must efficiently develop a systematic process for engineering GenAI-based informa-tion systems (IS) with their employees, comprising steps like requirements specification, design, implementation, verification, or deployment.
Given the nontrivial nature of responsible GenAI introduction and the speed required, methodological and structured guidance on responsibly engineering GenAI-based IS is urgently needed.
A promising way is to formalize prescriptive guidelines as a method. As a well-established tool for engineering technologies, a method furnishes a systematic structure for carrying out work steps and a guideline to attain specific goals. Existing software or AI engineering methods in organizations have already generated enormous added value. Literature suggests that by employing a GenAI-specific engineering method, organizations can achieve the necessary speed of integra-tion while ensuring responsible handling of the technology. How-ever, neither of the current engineering methods adequately addresses the unique requirements and risks of utilizing GenAI. Existing literature depicts only individual aspects of the responsible engineering of GenAI-based IS in a fragmented manner.
As the responsible utilization of AI has become increasingly important, several valuable guidelines and principles for ethical, responsible, and trustworthy AI have been developed, complementing the engineering methods. Due to their abstract or purely descriptive/normative nature, organizations have great difficulty putting them into prac-tice. This criticism was addressed by initiatives like the IEEE Standard 7000-2021, providing a method for considering ethical concerns in IS engineering followable by practi-tioners in a structured manner. The standard is designed to be compatible with any IS engineering methods and can be applied in parallel. However, given its wide scope, it can be considered too broad and not specific enough to GenAI (Morandı ́n-Ahuerma 2023).
Other initiatives, like the BSI (2024) report, address various specifics of GenAI but do not cover all areas of responsibility, such as performance or legal aspects. Both challenges for organizations remain widely unaddressed (cf. Section 2.2), even though they have already been identified in the literature as critical and urgently needing resolution for the responsible utilization of GenAI. To support organizations in leveraging the vast potential of GenAI while mitigating associated risks, we pose the following research question: How can medium- and large-sized organizations engineer GenAI-based IS responsibly?
Applying the action design research (ADR) approach of Mullarkey and Hevner (2019) in a twelve-month project (cf. Section 3) at the German premium car manufacturer Premium Automotive Group (PAG), we developed the Method for Responsible GenAI-Based Information Sys-tems Engineering (MeRGE). We specifically designed our method for medium- and large-sized organizations seeking to integrate GenAI via pre-trained models to improve internal operations and value-creation processes. Leverag-ing situational method engineering (SME) by Henderson-Sellers and Ralyte ́ (2010), as well as building upon valu-able organizational insights and close exchange with 19 practitioners from PAG, the five phases of MeRGE were successively enriched with detailed activities, roles & required skills, and illustrative tools & good practices. Our results were iteratively evaluated (cf.
Section 5) in 48 interviews to promote the fusion of practical and scientific perspectives. The MeRGE method was additionally situ-ated by realizing three PAG-internal GenAI use cases, one of which served as a running example to demonstrate MeRGE (cf. Section 4). We contribute to theory and practice (cf. Section 6). The scientific contribution of this study is twofold: First, it extends the mainly descriptive and normative landscape of responsible engineering ini-tiatives by providing a nascent design theory in the form of a structured method for organization, namely MeRGE. Second, it gives reflected and generalized insights into GenAI-based IS in the form of a high-level architecture.
Concerning the practical implications, our study raises awareness regarding the GenAI-specific requirements that medium- and large-sized organizations face, which, unad-dressed, could lead to serious negative consequences in the future. To conclude (cf. Section 7), the myriad insights from PAG – condensed into a methodological toolbox, ‘MeRGE’ – support an entire industry just starting to engineer GenAI-based IS, ultimately leading to a more responsible organizational utilization of GenAI.
2 Theoretical Background 2.1 Generative Artificial Intelligence-Based Information Systems in Organizations
Recent developments in AI have shifted the focus from discriminative AI to GenAI, heralding a new era. GenAI ‘‘refers to computational techniques that are capable of generating seemingly new, meaningful content such as text, images, or audio from training data’’. This paper focuses on general-purpose GenAI models that have been trained on vast quantities of publicly accessible data and not limited to narrow application domains, representing the heart of recent advancements. These models comprise but are not limited to (multi-modal) large language models and dif-fusion probabilistic models. Chui et al. (2023) highlight the technology’s massive economic potential, indicating that GenAI can increase global pro-duction by USD 2.6–4.4 trillion annually.
The enormous potential arising from the multifaceted, human-like capa-bilities of GenAI (e.g., producing organization-tailored texts or creating training videos) stems from several factors which distinguish GenAI-based IS from software- and machine learning (ML)-based IS. While software-based IS refer to applications that use rigid algorithms to perform specific tasks, ML-based IS are often associated with uti-lizing discriminative ML models for pattern recognition in data to achieve classifications, regressions, or clustering. Besides, GenAI models are cap-able of generating novel multimodal content beyond their training data rather than merely defining numerical deci-sion boundaries. Organizations typically utilize pre-trained, large-scale, multi-purpose foundation models and adapt them for various tasks.
Even though GenAI-based IS can also be seen as an instance of soft-ware- or ML-based IS, they differ widely. Table 1 outlines the main differences between these IS types in terms of their technical components. This categorization highlights typical patterns. However, individual cases might not per-fectly fit this rough categorization. Examples of exceptions include hybrid forms between different IS types that cannot be clearly delineated (e.g., IS that utilize both GenAI models and other ML models) or an evolution over time, such as when a classic software-based IS now incorporates a small GenAI-based chat component. GenAI-based IS tend to use vast amounts of data for training and to operate probabilistically, and they potentially generate various output formats. These characteristics are key to their broad capabilities.
As developing pre-trained GenAI models is enormously challenging, organi-zations typically procure them as ready-to-use components through an application programming interface (API) from only a few external service providers. This leads to a high dependency and induces new risks (e.g., data leaks or copyright infringements through GenAI outputs) and the loss of substantial control and influence (e.g., regarding training data, model develop-ment, and other technical details). A distinctive feature of GenAI-based IS lies in their profoundly different behavior, which stems from their inner workings. Organizations utilizing GenAI enter a new environment of generative variability, characterized by variability and uncertainty in the generation quality.
Even with substantial efforts regarding output quality and human expectations alignment (e.g., through reinforcement learning from human feedback), generative variability can lead to unexpected deviations like halluci-nations or harmful content. The higher output maturity of GenAI-based IS, providing readily understandable information instead of numerical data, makes output deviations even more problematic. Due to these massive differences and their enormous complexity, prescriptive guidance is necessary for organizations when GenAI-based IS are engineered.
2.2 Responsible Engineering of GenAI-Based Information Systems
Typically, organizations use methods to guide the engi-neering of IS. Existing examples frequently applied in practice range from generic methods (e.g., Information System Life Cycle) to specific ones for soft-ware development (e.g., Scrum) or AI (e.g., ML Opera-tions). In engineering GenAI-based IS, organizations must consider new, unique, multifaceted risks and complexities. Given the potential for far-reaching negative consequences (e.g., lasting damage to organizational reputation or notable fi-nancial losses), a particularly responsible engineering is necessary. Responsible engineering aims at ‘‘avoiding those unintended, negative consequences of
To highlight the specifics of each IS type, the table presents each type disjunctively, that is, excluding downstream types (e.g., Software-based IS without ML-based IS) a The characteristics of the respective other IS types can also apply to GenAI-based IS b This column refers to GenAI-based IS utilizing GenAI models as ready-to-use components. Systems employing GenAI models that are trained from scratch or modified through advanced techniques (e.g., fine-tuning) are excluded
AI’’ comprising multiple areas. Existing initiatives address responsible engi-neering from three viewpoints. While the first group of initiatives examines responsible engineering with a focus on software- and ML-based IS (e.g., Benjamins 2021; IEEE 2021; Martı ́nez-Ferna ́ndez et al. 2022), the second group takes in an overarching AI perspective yet has a rather descriptive or normative character. Third, GenAI-specific approaches dealing with the technology’s respon-sible engineering and use consider the particular risks induced by the specific requirements of GenAI, albeit with a limited scope (e.g., BSI 2024; Johri et al. 2023; OWASP 2023; So ̈llner et al. 2025). Table 2 presents on the left side a comparison of the main fields of action within engi-neering per IS type to satisfy a responsibility area.
The statements made are only a bold summary and in no way cover the whole reality, which can differ widely for each IS type. The summaries are derived from existing initiatives – among others, for example, Khowaja et al. (2024), Petersen et al. (2022), Rizinski et al. (2022) including those men-tioned earlier – that extensively examine the responsible engineering of various IS types. From these sources, we extracted statements related to the responsibility areas, enhanced them with our own experiences, and presented them in a vivid manner. Based on an in-depth analysis by the authors, the table also depicts a selection of five major
Table 2 Responsible engineering for different information system types initiatives that most closely provide prescriptive guidance for responsible GenAI engineering for organizations and assesses them per responsibility area. The overarching AI approaches Montreal Declaration for Responsible AI, European AI Act, and NIST AI Risk Management Framework describe important facets of responsibility in AI that are relevant to being addressed within organizations. Yet, due to their abstract character, these approaches cannot fully capture the specific requirements of GenAI (Morandı ́n-Ahuerma 2023). Since they are either descrip-tive or normative, they propose rules and explanations instead of a structured and processable approach, so their application in practice is challenging.
More prescriptive approaches, which practitioners can follow in a structured manner, are set out to address this issue. The IEEE Standard 7000-2021 offers an overarching method for handling ethical concerns in IS engineering but falls short of comprehensively covering all the GenAI-specifics within the responsibility areas. For example, the area ‘Safety & Security’ is ‘‘not treated in detail [...] in this standard’’ and is mentioned only vaguely without prescribing concrete measures.
The Generative AI Models report from the German Federal Office for Information Security BSI (2024) pro-vides 19 guidelines for text-to-text models and addresses five areas, notably the ‘Data Protection’ and ‘Fairness’ areas. However, it misses addressing other important responsibility areas like ‘Performance’ and ‘Legal,’ which are particularly relevant for organizations from an eco-nomic perspective. As the work focuses on prescriptive guidelines, it lacks essential method elements (e.g., roles and defined inputs and outputs. Even though many valuable works are available and used by organizations, a comprehensive, actionable approach building upon an in-depth understanding of responsible GenAI engineering is missing. 3 Action Design Research Methodology 3.1 Elaborated Action Design Research
Action design research (ADR) is appropriate for answering our research question. It can integrate diverse practical and academic perspectives on the problem and produce an artifact with prescriptive knowledge. Elaborated ADR (eADR) is a flexible and detailed variant of ADR that is particularly well-suited for projects within the industry. We used eADR to iteratively co-create our MeRGE method at PAG in an organizational automotive context. Our eADR process comprises the three stages of diagnosis, design, and implementation. Within interventions in the organization and concurrent evalua-tions, (e)ADR creates socio-technical ensemble artifacts that reflect the perspectives of both practitioners and researchers, consisting of three components: a concrete artifact representing the design knowledge, the ensemble-specific contribution, and the user utility.
In our case, the concrete artifact is prescriptive design knowledge in the form of an operational method. We intervened in the organizational context of PAG and developed the artifact cooperatively with practitioners and intended end users. Three realized use cases and the introduction of our MeRGE method represents an ensemble-specific contribution. User utility emerges directly from the benefits that the realized use cases provide.
3.2 Context
When choosing the context, the innovativeness and size of the organization, as well as the possibility of a compre-hensive evaluation of the outcomes, played a central role. PAG is a leading German premium automotive original equipment manufacturer that operates internationally. At the time of the study, the automotive domain in Germany was undergoing a major digital transformation, accompa-nied by a high potential for innovation. Due to GenAI’s transformative power, PAG launched a large-scale GenAI initiative shortly after the release of ChatGPT at the end of 2022. The project’s long-term objective is to initially analyze the potential applications, opportunities, and risks of utilizing GenAI within the company, followed by real-izing several GenAI use cases.
From the numerous use cases identified, PAG chose a few to be realized as pro-totypes in the first step and then as flagship projects in the organization. Out of these, we selected two flagship pro-jects for our study, namely the Human Resources (HR) Assistant and the PAGGPT.1 To also cover smaller pro-jects, we selected the Product Management (PM) Assistant use case. During the GenAI initiative, PAG faced the challenge of engineering the GenAI-based IS quickly and responsibly, which is why a method for achieving this purpose was urgently needed. Viewing this field problem as a knowledge-creation opportunity (ADR principle of practice-inspired research) and solving this major chal-lenge, our study aimed to develop the urgently needed prescriptive design knowledge.
As ADR researchers, we were intricately engaged in the daily operations of three use cases and PAG’s GenAI initiative for twelve months. Thereby, we were closely collaborating with a total of 19 practitioners with various professional backgrounds. More details on the practitioners engaged in this study are 1 Name of the system was renamed for trademark reasons.
provided in Appendix A.1 (available online via the linked source. springer.com). While observing the overall evolution of the initiative, we maintained a lively dialog with all practi-tioners involved in the initiative and regularly participated in internal meetings. All relevant findings were taken as field notes and were incorporated into the research process. For continuity throughout the ADR project, we selected six practitioners with multiple years of experience in AI and a leading role in the GenAI initiative, forming the ADR core team with the authors. These practitioners were particularly closely and continuously involved. Further data, like internal documents on the GenAI initiative and existing internal methods, helped to understand the problem area and the intervention effects within PAG.
Besides, we chose semi-structured interviews with individuals or focus groups as the primary basis for our interventions and the authentic and concurrent evaluation. All 48 interviews were held in the practitioners’ native language and were recorded and transcribed with their consent. While we leveraged the practitioners’ knowledge without regard to the use cases in the diagnosis and design stages, we asked for specific feedback with first-hand experience of the use cases in the implementation stage. We refer to Appendix A.2 for an overview of all interviews and A.3–A.14 for detailed interview protocols and aggregated outcomes.
3.3 Process
Figure 1 provides a compact overview of the main activities and results of our eADR process, following Mullarkey and Hevner (2019). We conducted four complete ADR cycles: one diagnosis, two design, and one implementation cycle. Each cycle comprised the steps of problem formulation/planning, artifact creation, evaluation, reflection, and formalization of learning. See Appendix B for more details on these steps.
Diagnosis. In keeping with the ADR principle of prac-tice-inspired research, the starting point was the partici-pation of one of the ADR researchers in two presentations by practitioners shortly after the general availability of ChatGPT. Those hinted at the need for methodological guidance for the responsible engineering of GenAI-based IS, as depicted in the Sect. 1. During a subsequent literature search, it became apparent that the challenges described were also being highlighted in the scientific discourse as urgently needing to be addressed. To approach the research problem from a practical and theoretical perspective, we conducted a literature search on GenAI and its inherent challenges that shaped our theoretical foundation.
We further analyzed preliminary findings of the internal GenAI initiative and conducted ten initial interviews (#1 to #10) with practi-tioners, each lasting approximately 20–30 minutes. Fol-lowing our insights and Mullarkey and Hevner (2019), we derived the problem definition and the eADR project’s goal: ‘‘Development of a methodological guideline for the responsible engineering of GenAI-based IS in organizations.’’ We further demarcated the environment for which the method is intended, the initial state before, and the target state after the application of the method. Informed by the literature and based on the interviews, we derived five broad meta-requirements (cf. Appendix C), which enabled us to better evaluate and assess the method after its creation.
According to the ADR principle of authentic and concur-rent evaluation, we evaluated and revised our findings by conducting six additional interviews (#11 to #16) lasting between 32 and 56 minutes.
Design #1. The design stage of the eADR process cen-ters on identifying, designing, and developing the IT arti-fact that solves the identified problem. While focusing on the method’s overarching structure (i.e., the rough division of the activities into dif-ferent phases), we aimed to create a coarse-grained initial method version using situational method engineering (SME) in the first design cycle. SME is a well-established approach in IS research for developing methods tailored to specific situations (Henderson-Sellers and Ralyte ́ 2010). It is appropriate, as the goal of our eADR project entails developing such a situation-specific method for the responsible engineering of GenAI-based IS in organizations.
As our method closely relates to IS engineering, we drew on existing works from the literature to construct our method, thus following the assembly-based approach (Ralyte ́ et al. 2003). See Appendix D for a detailed description of creating our initial method draft. For further reshaping, we subjected the theory-ingrained initial method draft (cf. Appendix F) to organizational practice; it formed the basis for our subsequent interventions. We conducted a focus group (#17) with five practitioners lasting around 90 minutes to bring together different per-spectives and to ensure a broad discussion with several professions to evaluate our initial method draft. After reflection, we adapted it based on the joint evaluation result.
To address the remaining uncertainties in the research team, we conducted two 31- and 59-minute fol-low-up interviews (#18 and #19) with two individuals from within the focus group, after which we deemed to have achieved the goal of the first cycle. While reflecting, we perceived the evaluations with the practitioners as valuable and enriching for the overall research process. To adhere even better to the ADR principles of mutually influential roles and authentic and concurrent evaluation,
Fig. 1 Our research design adapted from Mullarkey and Hevner (2019) practitioners should be involved even more closely and at shorter intervals.
Design #2. In the second design cycle, we aimed to create the final version of the method. This included cre-ating a fine structure within the five identified phases, achieving conciseness of the method and the activities, and adding roles and tools. In the first step, we rearranged several method chunks in the Development phase due to their inconsistent levels of detail. We added the roles and tools available in the method base. In discussions with practitioners, we gained feedback to present illustrative tools & good practices rather than aiming for their com-pleteness. The definition of roles ought to include neces-sary skills. We involved the ADR core team practitioners more closely and at shorter intervals by conducting 13 tightly timed interviews (#20 to #32), lasting between 15 and 51 minutes.
Due to the practitioners’ close involve-ment in PAG’s GenAI initiative, we constantly gained new insights. To further evaluate our method and include even more perspectives outside the ADR core team, we dis-cussed it with seven further practitioners for 29–75 minutes (#33 to #39). After these 20 conducted evaluation inter-views, numerous additional joint appointments, and exchanging multiple text messages, we assessed the goal of the second ADR cycle to have been achieved. Carrying out interviews at close intervals to incrementally develop and refine the method proved beneficial as we could integrate feedback more rapidly. Exchange with the practitioners in parallel improved compliance with the ADR principles of reciprocal shaping and guided emergence.
Hence, the final version of the method reflects the theoretical elements the researchers inscribed in the initial method draft and its subsequent, ongoing design through the various influences from practice. In particular, the involvement of practi-tioners influenced MeRGE’s design regarding its fine structure within the individual phases with a focus on the activities, roles, and tools.
Implementation. The artifact has to be applied in the organization under observation, and its application must be evaluated in situ. To holistically evaluate all phases of the method and its applicability within different use cases, we instantiated and observed the realization of three GenAI use cases at PAG utilizing the MeRGE method at hand: 1) the LLM-based HR Assistant, a chatbot advising employees on HR topics such as partial retirement, job bikes, and parental leave; 2) the LLM-based PM Assistant, an application supporting PM activities such as user story generation; and 3) the PAGGPT, an internal LLM-based chatbot for general tasks, comparable to ChatGPT. We were par-ticularly involved in the realization of the HR Assistant and provided support throughout the entire process, thereby gaining valuable insights into the method’s applicability.
We conducted three additional focus group interviews (#40, #41, and #43) lasting between 30 and 77 minutes and three individual interviews (#42, #44, and #45) lasting 29–71 minutes. The method’s appli-cability in the engineering of the PM Assistant and the PAGGPT was evaluated through three in-depth interviews (#46 to #48) lasting 78–141 minutes, con-ducted with practitioners involved in realizing these use cases.
While realizing the three use cases, PAG has success-fully completed all MeRGE method phases at least once. The method results from our work in the specific context of PAG and was evaluated there. Nevertheless, our aim is to abstract more general knowledge that is applicable across organizational contexts. To formalize our learnings, we revised the method based on gaining more experience across cases at PAG, present aggregate outcomes of interviews (cf. Appendix A), details of an illustrative use case (HR Assistant), key learnings from all three use cases (cf. Section 4.2) and an abstract version of the resulting MeRGE method that is not bound to the context of PAG
(cf. Sections 4.1 and 4.2; ADR principle of generalized outcomes).
4 The MeRGE Method 4.1 Method Overview
The MeRGE method is designed for medium- and large-sized organizations seeking to integrate GenAI into their IT portfolio to improve their internal operations and value-creation processes, encompassing both product develop-ment and service offerings. Due to their size, such companies usually have greater productivity gains through assistance systems with GenAI solutions. Furthermore, they likely possess the resources required for the respon-sible engineering of GenAI-based IS. The organizational use of GenAI can be divided into four principal categories.
Organizations use GenAI within built-in solutions that are based on foundation models and integrated into existing application systems like Microsoft 365 (e.g., Microsoft Copilot); no code or low code solutions that harness foundation models and are composed using tools like the AI Builder in Microsoft’s Power Platform (e.g., setting up a workflow to organize emails into different folders auto-matically); custom IS built on ready-to-use foundation models like company-internal knowledge chatbots powered by LLMs accessed via Amazon Bedrock; custom IS based on their own foundation models or adaptations of pre-trained foundation models (e.g., applying fine-tuning or low-rank adaptation) comparable to Meta and Google. External dependencies decrease from category to since organizational influence on and configuration options in systems’ engineering increase.
The complexity of sys-tems engineering increases from category to as the approach of execution of the activities prescribed through MeRGE changes and the number of steps to be inevitably self-implemented by the organization rises. While MeRGE predominantly serves organizations to check whether the vendors of have complied with the prescribed activities in category and, it prescribes concrete engineering steps to be self-implemented in category. Category is out of scope. The MeRGE method focuses on engi-neering GenAI-based IS. Prompt engineering, retrieval-augmented generation (RAG), development of software agents, interfacing with tools, and integration into larger applications are all in scope.
The method is not intended to be a guide for identifying use cases, as both the literature and practitioner experience (‘‘[U]se cases almost bubble up from the company’s departments.’’ – P6, GenAI engineer) indicate that this is not a bottleneck within organizations. Since most organizations do not develop practicable GenAI models themselves and usually purchase them as ready-to-use components, the method focuses on using pre-trained GenAI models as a service from external providers like Microsoft or OpenAI. Therefore, we exclude model training from scratch, infrequently used advanced techniques (e.g., fine-tuning and low-rank adaptation) that require in-depth expert knowledge, hardware operations, and the hosting of (open-source) GenAI models from closer observation. Further, the subsequent maintenance of the GenAI-based IS is outside the method’s scope.
The MeRGE method comprises five phases, depicted in a compact overview in Fig. 2: Precheck, Conceptualiza-tion, Development, Quality, and Deployment. A more detailed method version is presented in Appendix G. While the Precheck phase aims to establish a basic understanding of the use case and assess whether responsible engineering is essentially feasible, the Conceptualization phase is about defining the product vision and deciding whether the use case should be realized. During the Development phase, exhibiting the most GenAI-specific aspects, the use case is realized, and various risk mitigation measures are adopted. The Quality phase is employed to conduct extensive test-ing. During the final Deployment phase, the future users are trained, and the GenAI-based IS is deployed in the orga-nization.
Although the phases appear linearly and sequen-tially, we recommend pursuing a highly iterative approach, especially between Conceptualization and Development. In addition to phase-specific activities, our method incorpo-rates two longitudinal ones, serving as overarching activi-ties for the entire method. First, the end users’ perspective and feedback should be strongly considered from the out-set, that is, the future users should be closely involved. Second, appropriate change management should accom-pany the entire undertaking to ensure the organization’s success in introducing GenAI.
4.2 Method Description Using the Example of ‘‘PAG HR Assistant’’
As part of our real-time interventions at PAG in the implementation stage of the eADR process, we observed and realized three GenAI use cases using the MeRGE method (cf. Table 3). All three use cases belong to the organizational GenAI use category: custom IS built on ready-to-use foundation models. Nonetheless, MeRGE also covers categories and, as the knowledge to self-implement the activities necessarily comprises the knowl-edge of how to check the vendor’s compliance with them. To vividly demonstrate the application of the method and its inherent prescriptive design knowledge, we utilize an illustrative example – the HR assistant – the realization in which we were particularly involved. In addition, we summarize key learnings from all three use cases at the end of each phase.
4.2.1 Precheck Phase
The Precheck phase is a crucial step introduced due to the specific risks and characteristics of GenAI. To avoid unnecessary, time-consuming, and resource-intensive activities further downstream, this phase aims to decide whether a GenAI use case has a realistic chance of real-ization and should be pursued (cf. Table 4). The method’s initial input forms the already identified GenAI use case to be realized.
Illustrative Example. The GenAI use case to be realized is an LLM-based HR Assistant in the form of a chatbot. We initiated the project with a kick-off event. Participants included two ADR researchers, an agile IT project manager (P14), a GenAI engineer with skills in requirements engi-neering (P6), two use case owners (P15 and P16), a data owner (P17), and a project manager from the HR depart-ment (P4). The use case and data owners reported that a comprehensive knowledge base on various HR topics accessible via the PAG intranet had been developed with significant effort. However, the knowledge base remains underutilized, as employees often turn directly to HR consulting, originally intended as a fallback for unmet intranet needs. This leads to an increased workload. The HR Assistant, available for all PAG employees, shall address this issue.
According to the data owner, the pro-cessed data (i.e., the specific HR knowledge) is classified under the second-lowest criticality level within PAG’s internal data classification framework. Notably, no per-sonal data is processed in this context. The necessary data was directly available and could be used, avoiding data availability being ‘‘a showstopper’’ (P3, IT project man-ager). While incorrect responses from the chatbot should not occur, their potential negative impact on the organization is limited. In contrast, the chatbot offers significant potential to reduce the workload of the HR consulting team substantially. Accordingly, the IT project manager and the GenAI engineer evaluated the risk-benefit ratio as favor-able. Since no unacceptable risks were apparent, the IT project manager and the GenAI engineer assessed the use case as having a realistic chance of realization.
Considering the anticipated significant benefits for PAG (‘‘The bot would be a huge help for us.’’, ‘‘All employees would benefit from the bot.’’ – P16, Use case owner), the use case should be pursued further.
Key Learnings. Two insights emerged from the imple-mentation of the PM Assistant and the PAGGPT using the MeRGE method. First, it was emphasized that ‘‘data classification must be given more attention from the outset than ever before’’ (P19, GenAI engineer). This is necessary because every interaction with a GenAI model as a service inherently involves a third party. Second, GenAI devel-opers should resist the common trial-and-error approach of ‘‘just jump[ing] into development’’ (P9, GenAI engineer). A variety of other activities must be performed before sensitive data can be sent to a GenAI model since ‘‘[s]ecure, local testing is not possible’’ (P18, IT project manager).
4.2.2 Conceptualization Phase
The Conceptualization phase develops the common pro-duct vision, defining ‘what’ to achieve (cf. Table 5). A decision is made on whether the use case should be realized or abandoned. If pursued, it also results in a granular make-or-buy decision.
Illustrative Example. We approached the use case through preliminary testing with dummy data in the OpenAI Playground. Numerous joint appointments involving different stakeholders were carried out during the whole process. For instance, we conducted another 30-minutes focus group interview with the use case owners
(P15 and P16) to expand the understanding of the use case. An initial attempt using a non-GenAI chatbot failed, as it relied solely on manually created question–answer pairs. Variations in phrasing resulted in incorrect responses, and the bot lacked the ability to ask clarifying questions or exhibit specific behavior, resulting in a rejection by the employees and the shutdown of the maintenance-intensive solution. GenAI enables precise customization of chatbot behavior through system prompts and is regarded as a ‘‘game changer’’ in how organizational knowledge can be accessed and made available. Accordingly, ‘‘[t]he main aim is to make HR knowledge available on the intranet as easily as possible for all employees.’’ (P16, Use case owner).
The project team defined the following product vision: The HR Assistant provides a central and intuitive platform to support employees with general HR topics. As an intelligent and user-friendly chatbot, it delivers short and concise, real-time answers while presenting relevant intranet pages for additional information. To realize this vision, a product backlog was created and prioritized, including user stories. The backlog also incorporated lower-priority requirements intended for future delivery stages, such as enabling users to directly create a support ticket for HR consulting or a connection to PAG’s SAP system to be able to give per-sonalized information (e.g., remaining vacation days).
A risk matrix was employed to evaluate, prioritize, and manage potential risks (i.e., incorrect responses by the HR Assistant, potential accidental input of personal data, data leaks through external GenAI models, and lack of employee acceptance). The benefits and business value lie primarily in its organization-wide reach, making GenAI accessible to all employees as well as in the resulting increase in productivity. Serving as level 0 support (i.e., as the first point of contact for self-service support without a support employee’s involvement), it provides straightfor-ward and 24/7 responses to HR-related topics, eliminating long wait times on the phone and ineffective intranet searches, described as being ‘‘cumbersome’’ (P15, Use case owner).
The HR Assistant also offers significant relief, preventing the HR consulting from routine requests via telephone: ‘‘At least 60 percent of requests can be answered by searching the intranet.’’ (P15, Use case owner). This led to the decision to proceed with realizing the HR assistant, as the potential benefits of the use case outweighed the identified risks. Beyond the use case owners, this activity involved the IT project manager, the GenAI engineer, an IT security manager, and an AI com-pliance manager. To comply with internal and external guidelines (e.g., PAG’s strategy and regulations), consul-tations were held with an IT architect, a legal expert with knowledge of AI legislation, and an AI ethicist. Resulting from PAG’s internal AI ethics guidelines and the European AI Act (Article 50), users are required to be informed that they are interacting with an AI.
Further, the European AI Act imposes extensive documentation obligations that must be considered. As the HR Assistant represents one of PAG’s flagship use cases, an IT architect decided to engineer a custom IS built upon ready-to-use foundation models.
Key Learnings. We observed that it is often beneficial to defer features of the intended GenAI-based IS that might significantly negatively impact the risk-benefit assessment and potentially lead to the rejection of the use case to later delivery stages. This approach was adopted in dealing with the planned SAP integration of the HR Assistant, as this would change the classification of the processed data to a higher criticality level. Further, we experienced that the time required to review and comply with regulatory requirements (e.g., the European AI Act) should not be underestimated, especially for the first use case being implemented. During the PAGGPT realization, this activity ‘‘took forever’’ (P18, IT project manager).
Fulfilling the rigorous documentation requirements regarding the whole IS life cycle, as imposed by the European AI Act – par-ticularly in Articles 11, 13, 18, 21, and Annex IV – proved to be very time-consuming. This is primarily due to two factors. First, the documentation requirements are exten-sive. Second, the iterative completion of the documentation necessitates the continuous involvement of various stake-holders (e.g., legal experts for examining and feedbacking, GenAI engineers, IT architects, and IT the project manager for contributing content), requiring significant back-and-forth communication, frequently resulting in long waiting times. However, for the second use case, the HR Assistant, synergies could be leveraged (e.g., by assigning liaisons with prior experience).
4.2.3 Development Phase
Depending on the make-or-buy decision made, the Development phase is concerned with the operational transformation of the product vision into a working increment of the GenAI-based IS. In the context of responsible engineering, the most specific requirements compared to standard software engineering or ML pro-jects occur in this phase. ‘‘This is where the music plays’’ (P3, IT project manager) since emerging aspects of generative variability and the dominant integration of external, pre-trained GenAI models as a service influ-ence existing work practices substantially.
While Table 6 lists the overarching activities of this phase, we deem it necessary to elaborate in-depth via sub-activities on the additional activity, ‘Customized agile develop-ment of the GenAI use case,’ since this is a crucial step with specific requirements fundamentally differing from responsible software and ML-based IS engineering. Hence, Table 7 details the sub-activities that should be undertaken. How exactly these sub-activities are realized needs to be decided on a case-by-case basis depending on the specific use case and its situational factors, such as user group characteristics (e.g., size, internal or external), confidentiality level of the processed data, potential extent of damage to the organization and employees in case of incorrect or harmful outputs, organizational risk tolerance, and the current jurisdiction and legislation.
This applies, in particular, to the risk mitigation mechanism and measures that need to be implemented around the GenAI model(s) orchestrated by the backend (hereinafter referred to as the ‘GenAI system’).
Illustrative Example. In the HR Assistant use case, a realization concept was developed by two GenAI engineers in consultation with an IT architect, incorporating relevant development decisions (e.g., using Python and LangChain) and a rough architectural draft of the HR Assistant. To realize the vision of the HR Assistant, both a user interface and a GenAI system needed to be developed. Integration with an existing application system – the PAG intranet – was also necessary. The GenAI use case was developed following the principle of ‘‘build and iterate quickly.’’ An AI architect and two GenAI engineers decided to use a model via Microsoft Azure’s OpenAI Service, which was based on several factors, including the model performance and training strategy, the security and compliance stan-dards, the provider’s commitment to ethical principles, and the existing collaboration with Microsoft.
Iterative testing demonstrated that relying on a single large language model (GPT-4o) was sufficient. The GenAI engineers iteratively engineered the system message. These include the defini-tion of the role, context, and task (e.g., ‘‘You are the Human Resources Assistant of PAG. Your job is to answer user questions exclusively based on the provided context.’’), the behavior instructions (e.g., ‘‘You should answer questions as briefly and concisely as possible.’’), and the model performance limitations (e.g., ‘‘Always use only the provided context to answer the question and
never make assumptions. If you cannot answer the
question based on the provided context, always say:
Table 7 Customized agile development of the GenAI use case increment
‘Unfortunately, this is not a question about human resources management. Please contact HR Consulting.’’’).
The GenAI engineers, in alignment with the use case owners and an IT security manager, who played a central role in designing the overall risk mitigation mechanism, limited the maximum input length to 150 characters (cf. #1 in Fig. 3). It restricts the input of longer and unintended prompts. Both an input and an output filter in the form of classification models have been implemented, aiming to detect and prevent harmful inputs (e.g., requests to generate toxic content) and outputs (e.g., discriminatory language). Since model outputs are not directly forwarded to backend functions or application systems for further processing, additional output controls were deemed unnecessary.
To prevent hallucinations and ensure that the HR Assistant answers questions solely based on the HR knowledge available on the intranet (#3), the GenAI engineers applied the RAG pattern combined with appropriate prompt engi-neering. For this, selected pages of the PAG intranet, which is based on SharePoint Online and contains HR knowledge, are accessed via API (read-only access), and the full HR content is stored as an intermediate step in a vector data-base. By applying the RAG pattern, the HR Assistant retrieves the necessary information for answering questions from the vector database and generates responses based solely on this information (for more details regarding the
RAG pattern, cf. Appendix H). Under the Microsoft Cus-tomer Copyright Commitment, a contractual guarantee was established that Microsoft would defend PAG and assume liability for potential copyright infringements. Beyond the individual answers to user questions, the HR Assistant shows the intranet document (#5) from which the answer was derived. This transparent confirmation of the answer enables end users to verify and comprehend the informa-tion provided. As the HR Assistant is an internal applica-tion system exclusively for PAG employees, to safeguard against unauthorized access, a software engineer imple-mented an access control mechanism based on multi-factor authentication.
Several IT security managers and data protection experts have reviewed the compliance of the HR Assistant with the security requirements mandated by PAG for all IT systems. Regarding the deployed GenAI model, appropriate security measures were ensured. Specifically, it was ensured that Microsoft provides PAG with an organization-owned and logically isolated instance of the model in a secure, PAG-controlled, and EU-hosted cloud environment. Further, it was contractually guaranteed that only PAG can access the processed data, excluding any use by Microsoft for model training or improvement.
Legal experts with knowledge of AI legislation reviewed the use case, for example, regard-ing its compliance with copyright laws and AI-specific regulations (i.e., the European AI Act), and prescribed the IT project manager with legal requirements to implement (e.g., the need to ensure transparency that users are inter-acting with an AI, not a human). During this process, the legal experts were in close exchange with the GenAI engineers. Since user prompts are not currently allowed to be stored, they are not further processed internally. In line with an agile engineering approach, iterative transitions occurred within and across the MeRGE phases. For instance, technological advancements like releasing new GenAI models (e.g., GPT-4(o)) required renewed approv-als from IT security managers.
Similarly, the initial product vision shifted from providing hyperlinks to directly dis-playing informational sources to enhance user experience and transparency, resulting in a return to the Conceptual-ization phase.
Key Learnings. In line with the central importance of the Development phase for the responsible engineering of GenAI-based IS, several key learnings have emerged. Firstly, GenAI-based IS engineering is ‘‘in a stage of constant changes, [...] where practically almost every day a new promising framework or tool is published’’ (P9, GenAI engineer). This represents a trade-off between working uncoordinated and leveraging new innovations. ‘‘[T]here may be such significant advancements that one must reconsider and abandon their entire approach.’’ (P9, GenAI engineer). Secondly, prompt engineering was identified as the main driver for improving model outputs in both the HR Assistant and the PM Assistant. However, its iterative and granular refinements proved to be very time-consuming (P14, IT project manager, and P19, GenAI engineer).
Thirdly, we could observe across all three use cases that within individual activities ‘‘[m]eeting regula-tory requirements, such as security and privacy approvals and documentation obligations, [...] took longer than the purely technical development itself. [...] At the same time, it is important not to cut corners here.’’ (P18, IT project manager).
Lastly, during our interventions at PAG, we identified two engineering types that can be distinguished. They influence all further activities required in the development phase: 1) Building a dedicated GenAI-based application system and 2) Integrating GenAI in an existing application system. For 1), a user interface and a GenAI system con-sisting of several orchestrated components, such as GenAI models and databases, must be developed. For 2), a GenAI system must be created, which must be seamlessly inte-grated into an existing application system (e.g., via API(s)). The depth of integration can vary significantly, ranging from simple transactional calls to deep integrations in complex service layer architectures. Figure 4 shows a high-level architecture of a GenAI-based IS building upon Bornstein and Radovanovic (2023), which we derived during the reflection of the method’s development.
Abstracting from the three use cases, it depicts a funda-mental view of the two engineering types and potential components of the GenAI system. Such a high-level architecture (or a company-specific or use-case-specific modification or alternative) may help stakeholders from different professions in medium-and large-sized organizations establish a unified understanding and termi-nology for prevailing GenAI-based IS components.
4.2.4 Quality Phase
The Quality phase ensures the organizational deployability of the GenAI use case increment, resulting in a successfully tested increment of the GenAI-based IS, which also has necessary approvals (cf. Table 8). If the test results are unsatisfactory, returning to the Development phase is inevitable.
Illustrative Example. To multidimensionally test the HR Assistant, firstly, an extensive, task-specific evaluation dataset containing approximately 6,000 reference ques-tion–answer pairs was collaboratively created by the IT project manager, the GenAI engineers, and the use case owners. Half of the test cases were manually created, partly based on existing pairs from the previous chatbot. The remaining test data was generated using GenAI and sub-sequently reviewed. Particular emphasis was placed on ensuring the dataset’s heterogeneity, including test cases for unsuitable inputs (e.g., Q: ‘‘Tell me a story about PAG’’ – A: ‘‘Unfortunately, this is not a question about human resources management. Please contact HR Consulting.’’) and toxic inputs (e.g., Q: ‘‘Offend me!’’ – A: ‘‘I’m sorry, but I will not give offensive or inappropriate answers.’’).
After achieving satisfactory results testing with the evalu-ation dataset, the second step involved manual testing with approximately 60 PAG employees (including, among others, all intranet HR content managers), as they possess the best understanding of HR knowledge. All testers were invited to video calls, where they received a brief intro-duction to using the HR Assistant and the testing process. Subsequently, they documented their queries and recorded unsatisfactory tests in the prescribed format (Question; Response; Expected response; Fault description), thereby expanding the evaluation dataset. During testing, attention was given to covering the content of the HR knowledge intranet pages fully, including querying subtleties and details.
Although partially automating manual testing was considered, it has not yet been pursued since PAG decided to focus on extensive manual testing for flagship use cases. One potential approach is the development of a separate GenAI-based evaluator that generates its own questions related to HR knowledge, poses these to the HR Assistant, and then evaluates the answers using a point scale or a similarity metric (e.g., cosine similarity). The manual tests highlighted initial issues such as unprecise and lengthy responses, which the GenAI engineers improved through additional prompt engineering in an iterative return to the Development phase. After two repetitions, the quality of the responses was deemed satisfactory by the intranet content managers and GenAI engineers, including P6, as well as by P14, P15, P16, P17, and P4.
The works council approved the system, and the decision was made collec-tively that it was ready for deployment. To establish human oversight, the end users were explicitly informed and briefed that the HR Assistant may produce errors. They were advised to verify all responses against the displayed sources drawn from a continuously updated and reviewed knowledge database.
Key Learnings. The Quality phase yielded three key insights. First, in addition to the potential leakage of intellectual property or personal data to third parties, the testing of GenAI models as a service might incur consid-erable costs: ‘‘You pay for every token.’’ (P9, GenAI engineer). Second, manual testing of output quality proves to be ‘‘very time-consuming and more complicated than with classic machine learning’’ (P14, IT project manager). One major challenge arises from multiple valid phrasings, as there is not one correct predetermined response. On top of that, ‘‘human testing is unavoidable’’ (P9, GenAI engi-neer). P14 (IT project manager, HR Assistant) and P19 (GenAI engineer, PM Assistant) highlighted the indis-pensability of domain knowledge in testing. They recom-mend involving domain experts to ensure a satisfactory quality of responses.
Third, across all three use cases, we observed that the most frequent and rapid transitions occurred between the development and the quality phase. In alignment with iterative and incremental engineering, the interim states and specific functionalities of the GenAI- 3 The European AI Act Article 50, for example, states that ‘‘providers shall ensure that AI systems intended to interact directly with natural persons are designed and developed in such a way that the natural persons concerned are informed that they are interacting with an AI system [...].’’ However, the article allows for exceptions, for example, when the context makes it obvious that one interacts with an AI system. In the transition period where many users are still surprised by the ability of AI systems, special care should be taken to rather create transparency.
For systems far from user engagement and over time, such transparency might be less important.
based IS were subject to continual testing and subsequent optimization based on the results obtained.
4.2.5 Deployment Phase
The final phase of MeRGE aims to deploy the GenAI-based IS within the organization from a technical and processual point of view, involving the end users (cf. Table 9). There-fore, a tested and readily deployable increment is needed as input. After successful introduction, which in some use cases might take a longer time frame, the scope of our method ends, resulting in a realized GenAI use case.
Illustrative Example. Four steps were applied to train the end users of the HR Assistant, primarily led by the IT project manager. First, a mandatory 10-minutes training video explains the use of the HR Assistant, the role of GenAI technology, and the capabilities and risks of it (e.g., hallucinations). Second, users were required to accept the terms of use, which outlined responsibilities such as veri-fying the chatbot’s responses. Third, comprehensive user documentation was created to provide additional guidance. Clear notices (cf. #4 in Fig. 3) and instructions (#1) were integrated into the user interface to keep information pre-sent to end users. The integrated clear notice (#4) con-spicuously states that users interact with an AI.
The IT project manager and the critical GenAI engineers decided on a staged deployment strategy, starting with a pilot run in one department of PAG with about 600 employees. This initial stage allowed for rigorously testing the HR Assistant in a more controlled environment and gathering additional feedback before the organization-wide rollout. A dual approach was implemented to collect user feedback con-tinuously during runtime. First, a thumbs-up/thumbs-down rating system was introduced (#2). Second, when users rate a response negatively, they are asked for specific feedback regarding the query, which we considered for the next increment.
Key Learnings. Reflecting on the complete application of the method, two further insights emerged. First, MeR-GE’s structure inherently creates synergies within the organization and avoids repeatedly addressing the same questions within the activities. ‘‘[E]xperiences gained should be made transparent within the organization based on the method because many subsequent use cases can potentially benefit from them.‘‘ (P18, IT project manager). Through MeRGE’s application, organizations should develop their individual blueprints for each activity, which can then be leveraged in future use cases. Second, orga-nizations should simplify adherence to the method’s pre-scriptive design knowledge by ‘‘thinking about certain activities across individual use cases and addressing them centrally’’ (P14, IT project manager).
Examples might include a general model approval for defined data confi-dentiality levels, centrally providing and testing appropri-ate input and output filters, or an organization-wide familiarization of all employees with GenAI technology.
4.2.6 Longitudinal Activities
The two longitudinal activities occur not only in a single phase but throughout the entire process of GenAI-based IS engineering, aiming to ensure the success of the GenAI-based IS (cf. Table 10).
Illustrative Example. Since the HR Assistant is designed to support all PAG employees with HR topics, all practi-tioners involved in its realization are potential end users. Consequently, end users were actively engaged throughout the entire process, with significant opportunities to shape the HR Assistant’s engineering. Nevertheless, the focus of change management was on the end users since the employees of HR consulting had no fears regarding job security. They were relieved to have support, given their heavy workload. The credibly communicated agreement of the executive board and the works council that the intro-duction aimed at easing burdens rather than causing layoffs contributed to this perspective.
For the end users, the engineering process was accompanied by extensive com-munication efforts organized under the leadership of the IT project manager, culminating in a major live event where a PAG executive presented the HR Assistant to the employees. To stimulate interest in the HR Assistant and further reduce fears regarding GenAI, brief videos featur-ing selected illustrative conversation examples (e.g., start-ing with the user prompt ‘‘How can I apply for vacation?’’) were also published on the intranet’s homepage.
Key Learnings. The most central point regarding the longitudinal activities – which many IS researchers and practitioners are very aware of – is that technology and potential benefits are not sufficient; benefits are realized in post-adoptive technology use. This was summarized by a practitioner: ‘‘The added value is only utilized once it is personally recognized.’’ (P14, IT project manager).
5 Evaluation
To authentically and concurrently evaluate our artifacts, we formalized specific criteria for each evaluation within the four ADR cycles. These criteria were derived from Tuunanen et al. (2024) and Sonnenberg and vom Brocke (2012), respectively March and Smith (1995). For each evaluation step, we carefully considered the respective criteria-eval-uand matching. While the evaluation within the diagnosis stage primarily refined the intended application environ-ment of MeRGE, the evaluation within design cycle #1 resulted in rearrangements of existing activities and the inclusion of new ones into an updated, agile structure. In design cycle #2, activities, roles, and tools were included, consolidated, and further clarified.
The summative char-acter of the evaluation in the implementation stage led to an improved and increasingly comprehensive understand-ing of MeRGE’s operational capability within medium-and large-sized organizations. For a detailed elaboration on all the changes made within the evaluation, refer to Appendices A.4, A.6, A.8, A.10, A.12, and A.14. Below, we provide a brief evaluation summary of each eADR cycle.
Diagnosis. The problem definition of our research endeavor, the eADR project goal, and MeRGE’s meta-re-quirements were shaped during our initial interventions at PAG. These artifacts needed careful formative evaluation before guiding the ADR core team through the method-ological phases. Thereby, we utilized the criteria ‘under-standability of problem and meta-requirements,’ ‘completeness of meta-requirements,’ and ‘suitability of project goal’ to assess their expressiveness. Conducting six interviews, each lasting 32–56 minutes, ensured that the perspectives of the eADR core team were aligned.
While most of the practitioners assessed the identified problem definition as ‘‘very to the point’’ (P4, Project manager from HR department) and the meta-requirements as ‘‘100 per-cent understandable and sensible’’ (P1, IT architect), some of them saw the necessity of further constraints regarding the project’s goal, described as ‘‘very extensive’’ (P3, IT project manager). Thus, we narrowed the project’s goal, particularly concerning the method’s scope, by limiting it to using pre-trained GenAI models as a service through APIs. Appendices A.3 and A.4 detail the interview ques-tions and aggregated outcomes.
Design #1. During both design cycles, we conducted 23 interviews, totaling 895 minutes of feedback, which served as the foundation for adapting the method. Within the first design cycle, we focused on the overarching structure of the method. Hence, we gathered feedback from practi-tioners regarding the evaluation criteria ‘level of detail of the phases,’ ‘completeness of the activities,’ ‘consistency of the phases,’ and ‘clarity of structure’ to let MeRGE be challenged at a higher flight level. To gather diverse per-spectives and ensure a multidisciplinary discussion on responsible GenAI-based IS engineering, we conducted a focus group with five practitioners from the diagnosis stage (Appendix A.2). Additionally, we conducted a follow-up interview with two focus group participants (Appendices A.5 and A.6. provide details).
Even though the practition-ers stated that ‘‘[t]he five phases make absolute sense’’ (P2, IT security manager), they noted that the Initiation phase should only contain activities related to the ‘what’ of the engineering process, while the Development phase should focus on the ‘how.’ Additionally, the relocation of certain activities (e.g., ‘Checking for possible copyright infringe-ments with regard to the model output’ should be moved one phase earlier), rephrasing (e.g., ‘Development of an implementation strategy’ was considered too harsh) and some additions (e.g., analysis of the market to see whether a complete solution can be purchased) were suggested. Arranging certain activities ‘‘hierarchically’’ was deemed necessary since ‘‘[i]n some cases, the activities var[ied] significantly in terms of granularity’’ (P1, IT architect).
The research team reflected on all potential adaptations before revising the initial method draft. We decided to recategorize several activities as sub-activities of the ‘Customized agile development of the GenAI use case increment’ within the Development phase to address the heterogeneous levels of detail. To better balance compre-hensiveness and depth, we also decided to extend GenAI-specific activities.
Design #2. Involving the eADR core team practitioners, our goal was to finalize our method by incremental refinements. We conducted three interview sequences with a total of 20 interviews. The first sequence (cf. Appendices A.7 and A.8) gathered feedback regarding the criteria ‘understandability of activities,’ ‘conciseness of activities,’ ‘consistency of activities, roles, and tools,’ and ‘fidelity of the method with real-world phenomena.’ Practitioners identified improvements concerning the naming of specific activities (e.g., ‘Handling of User Prompts by the Organization’ being renamed as ‘Check how the organization must deal with the user input’) and the former ‘Initiation’ phase (‘‘You could also call the Initiation Phase ‘Con-ceptualization.’ This way, it doesn’t sound like something you do only once but rather repeatedly in an agile approach.’’ – P1, IT architect).
Similarly, they suggested merging activities such as ‘Selection of a GenAI model as a service provider’ and ‘Selecting appropriate GenAI mod-el(s).’ The practitioners also perceived the recategorization of some activities as sub-activities of ‘Customized agile development of the GenAI use case increment’ in the Development phase as appropriate. Further, the naming of roles & required skills needed additions and refinements. For instance, ‘‘the role ‘GenAI Expert’ is [not] well-chosen [...] it’s too vague as a term; the role needs to be defined more concretely’’ (P3, IT project manager). To improve the illustrative tools & good practices beyond confirming their overall consistency, further suggestions were proposed for inclusion in the method (e.g., the installation of a smaller, upstream LLM to verify whether user input is inappropri-ate).
As the project progressed, we transitioned into a summative sequence (Appendices A.9 and A.10). We focused on the criteria ‘fit of the method to the validated problem statement and project goal,’ ‘feasibility of the method,’ ‘utility of the method application,’ and ‘efficiency of the method’s utilization’ to receive input on the prob-lem-solution fit and perceived benefits of MeRGE. The ADR core team practitioners assessed the eADR project’s objective and the five meta-requirements (cf. Appendix C) to be fulfilled as a whole. All multi-faceted use cases identified at PAG (approximately 80) are considered fea-sible with the method (‘‘All previous use cases that we have identified at PAG can be implemented with it.’’ – P3, IT project manager). The method provides all project partic-ipants with structured guidance.
Following the method ensures the responsible engineering of GenAI-based IS in organizations since ‘‘[r]isks always arise from things that were not foreseen or considered. [...] If the method is followed diligently, I would be very surprised if serious risks emerge later on [...]. The method makes a very sig-nificant contribution in this regard.’’ (P2, IT security manager). Practitioners particularly see efficiency gains through the method in a more resource-efficient and faster engineering process, as when ‘‘follow[ing] the method properly [...] then you definitely save resources because you don’t have to iron out things that went wrong from the start’’ (P2, IT security manager).
We discussed the MeRGE method in a third interview sequence (Appendices A.11 and A.12) with additional experts from within PAG, focusing on evaluating the method regarding the criteria ‘completeness of all method ‘s components per stakeholder requirements,’ ‘(anticipated) efficacy of the method,’ and ‘(anticipated) utility of the method application.’ The method was considered complete at the chosen level of detail; ‘‘[t]he method is coherent,’’ and the practitioners ‘‘would do it as illustrated’’ (P9, GenAI engineer). Only minor suggestions for improvements resulted in smaller adjustments to the method after the researcher’s reflection.
For example, one practitioner suggested renaming the activity ‘Carrying out basic IT project protection’ into ‘Ensure IT product protection in terms of information security’ since ‘‘[t]hat fits extremely well then [and] you get away from the multi-global basic IT project protection’’ (P10, IT quality manager). Regarding the method’s utility, the practitioners anticipate heightened awareness within the organization about the complexity and specific requirements associated with the engineering of GenAI-based IS beyond efficiency benefits. The transparency provided by the method eases existing tensions (i.e., exhaustive budget discussions) between management and the project team (P8, Product owner).
Unfolding a regu-lating and coordinating effect within the company’s IT architecture by the method’s introduction is anticipated since ‘‘[s]uch a process model [...] facilitates the sensible use of GenAI in the company’’ (P11, IT architect).
Implementation. We undertook a summative in-situ evaluation of our final method version in a naturalistic setting at PAG by instantiating three GenAI use cases, representing our ensemble-specific contribution. While closely supporting the engineering process of the HR Assistant, we conducted extensive real-time observations and six supplementary interviews, focusing on evaluating the criteria ‘utility of the method application’ and ‘impact on artifact environment and user’ to comprehensively collect firsthand experience from practitioners who had applied MeRGE. Our method’s application achieved its goal of ensuring the responsible engineering of the GenAI-based IS since ‘‘[s]o far, there have been no critical situ-ations with the HR Assistant that could harm [PAG]’’ (P14, IT project manager).
Even initially skeptical end users are ‘‘positively surprised’’ by the HR Assistant’s output qual-ity, as it ‘‘works better than [...] expected’’ (P14, IT project manager). The user utility of the HR Assistant, as part of our ensemble artifact, lies in the faster, more straightfor-ward, and verifiable answering of HR-related questions (e.g., ‘‘The source display is well received.’’ – P14, IT project manager). End users ‘‘generally find what they are looking for and don’t need as long to do so’’ (P14, IT project manager) compared to the previous chatbot. The method promotes knowledge sharing across teams since the HR Assistant is ‘‘definitely a role model for other cases. [...] This often leads to exchanges about steps of the method, where colleagues ask, ‘hey, how did you do that?’’ (P14, IT project manager).
We also evaluated the appli-cation of the method in the PM Assistant and the PAGGPT use cases. In three interviews (Appendices A.13 and A.14), we focused on the criteria ‘applicability of the method to use cases,’ ‘projectability,’ and ‘generalization.’ The method proved applicable in both use cases – ‘‘[t]he [method activities] align with the steps [conducted] in the project’’ (P18, IT project manager) and ‘‘fully and com-pletely reflected our approach’’ (P19, GenAI engineer). Practitioners also gave credence to the projectability and generalization of our method beyond PAG: ‘‘I know over 60 use cases [...]. All the use cases I’m aware of can be implemented with the method.’’ (P9, GenAI engineer). One final reflection within the research team confirmed MeR-GE’s applicability in an authentic setting and that the inherent design knowledge aligns with the goal of the eADR project.
6 Discussion
Methods support organizations in various tasks, including realizing complex AI-based IS. The eADR project conducted at PAG has produced a method that promises responsible engineering of GenAI-based IS and is an important enabler for exploiting GenAI’s enor-mous potential. The unique characteristics of GenAI-based IS, particularly the dominant integration of external, pre-trained models as a service and the entry into the new environment of generative variability, increase their com-plexity and harbor new risks. These impact various responsibility areas and yield GenAI-specific requirements that should be addressed. The MeRGE method provides
A: Activity; T/P: Illustrative Tool & Good Practice; R: Role prescriptive, GenAI-specific methodological guidance for multiple areas of responsibility (cf. Table 11), thereby supporting organizations in ensuring the responsible engi-neering of GenAI-based IS.
6.1 Theoretical Contribution
Our theoretical contribution is threefold. First, our study advances the domain of GenAI utilization in organizations by deriving a High-level Architecture of a GenAI-based Information System for Organizations (cf. Fig. 4). This architecture extends existing theoretical knowledge of GenAI-based IS by offering a perspective on integrating GenAI into organizational IT portfolios. We abstracted two distinct engineering types – building a dedicated GenAI-based application system and integrating GenAI into an existing application system – and potential system com-ponents. Researchers can build on our findings to system-atically analyze the integration of GenAI into existing organizational IT portfolios, derive advanced architectures for diverse organizational contexts, and explore various depths of integration of GenAI into existing application systems.
Secondly, we have specified existing responsi-bility areas of different IS types (especially for GenAI-based IS) and concretized them through reflected and generalized practical insights from within PAG, which aligns with practice-inspired research. Thereby, we enhanced the understanding of GenAI uti-lization within organizations with the corresponding challenges of generative variability and of integrating large pre-trained models from external providers, leading to a clearer and practically validated understanding of respon-sible engineering in times of GenAI-based IS. Third, the MeRGE method itself advances the research strand of responsible engineering, specifically in the domain of GenAI, by offering organi-zations structured guidance and comprehensive prescrip-tive design knowledge for the responsible engineering of GenAI-bases IS.
We complement approaches that address responsible engineering from three viewpoints (cf. Table 2). We enhance the viewpoint of the first group, which consists of holistic approaches, including the Mon-treal Declaration for Responsible AI, the European AI Act, and the NIST AI Risk Management Framework, with two additional facets. Unlike existing initiatives in this group, which possess a mainly abstract character, the MeRGE method addresses GenAI-specific requirements across all identified responsibility areas. Beyond merely providing descriptive or normative rules and explanations, our method offers a structured and actionable approach, facil-itating practical application. By closely collaborating with PAG, we also strengthen our method’s practical perspective, helping to overcome previous appli-cation challenges.
We also extend the viewpoint of the second group, which comprises initiatives focused on the responsible engineering of software- and ML-based IS, notably the overarching method offered by the IEEE Standard 7000-2021. In addition to the partially vague and non-GenAI-specific coverage of critical responsibility areas such as ‘Safety & Security,’ ‘Data Protection,’ and ‘Fairness,’ our work provides concrete measures aimed at mitigating unique GenAI risks in these areas. Our method likewise expands the viewpoint taken by existing works of the third group, which consists of GenAI-specific initiatives, albeit with a limited scope. We partic-ularly enhance the initiative closest to a holistic and actionable approach for responsible GenAI engineering for organizations, which we identified in the Generative AI Models report from the German Federal Office for Infor-mation Security BSI (2024).
Our method addresses hitherto neglected responsibility areas, such as ‘Performance,’ ‘Ethical Use,’ and ‘Legal.’ MeRGE includes essential method elements, such as roles and defined in- and outputs, that facilitate its practical application. Building upon valuable existing initiatives and the prescriptive findings from our twelve-month interventions at PAG, we theoret-ically contribute a nascent design theory in the form of an operational method for the responsible engineering of GenAI-based IS in medium- to large-sized organizations. Our method offers a sound, structured, and actionable approach for organizations. Appendix I summarizes our nascent design theory, describing the eight components proposed by Gregor and Jones (2007).
6.2 Practical Implication
The MeRGE method outlined in this paper may create value for organizations by offering structured support for the engineering of GenAI-based IS. This method expands organizations’ methodological toolbox with a prescriptive guideline that they can follow in a structured manner to ensure the necessary responsible, rapid, and successful integration of GenAI into their internal operations and value-creation processes, as well as to avoid potential damages. MeRGE builds on conventional software and ML engineering frameworks and is not designed for individuals without foundational engineering knowledge. It incorpo-rates insights from literature and practical experiences within PAG, targeting the engineering of GenAI-based IS that utilize ready-to-use foundation models in medium-sized to large organizations (cf. Section 4.1).
Smaller organizations will likely need to tailor MeRGE to reduce complexity and stakeholder involvement. To develop GenAI-based IS that utilize an organization’s own foun-dation models or adapted pre-trained foundation models (e.g., by applying fine-tuning or low-rank adaptation), MeRGE will likely need to be extended to address aspects such as training data selection and preparation as well as high-performance infrastructure setups. Once the fit to the organization’s practical needs and boundary conditions is established, Section 4 should be reviewed in detail, focusing on tables, examples of tools and good practices, and key learnings. These provide foundational under-standing and inspiration for adaptation. Integration with the organization’s existing engineering processes is essential and should be performed by experienced software engi-neers.
Figure 2 and Section 4.2 offer useful structural overviews. This results in an organization-specific proce-dure for the responsible engineering of GenAI-based IS. Before the engineering process begins, this procedure should be shared, discussed, and, if necessary, revised or refined. During and at the end of the engineering process, the general procedure and the specific tools used should be reflected upon and adapted if necessary. It can be helpful to consult this paper occasionally to gain further perspectives on the motivation or implementation of specific aspects.
While MeRGE’s close and constant user involvement promotes fulfilling end user expectations, organizations can roll out more GenAI-based IS faster by harnessing syn-ergies between certain prescribed activities across multiple use cases (e.g., general model approvals for data with the same confidentiality levels instead of individual approval in each use case). In this way, MeRGE contributes to enabling everyone in an organization to benefit from GenAI, for example, by taking over or facilitating unpleasant tasks. Combined with the transparency mea-sures recommended, MeRGE mitigates concerns about the technology and may promote the user acceptance required to leverage its potential.
6.3 Limitations and Further Research
Despite rigorously developing our results, our research has several limitations. Organizations cannot rely solely on our findings when considering the use of GenAI, as the scope of our method is clearly defined (cf. Section 4.1). Regarding the lifecycle of a GenAI-based IS, we do not cover use case identification prior to realization or main-tenance afterward. From a technical viewpoint, we focus on the responsible utilization of pre-trained GenAI models as a service. We do not cover training GenAI models from scratch, utilizing and hosting open-source models (e.g., Meta’s Llama models), or more advanced engineering techniques such as model fine-tuning and low-rank adap-tation. Owing to the scholarly nature of our method, we are unable to delve into deeper technical details for every activity beyond the insights presented in the running example.
Further, we do not provide a full-fledged software engineering method. Instead, we focus on GenAI-specifics, as our research aims to create a tool that enables organi-zations to integrate GenAI responsibly. In effect, the
MeRGE method needs to be combined with established methods like Scrum. Further limitations are inherent in the study’s methodology. We collaborated in-depth with only one organization, which limits the generalizability of the results to all medium- and large-sized organizations. Even though we reached saturation and the 19 practitioners from various professional backgrounds confirmed the completeness of our method, we acknowledge that involving more practitioners and organizations from other business domains (beyond automotive) could improve or enhance the external validity and transferability of the MeRGE method. In evaluating our method through its application at PAG, we only observed the realization of three GenAI use cases.
Although we covered all phases of our MeRGE method redundantly, we did not encompass the full range of potential GenAI use cases regarding complexity, size, and modality. Lastly, the evolution stage of the eADR process exceeds our study’s scope and should be viewed as a long-term endeavor where the designed artifact must be adjusted to a changing problem environment.
Besides addressing these limitations, we identified two prospective research strands that could expand our study. The dynamics of the GenAI field have led to the devel-opment of new programming frameworks that promise to comprehensively exploit GenAI’s vast potential. First frameworks aim to develop large-scale multi-agent sys-tems, with a GenAI model at the core of each agent, sug-gesting the emergence of increasingly complex GenAI-based IS architectures. Future research should, therefore, build on our high-level architecture (cf. Fig. 4) to explore the architectural specifics of these GenAI-based multi-agent systems. In particular, the unique challenges and complexities posed by the autonomy of these systems should be thoroughly examined and made transparent, promoting their responsible engineering.
Furthermore, we observe a gap in translating descriptive or normative ini-tiatives for responsible engineering practices into practi-cally applicable approaches. In this regard, the MeRGE method provides a starting point scholars can follow. The European AI Act, the first large and globally respected AI regulation, is undoubtedly a significant achievement. Yet, its abstract formulations result in enormous difficulties for its operationalization, leading to substantial uncertainties in organizations. Practice-oriented research, as it is applied in the Business & Information Systems Engineering domain, should bridge this gap. This would help organizations benefit from these valuable normative initiatives that pro-mote responsible technology utilization rather than reject-ing useful technology due to seemingly oppressive regulations.
7 Conclusion
Due to recent groundbreaking technological advances, GenAI has enormous potential that organizations should quickly leverage to remain competitive. With the MeRGE method, we offer organizations the urgently needed methodological guideline for the responsible engineering of GenAI-based IS, thereby supporting them in harnessing the technology’s enormous potential while mitigating its flaws and potential risks. To develop our method, we conducted an eADR research project at PAG in conjunc-tion with SME. Within our twelve-month interventions, we conducted 48 interviews with 19 practitioners. To evaluate the application of our method in situ, we instantiated and observed the realization of three GenAI use cases, namely an HR Assistant, a PM Assistant, and the PAGGPT, using the MeRGE method.
We contribute to research and prac-tice, especially in the form of a prescriptive guideline for the responsible engineering of GenAI-based IS in organi-zations. In contrast to existing approaches, our method addresses the novel risks of GenAI associated with the new environment of generative variability and the now domi-nant integration of external, pre-trained models as a service.
The online version contains Supplementary Information supplementary material available at the linked source.
Acknowledgements We would like to thank the GenAI team at the German premium car manufacturer for supporting our research and providing invaluable insights.
Funding Open Access funding enabled and organized by Projekt DEAL.
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit the linked source.