You’re listening to “Positioning public sector practitioners as ‘moral crumple zones’: Mechanisms in the early use of generative AI work support tools,” by M. Choroszewicz. Published in 2026. Contents lists available at ScienceDirect Government Information Quarterly journal homepage: the linked source Positioning public sector practitioners as ‘moral crumple zones’: Mechanisms in the early use of generative AI work support tools Marta Choroszewicz Department of Social Sciences, University of Eastern Finland, Joensuu Campus, P.O. Box 111, FI-80101 Joensuu, Finland 1. Introduction. The release of large language models (LLMs) marked the beginning of a rapidly intensifying era of generative artificial intelligence (AI) innovation. In Finland, public organizations across various welfare domains have begun exploring the initial use of work support tools built on generative AI to streamline workflows and enhance efficiency. This includes intensive experimentation with both internally developed and commercially available generative AI tools such as Microsoft 365 Copilot (M365 Copilot) and other AI assistants. While the current impact of these developments remains unclear due to insufficient empirical research on their use in public sector, such tools are often portrayed as enablers of increased productivity, greater efficiency and even improved worker well-being. Research to date has shown that emerging technologies' innovation and integration in public organizations often reflect underlying power dynamics and perspective disparities. Managers tend to embrace them as catalysts for change, while frontline workers tend to perceive them as impositions that threaten their professional autonomy and control over tasks. Meanwhile, future users often have limited impact over shaping new technologies. Furthermore, emerging technologies have sparked concerns about their potential for eroding the skills and expertise that frontline workers have cultivated over time. New digital systems have been observed to standardize work practices in ways that may hinder the provision of welfare support tailored to individual citizens' needs. Many critical welfare services demand nuanced human judgement; therefore, it has been argued that human workers cannot be easily replaced or supplemented by machines. Furthermore, overreliance on AI systems can impair workers' ability to respond to unforeseen challenges and the development of critical thinking skills, an essential aspect of many public sector roles. Yet, despite growing institutional enthusiasm towards generative AI tools, there remains insufficient empirical insight into public sector practitioners' experiences of the work-related use of these tools across multiple welfare domains. In the rapidly evolving landscape of generative AI, public sector 0740-624X/© 2026 The Author. Published by Elsevier Inc. This is an open access article under the CC BY license (the linked source). practitioners such as claims specialists, frontline workers and administrative officers may find themselves at the intersection of technological innovation, professional practice and human welfare, as they are increasingly being held accountable for the use of LLM-based work support tools. The deployment of these emerging tools in public sector raises profound questions about who is responsible, especially when human oversight is both required but also taken for granted. Before these tools start to wield undue influence over the functioning of public sector and the delivery of public services and benefits, we need to address the challenge of determining not only who is accountable when things go wrong but also who is responsible for the adequate functioning and use of these tools. This issue is highly topical, as both public sector practitioners and everyday citizens may soon be interacting with LLM-based AI assistants. Furthermore, the issue is specifically concerning given the growing body of literature on the opacity of AI systems. In light of these concerns and inspired by Madeleine Clare Elish's (2019) concept of the ‘moral crumple zones’ (MCZ), this article explores the current practices of responsibility allocation and perspectives on human oversight of generative AI tools in public sector.1 Specifically, it asks how public sector practitioners are positioned as MCZ for the effective and ethical use of LLM-based tools in public sector? Elish (2019) borrowed the MCZ term from automobile safety design, where a crumple zone refers to the part of a car designed to absorb impact during a crash, thereby protecting its occupants from harm. In her article on technology accidents, Elish illustrates how proximate human operators in complex human-machine systems often absorb moral and legal responsibility when failures or accidents occur, despite their having little actual control over the system's design or functioning. Such failures typically arise from design flaws, programming errors, inadequate training, systemic complexity and opacity or even a series of unfortunate coincidences. Nevertheless, human operators tend to become the moral buffers between the public and these systems, absorbing blame and shielding the systems' designers, developers and providers from accountability. Drawing on Elish's study, this article illustrates how the MCZ phenomenon also occurs in the use of generative AI work support tools in public sector by highlighting the prevalent narrative that takes human oversight and responsibility for granted, while testers tend to absorb blame for poor-quality AI outputs or for falling to find the tools beneficial. Furthermore, it identifies challenges in responsibility allocation to testers by capturing blind spots in generative AI innovation. The analysis uses ethnographic data (fieldnotes and interviews) covering experimentation with M365 Copilot and the development of two internally created AI assistants in three welfare domains: administrative work, client service in social work and decision-making. The article thus highlights how public sector practitioners are often positioned as MCZ, shielding generative AI work support tools from broader scrutiny and criticism. These tools are often added on top of existing digital infrastructures, thereby reflecting broader tendencies of digital layering in public sector. In addition, LLM-based tools in public sector are frequently positioned as solutions to a persistent administrative challenge: the need to manage, retrieve, process and interpret vast amounts of complex, fragmented and constantly changing information from multiple sources. However, as the findings reveal, the very problems these tools were meant to solve persisted and became central to their use. This paradox resulted in three forms of opacity for users—epistemic, interpretative and operational. This article makes three key contributions to the research on public sector in the age of generative AI work support tools. First, it demonstrates how the MCZ concept emerges as relevant for describing the position of public sector practitioners who are the testers of emerging LLM-based tools; these practitioners are perceived as responsible for AI output generation, inspection, appropriate interpretation and use. This study extends the MCZ concept to the context of public sector, focusing specifically on the early stages of integrating generative AI tools. Second, the study captures hidden pressures and unrealistic expectations placed on testers, echoing wider critiques of AI innovation as a form of human experimentation lacking adequate oversight. The findings underscore the need for critical assessment of the technical utility of these tools and their ethical and professional implications for sustainable professional practice. Third, the article offers empirical insights into how the technical opacity of generative AI work support tools manifests in the everyday work of those using these tools and how the tools may contribute to reinforcing rather than transforming traditional hierarchies in public sector. 2. Theoretical framework: Operationalizing MCZ for LLM-based. AI tools in public sector This study is anchored in Elish's (2019) MCZ concept, which describes how proximate human operators absorb moral and legal responsibility for failures in complex human–machine systems, despite having limited control over these systems' design and functioning. The concept foregrounds misaligned accountability and advocates distributing responsibility among stakeholders, including designers, operators and institutions. Furthermore, Elish highlights that—often symbolically—placing humans ‘in the loop’ legally and ethically shields other actors responsible for design choices, technology development and regulatory frameworks that prioritize technological progress over clear accountability pathways. This study extends the application of the MCZ concept from early automation contexts to LLM-based tools in public sector to examine how public sector practitioners are positioned as MCZs when obtaining, evaluating, interpreting and integrating high-quality AI outputs into administrative work, decision-making and client counselling. While the MCZ concept is a powerful diagnostic tool, it needs to be combined with other research strands to build an analytical framework (see Fig. 1) that enables capturing obstacles to meaningful human oversight and the execution of responsibility as liability in the context of LLM-based tools in public sector. In this paper, MCZ positioning of public sector practitioners is conceptualized as the disproportionate allocation of responsibility to LLM-based tools' users relative to their actual authority, control and capacity to influence these tools' design, functioning and quality of outputs. Consequently, failures to obtain high-quality and useful AI outputs, or their improper or ethical use, can be framed as user errors. To further specify how MCZ positioning emerges in the use of LLM-based tools in Finnish public sector, this study draws on two research strands: 1) AI augmentation of professional work and 2) the opacity of AI systems and LLMs. Research on the AI augmentation of professional work helps reveal the masking of a responsibility shift to users and highlights the ‘human-in-the-loop’ approach: the assumption that users have a real understanding of the tools, enjoy decision-making authority and possess sufficient means and resources to be actively involved in the process and intervene at checkpoints, rather than being merely symbolically present. Furthermore, research on the opacity of AI systems and LLMs helps capture the epistemic risks they pose and the limited capacity users have to exercise meaningful influence over them. Ensuring a human-in-the-loop approach in the context of LLM-based tools means that those tools must be technically mature enough for specific work tasks or functions. They must also be adequately transparent in use so that users can engage with them in informed ways; the information base for output generation should be clear (and possibly accessible), links to sources for input generation and reasoning behind outputs should be provided, and users should understand how the tools work, what they do, and why they produce certain outputs during interaction. Such transparency aligns with Amoore's (2020) discussion of the ethical and practical capacity required to engage with algorithmic outputs, not merely by receiving them but by being able to interpret, question and act upon them. Such engagement demands situational judgement, attentiveness and the willingness to intervene meaningfully in algorithmic processes, which resonates with the broader scholarly notion of ‘response-ability’ (see e.g. Barad, 2007; Kleinman, 2012). 2.1. AI augmentation of professional work. The use of AI is increasingly framed through the augmentation perspective, which emphasizes supporting and enhancing human expertise rather than replacing it. The augmentation view promises superior performance and enhanced efficiency stemming from the combination of the complementary strengths of humans and AI. This perspective emphasizes the human-in-the-loop approach, where human judgement and oversight play a central role in accomplishing complex tasks. Scholars describe this collaboration as a dynamic interplay where both parties influence the other, which transforms knowledge by requiring an understanding of others' inputs and a willingness to adapt one's initial stance. Thus, effective augmentation involves not just adding AI inputs but professionals actively challenging and refining machine outputs. For instance, a U.S. study on AI use in diagnostic radiology illustrated how radiologists across three specializations respond to the opacity of AI tools when forming diagnostic judgements. Across three departments, AI initially increased uncertainty due to its lack of transparency. Only with lung cancer diagnosis did professionals achieve engaged augmentation, actively integrating AI outputs with their own expertise through deliberate interrogation practices. This required significant effort and resources. By contrast, in breast cancer and bone age diagnostics, professionals exhibited unengaged augmentation, either ignoring AI input or accepting it passively. These findings highlight that meaningful human–AI collaboration depends on a professional's ability and willingness to engage critically with opaque AI systems. Yet, such engagement might be constrained by time, workload and system design. 2.2. Opacity of AI systems and LLMs as sources of obstacles to users'. control and capacity Concerns are growing about the difficulty associated with understanding how AI systems produce outcomes. These emerging technologies often lack transparency in use, making it difficult for users to interrogate the logic and the processes behind AI output generation (see e.g. Burrell, 2016; Kellogg et al., 2020). This opacity exacerbates the risk of bias reinforcement, a phenomenon in which AI systems mirror and amplify human cognitive biases. Researchers have pointed out that AI systems operate through complex, multilayered algorithms that transform vast datasets into outputs that are often impenetrable to both users and developers. The term ‘black box’ captures this phenomenon, denoting when decisions are made without clear, interpretable reasoning. Even when technical details are available, constraints like time and cognitive capacity limit users' ability to engage meaningfully with the logic of AI systems. Burrell (2016) identified three sources of opacity in AI systems that make their outputs difficult to influence and scrutinize: intentional secrecy, technical illiteracy and machine learning opacity. These sources of opacity are exacerbated in the case of LLMs due to their immense size, stochastic processes and proprietary architectures. Because the internal statistical pathways leading to outputs are inaccessible, it is difficult to interpret and explain those outputs. Recent scholarship has specifically highlighted the epistemic risks posed by LLMs that are not knowledge systems. They generate plausible-sounding text by predicting likely word sequences without a grounding in truth conditions. This design creates epistemic risks related to production of coherent but incorrect or fabricated content that can be difficult to attribute, explain or audit. Tracing a causal chain between an agent's actions and their consequences is nearly impossible because LLMs produce text based on probabilistic associations across vast datasets without any embedded communicative intent or deterministic logic. Therefore, LLMs challenge traditional evaluation metrics and pose risks for oversight and accountability. In this sense, LLM opacity is structurally built into the model design, data scale and training pipelines; it is not merely a temporary transparency deficit. Hannigan, McCarthy, and Spicer (2024) refine this point by reframing opacity as an epistemic problem rather than solely a technical one. Because chatbots predict content rather than know or understand it, their outputs can become what the authors call ‘botshit’: untruthful content that is taken up uncritically by users. Therefore, they propose a task-sensitive framework that classifies chatbot use according to the importance of factual accuracy and the verifiability of responses. This yields four modes—authenticated, autonomous, automated, and augmented—each tied to a specific epistemic risk (i.e. miscalibration, black-boxing, routinization, ignorance) and highlights that even when verification is feasible, users may fail to verify, thereby amplifying epistemic risk. Read together with work on automation bias and selective adherence in public decision-making, LLM opacity is shown to be co-produced by model design and human behaviour. LLM opacity has direct consequences for user control, evaluation and accountability. First, when outputs are synthetic, coherent and persuasive but their provenance and generative pathways remain obscure, users struggle to interrogate claims or contest recommendations, especially under time pressure. Second, opacity complicates auditability and attribution; responsibility can drift towards end users as de facto risk absorbers, even though they lack the practical means to diagnose or correct model errors. Third, established expectations in public sector ethics—explainability and human oversight as preconditions for legitimate deployment—are hard to satisfy when explanations cannot be grounded in tractable causal accounts of model behaviour. This generates what some describe as a ‘responsibility vacuum’, wherein outcomes are systemically difficult to tie to accountable actors. From a governance perspective, these dynamics imply that challenges arise not only from poor implementation or immature processes, but from the intrinsic nature of current LLMs. Despite this opacity, generative AI work support tools are increasingly the object of experimentation, with the intention of being implemented in critical decision-making contexts where factual precision is essential and public sector practitioners are expected to integrate opaque outputs into their judgements. Yet, the probabilistic nature of LLMs, the insufficient transparency of their embeddings in the vast source materials and their technical development can cause them to produce inaccurate, misleading, or hallucinated outputs, obscuring the logic behind administrative recommendations. This raises concerns about the feasibility of meaningful human–AI collaboration and knowledge transformation, especially when users face obstacles to opportunities to interrogate AI-generated outputs. 3. Research design. This article is based on a multi-site digital ethnographic study conducted for 2.5 years in two Finnish public organizations—nearly 1.5 years in one site and nearly two years in the other—as part of “Perpetual piloting and invisible work of automating public services in Finland”. Both organizations are recognized as pioneers in the development and implementation of new technologies. They also play a central role in delivering public services and benefits in Finland. Each has one or more multiprofessional innovation teams responsible for enhancing and coordinating innovation activities. Since 2023, these activities have increasingly focused on generative AI work support tools. My research coincided with the rapid proliferation of generative AI following the release of ChatGPT by OpenAI in late 2022. A significant number of the observed innovation activities were linked to generative AI tools, the development of which tended to move more quickly towards experimentation compared to systems based on traditional AI methods. 3.1. Fieldwork overview: Studying generative AI work support tools. My fieldwork began with observations of weekly innovation team meetings in both organizations to gain an overview of ongoing innovation projects. Innovation activities were typically carried out in smaller groups that included internal and external actors, such as consultants, middle management, lawyers and experts from frontline units. With the first organization, my access to innovation activities focused on a generative AI decision support tool designed to assist claims specialists in retrieving and using relevant information for decision-making. This project began three months into my fieldwork. The pace of development accelerated significantly after the decision to use an LLM-based technical solution. From August 2023 to September 2024, I closely followed the team's intensive innovation work around this tool, including their formal and informal discussions, design workshops, discussions and demonstrations of technical developments and troubleshooting, the conducting of five rounds of user experiments and internal and public presentations about the project.2 With the second organization, my observations took place between February 2024 and December 2025. During this period, I gained broader access to innovation activities that included various AI and other digitalization projects involving other organizational and external actors, including smaller innovation teams across different sites of the organization. Additionally, I observed the innovation team's work in implementing AI regulation, cultivating experimentation-driven innovation and strengthening AI literacy within the organization. Several methods were used during this two-and-a-half-year ethnographic fieldwork, which included observation, qualitative interviews and document collection (e.g. PowerPoint presentations, experimentation plans and schedules, innovation process descriptions and innovation frameworks). The fieldwork resulted in research materials that include the following: fieldnotes from a total of 476 h of observing innovation activities related to particular AI and other technology projects, the nascent implementation of AI regulation, fostering innovation culture, the development of AI literacy among employees, co-creation sessions held in collaboration with external stakeholders and information sessions for employees on issues related to AI. Except for a few public events, all other observations were made online, as most of the meetings were organized either as online-only sessions or in a hybrid format. In addition, I conducted 109 interviews with members of innovation teams, consultants assisting innovators' work, test users and other organizational employees involved in innovation activities. All interviews were conducted over Microsoft Teams and lasted from 15 to 95 min. This article focuses on three key generative AI work support tools observed during the period of study: 1) an AI decision support tool, an internally developed LLM-based tool for claims specialists to assist them in the retrieval and use of relevant information for decision-making; 2) an AI assistant for client service in social work, an internally developed LLM-based tool designed to assist social workers in searching service guidance information; and 3) M365 Copilot, a commercial LLM- Overview of the three technologies in focus. powered productivity assistant embedded in Microsoft 365 applications for administrative officers (see Table 1). Each tool was designed to address the challenge of navigating vast, complex, scattered and constantly evolving information from various sources. The AI client service assistant and AI decision support tool used specific databases and frameworks like retrieval-augmented generation (RAG) to improve response relevance and accuracy. They were developed in collaboration with external consultants. For the AI decision support tool, the database consisted of benefit guidelines, while the database for the AI client service assistant included a list of available services and links to their websites. An M365 Copilot licence was purchased from Microsoft and tested by administrative officers across welfare domains to identify its most beneficial use cases. It was reported to be integrated with organizational data, enabling it to retrieve and analyse content from users' M365 applications (e.g. Word, Excel, Teams meeting transcripts and Outlook). These tools functioned as conversational AI bots operating on specific datasets. The internally developed tools generated outputs based on user queries and system prompts, often accompanied by links to source materials for validation or further exploration. An underlying idea shared across all three projects was that testing these tools involved not only new technologies but also new ways of working. Innovation teams aimed to understand how these tools fit into users' work realities, foster learning and support organizational change. 3.2. Data collection and analysis. This article focuses on data from the three generative AI tools (see Table 2), as these cases provided direct access to observing the perspectives of both the innovators and the test users on the tools. This included feedback sessions where users shared their experiences with the innovators and my interviews with some users. Unlike other AI projects I observed in my fieldwork, these three cases offered a rare opportunity to closely examine the interactions between innovation teams and end users, particularly their often-diverging perspectives. Although the broader dataset from the project was not analysed for this article, it informed my understanding of a recurring issue: the MCZ positioning of testers who bear the responsibility for ensuring the effective and ethical use of LLM-based tools. Data collection for the AI decision support tool lasted for over a year, resulting in fieldnotes from 212 h of observation of the innovation team's multifaceted work. Data collection included the observation of team meetings, co-creation sessions with external consultants, internal information sessions and workshops, small group discussions, the planning of six and organizing of five testing rounds, user feedback sessions, design workshops and technical demonstrations and internal and public presentations. User interviews were conducted during and after each testing period. In total, the dataset includes insights into 25 out of 30 officially registered testers. Data collection for the AI client service assistant lasted 13 weeks. This period covered my observations of the innovation team's weekly meetings, a user kick-off session, feedback sessions, a demo video issued Data collection overview. for users and three PowerPoint presentations about the project. User interviews were conducted after the testing period ended. The dataset includes insights from 18 of 32 officially registered testers. Data collection for M365 Copilot lasted 11 months. My involvement began with observing internal meetings of the innovation team responsible for planning and monitoring the tool's testing. These meetings covered systematic user feedback analysis, discussions on licence distribution and training participation, updates on the tool, testing progress and governance planning. I also observed five training sessions and received transcripts from additional sessions. I was invited to observe 10 interviews with early-stage testers conducted by the project manager and attended two sessions focused on developing a governance model for the tool. My own user interviews were conducted at the end of the testing period. In total, the data includes insights from approximately 43 of the approximately 900 officially registered testers. Data analysis was guided by the following research question: How are public sector practitioners positioned as MCZ for the effective and ethical use of LLM-based tools in public sector? To investigate this question, I employed an abductive analytical approach characterized by iterative movements between empirical data and the theoretical framework (van Hulst and Visser, 2024). This approach allowed me to remain open to emerging insights while grounding the analysis in the theoretical framework. The analysis process began during my fieldwork (see Fig. 2), particularly while observing the innovation work around the AI decision support tool. I was struck by how responsibility for realizing the tool's potential was implicitly placed on users, who were also expected to adapt to its limitations. This initial puzzlement became the starting point for a more systemic inquiry. As I continued observing, similar patterns emerged in the other two cases, confirming that this was not an isolated phenomenon. I began to track this practice more deliberately, identifying key actors, events and interactions where both responsibility and human oversight were discussed, assigned or assumed. The MCZ concept was introduced here to contextualize the emerging pattern of responsibility being assigned to users and was complemented by a literature review on AI augmentation of professional work and the challenges of AI opacity. After the conclusion of the three AI projects, I synthesized my fieldnotes, interview transcripts and observational data to construct a comprehensive overview of how responsibility was allocated and interpreted in each case. Initial coding of fieldnotes and transcripts focused on moments where responsibility, human oversight and user adaptation were discussed explicitly or implicitly. These codes were then grouped under emerging themes such as responsibility allocation and deflection, tool immaturity, user burden, technical complexity, insufficient trainings, learning requirements, limited transparency, persuasive fallibility and the burden of vigilance. Afterwards, I engaged more closely with the MCZ concept and research on AI augmentation and liability to examine the theme of responsibility allocation and deflection. Next, I focused on the analysis of users' experiences with the tools, especially emerging opacities and challenges to users' capacity and control. 4. Mechanisms producing moral crumple zones in public sector. The findings section discusses five mechanisms that position public sector practitioners as MCZ: the first mechanism explores innovators' practice of taking human oversight for granted and assigning human users' responsibility in the use of LLM-based tools, and the remaining four present obstacles to holding users liable for the effective and ethical use of LLM-based tools in public sector. 4.1. Human oversight and liability taken for granted. The analysis revealed that despite practical ambiguity and experimental improvisation around LLM-based AI tools, human oversight and responsibility as liability were explicitly codified as a user burden. This aligns with the augmentation argument that highlights human capacity in interrogating and integrating AI-generated outputs when striving for superior performance and enhanced efficiency. Across events, test users were systematically held responsible for the effective use of AI tools in terms of their output quality, validation and appropriate use. This delegation of responsibility to users rested on the taken-for-granted expectation that test users as domain experts are capable of overseeing and influencing output generation. The common assumption seemed to be that domain expertise—rather than technical or other expertise—would be central to or sufficient in overseeing and effectively users these tools. Thus, users were perceived as liable for the generated outputs and capable of obtaining high-quality outputs while navigating and interpreting the tools' opacity. Additionally, users were given extra tasks like identifying and reporting effective prompting strategies even though the users had received minimal or no training: The quality of the tool's outputs depends on how we, as users, learn to interact with it. I recommend experimenting with asking questions in various ways. It would be beneficial if you could identify patterns in the types of questions that would yield the best results. We [innovators] are unable to guide you on how to formulate effective questions, which is why we are experimenting with the tool. (...) We don't have the expertise to do it ourselves; you are the domain experts. (Fieldnotes on the innovators' comments during the feedback sessions with users in autumn 2023, AI decision support tool). In line with this rhetorical practice, flawed and low-quality outputs tended to be portrayed as the result of poor user interaction rather than issues with the tool's design and development. Users were also expected to have the skills and contextual understanding to assess whether outputs were appropriate or meaningful. Thus, testers were also held responsible for verifying accuracy and, when possible, cross-referencing with original sources to ensure that the information presented in outputs was correct: ‘It is always users’ responsibility to correctly interpret outputs. If needed, you can check the source materials for whether AI has understood or summarised them correctly’ (Innovators during the feedback sessions with users in autumn 2024, AI client service assistant). This rhetorical practice was particularly evident with internally developed tools that received significant critical feedback from users—feedback that often left innovators puzzled. When no quick technical fixes or clear explanations were available, the innovators' focus shifted to how users could interact with the tools to obtain more satisfactory outputs: At the feedback sessions, social workers testing the AI client service assistant were puzzled by the amount of incorrect and missing information in the outputs. The innovators and the tools' developers were puzzled too. Instead of trying to explain the reasons behind these issues, they encouraged users to actively challenge the AI assistant by saying, “It is worth challenging the AI if you notice crucial information missing from the outputs by asking follow-up questions or directly challenging it with questions like, ‘Why do you claim this? Are you absolutely sure?’ (...) Just ask more questions and challenge it. It never gets offended. It just apologises for being wrong and then tries again”. (Fieldnotes on the feedback session, autumn 2024, AI client service assistant). This tendency to encourage users to remain engaged with the tools until they obtained more satisfactory outputs aligns with the basic design principles of LLMs, which rely heavily on human input, continuous engagement and affective awareness. While prolonged human engagement may improve output quality, reliability remains doubtful without additional auditing mechanisms. This expectation of prolonged engagement can function as users' invisible work to compensate for the limitations of AI tools. This expectation and, more broadly, allocation of responsibility can also translate into management policies and an organizational culture that over-rely on individual responsibility and performance. However, as scholars highlight, LLMs are designed for plausibility and user engagement but not truthfulness. This raises concerns about the reliability of outputs, especially when tools operate on manually curated and incomplete databases. Such demands are particularly problematic in time-pressured environments like street-level bureaucracy, where users often lack the time, skills and technical literacy needed to effectively prompt, verify and interpret AI outputs. Ultimately, the delegation of responsibility for ensuring the effective and ethical use of LLM-based tools to test users—who are without adequate support—makes this model of oversight especially problematic, particularly when it serves as a way to manage the classic opacity problem of LLMs (see e.g. Bender et al., 2021). 4.2. Immaturity of prototypes marketed as ‘sparring buddies’. The argument for AI augmentation was specifically reflected in the innovators' promise to deliver tools that would support cognitive offloading, freeing users to focus on more meaningful and demanding tasks. These three tools were envisioned as tireless ‘sparring buddies’, always available and designed specifically to assist users in retrieving and managing relevant information to complete their work tasks. They were intended to address a well-known challenge in public sector: the limited availability of busy colleagues for consultation and the overwhelming volume of complex, scattered and constantly changing information considered to lower work quality, work efficiency and productivity. As a result of this promotional framing, these tools were warmly welcomed by users: I have tried to think with it even though I may not have personally experienced it as a very significant help or tool for me. I still wanted to try it because the opportunity existed, and this type of tool might be a big advantage when it gets improved further. (...) That's why I wanted to at least try it out. (Interview with tester, autumn-winter 2023, AI decision support tool). However, testing these tools revealed a more complex reality in public sector, with their immaturity emerging as a defining characteristic. These tools were rapidly developed prototypes given to users for immediate testing. Despite their potential technical sophistication and the considerable time and expertise invested in their development, they were frequently introduced into user testing before they had been rigorously vetted and reached maturity. Yet, that immaturity was rarely addressed directly; when it was, it was typically accompanied by optimism about future improvements. For example, the use of M356 Copilot was significantly hindered for half the testing period because it did not function in Finnish, the operating language of Finland's public sector. Users' experiences improved in the latter half as the tool gradually became functional in that language. Consequently, the tools' test users encountered and puzzled over fundamental flaws, including incomplete, incorrect and outdated information in outputs. Sometimes, the tools did not generate outputs even though they should have had adequate information from the database to do so. This undermined the reliability of the testing phase, as the tools were frequently too immature to accurately measure what was being tested. Specifically, users reported their inadequacy in performing work-related tasks: Our clients have very complex problems, which makes it difficult to find simple keywords or phrases to put into it [the tool] so that it can retrieve and offer suitable services to each individual client. (...) I ran some searches, and the results were mixed—some returned error messages or blank spaces, others were clearly incomplete or incorrect. (...) Regarding this tool, I'd say it's still in its infancy stage. (Interview with tester, autumn-winter 2024, AI client service assistant). Another shared the following: To be honest, it's a very poor substitute [for sparring with human colleagues]. [In this line of work] the questions you want answers to are so difficult that it [the tool] clearly doesn't provide them. In line with the instructions, I have tested it quite a lot, asking the same question in different ways or asking follow-up questions. For inexperienced workers, it [the tool] can even be dangerous when it gives completely incorrect answers, even though it doesn't actually know the answer. (Interview with tester, autumn-winter 2023, AI decision support tool). A shift in the narrative surrounding LLM-based tools marketed as sparring buddies occurred as their limitations became apparent through users' critical feedback. The narrative gradually moved towards a stronger emphasis put on the human role in an epistemological partnership between humans and machines. Yet, as Lebovitz et al. (2022) have argued, such partnerships require AI systems that are mature and adequate as well as additional time for reflexive interrogation—resources that are not always available to testers and do not always yield a return on investments. Furthermore, this evolving framing was often reinforced by general claims about the tools' usefulness. For example, statements such as ‘As a rule, it [the tool] does take you in the right direction’ were rarely elaborated on but served to affirm the tools' perceived value. The notion of the sparring buddy was rarely questioned in terms of how LLM-based tools, described by some researchers as ‘stochastic parrots’, fit into public sector, which relies on facts, nuance and public servants' discretion. 4.3. Trial-and-error learning offloaded to users. The tools' experimentation was characterized by offloading the burden of learning and testing onto users who often lacked sufficient training or technical support. This took place in an atmosphere of technology enthusiasm, where AI was seen as a catalyst for change and a solution to productivity traps (see Choroszewicz, 2025; Scarbrough et al., 2024). Users were repeatedly reminded that benefits would only emerge through active use, reinforcing a discourse that framed successful harnessing as dependent on user initiative and creativity rather than tool maturity and adequacy: It's definitely worth experimenting with what it can do specifically for you. (...) All you have to do is actively experiment with it and this way we [as an organization] will get competent in AI. (...) The question we will soon be answering is: What should we do with the saved time due to these technologies? (Innovators' quotes from information sessions on the testing results of M365 Copilot, spring 2024). In the case of the commercial tool, M365 Copilot, the organization made notable efforts to ensure compliance with data protection policies, select diverse testers, provide basic training and establish communication channels for updates, support and reporting AI incidents. It also invested notable effort in acquiring feedback on what kind of AI training its employees would need in the future. As part of the agreement with Microsoft, testers were also offered introductory training from a consulting firm. While testers appreciated the organization's efforts to carefully plan the experimentation around M365 Copilot, they often criticized the provided training sessions as overly promotional and lacking practical guidance: I attended some training sessions, but in my opinion, they were very general and of a promotional character. (...) Then we had the Teams channel to get support and ask about principles of use and such. (...) One should be open about using Copilot. (...) At least everyone should understand the rules of the game, which is not always the case. (Interview with tester, autumn-winter 2024, M356 Copilot). By contrast, much less training and guidance were provided for testers of internally developed tools. The tools were presented as open-ended, with potential to be harnessed individually: Well, not really [no training on using the tool]. It was more like, here it is. Ask it. (...) We [users] always discussed and developed how to ask questions and how to structure them. But we didn't get much technical guidance (Interview with tester, autumn-winter 2023, AI decision support tool). Across the tools' testing, significant attention was given to collecting user feedback through sessions and surveys. However, this feedback was often interpreted through an optimistic lens, emphasizing potential benefits while downplaying current limitations, especially for internally developed tools. By contrast, criticism of M365 Copilot was more openly expressed, perhaps because it was not an internal project. Negative feedback about internally developed AI assistants was frequently dismissed as temporary, cited as a reason for the tools' further development or redirected back to users as a need for their additional effort: So, in a situation where the list of services is not quite what you would hope for, but still somewhat in that direction, you can ask it [the tool] additional questions like ‘What else would you suggest?’ If there is any additional context that you feel might be relevant, you can try to provide it and ask it if there is anything else to add. Just challenge it and its output and ask it to come up with a better one. (Innovator during the feedback session, autumn 2024, AI client service assistant). One tester shared the following: Well, at least when I've talked to a couple of my coworkers, they seem to feel the same way I do, that the test questions were answered completely wrong. (...) I didn't feel like I was getting any useful information from it. (...) It doesn't really help with this job yet. (Interview with tester, autumn-winter 2023, AI decision support tool). Users who reported faulty outputs were asked to explain what was wrong and what would have satisfied them, subtly shifting the focus from present limitations to aspirational visions surrounding these tools. For users, this seemed to reinforce the belief that flaws could be overcome with time, their persistence in user engagement and continued technical development. Ultimately, users absorbed responsibility for navigating technical shortcomings and epistemic risks. This sustained a narrative in which human oversight, engagement and adaptability were central and served to justify further development and testing of generative AI tools across the public sector's welfare domain. Despite being marketed as tools to ease workloads, testers often found them confusing and ineffective. 4.4. Original opacity of LLMs compounded by limited technical. transparency The inherent opacity of LLMs appeared to be further compounded by limited technical transparency during the tools' further development and testing in public sector. Users often lacked access to meaningful explanations of how they functioned, which hindered users' ability to understand, trust or benefit from them. Many users, particularly those testing internally developed tools, had no prior experience using LLMs for work-related tasks. Despite widespread enthusiasm about AI, many users remained uncertain about the logic behind AI-generated outputs, especially when results were unexpected, incomplete or incorrect. This lack of clarity stemmed not only from the technical opacity of these tools but also from how they were introduced—without sufficient contextualization, interpretative guidance or explanation of their limitations. Although many found the tested tools easy to use, making sense of the outputs proved challenging, even for experienced workers. This issue was particularly pronounced among claims specialists whose work involves interpreting both publicly accessible and internal guidelines—some of which could not be fed into the tool—and conducting holistic assessments of claimants' life situations. By contrast, users of the AI client service assistant faced fewer difficulties identifying faulty outputs as testers were experienced social workers familiar with the available services. Nevertheless, even they reported puzzling over incomplete, outdated or inaccurate information that required time to inspect and verify. Specifically, the contents of databases created separately for each tool to anchor AI outputs in relevant public sector information remained unclear to testers: I don't know what's in there right now. I don't really know because I don't understand the process behind it, how it works, so it's difficult to take a position on it. (...) Sometimes I noticed that when I asked something, the answer from it was a little off topic. I don't really know how the source material was provided to it. (Interview with tester, autumn-winter 2023, AI decision support tool). For internally developed tools, the initial opacity of LLMs was compounded by a lack of transparency around their integration into the RAG framework. The problem was further exacerbated by the incomplete databases provided to each tool. It also remained uncertain whether it would ever be legally or technically feasible to expand these databases to cover all necessary information. An additional challenge was the constantly changing nature of public sector information. While some experienced users could detect outdated or incorrect information, others were confused by outputs based on obsolete data: My test question involved a situation where there had been a change in instructions. (...) However, this change had not been taken into account, and it [the AI tool] provided outdated information. I had thought about this issue too because our instructions change very often. How does it keep up with these changes and maintain its reliability? How does it stay up to date with all the changes in instructions? (Interview with tester, autumn-winter 2023, AI decision support tool). Unlike the internally developed tools, M365 Copilot operated directly on user data available through Microsoft applications. Nevertheless, users expressed concerns about the clarity and trustworthiness of the source material it used. Some reported difficulties with information retrieval and noted the tool's inadequate communicative capabilities: Copilot is really limited here now. When Microsoft pushes it [Copilot] into every single one of its applications, in some cases, it's good. In others, it's not. It does find everything from SharePoint pretty well from within Teams, but not reliably. You can't really trust it to do the work you need it to do. It's still very stiff. (Interview with tester, autumn-winter 2024, M365 Copilot). At the beginning of testing, users were often encouraged to test the tools actively, based on exaggerated promises about their capabilities and potential for improvement while the inner workings of these tools remained largely unexplained to testers. Only during later feedback sessions did testers receive additional information crucial to making sense of faulty outputs: At the kick-off sessions, the innovators and developers encouraged users to actively test the tool, often drawing on exaggerated promises about how well the tool worked: ‘In this tool, the language model is instructed to search only for services from the service database our team curated. (...) Yet, the list does not include all services. Please provide feedback if something is missing or if you receive strange responses. (...) Drawing on your feedback, we will be improving the tool during the testing period. (...) When, at the feedback session, users reported confusion over faulty outputs, the developers and innovators seemed to be equally puzzled. After a moment of collective confusion, they (finally) revealed to users that if they ask questions outside of the curated service database, the answers may be partly fictional because the LLM was not instructed to respond to them at all. (...) We also decided not to make any changes to the tool during the testing period’. (Fieldnotes on the kick-off session and feedback sessions, autumn 2024, AI client service assistant). This example illustrates how limited transparency and inconsistent communication about the tools' development and functioning kept users in a state of epistemic uncertainty. Such conditions undermined opportunities for human oversight and explainability, which are principles of emphasis in ethical AI governance frameworks. While technical transparency alone cannot resolve accountability concerns around complex technologies, its absence left users unsure whether errors stemmed from their own prompting, the tool's design, an incomplete database, the limitations of the LLM and its settings, the RAG framework or the latest technical changes introduced by innovators. It also remained unclear to users how much influence they had over a tool's performance or improvement. This lack of transparency not only hindered the effective and ethical use of the tools but also placed an undue burden on users to act as testers, AI outputs' quality verifiers and interpreters—roles they were not trained for. Furthermore, technical shortcomings were often reframed as solvable through better prompting or user behaviour, sustaining the belief that the outputs' quality could be improved with limited addressing of the tools' structural opacity. This belief was often reinforced by vague information sessions and the absence of clear mechanisms for measuring outputs' reliability. 4.5. Persuasive fallibility and the burden of vigilance placed on users. The tools produced affective turbulence for users as they tried to make sense of the outputs—deciding whether they were correct or wrong, trustworthy or misleading, helpful or not. A recurring issue across three tools was persuasive fallibility: users were unaccustomed to the confident, plausible and authoritative tone of AI-generated outputs, even when those outputs were factually inaccurate or completely incorrect. This created a usability paradox, blurring the line between surface-level and actual correctness. This issue was especially pronounced with more complex questions or tasks. Even experienced users consistently reported confusion when faced with neatly formulated and convincingly sound outputs. Epistemic and interpretative opacity appeared pervasive. Some users reported having to read outputs multiple times to determine their correctness, while also seeking peer validation. One experienced worker explained the issue as follows: I got such a response that I had to think about whether I had given the caller [colleague] the wrong answer, and then of course I started looking for the correct answer. Since the bot was so sure about the answer (...), I had to consult a colleague about this when I realized that I could no longer find the answer to the question directly from the benefit guidelines. The bot was of one opinion, and I was of the other. Later, it turned out that it was wrong, but it had answered so confidently. (Interview with tester, autumn-winter 2023, AI decision support tool). The quote illustrates a distinctive risk in LLM-based work support tools; persuasiveness may mask fallibility, increasing the likelihood of undetected errors, especially for insufficiently experienced workers (as noted by users of all three tools). The persuasive and authoritative tone placed an emotional and cognitive burden on users, deepening both ethical concerns and practical implications. While most users were experienced enough to detect inaccuracies and incorrect information or disregard misleading outputs, this may not hold true for new or less experienced workers. This is particularly relevant given that these tools were often designed to enhance job orientation and efficiency for just such workers, who may lack the time, opportunity or expertise to verify outputs. Under these conditions, users worried that the AI's fluency and confidence could lead to tacit acceptance or overreliance: ‘It misleads you, as if it knew everything for sure. That's why you have to be sceptical’. (Interview with tester, autumn-winter 2023, AI decision support tool). Users also expressed uncertainty about how to approach problem-solving tasks with these tools. Crucially, the opacity of output generation amplified this uncertainty rather than resolved it. This led to a persistent sense of epistemic doubt—a ‘nagging feeling’ that required users to reread and cross-check outputs repeatedly to distinguish facts from fabrications or inaccuracies. This vigilance became a form of continuous cognitive and emotional labour (see e.g. Carboni, Wehrens, van der Veen, and de Bont, 2023; Choroszewicz, 2022), especially in rule-bound environments like public sector, where even minor errors can have significant consequences. In the long term, as Elish (2019) noted, this dynamic may lead to deskilling. When humans are deprived of routine tasks and assigned only cognitively demanding ones, it does not necessarily sharpen their expertise. Instead, it can result in fatigue and disengagement. Constant alertness and scepticism can become burdensome. Thus, surface usability masks deeper epistemic risks. Ease of use and a persuasive tone may create a false sense of reliability, while the actual burden of ensuring correctness remains with users who may lack the time, skills or experience to carry out this task. Over time, routine use without critical engagement can turn AI from a helpful assistant into a source of risk. 5. Discussion and conclusions. This article has examined the mechanisms that positioned public sector practitioners as moral crumple zones for the effective and ethical use of large language model-based tools, focusing on experimentation with three LLM-based tools in two Finnish public organizations. By doing so, the study extended the MCZ concept, which was originally used to describe malfunctions in complex human–machine systems, to public sector practitioners involved in testing LLM-based tools offered by their employers. The findings captured five mechanisms that produce MCZ in the use of generative AI tools for administrative and frontline work: taken-for-granted human oversight and liability, the immaturity of tested tools marketed as sparring buddies, trial-and-error learning offloaded to users, limited transparency of AI tools' development and persuasive fallibility and a burden of vigilance placed on users. Through these mechanisms, end users were ultimately positioned as performing three roles—testers, AI outputs' quality verifiers and interpreters—despite the tools' technical immaturity and opacity. As testers, they were expected to experiment actively with underdeveloped tools, identify potential benefits and provide feedback. As quality verifiers and interpreters, they were held responsible for assessing, confirming, decoding and ensuring the correctness of AI outputs, even though they lacked adequate guidance, resources, support and evaluation frameworks. This dynamic shifted the burden of navigating the tools' technical opacity and epistemic risks from developers and innovators to end users, creating a paradox where tools intended to save time instead generated additional workload. 5.1. Theoretical and empirical extension of MCZ to the context of AI tools. in public sector The issue of human–AI collaboration has become highly topical in public sector, where workers are expected to do their jobs in a manner that is explainable and justifiable to citizens, fostering and maintaining trust in public services and institutions. As ethical considerations related to human oversight and responsibility are not yet fully integrated into the development of emerging AI tools, this study highlights a growing trend; the burden of effective and ethical use and oversight is increasingly being placed on users, especially when generative AI tools are intended to augment human judgement. In this article, such responsibility was evident in the expectation that end users possess the capability to unlock the potential of these tools and apply them ethically, preventing breaches of data protection policies and avoiding reliance on erroneous outputs. In practice, users were made accountable for producing high-quality and beneficial AI outputs, critically inspecting, validating and interpreting them, and knowing when and how to appropriately deploy them. The study indicated that, like cases described by Elish (2019), responsibility pertaining to the tools' designers, developers and providers can be displaced onto individual users of generative AI tools even though they might be structurally disempowered to exercise this responsibility in a meaningful fashion. The identified mechanisms producing MCZ highlight the paradox related to the AI augmentation argument, which ignores users' emotional and cognitive burden linked to the three forms of opacity: epistemic, interpretative and operational. This burden hinders the informed and meaningful use of LLM-based tools. While epistemic opacity around AI-generated outputs has already been widely recognized (see e.g. Bender et al., 2021; Hannigan et al., 2024), this study has provided additional empirical insights into how it is relevant and troubling in public sector. By contrast, interpretative and operational opacity have received insufficient attention. As the findings showed, interpretative opacity involves users' challenges in making sense of AI-generated outputs, including ambiguity in meaning or intent. Operational opacity refers to the practical application of outputs. It was especially pronounced in the expectation that users integrate AI recommendations into their own work practices and judgements without clear guidance on appropriate use cases or on how to ensure alignment with professional and ethical standards. These forms of opacity were intensified by the experimental nature of the tools' use in public sector, where trial-and-error learning was prevalent and governance structures were still emerging. Nevertheless, across the three cases, notable variations emerged based on both the development context and the intended use of the tools. Internally developed tools were typically narrower in scope and designed for specific purposes requiring high levels of accuracy, precision, and transparency. These functional constraints also limited testers' engagement. For example, the AI client service assistant did not support actual client counselling, though it had potential to assist in preparatory work if it had been technically more mature and technically transparent. Similarly, the AI decision support tool proved extremely laborious to use, as overcoming epistemic, interpretative and operational opacities appeared challenging, partly due to regulatory requirements as well as strict and often contradictory benefit guidelines. By contrast, M365 Copilot—despite its initial technical immaturity—was a commercially available, general-purpose tool supported by larger databases and more advanced architectures. This enabled broader creative and exploratory use, even with its inherent opacity. These distinctions significantly shaped testers' ability and agency to explore the usability of the tools. 5.2. Practical implications: Addressing the misalignment between. influence and accountability The findings point towards a troubling misalignment; despite not designing or fully understanding these tools, public sector practitioners in their role as users were held accountable for the AI tools' output quality, interpretation and use. This misalignment between influence and accountability exemplifies MCZ, where users absorb the burden and consequences of testing immature AI tools without having sufficient control or the right resources to shape them (aside from providing feedback). This issue was further complicated by users' limited abilities and resources to understand and interrogate the obtained outputs. Undoubtedly, this misalignment is particularly troubling in high-stakes domains like the public sector. As Lebovitz et al. (2022) have shown, the successful integration of AI-generated insights into professionals' judgement—known as ‘engaged augmentation’—requires deliberate interrogation of AI outputs. By contrast, unengaged augmentation occurs when users either ignore or accept AI outputs without reflection, often due to a system's lack of transparency. Codifying responsibility as a user burden insulates technologies and their developers from accountability, reinforcing the notion that AI-generated outputs are inherently human-mediated and risk-bearing. This practice of responsibility allocation aligns with current legal standards of practice in Finnish public sector, which emphasize human oversight and decision-making. The results presented in this article highlight that revising this current practice of responsibility allocation in the early stages of AI projects may help reveal systemic flaws and epistemic risks, establish clearer accountability pathways and prevent the premature institutionalization of AI systems that are inadequate. In line with emerging discussions on responsibility in the context of LLM-based tools, scholars have suggested that relational and distributed models of responsibility may offer appropriate conceptual frameworks. For instance, Amoore (2020) and Wachter et al. (2024) argue that responsibility around AI systems emerges from the capacity and willingness to respond to the effects of AI systems—not from the ability to control or predict them. They also critiqued the current framing of algorithmic governance, which assumes a legible world of rights and duties centred on rational human subjects. Their perspective further advocates for attention to the socio-technical arrangements surrounding AI systems: how they are designed, deployed and governed in institutional contexts. It also demands institutional infrastructures that support critical engagement, such as transparency standards, auditing mechanisms and participatory governance frameworks. Ultimately, this study underscores the need for more complex and transparent approaches to the innovation, use and governance of LLM-based tools in the public sector—ones that recognize the structural LLM opacity, the numerous epistemic risks, the limits of end user agency and the structural conditions under which these tools are deployed. The need for clear guidance, training and accountability pathways is particularly relevant given the rapidly proliferating AI technologies and risks associated with them. Accordingly, risk-sensitive deployment requires task-level calibration (matching use cases to accuracy needs and verifiability), verification protocols (ground-truthing and documentation), human-in-the-loop safeguards against automation bias and organizational accountability mechanisms (audits, incident reporting, and clear lines of responsibility) capable of operating under conditions of enduring opacity. In short, the governance of generative AI tools requires recognizing that opacity is deeply structural and behavioural—shaping what kinds of AI innovations are possible, for whom, and on what terms. 5.3. Limitations and future research directions. This study offers important insights into mechanisms positioning end users as MCZ in the public sector; however, certain limitations should be acknowledged. The analysis relied primarily on the MCZ concept to explore the current practice of responsibility allocation to end users and obstacles surrounding it. While this framework offered a useful lens, it may nevertheless oversimplify the complexity of varying levels of human involvement and control across experimentation with diverse and rapidly proliferating AI work support tools. Moreover, this research focused predominantly on similarities in trends and user experiences across early-stage cases of LLM-based tools, rather than exploring nuanced differences in design, development context and functional purpose, distinctions that may well shape how testers experience opacity, agency and responsibility in public sector. Future studies should more fully capture the relational and systemic aspects of accountability to understand how user agency, oversight and accountability operate under diverse conditions. Finally, although the study highlights the misalignment between responsibility and control, it does not fully account for emerging organizational strategies that might mitigate this tension. Future research should investigate emerging governance models, training practices and institutional infrastructures that support ethical AI use in public sector. In addition, evaluating the adequacy of generative AI work support tools for public sector tasks remains important, given their rapid proliferation and ongoing experimentation. CRediT authorship contribution statement Marta Choroszewicz: Writing – review & editing, Writing – original draft, Methodology, Investigation, Funding acquisition, Formal analysis, Conceptualization. Funding The author received the following financial support for the research: a two-year grant from the Finnish Cultural Foundation (Teresia and Rafael L ̈onnstr ̈om Fund) for the project “Perpetual piloting and invisible work of automating public services in Finland” and personal grants from the Ella and Georg Ehrnrooth Foundation and the North Karelia Regional Fund (Eini and Jorma Veikkolainen Fund) for the project “Navigating experimental AI tools within technology-enthusiastic context of Finnish public administration”. Declaration of competing interest The author declares to have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Acknowledgements. I would like to thank the anonymous reviewers for their thoughtful comments, which helped improve this manuscript. I am also grateful to the public organizations, their employees (organizational managers, innovation experts and testing users) and the collaborating consultants and other stakeholders for providing access to the observed meetings and events and for agreeing to be interviewed. I thank doctoral researcher Antti Rannisto for his collaboration in the collection of research material in one of the public organizations.