1 More Paper.
Full Reading02:45:10

Using LLM-generated draft replies to support human experts in responding to stakeholder inquiries in maritime industry: a real-world case study of Industrial AI

1 More Paper · Full Reading

Full Reading podcast cover
Listen to the Full Reading

About this paper

A full audio edition of this paper.

Authors: Tita Alissa Bach, Aleksandar Babic, Narae Park, Tor Sporsem, Rasmus Ulfsnes, Henrik Smith-Meyer, Torkel Skeie

Published in: Cogent Engineering

Publication date: 2026-07-29

Read the paper: https://doi.org/10.1080/23311916.2026.2704286

Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/

The authors and publisher do not sponsor or endorse this recording.

Brief episode

Transcript

You’re listening to “Using LLM-generated draft replies to support human experts in responding to stakeholder inquiries in maritime industry: a real-world case study of Industrial AI,” by Tita Alissa Bach and colleagues. Published in Cogent Engineering on July 29, 2026.

Cogent Engineering

ISSN: 2331-1916 (Online) Journal homepage: the linked source

Using LLM-generated draft replies to support human experts in responding to stakeholder inquiries in maritime industry: a real-world case study of Industrial AI

Tita Alissa Bach, Aleksandar Babic, Narae Park, Tor Sporsem, Rasmus Ulfsnes, Henrik Smith-Meyer & Torkel Skeie

To cite this article: Tita Alissa Bach, Aleksandar Babic, Narae Park, Tor Sporsem, Rasmus Ulfsnes, Henrik Smith-Meyer & Torkel Skeie (2026) Using LLM-generated draft replies to support human experts in responding to stakeholder inquiries in maritime industry: a real-world case study of Industrial AI, Cogent Engineering, 13:1, 2704286, DOI: 10.1080/23311916.2026.2704286

COMPUTER SCIENCE | RESEARCH ARTICLE

Using LLM-generated draft replies to support human experts in responding to stakeholder inquiries in maritime industry: a real-world case study of Industrial AI

Tita Alissa Bacha, Aleksandar Babica, Narae Parka‡, Tor Sporsemd, Rasmus Ulfsnesd, Henrik Smith-Meyerb and Torkel Skeiec aGroup Research and Development, DNV, Høvik, Norway; bMaritime Production Operations, DNV, Høvik, Norway; cTechnical Support Norway, DNV, Høvik, Norway; dDepartment of Software Engineering, Safety and Security, Sintef Digital, Trondheim, Norway

ABSTRACT.

This study investigated the utility of Large Language Models (LLMs) in supporting human experts in the maritime industry by generating draft replies to stakeholder inquiries. Using a mixed-methods approach, observations, interviews, a survey, and text similarity analysis, we examined how human experts perceive LLM drafts, how drafts align with final replies, and their potential impact on workflows and stakeholder trust. Findings show that while LLM drafts can improve efficiency, linguistic consistency, and time management, they frequently require manual adjustments to ensure accuracy and relevance. Case handlers’ attitudes were generally ‘skeptical but curious,’ reflecting cautious adoption and emphasizing the role of managerial support in fostering trust and engagement.

Text similarity analyses indicated moderate alignment between drafts and final replies, reinforcing the need for human oversight. Overall, LLMs are not yet suitable for independent use in safety-critical settings but can serve as valuable augmentative tools. By leveraging human expertise alongside LLM capabilities, organizations can enhance workflow efficiency while maintaining quality, accuracy, and stakeholder trust. These findings provide actionable guidance for implementing AI responsibly in specialized industrial contexts and for designing human-AI collaboration that balances automation benefits with critical human judgment.

1. Introduction.

ARTICLE HISTORY

Large language models (LLMs); generative AI; industrial AI; maritime industry; stakeholder communication; human-AI; retrieval augmented generation (RAG)

SUBJECTS

Artificial Intelligence; Psychological Science; Multidisciplinary Psychology; Industry & Industrial Studies; computation; Computer Science (General)

The maritime industry is a significant contributor to worldwide trade, transportation, and the global economy. It is a complex and dynamic sector characterized by multifaceted operations, increasingly more stringent regulations, and a diverse array of stakeholders that include shipowners, port authorities, regulatory bodies, suppliers, and ship management organizations. In such a complex environment, effective communication is paramount, requiring timely, accurate, and contextually appropriate responses to stakeholder inquiries. This can often be a daunting task for human experts due to the volume, complexity, time-critical, and especially nicheness of the information involved. Artificial intelligence (AI) has the potential to alleviate these communication challenges in the maritime industry by assisting human experts with the demands of stakeholder interactions.

This setting can be understood as a form of human–AI teaming in a safety-critical context, where human operators interact with AI-based decision support under conditions of uncertainty and operational risk (Tita A

ß 2026 DNV AS. Published by Informa UK Limited, trading as Taylor & Francis Group

Bach, Kristiansen, et al., 2024). Prior research has shown that in such settings, effective collaboration depends not only on system performance but also on appropriate trust calibration between human users and automated systems. However, how these dynamics play out in the context of Large Language Models (LLMs) in industrial environments remains underexplored.

In recent years, advancements in Industrial AI-enabled systems have introduced promising tools to support and enhance human capabilities in various professional domains. Here, we define AI-enabled systems as any system that contains or relies on one or more AI components, distinct units of software that perform a specific function or task within an AI-enabled system and consist of a set of AI models, data, and algorithms. We use ‘Industrial AI’ here to refer to the specialized application of AI-enabled systems within specific industries – such as energy and maritime – to optimize processes, enhance decision-making, and drive innovation. Unlike general AI applications such as customer AI, which often prioritizes broad adaptability and versatility across various domains, Industrial AI is tailored to address the unique challenges and requirements of a particular industry.

In this paper, we distinguish three related terms that differ in scope, output type, and deployment context. First, AI-enabled systems represent the broad umbrella, referring to computational approaches that enable systems to perform tasks that would otherwise require human intelligence (e.g., classification, prediction, language understanding). Second, generative AI is a subset of AI-enabled systems defined by its output, consisting of models that generate new content (e.g., text in our study) rather than only predicting labels or scores; LLMs are a prominent class of generative AI. Third, industrial AI is defined less by model type than by where and how AI is applied, referring to AI-enabled systems deployed in industrial domains (e.g., maritime), where performance must align with operational constraints, integration into established workflows, and safety and regulatory requirements.

Importantly, this distinction also reflects different theoretical perspectives. General-purpose LLMs are often conceptualized and evaluated as standalone tools based on intrinsic model performance, whereas industrial AI is inherently socio-technical, embedded within organizational processes, regulatory environments, and safety-critical contexts. As a result, evaluation in industrial AI settings must extend beyond model capability to include reliability, accountability, and alignment with operational constraints and human expertise. Accordingly, our study case is an example of industrial AI that uses generative AI (LLM-based textual drafting) within a safety-critical, highly specialized work process.

Industrial AI must be deployable with acceptable risk to be robust and trustworthy. This requires an understanding of both conventional industry risks and new AI-specific risks. Key characteristics of Industrial AI include: the AI’s technical performance; its role and agency within the system; its interaction with other components; and its technical, legal, and ethical impact on stakeholders. A thorough understanding of these risks ensures that Industrial AI systems align with industry-specific standards and deliver value while maintaining safety, reliability, and trust.

One advancement in Industrial AI-enabled systems that is pertinent to the maritime industry is the development of Large Language Models (LLMs), a type of Generative AI. These models, such as OpenAI’s GPT-4, can generate human-like text based on the prompts they receive. They have demonstrated potential in various applications, including drafting emails, creating content, and even coding. This could potentially address many of the challenges to effective communication in the maritime industry.

However, LLMs’ efficacy and usefulness in highly specialized, safety-critical fields, such as the maritime industry, remain areas for exploration. While LLMs have shown great potential in various industries such as healthcare, law, finance, and education, their reported weaknesses – such as factual inaccuracies (C. Wang et al., 2023), hallucinations, and a lack of contextual understanding, pose significant risks (DNV, n.d.c; Lappin, 2024). In the maritime industry, where precision and adherence to regulations are critical, the consequences of LLM inaccuracies could be severe. Erroneous information in compliance documentation, misinterpretation of international maritime laws, or incorrect advice on safety procedures could lead to regulatory penalties, operational inefficiencies, or even accidents that threaten both human lives and the environment.

These weaknesses highlight the need for robust validation processes and human oversight when deploying LLMs in these high-stakes contexts.

Fortunately, LLMs need not to be perfectly accurate to be valuable. They can, for example, serve as tools for generating drafts, offering preliminary suggestions, or enhancing communication by translating technical information into a more accessible language. As long as LLMs transmit the appropriate information, they can provide value.

Our real-world case study investigates the integration of Large Language Models (LLMs) into the workflow of maritime technical experts (i.e., case handlers) at DNV Maritime’s Technical Help Desk (THD). We focus on the role and impact of LLM-generated draft replies in assisting case handlers when responding to stakeholder inquiries. To achieve this, our study aims to comprehensively evaluate the utility of LLM-generated draft replies through two complementary perspectives: the subjective experiences of human experts and the objective textual similarity between LLM drafts and case handlers’ final responses. These dual objectives lead to the following two research questions, each investigated through distinct methodological approaches.

RQ 1: How do case handlers perceive the usefulness of LLM drafts, and how can these drafts support their workflows?

To address the first question, we adopt a mixed-methods approach combining qualitative preliminary research (observations and interviews) with quantitative survey data from case handlers. This question emphasizes the subjective experiences and insights of case handlers, capturing both the perceived usefulness and the various ways LLM drafts may assist their tasks (e.g., summarizing inquiries, standardizing replies, or providing reference points).

RQ2: How similar are the texts generated by the LLM to the final replies sent by case handlers, and what does this similarity imply about the drafts’ potential utility?

To address the second question, we employ a computational approach that quantitatively assesses the alignment between LLM drafts and case handlers’ final responses using text similarity metrics. The degree of similarity between LLM drafts and case handlers’ final responses can serve as a proxy for the drafts’ practical utility. Higher textual similarity may suggest that the drafts were useful as efficient tools requiring minimal modifications and aligning well with case handlers’ intentions. However, recognizing that similarity may not necessarily correspond to utility (e.g., cases where LLM drafts are useful as a starting point even when sent replies differ significantly), we broadly explore whether textual similarity between LLM drafts and sent replies can serve as a proxy for the drafts’ practical relevance.

This study contributes to the literature on human–AI collaboration and Industrial AI in three ways. First, it provides a rare real-world evaluation of LLM-generated draft replies in a highly specialized and safety-critical industrial setting, extending existing research beyond laboratory experiments and general-purpose productivity applications. Second, the study highlights that the value of LLM support cannot be understood solely through traditional technology acceptance perspectives. Our investigation examines situations in which users continue to engage with and recommend LLM-generated drafts despite perceiving important limitations in their accuracy, suggesting that utility may arise not only from automation, but also from functions such as brainstorming, knowledge retrieval, communication support, and workflow augmentation.

Third, the study contributes practical insights into human–AI collaboration by illustrating how LLMs can operate as augmentative tools that support, rather than replace, human expertise and critical judgment. These contributions provide implications for the responsible design and deployment of generative AI in specialized industrial contexts.

Specifically, we investigate how LLMs can augment human expertise and streamline workflows in the maritime industry as a case study for Industrial AI applications. We focus on their potential utility in drafting initial responses to inquiries, synthesizing complex information into accessible formats, and identifying patterns or insights from vast amounts of data. By evaluating these applications, we seek to understand how these draft replies (‘LLM drafts’) can serve as a starting point rather than final deliveries, while remembering that the role of human judgment is essential for refining and validating the AI-generated content. Through this lens, we aim to uncover practical use cases that balance the strengths and weaknesses of LLMs and align their capabilities with real-world needs in maritime operations (Tita A Bach, Kristiansen, et al., 2024).

The findings from our case study are intended to provide insights into the potential benefits and limitations of using LLMs to support communication in a real-world specialized industry setting. We aim to offer practical recommendations and future work for integrating such technologies to enhance operational efficiency and stakeholder engagement for industrial use cases. Our case study also contributes to the broader discourse on human-AI collaboration by showing that the utility of LLMs in safety-critical settings extends beyond direct automation and depends on how they augment human expertise, support critical judgment, and are integrated into existing socio-technical workflows (Tita A Bach, Kristiansen, et al., 2024).

2. Background.

2.1. The case study: Direct access to Technical experts (DATE) system.

We conducted our case study within DNV Maritime (n.d.b), a division of DNV that is responsible and specializes in classification, technical assurance, software, and advisory services for the maritime industry. DNV as an organization operates as an independent third party, providing services such as quality assurance and risk management in safety-critical industries and critical infrastructure. DNV Maritime plays a critical role in ensuring the safety, compliance, and technical integrity of ships and maritime operations worldwide. It provides classification, certification, and technical advisory services based on rigorous rules and international regulations.

DNV Maritime has a flagship service called The Technical Help Desk (THD) branded as ‘Direct Access to Technical Experts’ (DNV n.d.a). The DATE system is a specialized service designed to assist with most technical inquiries related to Fleet in Service (FiS), Certification of Materials and Components (CMC), and Newbuilding projects. The goal is to provide timely and expert support to DNV Maritime’s stakeholders (customers), ensuring their technical issues are resolved efficiently and effectively. In practice, the DATE system functions as a query-answering workflow: when a stakeholder (customer) submits a technical question, such as how to interpret a classification rule or whether specific equipment modification complies with regulations, a case handler (a DNV technical expert) is assigned to respond.

In brief, the query-answering process on DATE system works as the following in practice:

Submission of an inquiry. An inquiry from a stakeholder is submitted on the DATE system’s digital  platform.

Case assignment. Each inquiry is logged as a case and assigned to a case handler, a DNV technical  expert with relevant domain knowledge. Case handlers are located across global offices (e.g., Norway, Germany, Singapore, USA), enabling 24/7 coverage.

Initial Assessment. The case handler reviews the query, assesses its scope, and determines whether  they can respond independently based on, for example, relevant rules, procedures, and internal guidelines.

Expert Collaboration (if needed). Due to the highly specialized nature of maritime engineering and  regulation, spanning over 650 distinct fields of expertise, case handlers frequently consult internal subject-matter experts (SMEs). This collaboration ensures technical accuracy, especially for complex or novel cases. The SMEs may be engaged asynchronously via the DATE platform and discussions are documented within the case file for traceability.

Response Drafting and Validation. Based on their own knowledge and input from SMEs, the case  handler drafts a response. The response is reviewed if necessary by the SMEs before finalizing and sending it to the stakeholder through the DATE system. For Fleet in Service (FiS) cases as the most common type, DNV guarantees a response within 24 hours. The system handles approximately 60,000 FiS cases annually, with a 96% on-time response rate.

Knowledge Retention. All interactions, decisions, and expert consultations are stored in the DATE  platform. This creates an auditable trail and supports future case resolution through precedent.

Experts use the DATE system platform to ensure consistent and efficient handling of both internal and external inquiries. These inquiries are made primarily by the ship management organizations within the maritime industry (see Table 1 for examples of real-world inquiries and replies). Ship management organizations are companies that oversee and manage the operational, technical, and administrative aspects of ships on behalf of their owners, ensuring compliance with international and national regulations and efficient performance. Most of the inquiries come from ship management personnel. In some cases, the inquiries can also come from ship owners, captains, and technical personnel (in this case study, all are included in ‘stakeholders’).

The highly niche nature of the stakeholders’ needs makes responding to these inquiries challenging. Table 1 provides illustrative examples of real maritime stakeholder communications, specifically the inquiries and sent replies related to radar malfunctions. We include this table to demonstrate the highly technical and domain-specific nature of inquiries and replies in this industry, which motivated our case study, methodological choices and the need for LLM support. Case handlers often have to consult with colleagues and other experts and search for similar past cases to guide their replies. Historical records of previous inquiries can be invaluable for providing context and ensuring accurate, consistent, and timely responses to these complex and specialized requests.

Not all inquiries are highly complex or niche; some – such as requests for specific regulations that apply to a particular case – are more straightforward. Even in these simpler scenarios, referring to past cases can help ensure consistency and efficiency in responses. However, the sheer volume of accumulated cases makes it difficult, or at times impossible, to efficiently identify relevant precedents. The database contains over 600,000 past cases which pose significant barriers to extracting useful information when needed.

2.2. The LLM draft function in the DATE system.

In response to this challenge, the DATE system introduced LLM draft functionality as an experimental feature during the first quarter of 2024 (January–March). This is part of an exploratory initiative aimed at understanding whether, and in what ways, LLM drafts can assist case handlers. By producing initial drafts of replies, this functionality seeks to reduce case handlers’ workloads, allowing them to focus on more nuanced and critical aspects of communication to better support stakeholders.

The task of LLMs is to search through the DATE system’s extensive database of historical cases and quickly generate draft replies to new inquiries using relevant information from similar past cases. These drafts are generated using the Retrieval-Augmented Generation (RAG) approach, which enhances the model’s ability to create contextually relevant responses by integrating document retrieval with text generation. This approach ensures that the generated drafts are grounded in relevant historical records while being tailored to the specifics of the new inquiry.

The RAG method augments the input to the LLM with relevant information from knowledge sources. It is known to enhance the accuracy and reliability, and thereby trustworthiness, of its output by enabling the LLM to reference credible knowledge sources outside of its training data when generating text. Research suggests that RAG approach provides improved control over LLM-generated content while reducing hallucinations, a common issue where LLMs produce incorrect, false, or fabricated information (G. Gao et al., 2024; Bechard & Ayala, 2024; Shuster et al., 2021; Fan et al., 2024).

The process of generating draft responses to stakeholder inquiries using RAG in the DATE system is as follows (Figure 1): The DATE system maintains a vector database of 600,000 past inquiry cases. Existing inquiry cases are converted into text embedding vectors, numerical representations of the semantic aspects of the texts. When a new stakeholder inquiry is received, the system converts the inquiry text into an embedding vector and then searches the vector database to identify the most semantically similar cases using cosine similarity. The cases corresponding to these vectors are considered the most relevant to the new inquiry. The system then provides the LLM with a prompt that includes the three most similar inquiry-response pairs along with the new inquiry text. The LLM references these inquiry-response pairs as contextual examples to generate a response for the new inquiry.

This allows LLM to generate a more contextually appropriate and accurate response for the new inquiry. The LLM used for generating draft replies is OpenAI’s ChatGPT-3.5 model.

2.3. LLM Draft usage in practice.

The LLM draft generation feature is currently in trial and optional for case handlers to use on a case-by-case basis. Case handlers are encouraged to explore the function critically. As part of our study, we disclosed to case handlers that the drafts were AI-generated, since the LLM was being co-developed with them and their feedback was essential to its improvement. When an LLM draft is generated, case handlers are also presented with references and links to similar cases that inform the draft, allowing direct access to these cases for further review. To receive the LLM drafts, case handlers must click the ‘Generate LLM Drafts’ button in the user interface.

It is important to note that this feature is only available for responding to a new inquiry; if multiple exchanges occur (e.g., an email chain), the feature is no longer accessible (Appendix A shows the distribution of the number of exchanges between the stakeholders and the case handlers per case). Another limitation is that attached documents (e.g., PDFs or images) are not included in the LLM’s analysis. As a result, some crucial information may not be fully conveyed when generating drafts. In our setup, case handlers could, in principle, manually add relevant details from such documents into the inquiries if they wanted the LLM to take them into account, but this was not explicitly encouraged in the trial and was not a core part of the workflow we studied. Instead, the focus was on evaluating the usefulness of the draft generation as designed.

Case handlers can re-generate the LLM drafts as many times as they like, combining text from different versions before finalizing and sending the response to the stakeholders. Each time case handlers re-generate the LLM drafts, the newly generated text is displayed below the previous versions, allowing comparisons between versions.

3. Methodology.

Our case study employed both subjective and computational measures to evaluate the usefulness of LLM drafts by case handlers.

3.1. Preliminary study.

When the LLM draft functionality was first introduced, we conducted a preliminary qualitative study to understand early real-world use and to inform the subsequent survey instrument. This phase included direct, non-participant observations of case handlers’ work and semi-structured interviews with both case handlers and developers. In total, we conducted 42 hours of observations, four interviews with case handlers, and two interviews with developers. The observations focused on case handlers’ day-to-day work, particularly how they interacted with LLM-generated drafts, integrated them into replies, and navigated supporting systems. Field notes captured examples of challenges (e.g., alignment between draft suggestions and workflow) and opportunities (e.g., efficiency in searching for previous cases).

The interviews with case handlers provided reflections on the practical usefulness of the drafts, factors influencing adoption, and perceptions of accuracy and relevance. The developer interviews offered complementary perspectives on the design choices, system constraints, and expectations regarding draft uptake.

All notes and transcripts were coded thematically, with an emphasis on identifying recurrent patterns rather than statistical generalization. These findings are presented as illustrative and exploratory: their purpose is to shed light on possible mechanisms behind the patterns observed in the main analyses, rather than to serve as stand-alone results.

3.2. A Survey for case handlers (subjective measure).

The goal was to capture the case handlers’ perceptions of the LLM drafts (See Appendix B for the questionnaire). We disseminated the questionnaire between May and June 2024 to match the data collection timeline for our text similarity analysis.

3.3. A Text similarity analysis (computational measure).

We measured the textual similarity between LLM drafts and sent replies. Our underlying assumption was that smaller textual differences would indicate that case handlers were able to make more effective use of the LLM drafts. To evaluate textual similarity, we used two methods as a triangulation approach: semantic embedding similarity (SES), and LLM-as-a-judge (LAAJ) rating. We also analyzed text similarity scores between SES and LAAJ to evaluate the reliability of results across LAAJ and SES through cross-verification and assess their complementarity. Data used for the analysis was collected for inquiry responses between May and June 2024 to match the survey data collection timeline.

3.4. Ethical considerations for human participation.

Our case study was conducted in accordance with the guidelines provided by Sikt, the Norwegian Agency for Shared Services in Education and Research, a public administrative body under the Ministry of Education and Research in Norway. Based on Sikt’s criteria, our research did not require formal ethical review for approval or notification, as it exclusively involved the collection and analysis of anonymous data and no personal data was processed (Sikt - Norwegian Agency for Shared Services in Education and Research, n.d). This means that at no point during the study – which included observations, interviews, a survey, and text similarity analysis – was it possible to identify individual participants. Thus, our research complied fully with Sikt’s requirements and ensured the protection of participants’ data (See Supplementary Material).

In addition to not collecting personally identifiable information, all responses were recorded and stored in a manner that prevented the identification of individual participants.

All interview and observation participants gave informed consent prior to their participation and written informed consent was obtained. The process for interview and observation participants was thorough. Participants were first approached by email and received a detailed attachment outlining their rights as participants. At the start of every interview or observation session, participants were reminded that participation was entirely voluntary. Acceptance to participate was formally documented. Furthermore, participants were reminded at the conclusion of the session that they retained the right to withdraw any data gathered, at any time, by contacting the research team.

For survey participants, implied consent (a form of verbal consent) was obtained by their voluntary action of completing and submitting the questionnaire. This approach was chosen over written consent for the survey to ensure anonymity, maximize the convenience for a large sample, and maintain a high response rate, as the study presented no more than minimal risk to participants. Participation in the study was entirely voluntary, and all participants were informed that they could withdraw at any time without providing a reason. Participants were also informed that their data would be used exclusively for research purposes and handled in compliance with ethical and data protection standards. More detailed ethical considerations are also included in the next subchapters.

3.5. The survey study.

3.5.1. The development of the questionnaire.

To assess case handlers’ experiences with the LLM-generated drafts, we developed a questionnaire. Questionnaire development was informed by in-depth discussions among the research team, DATE developers, and several case handlers to identify key aspects of LLM draft use in this work setting. In addition, we drew on established validated instruments, including those by Brooke (1996), Davis (1989), and Venkatesh et al. (2003), selecting items and constructs that were relevant to the DATE context. We then used insights from the preliminary observations and interviews to contextualize and refine the selected items, adjusting wording to reflect case handlers’ terminology, tasks, and interaction patterns (e.g., referring explicitly to ‘LLM drafts’ and concrete work practices rather than generic ‘the system’).

Where the preliminary findings pointed to important context-specific issues not captured well by the validated scales (e.g., draft modification practices and reliance when in doubt), we added tailored items. The questionnaire was piloted with a small number of case handlers and iterated until all authors approved the final version (see Appendix B for the full questionnaire).

More specifically, the preliminary qualitative findings shaped the survey in three ways. First, they guided which constructs were prioritized for measurement in this setting (e.g., perceived accuracy/ relevance, usefulness of summaries, and intention to use). Second, they motivated the inclusion of context-specific items not well captured by generic acceptance scales, such as the extent to which case handlers modify drafts and what sources they rely on when in doubt (Table 2, items 6 and 13).

Third, observations of how drafts were used in practice (e.g., as starting points for brainstorming and for aThis questionnaire item is treated as an independent variable for statistical analyses. bThis questionnaire item is treated as a dependent variable for statistical analyses. cExtensively is scored as Strongly Disagree, Moderately as Disagree, Slightly as Agree, and None at all as Strongly Agree. standardizing tone) informed both the inclusion and contextual wording of items related to productivity, language quality, and recommendation/intention to use (Table 2, items 2–5, 8–11).

The questionnaire consisted of 21 questions, in which 7 were categorized as participant demographics and background characteristics (hereafter, ‘participant characteristics’) questions and 14 were measuring case handlers’ experiences and opinions with the LLM drafts. The last participant characteristics question measured how long a participant had been using the LLM drafts (i.e., a few days, a few weeks, a few months, or never). The survey was ended for those who responded with ‘never used the LLM drafts’ after asking them to provide a reason (i.e., unavailable to them, chose not to use them, or other reasons followed by a free text comment field). Participants who had used the LLM drafts received 14 questions (see Table 2). For these participants, the questionnaire was supposed to take 3–5 minutes to complete.

The questionnaire was designed this way to minimize the time taken from case handlers’ chargeable work hours.

3.5.2. Survey dissemination and participation.

To ensure maximum participation, the questionnaire (Appendix B) was distributed electronically to a randomly selected sample of 198 case handlers within the DATE systems department based in USA, Norway, Greece, Singapore, Germany, and Poland. The selected sample population included case handlers with diverse roles, experience levels, and geographical locations. This ensured that the survey results reflected a broad range of perspectives and were representative of the organization’s operational diversity. We adopted a sampling approach in alignment with organizational confidentiality policies to ensure that sensitive workforce data remained protected.

Prior to dissemination, we obtained formal approval from senior management to distribute the questionnaire. The link to the questionnaire was emailed directly to case handlers during May and June 2024 with a description and purpose of the questionnaire. We also sent three reminders during the data collection period. To encourage participation, the senior managers of DATE systems actively promoted the survey during department meetings, emphasizing its importance for improving case handling processes. No incentives were given to participants.

For the survey data collection, this case study complied with required ethical and legal standards relating to participant anonymity and confidentiality. Participation in the survey was anonymous and voluntary. The questionnaire was created and all responses were collected on Microsoft Teams Form with restricted access. Only the research team has access to the Teams channel and only the lead researcher of this case study had access to the questionnaire anonymous responses.

3.5.3. The analysis of the questionnaire responses.

The independent variables included participant characteristics questions and one question about trust (See Table 2 and Appendix B):

The participant characteristics questions: years of experience as a case handler, office location, age  group, gender, current role, time using the LLM drafts, and native language(s)

The native language(s) were converted into non-native English speakers and native English s speakers.

The questionnaire item #7 ‘Stakeholders’ trust in the organization’: ‘I believe that the LLM drafts will:  increase, decrease, or have no impact on our stakeholders’ trust in the organization’. Given the externally facing nature of DATE communications and the safety-critical context, we included here a measure of perceived impact on stakeholders’ trust in the organization as a potential factor influencing adoption and user perceptions.

The dependent variables were the questionnaire’s items number 1–8 and 11 from Table 2 (all used a 4-point Likert-type scale). See Table 2 for selection of independent and dependent variables. For transparency, we retain the original questionnaire item numbers when reporting results. Because the survey items were presented in a shuffled order, related items are not consecutive and may therefore appear out of sequence in Table 2.

To examine group differences in perceived usefulness of the LLM drafts, the questionnaire items of dependent variables were averaged for each participant as a total mean score rather than as individual questionnaire items. We compared the dependent-variable scores (i.e., the total mean scores of the nine questionnaire items) across groups defined by the independent variables (i.e., participant characteristics and the impact on stakeholders’ trust in the organization) using Kruskal-Wallis H test. This analysis determined whether there were statistically significant differences in the total scores between groups for: the current roles, years of experience being a case handler, office locations, age groups, gender, non-native vs. native English speakers, length of using the LLM drafts, and perceived impact on stakeholders’ trust in the organization.

We conducted post-hoc pairwise comparisons using Dunn’s test with a Bonferroni correction applied to control for multiple comparisons. This adjusted the significance level to p < 0.017.

To assess whether users and non-users differed in their characteristics, we tested differences in the composition of the participant characteristics between users (i.e., case handlers who had used the LLM drafts) and non-users (i.e., case handlers who never used the LLM drafts) using Pearson’s chi-square tests separately within each participant characteristic (i.e., the current roles, years of experience being a case handler, office locations, non-native vs. native English speakers, age groups, gender, length of using the LLM drafts, and the impact on stakeholders’ trust in the organization) as factors. If interactions between group and category were found (i.e., if the participant characteristics composition of users and non-users differed), we conducted follow-up pairwise Pearson’s chi-square tests between sub-groups.

Odds ratios (where an odds ratio of 1 indicates that the participant characteristics composition of users and non-users did not differ) calculated the effect size. We used a significance level of p < 0.05, the Bonferroni correction for each family of comparisons, where appropriate.

To examine relationships between questionnaire constructs, we computed Pearson correlation coefficients between the nine questionnaire items. This analysis was conducted separately from the group comparison analyses to explore associations among the items.

Finally, to analyze qualitative feedback, we conducted a thematic analysis following an inductive approach of the survey response free text for the questionnaire item #14 (i.e., General comments or suggestions about the LLM drafts: [free text]). We identified and extracted significant words and phrases from each transcript and interpreted their meanings in context. Codes expressing similar meanings were grouped into themes.

3.6. Text similarity analysis.

The initial dataset for text similarity analysis included all 13,794 stakeholder inquiries received through the DATE system between May and June 2024 (Figure 2).

The DATE system stores stakeholder inquiries and LLM-generated drafts in separate databases. We merged these using case IDs to create a dataset linking drafts with the corresponding replies sent by case handlers. This process yielded 7,061 cases after excluding 6,733 cases where either the reply or the draft was missing.

To ensure comparability, we further restricted the dataset to 2,414 cases in which only one LLM draft was generated, excluding 4,647 cases with multiple drafts (see Appendix C for distribution). Multiple draft generations risked case handlers combining text from different versions, making it difficult to identify the specific draft associated with a reply. Although this restriction may not capture all usage patterns, it enables clear draft–reply matching and improves the accuracy and consistency of the analysis.

We note, however, that this restriction is unlikely to be neutral with respect to our similarity findings. Re-generation of a draft is itself an indication that the case handler found the initial output insufficient; the excluded subset therefore plausibly represents cases in which the LLM’s first draft was less useful than in the retained single-draft cases. As a result, our reported similarity scores should be interpreted as an upper-bound estimate of LLM draft utility across the full population of inquiries handled in the DATE system during this period, rather than as representative of all usage patterns.

We considered a sensitivity analysis comparing first-draft similarity across the multi-draft subset, but because case handlers could combine text from multiple regenerated versions before sending their reply, no single version can be reliably identified as ‘the first draft’ actually used, the same ambiguity that motivated our single-draft restriction in the first place. We return to this point, together with a related source of potential upward bias in our LAAJ scores, in Section 4.3.3.

To protect personal information and focus on the inquiry content, personally identifiable details (e.g., sender and recipient names, salutations, affiliations, and contact information) were removed using rule-based preprocessing such as pattern matching and keyword-based filtering. To standardize text formats and prevent non-content elements (e.g., HTML tags, line breaks, and consecutive spaces) from affecting text comparison, these elements were also removed from the dataset using text cleaning procedures, including regular-expression-based parsing and normalization. After cleaning, 108 cases where inquiries or replies were left without content were excluded, resulting in a final dataset containing 2,306 cases.

3.6.1. Text similarity measures.

To computationally compare LLM drafts and replies sent by case handlers, we employed two measures of text similarity: semantic embedding similarity (SES) and LLM-as-a-judge (LAAJ). We then compared the similarity scores produced by these two measures to evaluate their consistency and reinforce the reliability of our findings.

3.6.1.1. Semantic embedding similarity (SES).

Semantic embedding represents text as a vector in a high-dimensional space, where each dimension corresponds to a latent semantic feature learned by the model during its training on vast text corpora. This vector captures semantic nuances, allowing for the quantification of semantic similarity between two texts by measuring the closeness between their respective vectors. Cosine similarity is a common metric used for this purpose, quantifying how close two text embeddings are by calculating the cosine of the angle between their vectors. For text embeddings, cosine similarity scores typically range from 0 to 1, where 1 indicates perfect similarity and 0 indicates no similarity.

For each of the 2,306 cases in the final dataset, we converted both the LLM-generated text (the LLM draft) and the case handler-generated text (the sent reply) into semantic embedding vectors and calculated the cosine similarity between them. This process yielded a semantic embedding similarity score for each case, as well as the overall similarity score distribution across the dataset. For converting text into semantic embedding vectors, we used two OpenAI text embedding models to ensure the reliability of the results: text-embedding-ada-002 and text-embedding-3-large1.

3.6.1.2. LLM-as-a-judge (LAAJ).

LAAJ leverages the advanced language comprehension capabilities of

LLMs to evaluate the quality of texts based on specific instructions (prompts). This method is grounded in LLMs’ extensive language knowledge and contextual understanding abilities, enabling detailed and specialized evaluations across various dimensions of language. Prior work has examined LLM-based evaluation as an emerging approach for assessing open-ended text outputs, including whether LLM judgments can approximate human evaluations, how such judgments can be structured through criteria-based prompting, and under what conditions they remain reliable At the same time, previous studies also highlight important limitations of LLM-based evaluation, including task-dependent variability, potential judge biases, and limited reliability for some evaluation dimensions such as factuality. We situate our use of LAAJ within this evolving methodological literature.

Drawing on its demonstrated utility, we adopt LAAJ as a measure of textual similarity between LLM drafts and case handlers’ sent replies; in light of its known limitations, however, we apply it not as a standalone substitute for human evaluation but as one of two complementary measures – together with semantic embedding similarity (SES).

The evaluation dimensions, comparability of language and similarity of interpretability, were directly adopted from the validated framework of Sperber (2004), originally developed to assess functional equivalence between translated versions of research instruments. This grounding in a published, validated method provides a principled basis for the evaluation criteria.

We adapted this framework in four ways for our computational setting: the scale direction was reversed relative to Sperber’s original so that higher scores indicate greater similarity, for more intuitive reporting of results; explicit scale range anchors were embedded directly within the prompt (1–2 1⁄4 not at all comparable/similar; 3– 5 1⁄4 moderately comparable/similar; 6–7 1⁄4 extremely comparable/similar), since an LLM rater cannot be assumed to intuitively interpret an unlabeled scale consistently; an overall similarity dimension was added as an integrative, holistic measure capturing aspects of textual equivalence not fully reducible to the two individual dimensions; and structured JSON output with mandatory explanations to ensure transparency and support verification of the evaluations.

Comparability of language refers to the degree of overlap in the surface form of the texts, including similarities in words, phrases, sentence structures, and stylistic expressions. A high score on this dimension indicates that the texts use similar wording and phrasing, whereas a low score reflects substantial differences in linguistic expression even if the underlying meaning may be similar. For example, the draft ‘The vessel must undergo inspection before departure’ and the reply ‘The vessel must be inspected prior to leaving port’ would be considered highly comparable in language, while a version such as ‘Inspection is required before sailing’ would reflect lower language comparability despite conveying a similar requirement.

Similarity of interpretability captures semantic and pragmatic equivalence between the texts: whether the LLM draft and the sent reply convey the same underlying meaning, intent, and actionable outcome. A high score indicates that both texts would lead to the same understanding, conclusion, or decision by a reader, whereas a low score reflects differences in meaning, omitted or additional information, or divergent interpretations. For example, ‘A survey is required before approval’ and ‘Approval cannot be granted until a survey is completed’ would score high on interpretability despite wording differences, whereas a draft stating ‘A survey may be required’ compared to a reply stating ‘A survey is mandatory’ would reflect low interpretability similarity due to a difference in obligation.

This dimension is therefore distinct from language comparability, as texts can differ linguistically while remaining interpretatively similar. Overall similarity represents a holistic assessment that integrates both linguistic and interpretative aspects, capturing the extent to which the two texts are comparable as complete responses. For example, two texts may have low language comparability but still receive a high overall similarity score if they convey the same conclusions and recommendations, whereas texts that share wording but differ in key details or conclusions would receive lower overall similarity scores. Importantly, the overall similarity dimension is not simply a mechanical average of the other two dimensions.

Rather, it captures an integrative judgment of how the two texts compare as complete responses – one that may diverge from what either individual dimension would predict, particularly in cases where language comparability and interpretability pull in opposite directions. For example, a draft that is linguistically distant from the sent reply but conveys the same actionable conclusion would score low on language comparability but high on interpretability; in such cases, overall similarity allows the LLM to reflect on how comparable the texts are in practice, taking into account how the two dimensions interact rather than treating them as independent. Including overall similarity in the composite score therefore adds a holistic, context-sensitive component that the two specific dimensions alone cannot fully capture.

The composite score thus reflects both dimension-specific assessments and this broader integrative judgment, which is particularly relevant for assessing practical equivalence in real-world use.

To enhance replicability and transparency, the LAAJ prompt was developed through an iterative prompt engineering process. This involved grounding the evaluation criteria in an established framework, translating these criteria into structured instructions suitable for LLM interpretation, and refining the prompt to ensure consistent scoring behavior through explicit scale anchors and output constraints. Particular attention was given to minimizing ambiguity by embedding clearly defined scale ranges and requiring structured JSON outputs with mandatory explanations. The full prompt is provided in Appendix D to enable direct reuse. To support validation of the prompt design, we assessed the stability and reproducibility of LAAJ outputs by conducting three independent runs and evaluating inter-run agreement (see Appendix F).

The high consistency across runs indicates that the prompt yields stable evaluations despite the probabilistic nature of LLM outputs. These design choices align with prior work emphasizing that explicit criteria and structured outputs improve the reliability of LLM-based evaluation.

The LLM’s evaluations were conducted using a 7-point Likert scale, where 1 represented ‘not at all comparable/similar’ and 7 represented ‘highly comparable/similar.’ Additionally, to enhance transparency and enable researchers to verify the evaluations, we instructed the LLM to provide explanations justifying its assessments. The complete prompt provided to the LLM, including these instructions, is provided in Appendix D. We used OpenAI’s GPT-4o (version 2024-05-13) as the judge LLM, accessed via the OpenAI API at temperature 1⁄4 0 to maximize output determinism and minimize stochastic variability across evaluations. We utilized Microsoft Azure OpenAI services to implement our evaluation framework and designed the system to generate and store structured responses in JSON format for easier analysis.

For each of the 2,306 cases in the final dataset, we provided the LLM with the prompt along with the LLM draft and the sent reply to evaluate textual similarity. As a single metric for comparison with other similarity measures like SES, we created a composite score by taking the mathematical average of the three dimension-specific scores obtained (i.e., overall similarity, language comparability, and similarity of interpretation), representing combined evaluations across all three dimensions. Furthermore, to verify output stability, since temperature 1⁄4 0 produces near-deterministic rather than strictly deterministic outputs due to floating-point variability across API calls, we conducted three independent evaluation runs and averaged these scores to obtain more stable and representative final similarity scores.

While we report a composite score for comparability across cases, the inclusion of the overall similarity dimension ensures that this aggregation incorporates both dimension-specific assessments and a more holistic evaluation reflecting practical equivalence.

To further support the robustness and replicability of the LAAJ approach, we assessed inter-run consistency across the three evaluation runs using statistical measures (ANOVA and Kendall’s W; see Appendix F). The absence of significant differences between runs and the high agreement levels indicate that the prompt produces stable and reproducible evaluations despite the probabilistic nature of LLM outputs. This combination of a validated evaluation framework, structured prompt design, and empirical consistency testing strengthens the LAAJ-based approach.

3.6.1.3. Comparison of SES and LAAJ scores.

As a methodological triangulation approach for text similarity assessment, we conducted comparative analysis of the similarity scores obtained through both SES and LAAJ methods. We calculated the correlations between SES and LAAJ similarity scores across the entire dataset using Pearson and Spearman correlation coefficients. We also examined the statistical characteristics and density distribution plots of both similarity score distributions to thoroughly compare and analyze the distribution patterns of the two similarity scores.

4. Results.

For ease of navigation, we present the results in the same phase sequence as the Methodology: preliminary qualitative themes, survey findings, and text similarity analyses.

4.1. The preliminary findings.

This subsection corresponds to the ‘Preliminary study’ phase as described in the Methodology and provides contextual themes that help interpret the survey and text-similarity findings reported below. We conducted an exploratory qualitative phase (observations and interviews) to provide contextual insight into early use of the LLM draft function and to inform the questionnaire development. Here we summarize the most salient preliminary themes that help interpret the survey and text-similarity findings.

The observations focused on case handlers’ day-to-day work, particularly how they interacted with LLM-generated drafts, integrated them into replies, and navigated supporting systems. Field notes captured examples of challenges (e.g., alignment between draft suggestions and workflow) and opportunities (e.g., efficiency in searching for previous cases).

The interviews with case handlers provided reflections on the practical usefulness of the drafts, factors influencing adoption, and perceptions of accuracy and relevance. The developer interviews offered complementary perspectives on the design choices, system constraints, and expectations regarding draft uptake.

All notes and transcripts were coded thematically, with an emphasis on identifying recurrent patterns rather than statistical generalization. These findings are presented as illustrative and exploratory: their purpose is to shed light on possible mechanisms behind the patterns observed in the main analyses, rather than to serve as stand-alone results.

4.1.1. Levels of complexity of inquiries.

Stakeholder inquiries present an almost limitless range of variations, scenarios and combinations. These include factors such as vessel type, construction site, operational timeline, sailing conditions, geographical location, local regulations, crew history, vessel maintenance records, engineering culture, maintenance decisions, shipping tax regulations, onboard equipment types, and the diversity of equipment suppliers, among others. This made users question whether an LLM could have been accurately trained on such complex data. We also found that the level of complexity of stakeholder inquiries significantly influenced how case handlers perceived the usefulness of LLM drafts. A case handler’s quote illustrates perceived challenges for the LLM to address complex inquiries:

In the maritime industry, every ship is different, even sister ships have differences. The problems they encounter are variable. This is because different people manage the ship, different management styles, different operational areas, cargoes, and flags.... So probably [the LLM] would take the answer from an already answered question, which is of similar age, vessel, similar type of vessel, similar flag, but not necessarily the same management. – A case handler

Conversely, simpler inquiries (e.g., those involving clear procedures and straightforward instructions with readily available explicit knowledge) tended to inspire greater trust in the LLM drafts. In such cases, the required information was often readily available in past cases or documented sources like regulations and procedures, making it easier for the LLM to generate more accurate LLM drafts. Case handlers also suggest that these low-complexity cases would be easier for them to verify.

4.1.2. Knowledge sharing.

Case handlers dealt with the level of complexity presented by the stakeholder inquiries by discussing with colleagues, calling surveyors on board for field-based information, and using their experiences and domain knowledge. Our observations also suggest a high level of collaboration and teamwork between and among case handlers and other colleagues, especially when responding to complex inquiries. Often, handlers made decisions collaboratively, rather than alone, to raise the confidence and quality of the replies to the stakeholder inquiries. Case handlers reported that collaboration and teamwork might also lead to the creation of new knowledge or new ways of dealing with cases. Lastly, case handlers felt that the usual way they responded to inquiries, especially complex ones, could not be incorporated into the LLM since this involved tacit knowledge.

4.1.3. Brainstorming.

Although the LLM drafts were created to help case handlers respond to the inquiries, we observed that the LLM drafts were used in other ways as well. Case handlers leveraged LLM drafts to generate ideas for addressing new inquiries. They reflected on aspects, such as necessary information, relevant regulations, and applicable domain expertise, to help shape their replies. In addition, case handlers tested a filtering function within the user interface (UI) that enabled them to collaborate with the LLM in retrieving cases used to generate drafts, based on criteria such as vessel type, flag, or age. Case handlers explored how different filters changed the content of the LLM drafts to generate different ideas and evaluate the specificity of the content of the LLM drafts.

4.1.4. Skeptical but curious.

Initially skeptical, case handlers carefully evaluated the LLM drafts to determine alignment with their knowledge. When discrepancies arose, they often dismissed the system as immature. However, they revisited it after updates or improvements were announced, demonstrating a cautious openness to innovation.

4.1.5. Humanizing.

In the cases when case handlers did not understand the LLM drafts, they tended to humanize the technology by making comments such as calling it a toddler, that ‘it’ did not understand the context of the inquiry, and ‘it’ made up content.

4.1.6. Trust.

We observed that the trust of case handlers in technology increased when they recognized phrases from prior cases. In contrast, LLM drafts perceived as drawing primarily from generic internet sources were met with greater skepticism.

4.2. The survey results.

This subsection reports results from the ‘Survey study’ as described in the Methodology, presented in the same logical sequence: (i) dissemination and participation, (ii) respondent characteristics (users vs. non-users), and (iii) descriptive item responses and group comparisons. A total of 74 of 198 invited, selected case handlers responded to the questionnaire (response rate 1⁄4 37.37%), in which 23 selected ‘never used the LLM drafts,’ leaving 51 participants (response rate 1⁄4 25.76%) to respond to the 14 questionnaire items measuring their experiences and opinions of the LLM drafts (Figure 3). In this case study, we designated participants who reported to have used the LLM drafts as ‘users’ and those who reported to never use the LLM drafts as ‘non-users’.

Non-users commonly reported their reasons for never using the LLM drafts were that it was not fit-for-purpose and unavailable (Table 3). Unavailability to non-users was probably because the non-users moved to another unit or team without access to the LLM draft function on trial. It might also be possible that non-users were not aware of the new functionality.

Table 4 shows the participant characteristics of all the participants, users, and non-users. The most common role of the case handlers across the three groups was engineering. All participants and users most commonly had 5–10 years experience as case handlers, whereas non-users most commonly had longer than 15 years’ experience. Norway and Germany were the most common office locations for the participants across the three groups. The most common age group for participants across the three groups was older than 50 years, followed by 41–50 years old. Most of the participants identified themselves as male across the three groups. German and Norwegian were the most common native languages across the three groups. Similarly, most of the participants across the three groups were aNon-users who selected this response received a follow-up question to provide their reasons in free texts.

categorized as non-native English speakers. A Pearson’s chi-square analysis (N 1⁄4 51) revealed that there were no statistically significant differences in the composition of the participant characteristics between users and non-users (Table 5).

Over 60% of the users (N 1⁄4 34) did not provide any responses when asked how the LLM drafts improved their responses to stakeholder inquiries. Of the 17 users (33.33%) who did respond (Figure 4), the majority indicates that the LLM drafts enhanced their efficiency. Almost all users (92.16%) stated that using the LLM drafts did not affect their frequency of consulting with their colleagues (Figure 5). Users relied mostly on their knowledge and colleagues rather than on the LLM drafts when in doubt (Figure 6). The nine users who selected ‘Other,’ either alone or in combination with other options, provided free-text responses indicating that, when in doubt, they relied on previous replies to similar cases and on rules and regulations. One user mentioned that they did not rely on the LLM drafts until they believed the technology had become more mature.

More than half of the users (50.98%) believed that the LLM drafts would have no impact on the stakeholders’ trust in the organization (Figure 7).

aExtensively was scored as Strongly Disagree, Moderately as Disagree, Slightly as Agree, and None at all as Strongly Agree.

In general, users responded more negatively to items 1, 6, 8, and 11 (Table 6). This shows that they perceived that the LLM-generated draft replies’ quality still needed significant improvement. Users responded more positively to the items 2, 3, 5, and 9, indicating that users intended to continue using the LLM drafts and recommend others use them as well. Users also perceived the language as of high quality. This shows that although the quality of the LLM drafts was perceived as an area for improvement, users were willing to continue using the LLM drafts. A split was observed whether the summary of the LLM drafts was perceived as useful.

The results from statistical analyses show that the only significant difference is the impact the LLM drafts have on stakeholders’ trust in the organization (Table 7). We conducted a Dunn’s post-hoc test with Bonferroni correction to follow up on the significant Kruskal-Wallis test results for the variable ‘stakeholders’ trust in the organization’ on ‘responses.’ The results indicate significant differences between the ‘decrease’ and ‘increase’ groups, Z 1⁄4 −4.29, padj < 0.001, and between the ‘decrease’ and ‘no impact’ groups, Z 1⁄4 −4.07, padj < 0.001. There was no significant difference between the ‘increase’ and ‘no impact’ groups, Z 1⁄4 0.96, padj 1⁄4 1.00. The ‘decrease’ group’s mean score was 1.69 of 4.00, the ‘increase’ group’s was 2.00, and the ‘no impact’ group’s was 2.15.

Based on these mean scores, the ‘decrease’ group seems to have the most negative perception of the overall score of the nine dependent variables.

We conducted correlation analysis on the nine questionnaire items, resulting in 36 pairwise correlations (Appendix E). Of these, 20 correlations were statistically significant (p < 0.05), indicating meaningful relationships between the corresponding variables. All the significant correlations showed positive relationships, suggesting that as one variable increased, the other also tended to increase. In contrast, 16 correlations were not statistically significant (p > 0.05), indicating no strong evidence of a relationship between those pairs of items. These findings suggest several patterns in users’ perceptions.

First, adoption-related items (recommendation, intention to continue, frequency of use, and perceived productivity) were strongly inter-correlated (r 1⁄4 0.59–0.72, p < 0.001), indicating that willingness to adopt the LLM drafts co-varied closely with perceived productivity gains. Second, perceived usefulness of the summary and perceived accuracy was moderately associated with both productivity and adoption intentions (e.g., summary usefulness with intention to continue r 1⁄4 0.63, p < 0.001; accuracy with intention to continue r 1⁄4 0.49, p < 0.001), suggesting that perceived content value is linked to adoption partly through perceived efficiency. In contrast, language quality showed weak and largely non-significant correlations with other items, suggesting it was evaluated as a relatively distinct attribute rather than a primary driver of continued use.

Finally, perceived accuracy was associated with lower reported modification (r 1⁄4 0.40, p < 0.01; higher scores represent less modification), consistent with the descriptive finding that extensive editing was common when drafts were perceived as inaccurate.

The general comments from the questionnaire item #14 from the users about the LLM drafts were summarized as follows: 1. Accuracy and tailoring of the LLM drafts Users often perceived the LLM drafts to lack technical accuracy and relevance, with some comments noting that replies were either too general, missing the point of the stakeholders’ questions, or repeated the same information. Although users perceived the language of the LLM drafts as high-quality, they suggested that the content lacked substantive content or even provided incorrect information. The need for the LLM to handle cases individually, rather than applying a one-size-fits-all approach, was a recurring theme.

Consequently, users expressed a need for the LLM drafts to be more precise and better tailored to the specific details of each case. They also wanted to allow for customization based on specific stakeholder needs and the complexity of the inquiries. This might be achieved by allowing the LLM access the documents attached to stakeholder inquiries and by considering regional specifics such as differences in handling, fees, and regulations in different countries. Importantly, several comments emphasized the need for the LLM to integrate data from other sources and databases to provide more informed and contextually accurate LLM drafts. Users reported that the frequent inaccuracies of the LLM drafts gave them concerns about the reliability of the LLM drafts, with some expressing that the current system could not be fully trusted.

There was a suggestion that the LLM should refrain from generating an LLM draft if it was unsure, rather than providing potentially misleading information.

2. User experience and feedback Some users appreciated the concept and potential of the LLM drafts, finding it useful as a starting point or for generating initial insights (e.g., summarizing inquiries). However, they acknowledged the LLM drafts still needed significant improvement as they did not necessarily solve the problems posed in the inquiries. Users suggested having more direct involvement in the improvement of the LLM drafts, such as the ability for users to mark correct and incorrect drafted replies, as feedback to the LLM and the developers. Additional training and system refinement were suggested to improve the overall effectiveness of the system.

4.3. Text similarity analysis.

This subsection reports the ‘Text similarity analysis’ as described in the Methodology, including semantic embedding similarity (SES), LLM-as-a-judge (LAAJ), and the triangulation comparison between measures.

4.3.1. Semantic embedding similarity (SES).

We calculated the Semantic Embedding Similarity (SES) score for each case by converting the LLM draft and the sent reply into semantic embedding vectors and then computing the cosine similarity between them. A score closer to 1 indicates higher semantic similarity, while a score closer to 0 indicates semantic irrelevance. Figure 8 shows the distribution of SES scores obtained from the final dataset of 2,306 cases. We employed two text-to-embedding models – text-embedding-ada-002 (ADA) and text-embedding-3-large (LARGE) – to verify whether the SES score distribution was consistent across embedding models, rather than being driven by a single model choice. Panels A and B display the results from the ADA and LARGE models, respectively. Each distribution has been adjusted to its valid score range.

Most notably, both panels exhibit a bimodal distribution pattern. The distribution reveals two distinct groups: one with higher similarity scores where LLM drafts showed moderate alignment with sent replies, and another with lower similarity scores where LLM drafts and sent replies diverged significantly. From a utility perspective, the higher similarity group suggests that LLM drafts captured the semantic content of the sent replies to some extent, allowing case handlers to utilize them with moderate modifications. Conversely, the lower similarity group indicates that LLM drafts and sent replies were semantically divergent, requiring case handlers to undertake extensive rewriting. This polarized distribution suggests that the LLM’s performance was inconsistent, being effective in some cases but inadequate in others.

A manual inspection of cases with low similarity scores revealed that they often included generic content not directly relevant to the case, such as requests for more context, clarifications of the question, or suggestions to contact the relevant office or expert. Such generic text was dissimilar to the specific, contextualized answers required by the inquiries, which resulted in very low similarity scores.

The two models differed in the range of their score distributions. The ADA model (Panel A) produced a compressed distribution of scores mainly ranging from 0.65 to 1, whereas the LARGE model (Panel B) covered the full 0-to-1 range. However, to assess whether the two models evaluated case-level similarity in a consistent way despite this difference in score range, we examined the correlation between the case-level SES scores produced by the two models. The two sets of scores were highly correlated (Pearson r 1⁄4 0.9604, p < 0.001; Spearman q 1⁄4 0.9627, p < 0.001), indicating that the main SES patterns and case-level rankings were highly similar across the two models. This cross-model convergence suggests that the SES results were not dependent on a specific embedding model. Accordingly, to avoid redundancy in the subsequent analyses, we used LARGE as the representative SES measure.

This choice was also supported by its reported stronger performance on standard embedding benchmarks, and by the fact that its scores were distributed across the full 0-to-1 range, which facilitates interpretation and comparison.

The descriptive statistics of semantic embedding similarity scores in Table 8 provide additional insights into the SES results. The standard deviation (0.2260) highlighted significant variability in text similarity, reflecting a wide range of LLM’s performance across cases. This variability corresponded to the observed bimodal distribution. The bottom 25% of cases had low similarity scores below 0.2704, suggesting that drafts and replies differed significantly. The top 25% of cases showed similarity scores above 0.6717, representing cases in which the LLM drafts and sent replies exhibited relatively high semantic similarity. With a median of 0.5331 and the top 25% boundary at 0.6717, approximately one-quarter of all cases were concentrated between scores of 0.53 and 0.67. This suggests that a substantial number of cases exhibited moderate similarity between LLM drafts and sent replies.

The mean of 0.4879 being slightly lower than the median of 0.5331 reflected the influence of cases with considerably low scores.

This score distribution suggests that LLM drafts have not yet reached a level where case handlers can use them without significant modifications. This underscores the high standards of accuracy and domain expertise required in maritime industry communications.

4.3.2. LLM as a judge (LAAJ).

We had the LLM evaluate the similarity between LLM drafts and sent replies across three criteria: overall similarity, language similarity, and interpretative similarity, and also calculated a composite similarity score by averaging these results. The evaluation was conducted on a 1–7 Likert scale, where 1 indicated low similarity and 7 indicated high similarity. Figure 9 shows the LAAJ score histograms for these evaluation criteria. Since the scores represented averages from three independent runs for each case, they appeared as discrete distributions with some variation.

Overall similarity exhibited, in general, a bimodal distribution with high frequencies at scores 1–2 and at score 6. Interpretative similarity showed a generally similar pattern, also having high frequencies at scores 1–2 and at score 6. Language similarity displayed a somewhat different pattern, showing high frequencies at scores 2 and 4. The similar patterns between overall similarity and interpretative similarity suggested that overall similarity largely reflected semantic and interpretative aspects of the text. The relatively moderate levels in language similarity indicated that case handlers frequently modified specific expressions and language, while the bimodal distribution in interpretative similarity suggested that the LLM either properly grasped the core meaning of inquiries or failed to do so.

The composite score distribution also showed a broadly bimodal pattern, with peaks in the lower range (around scores 1–2) and in the upper range (around scores 5–6), and less mass in the moderate range (scores 3–4). Scores at the very top of the scale (around 7) were uncommon, which mirrored the SES results, where scores near the upper bound (close to 1) were likewise rare. Because each composite value was derived by averaging three criterion scores, each itself an average of three independent runs on the 1–7 integer scale, the scores fell on a finer grid of fractional values. These fractional values reflected the averaging of adjacent integer ratings rather than additional modes in the underlying distribution. This concentration of cases in the lower and upper ranges suggested that the LLM’s performance was inconsistent, a pattern similar to that observed in the SES results.

For comparison, a system demonstrating consistently higher similarity would show a markedly different pattern: a unimodal distribution heavily skewed toward higher similarity of scores 6–7, with minimal cases in the lower similarity ranges. The actual bimodal distribution we observed, with substantial mass at both ends of the scale, indicated room for improvement in the context of professional communication where accuracy and precision are crucial. Despite the presence of some highly similar responses, this pattern indicated the system had not yet achieved the consistent, reliable performance required in professional maritime communications.

To ensure the reliability of these similarity assessments, we conducted three independent evaluation runs, and the consistency of these results was statistically confirmed. ANOVA results showing no statistically significant differences between runs, very high Kendall’s W values, and visually consistent score distributions (see Appendix F for detailed information) suggested that the LLM applied consistent evaluation criteria across all runs, supporting the reproducibility and reliability of LAAJ.

The descriptive statistics of the composite scores in Table 8 provided complementary insights into the LAAJ results. The LAAJ evaluation yielded a mean composite score of 3.63 and a median of 3.55 on the 1–7 scale, suggesting a moderate level of similarity between LLM drafts and case handlers’ sent replies. While moderate similarity scores might appear encouraging, taking a conservative interpretation appropriate for safety-critical maritime industry communications, these scores indicated insufficient evidence to claim high textual similarity between LLM drafts and case handlers’ sent replies.

4.3.3. Additional analysis (triangulation): comparing similarity scores between SES and LAAJ.

We sought to establish the reliability of our textual similarity assessments through methodological triangulation using SES and LAAJ. This additional triangulation analysis was to emphasize its role in validating the reliability of the textual similarity assessments, highlighting the complementary perspectives and strong correlation as evidence of robustness. While both methods indicated moderate average similarity levels between LLM drafts and case handlers’ replies (SES mean 1⁄4 0.4879; LAAJ mean 1⁄4 3.63 on 1–7 scale), they offered complementary perspectives on the similarity patterns. The strong correlation between these methods (Pearson r 1⁄4 0.8128, p < 0.001; Spearman q 1⁄4 0.8025, p < 0.001) provided robust evidence that both methods were capturing related aspects of text similarity, despite their different methodological foundations.

The comparison of score distributions (Figure 10) revealed both commonalities and differences in how these two methods characterized similarity patterns. The SES method produced a distinct bimodal distribution, clearly separating cases into higher and lower similarity groups. In contrast, the LAAJ method yielded a more nuanced multi-modal distribution, suggesting finer distinctions across similarity dimensions. This difference reflected the inherent characteristics of each method: SES provided a continuous measure of semantic similarity, while LAAJ offered discrete, multi-faceted evaluation capturing distinct aspects of text comparability.

To validate the reliability of this comparative analysis, we examined the statistical properties of both methods (Table 8). Table 8 presents summary statistics for SES and LAAJ scores on their respective scales: SES on a 0–1 cosine similarity scale and LAAJ on a 1–7 Likert-based composite scale. Because these scales differ, the values are not directly comparable; the table instead describes relative patterns such as distribution and variability. Both methods demonstrated considerable range in their measurements (SES: 0.0089–0.9969; LAAJ: 1–7), with similar patterns of central tendency and spread relative to their respective scales. The consistency in these patterns, combined with the strong correlations between methods, provided strong evidence for the robustness of our similarity analyses.

To further explore the relationship between SES and LAAJ, we attempted to conduct a manual, human expert evaluation on a stratified subset of cases (N 1⁄4 60), sampled across four groups representing extreme combinations of similarity scores between SES and LAAJ (i.e., 15 cases each for low–low, high–high, high–low, low–high). This sampling was supposed to enable an examination of cases where the two measures align (related cases) and diverge (non-related cases). We found that interpreting these cases was highly non-trivial, as similarity assessments depend on nuanced contextual, regulatory, and domain-specific considerations that are not fully captured by either metric alone.

In particular, cases with divergent scores (i.e., high SES but low LAAJ, or vice versa) often required specific human expert judgment to determine whether differences reflected meaningful discrepancies in content or acceptable variation in phrasing and interpretation. Consequently, we did not manage to complete the human expert evaluation of the 60 cases, let alone scaling such analysis to the full dataset (N 1⁄4 2,306). Manual evaluation required careful inspection of inquiries, LLM drafts, and final replies, and was extremely challenging due to the nuanced and highly specialized nature of the content. In many cases, meaningful interpretation depended on deep maritime domain expertise, including familiarity with regulatory requirements and context-specific decision-making practices.

Given that the inquiries span over 650 distinct fields of expertise, a comprehensive evaluation of all cases would require extensive expert involvement across domains. Instead, we interpret SES and LAAJ as complementary measures that capture different aspects of similarity, using them in a triangulation framework rather than as directly comparable or interchangeable metrics.

Note.: Diagonal cells (SES low/LAAJ low; SES high/LAAJ high) represent concordant classifications (1,806 cases, 78.3%). Off-diagonal cells represent discordant classifications (500 cases, 21.7%). Thresholds: SES median 1⁄4 0.53; LAAJ median 1⁄4 3.56.

This triangulation of methods suggested that while SES and LAAJ should not be used interchangeably due to their distinct measurement characteristics, their strong correlation and complementary insights strengthened our overall analyses of text similarity between the LLM drafts and case handlers’ sent replies. The convergence of results across these methods reinforced our conservative interpretation regarding the moderate levels of text similarity observed.

Nevertheless, we note two potential sources of bias in our text similarity pipeline that merit explicit discussion. First, because drafts were generated by ChatGPT-3.5 and evaluated by GPT-4o, our LAAJ scores may be subject to self-enhancement or family bias, whereby a judge model assigns systematically higher ratings to outputs from related models. Spiliopoulou et al. (2025) report this effect specifically for GPT-4o. Our study, however, is based primarily on pairwise preference judgments between candidate LLM outputs; whether the same mechanism extends to similarity scoring against an independently human-authored text, as performed here, remains untested. The choice of GPT-4o as judge was constrained by data sensitivity requirements restricting our evaluation pipeline to Microsoft Azure-hosted OpenAI models.

Second, our attempt to calibrate LAAJ against manual human judgment on a stratified subset of 60 cases did not yield interpretable results, reflecting the same domain-expertise barrier, spanning over 650 maritime fields, that motivates automated evaluation in this setting, rather than indicating that LAAJ scores are unreliable. The first of these, self-enhancement bias, has a known direction: it would tend to inflate rather than deflate measured similarity. The second, the failed human calibration, does not itself have a direction; rather, its absence means we cannot independently verify whether such inflation occurred.

Taken together, we therefore adopt a conservative interpretive stance, treating our reported similarity scores as an upper bound on the true alignment between LLM drafts and sent replies rather than a validated point estimate, reinforcing rather than undermining our central conclusion that substantial human oversight remains necessary. We further note that our triangulation between LAAJ and SES provides partial, though not complete, independent corroboration: while SES embedding models (text-embedding-ada-002, text-embedding-3-large) are also OpenAI products, the self-enhancement/family-bias mechanism described above is specific to generative evaluative judgments, and cosine similarity between fixed vector representations has no clear analog to this stylistic-preference mechanism.

This is further supported by a case-level concordance analysis presented in the next paragraph: 78.3% of cases are classified consistently as low or high similarity by both measures (Table 9), a categorical result independent of, and consistent with, the continuous-scale correlation reported here.

We examined this concordance by classifying all 2,306 cases independently under each measure as ‘low’ or ‘high’ similarity, using each measure’s median as the classification threshold (SES median 1⁄4 0.53; LAAJ median 1⁄4 3.56). Table 9 presents the resulting cross-tabulation: 78.3% of cases (1,806 of 2,306) were concordant across SES and LAAJ, while 21.7% (500 of 2,306) were discordant. This level of case-level agreement between two methodologically independent measures, a deterministic embedding computation on one hand and a generative judgment on the other, indicates that the bimodal distribution reflects a property of the underlying cases rather than an artifact of either measurement approach individually, and provides a categorical complement to the continuous-scale correlation reported above (Pearson r 1⁄4 0.81, Spearman q 1⁄4 0.80).

Together with the argument above regarding self-enhancement/family bias, this concordance is difficult to reconcile with the two measures’ agreement being driven by a shared measurement artifact.

5. Discussion.

Our real-world case study explores how case handlers, human experts in the maritime industry, perceive the usefulness of LLM-generated draft replies (i.e., LLM drafts) in responding to stakeholder inquiries. Our research questions were:

1. How do case handlers perceive the usefulness of LLM drafts, and how can these drafts support their workflows?

2. How similar are the texts generated by the LLM to the final replies sent by case handlers, and what does this similarity imply about the drafts’ potential utility?

In this section, we discuss the key findings derived from a preliminary qualitative study, a survey study, and a text similarity analysis, while also identifying potential directions for future work.

It is important to note that the LLM draft function had been operational for only a few months at the time of our study. While this might be considered premature especially for assessing long-term impacts, the goal of our research was to establish a baseline. This baseline serves not only as a reference point to evaluate future advancements in technology and user perceptions but also as a foundation for tailoring improvements to the LLM drafts.

By analyzing our findings, we aim to start addressing the broader challenge of bridging the gap between user expectations and the current LLM functionality in general, and for the DATE system specifically. The findings can inform targeted improvements to the DATE system, helping it meet the needs of case handlers more effectively. This approach emphasizes using the findings proactively, not just as benchmarks but as actionable guidance for developing the technology further, enhancing its relevance, and addressing areas where user expectations remain unmet.

5.1. Case handlers’ perceptions of the LLM drafts.

Our findings on perceptions of the LLM drafts can be grouped into three interrelated themes: perceived accuracy and relevance, attitudes toward adoption, and perceived impact on stakeholder trust. It is important to note that the survey only indicates how many case handlers had tried the LLM, not how frequently they used it, as the system was still in its early experimental phase at the time of the study. Moreover, the survey does not capture the perspectives of non-respondents or respondents who had not yet tested the drafts. Finally, prior research has shown that humans often hold biases against LLM-generated texts, and such biases should be considered when interpreting our findings.

5.1.1. Perceived accuracy and relevance.

In general, the preliminary study and survey findings suggest that case handlers who tested the experimental LLM draft functionality feel that it requires significant improvement in terms of accuracy and relevance. Survey results indicate that case handlers still rely heavily on their own expertise or that of their colleagues, particularly when dealing with uncertainties (Figure 6). This reliance is likely influenced by the specialized nature of the maritime industry and the high level of expertise required for case handling. The complexity of stakeholder inquiries in this domain often necessitates niche expertise or a combination of specialized knowledge to provide accurate responses, including forms of knowledge developed through experience and collaboration.

For example, as reported in the preliminary findings on Levels of Complexity of Inquiries (Section 4.1), participants indicated that LLM drafts were often more useful for relatively routine inquiries involving clear procedures or readily available regulatory information. In such situations, the draft could provide a reasonable starting point that required only verification and limited adaptation before a reply was sent. In contrast, participants also described situations where an LLM draft suggested a generic response based on similar historical cases, while the final reply required additional contextual considerations derived from consultation with colleagues, previous experience, knowledge of the stakeholder, or vessel-specific operational circumstances.

These findings illustrate how LLM drafts may perform adequately when the required knowledge is explicit and well represented in historical cases, but become less useful when responses depend on contextual and experience-based knowledge.

This raises important questions about how an LLM can be effectively utilized to handle inquiries of such complexity. In the survey findings, case handlers provided several suggestions for improving the LLM drafts. These improvements include broadening the system’s data sources beyond historical cases to incorporate additional relevant materials – such as attachments – and integrating user feedback at the case level. For example, case handlers can mark text as correct or incorrect, enabling the system to refine its future LLM drafts. This feedback can then be used to improve the system’s ability to retrieve and present relevant information (G. Gao et al., 2024; Shankar et al., 2024), ultimately aiming for its enhanced accuracy and relevance across a broader range of inquiries.

Despite these challenges, the survey findings also highlight several positive perceptions of the LLM drafts. Case handlers rated the language quality of the drafts as high and noted that the drafts helped them work more efficiently and maintain consistency in their replies to the stakeholders. These findings indicate that, while accuracy and relevance remain areas for improvement, users already perceive other aspects of the LLM drafts, such as linguistic quality and time-saving potential, as valuable.

5.1.2. Skeptical but curious.

Our survey findings are in line with our preliminary findings, especially in identifying the attitudes of case handlers as ‘skeptical but curious’. This skepticism towards the LLM drafts is an important and healthy safeguard against overreliance while encouraging the advancement of the LLM’s utilization to create more accurate and relevant LLM drafts. Such a stance can be understood as a form of trust calibration in human–AI interaction, where users balance caution with engagement to avoid both overreliance and underutilization. At the same time, the observed curiosity is consistent with early-stage technology adoption dynamics, where users explore new systems despite known limitations to assess their usefulness in practice.

This openness to innovation is likely related to the fact that the LLM drafts are developed in-house and the case handlers are involved and encouraged to explore the new function critically.

Importantly, senior managers are supportive and engaged in the development of the LLM drafts. Such support is crucial; managerial endorsement plays a significant indirect role by building trust among employees, which in turn increases use of the AI-enabled systems (Korzynski et al., 2024). It is worth mentioning that the deployment of the LLM drafts follows an earlier implementation of two natural language processing (NLP) solutions based on Term Frequency-Inverse Document Frequency (TF-IDF) on the DATE systems, which had been well-received and perceived as effective. In addition, the LLM drafts’ deployment has been communicated as a collaborative and exploratory effort, emphasizing that it is an initial step rather than a finalized product.

The LLM drafts have not been overstated as a revolutionary AI-enabled system; rather, both the design and communication strategies have been deliberately cautious, acknowledging potential limitations and framing the deployment as a learning process. Such strategies are followed by the iterative design process, incorporating improvements based on real-world use and feedback. This strategies and framing may influence the creation of a more favorable environment for deploying the new system (i.e., the LLM drafts).

Our case study also highlights a broader organizational motivation for piloting the LLM functions: the recognition that AI-enabled systems are becoming ubiquitous and cannot be ignored. By actively engaging with these technologies, organizations can develop internal expertise, better understand the associated risks, and identify opportunities to harness AI-enabled systems effectively (Sarri & Sj€olund, 2024). This proactive approach positions organizations to embrace AI-enabled systems responsibly while mitigating potential drawbacks.

Having said that, the fact that only 74 of 198 case handlers responded to the survey and only 51 case handlers had tried the LLM drafts at the time of the survey can mean that those case handlers are the early adopters (a.k.a. champions). Early adopters are likely to have more optimistic attitudes towards the technology. This is a limitation in our survey and the findings should be read in this context. If the participants are indeed skewed towards more optimistic perceptions, the actual reception of the technology among the broader case handler base may be less favorable than the data suggests. We will need further investigation into barriers or hesitations among case handlers who did not respond and the non-user case handlers. Future work can also explore factors that influenced the decision whether to use the LLM draft.

This can provide valuable feedback on the technology itself and help us understand the socio-technical factors surrounding the perceived usefulness of the LLM drafts.

5.1.3. Stakeholders’ trust.

The only statistically significant difference we found in the responses is between the case handlers who believe that the new function of LLM drafts will decrease the stakeholders’ trust in the organization (the ‘decrease’ group) and those who believe it will increase (the ‘increase’ group) or have no impact (the ‘no impact’ group). The ‘decrease’ group having the lowest mean scores shows that these case handlers perceive the LLM drafts as less aligned with their expectations or less capable of accurately addressing stakeholder inquiries. This perception likely reflects concerns about the quality or appropriateness of the LLM drafts, which may lead to fear of misunderstandings or dissatisfaction among stakeholders ultimately impacting trust in the organization.

In contrast, the ‘increase’ and ‘no impact’ groups likely see the LLM drafts as being sufficient to support or maintain stakeholder trust. This difference in perception suggests that some case handlers have more confidence in the LLM than others, highlighting the need to address concerns about the quality and reliability of the LLM drafts. Addressing case handlers’ specific concerns can contribute to building trust within the organization and with stakeholders. Accordingly, future work can explore underlying issues in the different perceptions among case handlers to tailor improvement efforts.

5.2. Similarity between LLM drafts and sent replies.

5.2.1. Variability in text similarity.

Text similarity analyses between LLM drafts and case handlers’ sent replies provide valuable insights into the LLM’s performance. The similarity scores exhibit a bimodal distribution suggesting variability of similarity levels. This distribution highlights two primary patterns in case handlers’ use of LLM drafts: cases where the drafts are largely retained and cases where extensive revisions are required. Specifically, the group of highly similar texts suggests that the LLM performs well, enabling case handlers to use the LLM drafts without making substantial revisions. Conversely, the group with lower similarity may indicate instances where the LLM lacks sufficient contextual understanding or information, requiring case handlers to significantly adapt the LLM drafts to meet the inquiry’s specific needs.

However, the survey findings appear more weighted towards the lower similarity group and much less towards the higher similarity group. The survey suggests that most LLM drafts require significant changes (Table 6). The respondents who only modified the LLM drafts slightly or none at all, may correspond to the high-similarity group. The widespread perception of inaccuracy may help explain why extensive modifications are frequently needed, as the survey findings reveal that the majority of case handlers (86.28%) disagree or strongly disagree with the statement that the LLM drafts are accurate (Table 6).

These findings suggest that the practical utility of LLM drafts is case-dependent: they can align well with final replies in some instances but require substantial rewriting in others. This is particularly relevant in the complex and safety-critical maritime industry, which demands high levels of accuracy and precision in the communication of critical information. Therefore, we find that these LLM drafts are not yet mature enough for use in high-risk industries without human experts.

The variability in similarity between the LLM drafts and the case handlers’ sent replies may be influenced by several factors. One key limitation is that the LLM system cannot process attachments, which may contain critical information necessary to address the inquiries. Additionally, the LLM lacks access to vital resources available to case handlers such as internal reports, relational knowledge, historical context, and access to other experts. Outdated past cases – due to updated or new regulations, for example – also require case handlers to amend the LLM drafts to ensure compliance and accuracy. In these cases, case handlers leverage their more comprehensive understanding of the inquiries to refine the drafts. This directly impacts the text similarity results.

In addition, variations in writing and communication styles between the LLM and individual case handlers may further contribute to the observed variability.

Among these factors, inquiry complexity is of particular interest, though we do not have a methodologically sound basis, in this real-world operational setting, to quantitatively test whether it drives the observed similarity pattern.

The candidate proxies available to us are each confounded in identifiable ways: duration of LLM draft usage is conflated with case handlers’ workload prioritization under service-level constraints (Section 2.1) rather than reflecting time spent specifically on a given case; inquiry category does not capture within-category complexity variance, as nominally similar cases can range from routine to highly specialized (Table 1); and retrieval confidence, the RAG system’s own similarity score between an inquiry and its retrieved historical precedents (Section 2.2), reflects the availability of a comparable precedent rather than the complexity of the inquiry itself, since a rare but straightforward inquiry can show low retrieval confidence for lack of precedent while a recurring complex inquiry can show high retrieval confidence despite genuine difficulty.

We therefore do not present a quantitative test of this hypothesis. We do, however, note that it is grounded in convergent qualitative evidence from two independent sources: preliminary interview and observation data (Section 4.1) and anonymous survey free-text responses (Section 4.2), collected from different participants and at different points in our data collection. Both independently point to perceived inquiry complexity as a factor influencing LLM draft usefulness. We additionally find that the bimodal pattern itself is robust across two independent computational similarity measures (Section 4.3.3, Table 9), with 78.3% of cases classified concordantly by SES and LAAJ, supporting that it reflects a genuine property of the case population rather than a measurement artifact.

We flag the link between inquiry complexity and measured similarity as a hypothesis for future work, contingent on development of a validated complexity-rating instrument for this domain.

5.2.2. Pathways for improvement.

The wide variability in similarity scores highlights several areas for potential improvement in the system. For example, the tendency of the LLM to produce generic or insufficient responses, as also reported in the preliminary and survey findings, can be mitigated by improving the model’s contextual understanding. The inability of the current LLM system to handle attachments or multi-turn conversations that limits its applicability to more nuanced inquiries may be addressed through improved data integration and dialogue management. The improved data integration may also address the limitation that the LLM does not have access to the same resources as case handlers. In addition, removing outdated cases may improve the accuracy and relevance of the LLM drafts.

Looking ahead, several technical advancements may improve the consistency of the LLM draft responses to inquiries with different levels of complexity. For example, computation enhanced generation could integrate external computational tools and models with LLMs, enhancing their accuracy in tasks that require precise calculations or domain expertise. The expanded context window capabilities seen in newer models, including those that can handle up to 1–2 million tokens, could potentially reduce reliance on retrieval for handling complex inquiries. Multimodal capabilities would enable processing of various data types often present in maritime documentation, enhancing the system’s ability to comprehend and respond to diverse information formats.

The development of smaller, more efficient models that maintain performance – through techniques like distillation and quantization – could make the system more practical and cost-effective. Dynamic prompting that creates prompts in real-time based on user behavior and context could improve the system’s adaptability and responsiveness. Finally, better integration with operational workflows through LLMOps could enhance the system’s reliability and maintainability in production environments.

5.3. Beyond generating accurate and relevant LLM drafts.

The preliminary study and survey findings reveal that case handlers continue to place significant emphasis on the accuracy and relevance of LLM drafts when handling inquiries that are often highly specific and complex. Interestingly, despite the relatively negative perception of the LLM drafts’ accuracy and relevance, the case handlers express a willingness to recommend the drafts to their colleagues, engage in frequent discussions about them, and intend to continue using them over the next 12 months (Table 6). These results suggest that the value of LLM drafts is likely to extend beyond merely generating accurate and relevant replies. For example, the preliminary findings indicate that case handlers use the drafts as a tool to generate ideas and explore potentially necessary information for crafting responses.

This example highlights the potential of LLM drafts to serve as creative and exploratory aids in workflows. Below, we provide some known strengths and limitations of both the LLM drafts and case handlers, with the purpose of identifying potential use cases where the LLM drafts can provide value beyond generating accurate and relevant responses (Table 10).

By exploring these examples, we can find new avenues for the application of LLM drafts and expand their utility beyond generating accurate replies. Consistent with our findings, LLM drafts may support brainstorming, provide quick access to relevant precedents, and assist case handlers in reviewing and interpreting lengthy or complex inquiries. By processing and summarizing information from past cases, LLM drafts can help reduce effort spent on preliminary tasks and allow case handlers to focus on decision-making and tailoring responses. They may also provide templates or examples that support less experienced case handlers.

In addition to improving workflows, the LLM drafts enhance the overall quality of written communication. Survey findings suggest that case handlers’ perception of the language quality of LLM drafts is high, indicating their potential to set benchmarks for clarity and professionalism. By creating templates for common cases, LLM drafts can improve efficiency while also standardizing communication. This standardization ensures consistent, high-quality responses aligned with organizational standards, which can enhance stakeholder satisfaction and the organization’s reputation.

6. Lessons learned, identified gaps and directions for future work.

We have identified limitations in our case study, as well as lessons learned, gaps, and directions for future work. Although we validated the text similarity results within and across the LAAJ and SES, we acknowledge that a fully accurate evaluation remains difficult. We did attempt a manual/human evaluation of a very small subset of cases from the text similarity analysis to calibrate its results, including reviewing these cases together with a domain expert (i.e., case handler), focusing on the inquiries, LLM drafts, and sent replies. However, this proved extremely challenging due to the nuanced content. Even with domain expertise, it was difficult to consistently interpret whether differences between LLM drafts and sent replies reflected meaningful discrepancies or acceptable variation in phrasing and professional judgment.

Relatedly, our LAAJ scores may be subject to self-enhancement bias given the shared model lineage between the drafting and judging LLMs (Section 4.3.3), a risk our failed human-calibration attempt could not independently rule out; future work should pursue calibration with a judge model outside this lineage, alongside a validated complexity- or quality-rating protocol. In a related vein, our restriction to single-draft cases (Section 3.3) reflects a similar limitation: because case handlers could combine text across multiple regenerated draft versions, no principled way exists to isolate a single ‘first draft’ per case among multi-draft cases, precluding a sensitivity analysis comparing first-draft similarity across the excluded subset. In addition, scaling up such manual evaluation would prove even more challenging considering the volume of our dataset (N 1⁄4 2,306).

We concluded that completing a comprehensive evaluation for all cases was nearly impossible without extensive domain-specific expertise. This is because of the highly specialized nature of the inquiries and LLM drafts that necessitates specific domain expertise. The inquiries span over 650 maritime fields of expertise, reflecting the complexity and breadth of knowledge required to understand them. Although our dataset includes queries across these 650 categories, we did not analyze variance between categories, as this was beyond the scope of the study. The mention of these categories is intended solely to illustrate the diversity and complexity of the maritime domain. Our primary focus was instead on establishing baseline perceptions of the LLM drafts and examining overall similarity patterns.

Relatedly, a further limitation is that no validated instrument exists for measuring inquiry complexity in this operational setting; while our concordance analysis (Section 4.3.3) confirms that the bimodal similarity pattern is a robust property of the case population rather than a measurement artifact, it remains diagnostic rather than explanatory, and testing whether complexity specifically drives this pattern (Section 5.2.1) awaits development of such an instrument in future work. Accordingly, future work should explore more robust, mixed-methods evaluation approaches that incorporate domain expertise, case handlers’ reasoning, and qualitative insights to assess LLM-generated content more comprehensively. In relation to this, our focus on the maritime industry may limit the generalizability of the findings to other domains.

Future research is needed to explore LLM utility across different industries and use cases beyond maritime inquiries. For example, conducting comparative studies in other highly specialized domains to understand how LLMs perform across diverse contexts and tasks to enable cross-domain learning.

While similarity metrics can provide a useful starting point for evaluating the alignment between LLM drafts and sent replies, they may not be sufficient to fully capture the LLM’s utility. For example, lower similarity texts could reflect cases where the LLM drafts served as a starting point for generating nuanced replies, thus demonstrating some value in utility even though the similarity scores were lower. Additionally, higher text similarity scores may not always correspond to high utility.

Consider a scenario where a stakeholder inquires, ‘What documents do I need to submit for a cargo claim?’ The LLM draft responds, ‘What type of cargo is involved, which flag, is there damage or loss, which location, and under what circumstances?’ While this response aligns with the general process of gathering additional information and the case handler sends it with minimal edits, resulting in a higher similarity score, the inquiry remains incomplete. The response does not directly address the stakeholder’s request for a document list and instead prompts further clarification, necessitating follow-up interactions to resolve the inquiry.

Another limitation with text similarity metrics is ambiguity in dimensions and extent of similarity (e.g., language, interpretability, or contextual similarity). LLM drafts fundamentally operate as language tools, and the nuances of language can lead to complexities that computational text similarity metrics often fail to capture. A phrase can have the same underlying meaning but be expressed in entirely different ways. Conversely, phrases that appear linguistically similar may carry significantly different meanings depending on context. As an example, the LLM draft can respond to an inquiry of ‘What is the status of my shipment?’ with ‘Your shipment is currently in transit and will arrive on Thursday’, which the case handler might amend to ‘The shipment is on its way and should reach you by Thursday’.

While the LLM drafts and sent reply convey similar information, they are phrased differently, potentially resulting in a lower similarity score despite their equivalence in meaning and utility. Similarly, there might be ambiguity in interpretation: an inquiry that says, ‘Do I need to file additional forms?’ can be responded by the LLM draft as ‘No additional forms are needed,’ in which is amended by a case handler into ‘No, you do not need to submit extra forms at this stage’. While the LLM draft and sent reply are highly similar in language, the case handler’s reply adds the clarification ‘at this stage’, which acknowledges that further steps might require additional forms, an important distinction from the LLM draft.

Future work should explore complementary metrics, such as uncertainty measures or confidence levels, which can provide a more comprehensive understanding of the LLM system’s effectiveness. Importantly, future work should include a thorough investigation to explore the specific degrees and aspects of textual similarity represented by both higher and lower similarity scores, ideally through human-as-a-judge, to calibrate the computational evaluation metrics in real-world settings. This deeper understanding would enable more precise and meaningful evaluations of the LLM performance. Additionally, analyzing whether certain patterns emerge in the generated LLM drafts and examining the factors that influence varying levels of text similarity with domain experts could provide valuable insights for system improvements.

Our case study lacks detailed investigation into case handlers’ needs and interaction patterns with the LLM system. Incorporating user feedback at the inquiry and case level could provide valuable insights into where the system excels or struggles (G. Gao et al., 2024; Shankar et al., 2024). For example, analyzing the types of questions posed and how users engage with the LLM could help identify scenarios where the system performs effectively – such as responding to straightforward, simple inquiries – and areas where it is less effective, particularly with more complex cases. Our findings imply that the LLM is less suited to handling these complex inquiries, but unpacking its performance in these contexts requires a more targeted approach.

Collecting and mapping user feedback at the case level, as suggested by case handlers in the survey findings, could highlight the LLM’s strengths and weaknesses and also identify areas where human intervention is most and least critical (G. Gao et al., 2024). This would support more effective human-AI collaboration and guide future system improvements (Tita A Bach, Kristiansen, et al., 2024). Hence, future work should focus on incorporating detailed user feedback at the inquiry and case levels to better understand how case handlers interact with the LLM system and where it succeeds or struggles (G. Gao et al., 2024). This may involve categorizing cases by complexity to map the LLM’s performance and identify patterns in scenarios where it performs effectively or falls short.

In a broader sense, future work should also focus on shifting the emphasis from designing for automation to designing for enhancing human-AI collaboration. This involves developing frameworks and methods that embed collaboration-first principles into algorithms and interfaces that align AI-enabled systems with real-world needs and fosters systems that augment humans rather than displacing them.

Our case study does not directly or in detail investigate potential implications of the integration of the LLM drafts into case handlers’ workflows for expertise development, particularly concerning knowledge sharing and competence building among junior or less experienced and senior or more experienced case handlers. The survey findings indicate that most case handlers rely on their own knowledge or consult colleagues also when in doubt, suggesting that the LLM drafts currently augment rather than replace traditional knowledge-sharing practices. However, a minority of case handlers perceived the LLM drafts to have influenced how they consult colleagues. This signals potential shifts in collaboration dynamics as LLM technology evolves. It is worth noting here that the LLM draft function had been operational for only a few months at the time of our study.

As the LLM technology advances, future work should look more detail into how the LLM drafts influence collaboration between and among case handlers and other experts in the organization, and the implications to knowledge-sharing and competence building.

7. Conclusion.

LLMs represent a significant opportunity to transform workflows in specialized domains. Our study shows that while LLMs can streamline processes and support case handlers as human experts, they are not yet mature enough for independent use in safety-critical applications, and human oversight remains essential. LLMs can play a valuable role as an augmentation rather than a replacement of human expertise, and system improvements, such as broader data integration and iterative user feedback, can enhance performance. However, incorporating tacit knowledge gained through experience remains challenging, if not highly improbable, in real-world applications.

Our findings demonstrate that case handlers’ expertise and experience will remain to be critical for ensuring replies to stakeholder inquiries are relevant, accurate, and aligned with the regulatory and operational context of each vessel. In the maritime industry, stakeholder inquiries often require tailored, highly specialized knowledge. While LLM drafts were helpful in many cases, they still exhibited limitations, highlighting the indispensable role of human expertise for precision, contextual understanding, and relational knowledge.

We conclude that LLMs should augment, not replace, human decision-making. Final responsibility must remain with case handlers as the human experts. By leveraging LLMs thoughtfully and fostering human-AI collaboration, organizations can enhance efficiency while maintaining high standards of quality, accuracy, and relevance tailored to each case.

Note

1. The text-embedding-ada-002 model, released in December 2022, has been widely used for various natural.

language processing tasks. It generates 1536-dimensional embeddings and has demonstrated robust performance across a range of applications. The more recent text-embedding-3-large model, introduced in January 2024, features up to 3072 dimensions, enhancing performance compared to its predecessors (the linked source).

Acknowledgments.

Authors would like to thank Dr. Martin Høy for providing the data for text similarity analyses, Dr. Caryl de la Serna for conducting the R analyses and supporting the interpretation of the results, the DATE system’s team and management, and the participants. Authors would also like to thank the reviewers for their constructive feedback, which helped improve the quality, clarity, and readability of the manuscript. This manuscript involved the use of ChatGPT 4o, Claude 3.5 Sonnet, and Microsoft 365 Copilot (GPT-5) to assist in generating ideas, improving language clarity, and translating to English. All content was rigorously reviewed and refined by the authors, with additional proofreading performed by a professional proofreader to ensure readability, accuracy, and quality. The pre-print of this manuscript is available on arXiv (Tita Alissa Bach, Babic, et al., 2024).

TAB, AB, NP conceptualized the study, designed the research methodology, conducted analyses, and led the writing of the manuscript. TSp and RU conducted the preliminary method, contributed to the overall data analysis, and critically reviewed the manuscript. HSM as solution designer and TSk as process responsible provided technical expertise on DATE systems and case handlers, contributed to the interpretation of results, and reviewed the manuscript critically. All authors read and approved the final version of the manuscript.

Authors contributions

CRediT: Tita Alissa Bach: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Resources, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing; Aleksandar Babic: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing; Narae Park: Data curation, Formal analysis, Investigation, Methodology, Resources, Software, Validation, Visualization, Writing – original draft, Writing – review & editing; Tor Sporsem: Data curation, Formal analysis, Investigation, Methodology, Resources, Validation, Writing – original draft, Writing – review & editing; Rasmus Ulfsnes: Data curation, Formal analysis, Investigation, Methodology, Resources, Validation, Writing – original draft, Writing – review & editing; Henrik Smith-Meyer: Conceptualization, Data curation, Investigation, Methodology, Resources, Software, Validation, Writing – review & editing; Torkel Skeie: Data curation, Investigation, Resources, Validation, Writing – review & editing.

Disclosure statement

No potential conflict of interest was reported by the author(s).

Funding

This project is partially funded by the Norwegian Research Council (grant number: 309631).

About the authors

Tita Alissa Bach received a PhD in behavioral and social sciences from the University of Groningen, the Netherlands, in 2012. She received a diploma in Top Tech Executive Education from Haas School of Business, University of California, Berkeley, CA, USA, in 2017. In January 2023, she joined Digital Transformation Team, Digital Assurance, Group Research and Development, DNV, Høvik, Norway, as a Principal Researcher. From 2011 to 2022, she was a Principal Researcher in the Healthcare program at the same organization. Her research focuses on several areas including human behaviors with technology in various environments, safety culture and cybersecurity culture, trust in technology and AI, AI deployment and adoption, human-AI interaction, and responsible AI.

Aleksandar Babic received a MS degree in electrical engineering and a MS degree in computer science from the University of Belgrade, Serbia, in 1999 and 2012, respectively, and a PhD degree in biomedical engineering from the University of Oslo, Norway, in 2019. He is currently a Principal Researcher with the Healthcare Program, Group Research and Development, DNV, Høvik, Norway. He has over two decades of industrial research and development experience and has worked on various projects, such as machine vision for 3D cameras and image fusion and deep learning for the applications of cardiac ultrasound. His research interests include the challenges associated with the implementation of AI in healthcare, including developmental and regulatory aspects, trustworthy and explainable AI, data quality and access, privacy, and algorithm development and robustness.

Narae Park is a researcher with an interdisciplinary academic background, holding MSc degrees in Language Technology from the University of Oslo and Computer Science from UiT The Arctic University of Norway, and an MA in Aesthetics from Seoul National University. She previously worked in the Digital Transformation Team, Digital Assurance, Group Research and Development at DNV, Høvik, Norway. Her research interests include language models, generative AI, extended reality, and the societal impact of these technologies.

Tor Sporsem is pursuing a PhD in Computer Science at the Norwegian University of Science and Technology (NTNU), Trondheim, Norway. Since 2019, he has been employed as a Research Scientist at SINTEF in Trondheim. He specializes in sociological research methods, conducting qualitative studies through interviews and observations. His research focuses on the interaction between software developers and users, particularly the feedback loop during development, requirements engineering, and the creation of software incorporating AI components.

Rasmus Ulfsnes is a research scientist at SINTEF Digital and a PhD candidate at the Norwegian University of Science and Technology (NTNU). His research focuses on the intersection of organizations, knowledge work, and artificial intelligence, with a particular emphasis on how organizations implement and adopt AI for experts and knowledge workers. Rasmus aims to understand the more systemic implications of the use of AI in organizations, through both qualitative methods, and digital data. He has previous experience from industry as both a Security manager and an Enterprise Architect.

Henrik Smith-Meyer received an MSc in Computer Science and Knowledge Systems from Stanford University in 1992 and an MSc from NTNU in 1989, building on a BSc from the University of Utah. In 2000, Henrik completed the Program of Strategic Leadership at the Norwegian Business Institute. With extensive AI experience, Henrik has implemented four Industrial AI production systems (in Hydro, Telenor, and Statnett) and contributed to the first Norwegian commercial AI company, Computas, managing areas such as Methods and Tools and multiple project deliveries. Worked on a tool framework recognized in Gartner’s Magic Quadrant for AI-supported BPM and Enterprise Architecture, contributing patentable innovations.

Henrik contributed to three AI research projects for the European Space Agency (ESTEC), and several for NRC and in EU Programs, advancing knowledge activation, interoperability, model-based design, HCI, and perceived value. Key contributor to team awarded DNDs prize of Business Intelligence 2018 for ML dev platform and first NLP solutions in DNV and DATE 2017. Solution Architect/Designer for DATE since 2013, in effect acting as production tool manager.

Torkel Skeie holds a bachelor’s degree with Honours in Offshore and Mechanical Engineering from Heriot-Watt University, Scotland (1996). Since joining DNV Maritime in 1998, Torkel has worked extensively across various roles, including drawing approvals, quality assurance for survey and reporting data, and as a surveyor for new buildings, materials and components certification, and ships in operation. Since 2008, Torkel has been working as a technical expert, advising surveyors and customers on ship specific questions as well as being lead trainer for a course on ship classification. Since 2013, Torkel has worked as a team with Henrik to further develop the DATE system, focusing on user experience for both case handlers and stakeholders. Their work emphasizes intuitive design and agile development to deliver practical solutions.

Data availability statement

Raw or minimally processed data are not available due to proprietary, ethical, legal, and confidentiality restrictions, and due to terms of participant consent. Specifically, the following constraints apply:

Proprietary and security data: The text similarity data used for analysis includes sensitive, proprietary informa tion related to DNV customer inquiries and internal company data, which cannot be released for legal, privacy and security reasons.

Ethical, confidentiality, and privacy concerns: The survey responses and qualitative data (interview transcripts  and observation notes) cannot be shared publicly. Even after anonymization, the unique nature of the data, combined with potential cross-referencing (‘slicing’) of responses, carries a significant risk of re-identifying participants, thereby compromising the guaranteed anonymity and confidentiality of the DNV employees.

Download transcript ↗