CHORUS: Designing Human-AI Multi-Agent Collaboration for Professional Translators
1 More Paper · Full Reading

About this paper
A full audio edition of this paper.
Authors: Not supplied
Published in: arXiv
Publication date: 2026-02-22
Read the paper: https://doi.org/10.48550/arXiv.2602.19016
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “CHORUS: Designing Human-AI Multi-Agent Collaboration for Professional Translators,” by the authors. Published in arXiv on February 22, 2026.
Abstract.
Despite the widespread use of automatic AI translation systems in daily language tasks, their limitations become apparent in profes-sional translation contexts, where human expertise remains crucial. Professionals rarely rely on these systems in their practice due to a lack of detailed support for the translation process, matching professional styles, and accountability for the final outcome. To
∗These authors contributed equally to this work.
CCS Concepts.
• Human-centered computing → Interaction design; • Com-puting methodologies → Multi-agent systems.
ACM Reference Format.
George Xi Wang, Jiaqian Hu, Guande Wu, and Jing Qian. 2026. CHO-RUS: Designing Human–AI Multi-Agent Collaboration for Professional Translators. In Proceedings of INTERNATIONAL CONFERENCE ON MUL-TIMODAL INTERACTION (ICMI ’26). ACM, New York, NY, USA, 10 pages. the linked source 1 Introduction
Professional translation is a crucial part of high-stakes scenarios, such as medical, legal, and business communications, where errors are much less tolerated, and quality and style are more prominent than in daily translations. Large Language Models (LLMs) have become increasingly popular, but they often lack accurate mean-ing preservation, cultural nuance, domain specificity, and personal style for professionals. Moreover, professional translation needs to meet specific requirements imposed by clients, and such requirements are rarely available in public domains for LLMs to learn. Translators often need to discuss and resolve nu-anced details with clients of different cultural backgrounds, as they are solely responsible for the final outcomes. This responsibility is particularly critical in high-stakes domains where human assurance and accountability are still essential.
Plus, most existing systems use a single LLM interface, which often blends accuracy, terminol-ogy, and style into a one-shot translation. This makes it difficult for translators to revise any of these dimensions individually, resulting in additional time during post-editing. Consequently, existing sys-tems fail to provide the in-progress granularity needed to support professional translation, and guidance in the literature on how to build such systems is largely missing. As a result, we ask: how can we harness the power of LLMs to design an efficient tool to support professional translation work?
Through a formative evaluation with 6 professional translators, we gather design insights around how to better scaffold professional work by incorporating Multidimensional Quality Metrics (MQM), a framework widely used to describe translation quality across dimensions such as accuracy, terminology, fluency, and style. Rather than treating MQM as an evaluation taxonomy, we use it to structure support during revision and enable translators to inspect, adjust, and review each dimension.
We present CHORUS, a mixed-initiative translation system that adopts an MQM-inspired multi-agent design. We develop a live effort algorithm that uses multi-modal input and editing history to adjust agents’ focus during the working process. A Live Style Guide summarizes revision patterns and provides personalized feedback for idiosyncratic traits. This design goal raises three research ques-tions: RQ1: whether CHORUS reduces translators’ editing effort and perceived workload; RQ2: whether it improves final translation quality; and RQ3: how CHORUS supports professional translators in revision, reflection, and preference formation.
We evaluated CHORUS in a within-subject study with 30 licensed English–Chinese translators using WMT24 translation tasks. Compared with a chat-interface LLM baseline (GPT-5.3 web), CHO-RUS reduced completion time by 33.8%, significantly lowered cog-nitive workload, and improved translation quality. Participants also found that CHORUS made error inspection clear, which was useful for self-improvement and assessment. This work contributes: formative results on why current LLM translation tools fall short for professionals and insights for improvement; CHORUS, an MQM-inspired multi-agent system that dynamically translates and adapts through multi-modal behavioral signals while scaffolding the translation process and style; and empirical evidence that a multi-agent translation system reduces effort, improves translation quality and speed, and supports more accountable revision.
The CHORUS system, including all agent prompts, is publicly available at the linked source.
2 Related Work 2.1 LLM-Based Translation and Post-Editing
LLM-based translation work studies prompting, adaptation, and post-editing strategies, including zero-shot prompting, few-shot learning, fine-tuning, prompt design, demonstration quality, and example selection. These methods improve bench-marks and affect document-level behavior, but translation quality remains context-sensitive, and perturbing demonstrations can sub-stantially degrade output. Although LLMs can outperform supervised baselines in some directions, commercial accountabil-ity remains a concern, motivating WMT evaluations with professional translators and human evaluation beyond reference-overlap metrics.
LLM-based post-editing can improve MT scores and perceived trustworthiness, and MQM-derived anno-tations can improve automatic quality metrics such as TER (Trans-lation Edit Rate, number of edits needed to match a reference translation), BLEU (n-gram overlap with reference transla-tions), and COMET (a neural metric trained on human quality judgments). Yet hallucinated edits threaten high-stakes deployment, and it remains unclear how annota-tions should guide concrete revision decisions. This motivates human-in-the-loop workflows where translators retain oversight over quality and accountability. Iterative and multi-agent LLM work also shows both promise and limits: self-refinement and iterative translation refinement can improve outputs, fluency, and naturalness, but can complicate metric interpretation and amplify model self-bias unless external feedback is introduced.
Multi-agent systems use specialized roles for translation production or evaluation, and interactive MT emphasizes trans-lator control; however, prior systems often focus on agent specialization, agent-to-agent refinement, or traditional interactive MT rather than close human coordination during multi-dimensional revision. CHORUS builds on this space by combining LLM revision, specialized agent support, and translator oversight.
2.2 Quality Evaluation in Translation
Translation quality is assessed through automatic metrics and hu-man evaluation, but evaluation can be misleading without explicit error analysis, especially for strong MT systems where differences are subtle. Prior work therefore argues for struc-tured error types and severities. Automatic metrics in-clude BLEU, which depends on reference similarity; TER, which estimates revision effort; and COMET, trained on human judgments such as Direct Assessment, HTER/TER-style judgments, and MQM annotations. Human eval-uation protocols and quality frameworks include Direct Assess-ment, continuous scoring, the Dynamic Quality Framework, SAE J2450, the LISA QA Model, and Multidimensional Quality Metrics (MQM).
Following prior work, we use MQM because it provides a fine-grained taxonomy of translation errors and is widely used in professional, research, operational quality-control, and agent-based evaluation settings. In CHORUS, MQM defines the quality concerns available in the interface and the seven dimensions summarized in Table 2.
2.3 Adaptive Translation Systems
Adaptive translation systems treat translation as an interactive pro-cess: the system proposes translations, humans correct or rate them, and the system updates to reduce future effort. Mixed-initiative and interactive MT have long offered alternatives to pure post-editing through target-text-mediated interaction, completions compatible with translator input, links between human effort and machine learnability, and feedback datastores for later transla-tions. Recent adaptive work collects fine-grained edits, prefix constraints, accept/reject actions, segment-level post-edits, and structured error tags to reduce editing cost and make feedback reusable. LLM-enabled sys-tems can produce useful edits but still require validation and human oversight; iterative self-refinement can improve fluency and naturalness while complicating metric interpretation.
This points to a design gap for human-centered multi-agent trans-lation: adaptation should respond not only to user feedback, but also to the effort and commitment behind that feedback.
3 Formative Study 3.1 Participants and Data Collection
We interviewed six professional translators (E1–E6), evenly split between client-side and vendor-side roles. Each participant had at least 6 years of professional practice across diverse translation domains, including game localization, marketing, government, med-ical translation, and chip design (Table 1). All interviews were conducted remotely via Zoom, and sessions were recorded and transcribed. Interviews averaged 42 minutes. The study was re-viewed by our institution’s IRB and granted Exempt status, and all participants provided informed consent prior to the interview.
3.2 Procedure
We conducted semi-structured interviews consisting of three parts: Background and tool usage: Participants were asked about the AI tools they use and their satisfaction with these tools (including reasons). Think-aloud translation task: Participants were given a sentence to translate in real time while verbalizing their thought process and explaining their edits. Open-ended design reflec-tion: Participants were asked to imagine and describe their ideal translation tool while thinking aloud. The full interview protocol is provided in Appendix.
We analyzed the interview transcripts and notes using thematic analysis. Two authors conducted open coding to identify re-curring issues and breakdowns in current translation workflows. The authors then compared and discussed their codes, resolved disagreements through discussion, and grouped related codes into higher-level themes. This analysis led to three recurring challenges that informed the design of CHORUS.
3.3 Challenges in Current AI Translation Tools
C1. Single-LLM rewrites obscure distinct translation qual-ity dimensions. Participants described professional revision as a process of balancing multiple quality dimensions within the same sentence, including accuracy, terminology, fluency, and style (E1, E2, E5). However, current AI tools often return a single broad rewrite, making it difficult for translators to see which quality dimension the system addressed and whether the revision introduced trade-offs elsewhere. As one participant explained, “It often rewrites every-thing at once. I cannot tell if it is fixing terminology or just changing the style, so I end up going back and rechecking” (E2). Across inter-views, all six experts reported general-purpose AI chatbots (GPT web) as their primary day-to-day AI aid, while CAT tools were reserved for projects with established translation memories.
C2. Current AI tools do not carry forward translators’ qual-ity adjustments. Participants emphasized that translation deci-sions unfold through revision: a sentence may evolve through sev-eral versions, some segments may be repeatedly reconsidered, and translators may concentrate more attention on one quality concern than another (E1, E2). These process traces matter because not all edits carry the same weight. A terminology correction in medical content may reflect a domain constraint, while a fluency revision may matter more in consumer-facing text (E2, E6). However, cur-rent AI tools rarely preserve these signals across the workflow. Translators must repeatedly tell the system what to prioritize, as in E3’s example: “This part has been revised three times. Don’t change it, but you can still improve the surrounding sentences.”
C3. Current AI tools provide little support for review-ing and justifying translation decisions. Finally, participants wanted more structured guidance for checking what had been addressed during revision (E1, E5). They described needs such as re-viewing style guides (E1), identifying critical issues (E1, E3), check-ing domain-specific constraints (E4), and understanding why a suggestion was made (E5). These needs were tied to professional ac-countability: translators often have to justify and defend decisions to supervisors, clients, or other stakeholders (E3, E4, E5). Current AI tools offer suggestions, but provide limited support for seeing which quality dimensions have been checked, which concerns may still need attention, and how a final decision can be explained.
3.4 Design Rationale
D1. Operationalize MQM as separable agent roles. Because LLM-based translation support is context-sensitive, a single-LLM in-terface can merge several quality goals into one opaque rewrite and become difficult for translators to inspect. All participants acknowl-edged MQM as an authoritative framework for multi-dimensional thinking. We therefore decompose revision support into MQM-aligned agents, each focused on one quality lens, so translators can inspect dimension-specific suggestions and coordinate trade-offs before deciding on the final translation.
D2. Capture and use editing history to reduce repeated MQM-specific quality adjustments. Professional translators of-ten need to steer the balance among quality dimensions across many segments, such as preserving terminology, adjusting fluency, or maintaining a client-specific style. In current LLM workflows, these preferences often have to be restated through repeated prompts, which adds effort to the revision process. We should use the trans-lator’s interaction history to ease this work. We need to introduce a novel mechanism with confirmed edits and revision traces to help the system recognize which quality adjustments the translator has already made and carry them forward into later suggestions.
D3. Support structured review through revision history and quality coverage.
We should help translators review their own revision process without turning reflection into a separate task. The interface should surface useful traces of the work already done: recurring edits and preferences from revision history, how attention has been distributed across MQM dimensions, and which quality concerns may still deserve review. This gives translators material for checking coverage and explaining decisions while keeping final judgment with the user.
4 CHORUS System 4.1 Human-AI Multi-Agent Collaboration
Synchronization among multiple agents. Separating quality 4.1.1 dimensions introduces a coordination problem: suggestions from multiple agents must remain anchored to the same current draft. CHORUS therefore synchronizes agent outputs after each user edit. Each agent’s response is represented as a token-level differ-ence against the latest draft using the Longest Common Subse-quence algorithm. CHORUS applies the minimal patch and immediately re-renders agent suggestions (see Fig. 2), keeping dimension-specific feedback aligned with the translator’s current text.
4.1.2 Ranking agents. Since several dimensions may be relevant at once, CHORUS ranks agents to reduce attentional load while preserving access to the full MQM space. The system obtains an initial relevance score for each agent from the translator’s high-level goal and current context. Subsequent interactions update each agent’s score, so agents that better match the translator’s current focus become more visible over time. This serves as an attention-management mechanism: the most relevant agents are ranked according to the current context while keeping other MQM dimensions always available.
4.1.3 Error handling. Hallucinations are known to cause errors in LLMs’ responses. CHORUS employs a list of “bad examples” whenever the user identifies an error in LLM output. During regen-eration, CHORUS commands the LLM to avoid similar mistakes cached in the bad example list. Additionally, users can regenerate any single agent’s output if unsatisfactory.
4.2 Effort-Aware Memory
To help the system personalize toward users’ habits, we introduce an effort-aware memory mechanism that converts users’ editing behavior into weighted memory for prompt adaptation. Live Effort estimates which edits required more translator effort, Micro-Edits preserve what was changed and where the change occurred, and Memory uses these effort-weighted edit records to generate prompt guidance for the AI agents. Together, these components allow CHO-RUS to remember which edits matter the most and adjust future prompts accordingly.
Live Effort Following Krings and Stasimioti, CHORUS models live effort at three levels: temporal, technical, and cogni-tive. Temporal effort is defined by two metrics: an initial pause that reflects the time spent reading and understanding the source text and a total edit duration that captures the time to refine. Tech-nical effort is measured through keystroke counts (deletion and cursor movement). Cognitive effort captures the deeper reasoning processes, such as the number of redos and stress levels, which are inferred by ChatGPT-5.3 using appraisal theories of emotion (difficulty, ambiguity, risk, and controllability). To combine these three metrics into one live effort score, we use a linear effort model and join these metrics into the final score.
Micro-Edits capture deletions or replacements during editing. This design allows the system to preserve not only what the trans-lator changed, but also where the change occurred. It can also be used to infer why the change was made and, when combined with the system’s memory component, how much effort it required.
Memory CHORUS uses a weighted memory to reduce the need for repeated edits. Inspired by importance-aware retrieval, CHORUS uses seven memory buckets to match up with seven AI agents’ editing history. The memory contains a snapshot of micro-edits, target translation, and user prompt whenever a change is made. Non-agent edits such as modifying the output sentence di-rectly are stored in a general memory bucket. To personalize CHO-RUS, we use the stored memory to adjust the prompt template to fit the professional translator’s idiosyncratic traits. This begins with ranking micro-edits using the live effort algorithm. After ranking, CHORUS injects the top-five micro-edits as few-shot examples into a prompt template to generate a weighted memory, personalizing what matters most to users without restating their needs. Appendix shows a complete example and prompt templates.
4.3 Translational Scaffolding
We instrument Live Style Guide to show translational scaffolding by surfacing the learned preferences, recurring correction patterns, and dimension-level strengths as a spider graph using the weighted memory. This helps users to see which of the MQM dimensions are more frequently used and their personalized feedback on their performance. This feedback is created based on the LLM’s response using the MQM website’s suggestions. This way, the system helps professional translators to recognize their blind spots over time and reflect on their strengths.
5 Evaluation 5.1 Experiment Design
The evaluation used English-to-Chinese sentence-level data from WMT24. We sampled 10 sentences from each of four domains: literary, news, social, and speech. WMT24 reference translations were retained for quality evaluation. In the Baseline condition, par-ticipants revised translations while using the GPT-5.3 web as AI support. Chat history was cleared between trials. In the CHORUS condition, they worked with the CHORUS interface (see Figure 3), which contains the source text, an editable machine-translated draft, and access to seven AI agents. In both conditions, initial drafts were produced by GPT-5.3. Condition order was counterbalanced using a pre-generated table, and to avoid learning effects, no identical sentence appeared across conditions or trials.
CAT tools were not used because they rely on accumulated trans-lation memories and termbases unavailable for our tasks. Other sys-tems were also unsuitable as end-to-end baselines: CASMACAT predates LLM-based assistance and was not evaluated on English– Chinese, TranSmart is a closed commercial service, and MQM-APE and HiMATE provide evaluation methods rather than interactive translation systems.
5.2 Participants
We recruited 30 licensed professional English–Chinese translators who provided proof of certification. Participants were 21 to 50 years old (M =28.9, SD =6.1) and reported 3 to 21 years of transla-tion experience (M =5.4, SD =4.9). They reported frequent use of AI translation tools (median = 4/5) and moderate-to-high familiarity with CAT tools (median = 3.5/5). Each participant was assigned two out of the four possible WMT24 domains and paired with con-ditions. For each domain, participants edited 10 sentences (10 trials), and each domain contained the two conditions. As a result, each participant conducted 10 sentences per condition, 20 per domain, and 40 in total. The overall frequency of domains was balanced across all participants. For example, we assigned Pn domains A and B and Pn+1 domains C and D. This pattern was alternated.
5.3 Procedure and Measures
After participants signed the consent form, the experimenter ex-plained the tasks and walked them through the interface and how to perform the trial. Participants had five minutes to practice. Once ready, they were assigned one of the conditions and switched to the other condition upon completing the first. The NASA-TLX sur-vey was used to assess cognitive load across conditions. At the end of the experiment, a semi-structured interview followed to collect qualitative feedback and their impressions of the two condi-tions. Sessions lasted approximately 100 minutes, and participants were compensated ¥100.
We collected interaction logs, outcome translations, and self-report data. Timed performance was computed from active editing duration, excluding idle periods. Perceived workload was measured with NASA-TLX on the original 0–100 scale, and open-ended re-sponses were collected. To assess quality, we used COMET and BLEU automatic evaluation metrics as a means of automatic comparison as they come with “ground-truth” labeling. Further, we asked three professional translators with more than 6 years of experience to compare the translated sentence with the original.
6 Results 6.1 Efficiency, Effort, and Workload
A paired t -test on log-transformed completion time showed that CHORUS significantly reduced completion time compared with Baseline (t = −5.35, p <.001). On the original scale, the geometric-mean time ratio of CHORUS over Baseline was 0.662, 95% CI [0.567, 0.772], indicating 33.8% faster task completion on average. Domain-level comparisons followed the same direction, with the largest reductions in literary and speech tasks (Fig. 4).
CHORUS also produced significantly lower Live Effort scores than Baseline (t = −5.81, p <.001). The average effort score was 54.85 (SD = 19.14) for CHORUS and 65.84 (SD = 19.61) for Baseline, a reduction of about 11 points. Mixed-effects analyses found no significant omnibus domain effect for completion time, χ 2 = 5.45, p =.142, or effort, χ 2 = 6.00, p =.112, suggesting that the CHORUS advantage was not driven by a single domain.
Task-order analyses further suggested that participants settled into more efficient interaction patterns with CHORUS over time (Fig. 6). For completion time, the interaction between task order and CHORUS was significant, b = −0.146, z = −6.66, p <.001, with a significant mean slope difference against Baseline, −0.147, t = −6.12, p <.001. Live Effort showed the same pattern: the CHORUS interaction was significant, b = −2.444, z = −5.67, p <.001, with a significant mean slope difference of −2.448, t = −5.14, p <.001 (Fig. 5).
6.2 Translation Quality and Adaptivity
CHORUS improved final translation quality relative to Baseline. Measured against human-translated WMT references, mean BLEU increased from 34.90 under Baseline to 37.98 under CHORUS, a mean paired difference of 3.08 points (t = 3.16, p =.0036; Wilcoxon p =.0081; Hedges’ g = 0.56). Mean COMET increased from 0.837 to 0.852, a mean paired difference of 0.015 (t = 3.51, p =.0015; Wilcoxon p =.0019; Hedges’ g = 0.62). For both metrics, 73.3% of participants had higher scores under CHORUS than Baseline.
Quality trends over task order provided additional evidence of adaptation (Fig. 7). BLEU increased over rows under CHORUS (slope
= 0.297, p =.014), as did METEOR (slope = 0.00236, p =.024), while Baseline remained largely flat on these metrics. A composite quality index was flat for Baseline (slope = 0.00062, p =.923) and showed an upward trend for CHORUS (slope = 0.0119, p =.059). BLEURT also suggested that Baseline quality declined over rows (slope = −0.00830, p <.001), while CHORUS showed a smaller, non-significant decrease (slope = −0.00295, p =.101).
Human expert ratings showed the same direction. Three transla-tion experts rated whether translation quality improved over time. CHORUS received a mean rating of 3.53 (SD = 1.38), whereas Baseline received 1.70 (SD = 1.06), with Baseline ratings concen-trated toward disagreement and CHORUS ratings showing more agreement (Fig. 8).
6.3 Structured Revision and Reflection
Open-ended responses helped explain why CHORUS reduced ef-fort and improved quality. Participants attributed the lower effort to targeted support, clearer problem visibility, and less need to manually formulate prompts. They described CHORUS as mak-ing issue-specific feedback easier to inspect than broad chatbot responses, and several participants noted that focusing on one qual-ity dimension at a time reduced the burden of mentally tracking accuracy, terminology, fluency, and style together (Fig. 9).
Participants also described CHORUS as supporting more de-liberate quality decisions. They reported that dimension-specific suggestions made trade-offs more explicit, helped them catch is-sues they might otherwise miss, and supported verification before confirming a sentence. In contrast to a static chatbot, participants described CHORUS as increasingly reflecting their ongoing edits, preferred correction patterns, and habitual areas of focus, reducing the need to repeatedly restate the same priorities.
The Style Guide and radar chart supported structured review by making revision history and MQM coverage visible. Participants used the radar chart to see which dimensions received more atten-tion and which were underrepresented, and many reported using the visualization to rebalance their focus in later revisions. The Style Guide also helped participants understand MQM dimensions, identify possible misclassifications in their own edits, and justify decisions with more confidence (Fig. 10).
7 Discussion
The results suggest that CHORUS improves professional translation not by replacing translator judgment, but by reorganizing LLM as-sistance around how translators revise. MQM agents made quality concerns more visible and easier to inspect, reducing the need to repeatedly prompt, compare broad responses, and mentally track several revision goals at once. This shift may help explain the 33.8% reduction in completion time, the lower Live Effort, and the lower NASA-TLX workload (RQ1). Because CHORUS differs from the base-line in both its multi-agent backend and its purpose-built interface, we interpret these gains as the effect of the integrated workflow rather than of the agent architecture alone. It also aligns with prior work on mixed-initiative translation, where useful automation sup-ports professional control rather than removing it.
CHORUS improved translation metrics (BLEU +3.08, COMET +0.015) and expert ratings (RQ2). One explanation is that partic-ipants could evaluate whether a suggestion improved accuracy, fluency, or another MQM dimension before committing to it. Mak-ing trade-offs inspectable reduced the risk of accepting suboptimal rewrites and supported more deliberate final decisions.
CHORUS also supported reflection beyond sentence-level editing (RQ3), complementing the inspectability and reduced re-prompting that participants highlighted in their qualitative feedback. The Style Guide and radar chart made revision history and quality coverage visible, helping participants notice where their attention was con-centrated, identify possible blind spots, and explain decisions using shared quality categories.
More broadly, CHORUS illustrates a division of labor for human-centered multi-agent systems. Agents can surface dimension-specific issues, organize alternatives, and accumulate interaction history, while the professionals remain responsible for interpreting context, weighing trade-offs, and making the final decision. This pattern may generalize to other revision-heavy domains where people co-ordinate multiple specialized AI agents under shifting priorities.
8 Limitations and Future Work
Our evaluation assessed CHORUS as an integrated workflow be-cause the goal of evaluation was to holistically understand the efficacy of the system for professional translation. Efficacy of indi-vidual components remains the future work.
In addition, our evaluation was limited to sentence-level English– Chinese tasks from WMT24. Other language pairs, and domain-specific effort signals remain to be tested, and we hope to expand the user study through multilingual and document-level studies in future work.
9 Conclusion
This paper presented CHORUS, an MQM-aligned multi-agent workspace for professional translators. CHORUS separates revision into seven inspectable dimensions and integrates users’ interaction history to reduce repeated MQM-specific adjustments, and visualizes revision patterns through an MQM review interface. In a study with 30 pro-fessional translators, CHORUS reduced completion time, effort, and workload while improving translation quality evaluated via both
CHORUS: Designing Human–AI Multi-Agent Collaboration for Professional Translators expert ratings and existing methods. Findings suggest that CHO-RUS can better support professional work by preserving translators’ control and revealing quality concerns for inspections.
10 Safe and Responsible Innovation Statement
CHORUS explores a human-accountable paradigm for multi-agent interaction in professional translation, where AI support remains explainable, inspectable, and subject to human judgment. Responsi-ble deployment should protect confidential materials, audit uneven performance across users and domains, and prevent over-reliance on agent suggestions in high-stakes settings.