The efficiency-verification trade-off in large language model-assisted workplace risk assessment; [Tõhususe ja usaldusväärsuse tasakaal LLM-toega töökoha riskihindamisel]
1 More Paper · Full Reading

About this paper
A full audio edition of this paper.
Authors: A. Bosler, T. Koppel, K. Reinhold
Publication date: 2026
Read the paper: https://doi.org/10.3176/proc.2026.3.10
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “The efficiency-verification trade-off in large language model-assisted workplace risk assessment; [Tõhususe ja usaldusväärsuse tasakaal LLM-toega töökoha riskihindamisel],” by A. Bosler, T. Koppel, and K. Reinhold. Published in 2026.
Abstract.
Large language models (LLMs) are increasingly adopted in occupational health and safety (OHS) to accelerate workplace risk assessment, yet their integration raises a dilemma between efficiency and verification. The study examines this trade-off in an LLM-assisted workflow using data from 111 student assignment files from an educational setting, analyzed at the session level, yielding 96‒121 analytic sessions depending on measure availability. Efficiency was measured via perceived time savings versus an estimated manual baseline, while the verification-burden proxy was operationalized as an AI‒human disagreement in quantitative scoring (probability and impact on 1‒5 scales).
User acceptance was captured via the utility and satisfaction / intention to reuse Likert scales. the linked source Estonian Academy Publishers AI-ASSISTED DECISION-MAKING RESEARCH ARTICLE The results show substantial perceived efficiency gains: the median LLM-assisted time was 20 minutes versus the median manual time of 120 minutes (median saving 0.84). The verification-burden proxy remained non-trivial: median disagreement was 0.50, and 26.8% of sessions exhibited disagreement ≥1.0. Acceptance remained high (mean satisfaction 4.06/5). Notably, 40.6% of the sessions showed a possible overreliance pattern, with high satisfaction coexisting with high disagreement. In a robust ordinary least squares (OLS) regression, utility was the strongest predictor of satisfaction, while disagreement reduced satisfaction; time saving was not significant.
1. Introduction.
The digital transformation of corporate governance has positioned generative AI as a practical tool for improving organizational resilience. In the context of sustainable management in the digital era, sustainability increasingly depends on maintaining high occupational health and safety (OHS) standards while optimizing limited time and human resources. Workplace risk assessment is central to this challenge, and large language models (LLMs) are now being integrated to support hazard identification and mitigation planning, with the promise of faster and more standardized safety workflows.
Citation:
Bosler, A., Koppel, T. and Reinhold, K. 2026. The efficiency–verification trade-off in large language model-assisted workplace risk assessment. Proceedings of the Estonian Academy of Sciences, 75, 251–259. the linked source
A primary managerial incentive for adopting LLMs is a potentially substantial increase in process efficiency. Recent field experiments have documented substantial productivity gains; for instance, Martin et al. (2025) found that AI-assisted groups completed construction risk evaluations in an average of 36 minutes, compared to 157 minutes for traditional methods. However, these gains may introduce an efficiency‒verification trade-off: while LLMs can rapidly generate broad and plausible risk narratives, they may lack site-specific depth and struggle with precise quantitative scoring (e.g., probability and impact), which remains critical for prioritizing controls and allocating resources.
A relevant concern for sustainable safety governance is automation bias (AB) – the human tendency to over-rely on automated suggestions even when they contradict contextual evidence. Empirical studies suggest that the “mere knowledge” of a suggestion being AI-generated can induce overreliance, leading to inferior decision payoffs. This phe nomenon is exacerbated by a “gulf of impatience,” where users prioritize the low latency of AI over the time-consuming process of manual verification. Although users frequently report high satisfaction with AI integration in safety-related work, there remains limited evidence on the “assurance cost” of perceived efficiency – namely, whether perceived time savings coincide with systematic AI‒human divergence in quanti tative risk metrics.
To address this gap, the present study analyzes structured workplace risk assessment data from an educational setting, linking perceived time savings to AI‒human disagreement in quantitative risk scoring and user acceptance. The aim of this study is to examine the efficiency–verification trade-off in an LLM-assisted workplace risk assessment across three dimensions: perceived time savings relative to a manual baseline, quantitative disagreement between AI-generated and human-finalized risk scores, and the extent to which high user acceptance may coexist with substantial scoring divergence. This approach allows us to explore a possible overreliance pattern in which satisfaction and intention to reuse remain high despite substantial scoring disagreement.
To investigate this efficiency‒verification trade-off, this study pursues three research questions:
1. What is the magnitude of perceived efficiency gains (time savings) when students use LLMs to.
assist in workplace risk assessment compared to their expected manual baseline?
2. How significant is the quantitative disagreement between initial AI-generated risk scores.
(probability and impact) and the final scores validated by human users?
3. To what extent do high user satisfaction and intention to reuse persist despite substantial.
disagreement with the AI’s quantitative output, suggesting a possible overreliance pattern?
2. Background.
LLMs are increasingly evaluated as safety support tools in workplace risk management, with evi dence showing strong performance in early-stage risk identification. For example, Nyqvist et al. (2024) reported that GPT-4 produced more comprehensive hazard identification outputs than human experts (overall score 8.6 vs. 5.7). Similarly, Oral et al. (2026) found high internal consistency in ChatGPT-generated assessments (similarity 0.90–0.99), suggesting stable output structure under repeated prompting. At the same time, technical specificity and numerically grounded scoring remain weak points: in failure mode and effects analysis (FMEA)-style risk assessments, Collier et al. (2024) reported that experts rated AI-generated severity and likelihood scoring as “poor” or “fair” in over half of cases, indicating limitations in decision-critical quantification.
A major driver of adoption is the prospect of process efficiency gains. In construction risk assessment, Martin et al. (2025) documented substantial time reductions, with AI-assisted groups completing evaluations in 36 minutes compared to 157 minutes using traditional methods. However, these gains may introduce an efficiency‒verification trade-off. Prior work argues that filtering irrelevant or inaccurate AI content can impose a non-trivial verification burden, including iterative prompting and “debugging” of AI artifacts. From a governance perspective, such acceleration may also increase the risk of expertise erosion if users reduce reflective deliberation in safety-critical decisions.
This trade-off may become particularly consequential under AB – the tendency to over-rely on automated suggestions despite contradictory evidence. Parasuraman and Riley (1997) first systematically described how misuse of automation through overreliance leads to monitoring failures and decision errors. Subsequent reviews confirmed that erroneous automated advice increases incorrect decision-making by over 25% compared to unaided conditions, and that verification complexity is a central factor determining susceptibility to AB – the harder it is to verify automation correctness, the greater the risk of overreliance. In AI-assisted workflows, recent evidence confirms that merely knowing a recommendation was AI-generated can induce overreliance, leading to decision payoffs roughly 22% lower than control conditions.
Low-latency outputs may further intensify this effect through a “gulf of impatience,” where users accept “fast” results to avoid cognitive effort. While satisfaction with LLM integration is often high, there remains limited evidence linking perceived time savings to objective AI‒human disagreement in quantitative risk metrics. This study addresses this gap by analyzing educational workplace risk assessment data and examining whether a possible overreliance pattern emerges when high acceptance coexists with substantial scoring divergence.
3. Methods.
This study examines the efficiency–verification trade-off in an LLM-assisted workplace risk assessment using data from a structured educational workflow. The dataset comprised 111 submitted student assignment files collected across multiple course cohorts. Although all assignments involved AI-assisted workplace risk assessment, some task formats required students to complete more than one tool- or interface-specific workflow. Accordingly, one file could contribute one or more analytic sessions. The 111 submitted assignment files yielded 121 analytic sessions in total. For analysis, the primary unit was an assessment session, defined as one completed risk assessment workflow under a given tool condition. Within a session, time measures were recorded separately for each workplace and then aggregated to session totals.
Hazard-level AI‒human scoring differences were averaged to obtain session-level disagreement. Likert-based perception measures were recorded once per session. Analytic sample sizes then varied by outcome due to incomplete time, scoring, or questionnaire data. Specifically, 113 sessions had valid time data, 97 had complete AI‒human scoring pairs, 120‒121 had questionnaire data depending on the scale, 96 had complete data for the overreliance zone analysis, and 91 had complete cases for regression analysis. The study does not aim to establish ground-truth occupational risk levels; instead, it analyzes perceived efficiency gains, AI‒human quantitative scoring disagreement as a verification-burden proxy, and user acceptance.
Time measures were self-reported for each workplace within a session: time spent using the LLM-assisted workflow (Tai) and the estimated manual baseline without AI support (Tmanual). Session totals were computed by summing across workplaces where both values were available. Perceived time saving was operationalized as:
SavingPct= Tmanual −Tai. Tmanual
For each hazard, both the AI and the participant provided numeric scores on 1‒5 scales for probability and impact. Session-level disagreement was computed as the mean absolute divergence across hazards:
Disagreement is interpreted as a verification-burden proxy and does not imply that either AI or human scores represent a ground-truth reference.
The key measures used in this study should be interpreted as analytically useful proxies rather than direct measures of the underlying constructs. First, the disagreement index captures divergence between AI-generated and human-assigned scores but does not establish which party is more accurate; human scores may themselves reflect miscalibration, and high disagreement does not necessarily indicate AI error. Second, the manual-time baseline (Tmanual) is a subjective self-estimate rather than an observed measurement, meaning that perceived time savings reflect the user’s internal valuation of the tool rather than objective efficiency gains.
Third, the overreliance zone is defined operationally as the co-occurrence of high disagreement and high satisfaction and should be interpreted as a possible overreliance pattern consistent with AB-related concerns rather than a direct measurement of overreliance or decision degradation.
Participants completed a 17-item Likert questionnaire (1‒5) measuring risk-metric realism, competence/accuracy, utility, and satisfaction / intention to reuse. Scale scores were computed as mean item ratings, and internal consistency was assessed using Cronbach’s alpha. No questionnaire items were dropped from the predefined scales. An overreliance zone was defined as the intersection of high disagreement (D ≥ 0.5) and high satisfaction (S ≥ 4.0) and is interpreted as an operational indicator of a possible overreliance pattern.
Descriptive statistics were reported using medians and interquartile ranges for time and disagreement measures, and means for perception scales. Associations were assessed using Spearman correlations. Group differences were tested with the Mann‒Whitney U test, reporting Cliff’s delta as the effect size. Finally, an ordinary least squares (OLS) regression with heteroskedasticity-robust standard errors (HC3) examined whether satisfaction could be jointly explained by perceived efficiency gains (SavingPct), disagreement, and perceived utility.
The study was conducted in an educational setting using anonymized student-generated materials, with results reported only in aggregated form.
4. Results.
4.1. Perceived efficiency gains and AI‒human scoring disagreement (RQ1‒RQ2).
Across sessions with a valid manual baseline (n = 113), participants reported substantial perceived efficiency gains when completing workplace risk assessments with LLM assistance. Median LLM-assisted time was 20 min, compared to 120 min for the estimated manual baseline, corresponding to a median proportional time saving of 0.84. Positive time savings were reported in 97.3% of sessions with valid baselines (Table 1). Figure 1 visualizes the strong separation between the distributions of LLM-assisted and manual baseline times.
As a verification-burden proxy, a session-level disagreement index was computed from absolute AI‒human differences in probability and hazard impact scoring. Among sessions with valid scoring pairs (n = 97), median disagreement was 0.50, and 26.8% showed mean disagreement ≥ 1.0 (Table 1), indicating non-trivial divergence between AI-generated and human-assigned scores in a meaningful share of cases. At the hazard level, exact agreement was lower for probability (43.9%) than for hazard impact (59.9%), with visible off-diagonal dispersion in the mid-to-high scoring range (Fig. 2).
4.2. User acceptance and possible overreliance patterns (RQ3).
Despite scoring disagreements, participants reported high acceptance of the LLM-assisted workflow. All perception scales averaged above 4.0, including utility and satisfaction / intention to reuse (Table 1). Internal consistency was acceptable to high across the four composite scales: metrics realism (α = 0.896, 7 items), competence/accuracy (α = 0.799, 3 items), utility (α = 0.902, 4 items), and satisfaction / intention to reuse (α = 0.862, 3 items). Disagreement was negatively associated with satisfaction (Spearman ρ = ‒0.26, p = 0.012), whereas the relationship between time saving and satisfaction was small and marginal (ρ = 0.20, p = 0.053).
Figure 3 maps disagreement against satisfaction using the median disagreement (0.50) and a satisfaction cutoff of 4.0. Among sessions with complete data (n = 96), 40.6% fell into the high-disagreement/high-satisfaction quadrant (overreliance zone), indicating frequent cases of strong tool acceptance even when participants assigned scores that diverged substantially from AI-generated values.
Finally, an OLS regression model predicting satisfaction (Table 2) showed that perceived utility was the strongest positive predictor (β = 0.98, p < 0.001), while disagreement remained a significant negative predictor (β = ‒0.40, p = 0.002). Perceived proportional time saving was not statistically significant once the other predictors were included (p = 0.35). The model explained 72.9% of variance in satisfaction (R2 = 0.729). Standardized coefficients confirmed that utility was the dominant predictor (β = 0.830), followed by disagreement (β = ‒0.211), while time saving was negligible (β = ‒0.041). A supplementary check for nonlinearity revealed no significant quadratic terms for any of the three predictors (SavingPct2, mean disagreement2, utility2; all p > 0.30), supporting the adequacy of the linear specification.
5. Discussion.
This study mapped the efficiency–verification trade-off in an LLM-assisted workplace risk assessment. The workflow appears to offer substantial perceived speed advantages, but the findings also suggest an assurance cost: quantitative risk scores often showed substantial AI‒human disagreement, while acceptance remained high even when disagreement was non-trivial. This creates a governance dilemma for sustainable safety management: perceived time efficiency is attractive, but it does not automatically imply reliability or well-calibrated trust.
5.1. Perceived efficiency gains (RQ1).
The median perceived time saving of 84% documented in this study indicates a substantial perceived shift in OHS process management. A reduction from a 120-minute manual baseline to a 20-minute AI-assisted session mirrors the “efficiency leap” observed in construction risk assessment by Martin et al. (2025), who recorded a move from 157 to 36 minutes.
However, our findings suggest that this efficiency is primarily perceived. By using Tmanual as a subjective baseline, we captured the user’s internal valuation of the tool. This strong perceived reduction in effort may serve as an important driver of technology adoption, but as Sun and Kalar (2025) note, such gains often come with hidden costs. The lack of statistical significance for SavingPct in our regression suggests that while speed may act as an initial adoption driver, it is not the primary driver of long-term satisfaction once the user begins interacting with the tool’s output.
5.2. AI‒human disagreement as a verification-burden proxy (RQ2).
Despite strong efficiency perception, scoring disagreement suggests a non-trivial verification-burden proxy. Median AI‒human disagreement was 0.50 (1‒5 scale), and 26.8% of sessions showed mean disagreement ≥ 1.0, suggesting that scoring divergence – and by implication, potential verification demands – remains substantial in a meaningful subset of cases. Probability judgments showed lower exact agreement than hazard impact (43.9% vs 59.9%), which is consistent with probability being more context-dependent (exposure, local conditions) and therefore harder for LLMs to infer reliably from limited textual input. This aligns with the view that LLMs can support structured hazard narratives while still requiring active human calibration for numeric prioritization.
5.3. The overreliance zone and a possible overreliance pattern (RQ3).
One notable finding is the identification of the overreliance zone, encompassing 40.6% of sessions. In these cases, users reported high satisfaction alongside substantial AI‒human scoring disagreement. This pattern may pose a challenge for sustainable safety governance and may be interpreted through the lens of automation complacency. When the perceived benefit of the system is high, partly due to the large perceived time saving observed here, users may, as Kücking et al. (2024) suggest, develop a “blindness to potential errors.” This may reflect the “gulf of impatience”, where the desire to finalize the assessment quickly outweighs the cognitive strain of rigorous auditing. For sustainable management, this implies that high user satisfaction may be an insufficient proxy for process quality; a satisfied user is not necessarily a vigilant one.
5.4. What drives satisfaction: utility matters most, disagreement still matters (RQ2‒RQ3).
The regression results are consistent with the interpretation above. Perceived utility was the strongest positive predictor of satisfaction, while higher disagreement significantly reduced satisfaction. Meanwhile, the time saving variable did not remain significant once utility and disagreement were included.
This suggests a plausible behavioral pattern. Perceived efficiency may support initial adoption, utility may sustain intention to reuse, and disagreement may act as a friction cost that reduces satisfaction but often does not eliminate acceptance – helping explain why a possible overreliance pattern may persist despite recognized scoring divergence. This is important because it helps clarify how such patterns may arise in ways consistent with AB-related concerns: people may recognize scoring divergence (hence, disagreement hurts satisfaction) but still judge the tool as “worth it” because it remains useful overall.
5.5. Implications for sustainable digital safety governance.
From a management perspective, the results support a practical recommendation: organizations should treat LLM-based risk assessment as a potentially speed-enhancing system that still requires structured assurance controls, especially for numeric risk metrics. Practical implications include:
● Do not equate satisfaction with assurance. Acceptance can remain high even when quantitative outputs show substantial AI‒human disagreement.
● Track AI‒human disagreement as a verification-burden indicator. Monitoring scoring divergence or disagreement rates can approximate verification burden and surface reliability issues.
● Design and train for verification. Workflow design should include explicit review steps (e.g., “why this score?” prompts) and user training focused on the most fragile component – probability calibration in context-dependent environments.
5.6. Limitations.
This study has several limitations. First, the dataset originates from an educational setting and reflects student-based validation, which may differ from professional OHS practice in important ways. Compared to students, expert assessors likely bring domain-specific contextual knowledge that could reduce AI‒human scoring divergence, as they may be better positioned to evaluate site-specific probability and impact estimates. At the same time, experts may exhibit different trust calibration patterns ‒ either greater skepticism toward AI outputs due to professional standards, or conversely, higher complacency if time pressure is a dominant factor in real workplace settings. Whether the overreliance zone identified here would persist, shrink, or shift in composition among experienced practitioners remains an open empirical question.
Second, the manual-time baseline was self-estimated rather than experimentally measured, so efficiency results should be interpreted as perceived savings rather than objective time-and-motion evidence. Third, the disagreement metric captures AI–human divergence, not ground-truth correctness; therefore, it should be interpreted as a proxy for verification effort and calibration rather than an accuracy benchmark. Finally, sessions varied in the number of workplaces and hazards assessed, which introduces heterogeneity in workload and opportunities for AI‒human scoring divergence.
5.7. Future work.
Future research should replicate the analysis in professional OHS contexts using a comparative design that directly contrasts student and expert assessors on disagreement magnitude, satisfaction, and overreliance patterns under matched task conditions. Such studies should also include observed manual baselines rather than self-estimated ones, to provide objective validation of time-effect estimates. Additional work could examine whether disagreement patterns differ by hazard type, workplace context, or user expertise, and whether structured interface interventions (e.g., mandatory justification steps) reduce possible overreliance patterns. Finally, qualitative analysis of edits and reasoning could clarify why users accept AI outputs even when numeric scores diverge from human-assigned values.
6. Conclusion.
This study suggests that the LLM-assisted workplace risk assessment delivers strong perceived efficiency benefits but introduces a potential verification burden, reflected in consistent AI‒human disagreement in probability and impact scoring. A substantial share of sessions exhibited high satisfaction despite high disagreement, indicating a possible overreliance pattern consistent with AB-related concerns that challenges the use of acceptance metrics as a proxy for process reliability. For sustainable digital safety management, the results highlight that speed gains should be paired with explicit verification structures so that efficiency does not come at the expense of risk governance quality.
Data availability statement
The data supporting the findings of this study are available from the corresponding author upon reasonable request.
Acknowledgment
The publication costs of this article were partially covered by the Estonian Academy of Sciences.