AI at the Front Lines of Platform Governance: Using LLMs to Support Illegal Content Reporting under the Digital Services Act
1 More Paper · Full Reading

About this paper
Illegal Content Reporting under the Digital Services Act MARIE-THERESE SEKWENZ∗, Delft University of Technology, Delft, The Netherlands SHREYAN BISWAS∗, Delft University of Technology, The Netherlands RITA HERMANN-GSENGER, Weizenbaum Institute, Germany UJWAL GADIRAJU, Delft University of Technology, The Netherlands Illegal content reporting mechanisms are a key technical and organizational measure through which online platforms address the dissemination of illegal content under European Union law. Under the Digital Services Act (DSA), user notices submitted pursuant to Article 16 must be sufficiently substantiated and provided in good faith, requiring users to interpret legal and procedural language and translate it into legally meaningful categories and reasons. In practice, however, reporting illegal content remains cumbersome across major social media platforms, placing substantial cognitive and legal demands on users. Without effective support at the reporting interface, operationalizing Article 16 in practice remains challenging. We investigate how large language model (LLM)–based assistants can support illegal content reporting. In a controlled user study (N = 450) using an interface modeled on a major platform’s reporting workflow, we compare three conditions: a conventional explainable AI assistant (XAI) that suggests a single legal category with a rationale, an evaluative AI assistant (EvalAI) that presents balanced pro and con arguments across candidate legal provisions for user deliberation, and a baseline reflecting unaided reporting (Baseline). We further examine these assistance forms under systematically varied AI error regimes. Our results show that EvalAI improves provision-level accuracy under AI error regimes and reduces misclassification distance relative to conventional XAI, particularly for near-miss and overbreadth errors. In contrast, conventional XAI does not improve—and can degrade—the quality of users’ rationales relative to unaided reporting, despite enabling faster decisions when the AI output is correct. We discuss implications for the design of compliance-oriented reporting interfaces, highlighting trade-offs between accuracy, deliberation, and vulnerability to misleading AI output.
Authors: M.-T. Sekwenz, S. Biswas, R. Hermann-Gsenger, U. Gadiraju
Publication date: 2026
Read the paper: https://doi.org/10.1145/3805689.3812301
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “AI at the Front Lines of Platform Governance: Using LLMs to Support Illegal Content Reporting under the Digital Services Act,” by M.-T. Sekwenz and colleagues. Published in 2026.
MARIE-THERESE SEKWENZ∗, Delft University of Technology, Delft, The Netherlands SHREYAN BISWAS∗, Delft University of Technology, The Netherlands RITA HERMANN-GSENGER, Weizenbaum Institute, Germany
UJWAL GADIRAJU, Delft University of Technology, The Netherlands
Illegal content reporting mechanisms are a key technical and organizational measure through which online platforms address the dissemination of illegal content under European Union law. Under the Digital Services Act (DSA), user notices submitted pursuant to Article 16 must be sufficiently substantiated and provided in good faith, requiring users to interpret legal and procedural language and translate it into legally meaningful categories and reasons. In practice, however, reporting illegal content remains cumbersome across major social media platforms, placing substantial cognitive and legal demands on users. Without effective support at the reporting interface, operationalizing Article 16 in practice remains challenging. We investigate how large language model (LLM)–based assistants can support illegal content reporting.
In a controlled user study (N = 450) using an interface modeled on a major platform’s reporting workflow, we compare three conditions: a conventional explainable AI assistant (XAI) that suggests a single legal category with a rationale, an evaluative AI assistant (EvalAI) that presents balanced pro and con arguments across candidate legal provisions for user deliberation, and a baseline reflecting unaided reporting (Baseline). We further examine these assistance forms under systematically varied AI error regimes. Our results show that EvalAI improves provision-level accuracy under AI error regimes and reduces misclassification distance relative to conventional XAI, particularly for near-miss and overbreadth errors.
In contrast, conventional XAI does not improve—and can degrade—the quality of users’ rationales relative to unaided reporting, despite enabling faster decisions when the AI output is correct. We discuss implications for the design of compliance-oriented reporting interfaces, highlighting trade-offs between accuracy, deliberation, and vulnerability to misleading AI output.
CCS Concepts: • Human-centered computing → Interactive systems and tools; • Social and professional topics → Computing / technology policy; • Applied computing → Law.
ACM Reference Format.
Marie-Therese Sekwenz, Shreyan Biswas, Rita Hermann-Gsenger, and Ujwal Gadiraju. 2026. AI at the Front Lines of Platform Governance: Using LLMs to Support Illegal Content Reporting under the Digital Services Act. In The 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’26), June 25–28, 2026, Montreal, QC, Canada. ACM, New York, NY, USA, 37 pages. the linked source
∗Authors contributed equally to this research.
Authors’ Contact Information: Marie-Therese Sekwenz, the email address, Delft University of Technology, Delft, Delft, The Netherlands; Shreyan Biswas, the email address, Delft University of Technology, Delft, The Netherlands; Rita Hermann-Gsenger, rita.gsenger@ weizenbaum-institut.de, Weizenbaum Institute, Berlin, Germany; Ujwal Gadiraju, the email address, Delft University of Technology, Delft, The Netherlands.
Online platforms have become core societal infrastructure and de facto public spaces shaping political discourse, cultural participation, and access to information. As a result, governments increasingly regulate how platforms address harms such as the dissemination of illegal content. The European Digital Services Act (DSA) harmonizes content governance rules across the European Union, aiming to create a “safe, predictable and trustworthy online environment” (Rec. 3 DSA).
Article 16 mandates that platforms provide user-facing mechanisms for reporting illegal content. What is ‘illegal’ is defined in Article 3 (h) DSA. 1 The requirements of Article 16 DSA are relevant to a wide range of platforms including hosting intermediary services (Art. 3(g) DSA), online platforms (Art. 3(i) DSA), online search engines (Art.3 (j) DSA), and Very Large Online Platforms (VLOPs) and Very Large Search Engines (VLOSEs) according to Article 33 DSA.2 Reports must include (i) a sufficiently substantiated explanation of the reasons (lit. a), (ii) a URL, (iii) identifying information of the reporting user, and (iv) a good-faith statement. While other DSA provisions impose systemic risk-management obligations for VLOPs and VLOSEs (Arts.
34–35, 37), Article 16 (which is applicable to hosting intermediary services, online platforms, search engines, VLOPs/VLOSEs), explicitly foregrounds the role of end users in initiating enforcement processes. Non-compliance can result in fines of up to 6% of global annual turnover for providers of VLOPs and VLOSEs (Art. 74). Recent enforcement actions against major platforms, including a e 120 million fine against the platform X, underscore the regulatory importance of DSA compliance.
The DSA further expands platform responsibility to content that is blurring the boundaries of ‘illegality’ in cases of disinformation3 which is at times referred to as “lawful but awful” content. Regulating speech, therefore, entails risks of over-removal and requires trust in regulators, platform participation, and avenues for user challenge and involvement. However, accurately distinguishing illegal content from harmful but lawful expression remains difficult even for experts, let alone ordinary users. Reporting requires mapping ambiguous posts to specific legal provisions under uncertainty. Errors at this stage can lead to under-reporting, over-reporting, or misrouting, with downstream consequences for both users and platforms. While platforms have no general monitoring duty (Art. 8), intermediary services (Rec. 20) and hosting providers (Rec.
22), however, do not have a carte blanche for liability questions mitigation as “other actors [...] should also help to avoid the spread of illegal content online, in accordance with the applicable law" (Rec. 27) and “it is important that all providers of hosting services, regardless of their size, put in place easily accessible and user-friendly notice and action mechanisms that facilitate the notification of specific items of information that the notifying party considers to be illegal content" (Rec. 50).
VLOPS and VLOSEs, furthermore, are obliged to mitigate the dissemination of 1The definition of ‘illegal content’ states: “any information that, in itself or in relation to an activity, including the sale of products or the provision of services, is not in compliance with Union law or the law of any Member State which is in compliance with Union law, irrespective of the precise subject matter or nature of that law." Recital 12 further clarifies: “‘illegal content’ should broadly reflect the existing rules in the offline environment. In particular, the concept of ‘illegal content’ should be defined broadly to cover information relating to illegal content, products, services and activities.
In particular, that concept should be understood to refer to information, irrespective of its form, that under the applicable law is either itself illegal, such as illegal hate speech or terrorist content and unlawful discriminatory content, or that the applicable rules render illegal in view of the fact that it relates to illegal activities."
2The VLOPs and VLOSEs have additional reporting and assessment duties compared to other platforms within the scope of the DSA. Such platforms have e.g., additional transparency reporting duties (Art. 42 DSA), or the obligation to conduct Systemic Risk Assessments (Art. 34 DSA) and be externally evaluated through Independent Audits (Art. 37 DSA). To qualify as a VLOP/VLOSE platforms must reach a threshold linked to "number of average monthly active recipients of the service in the Union equal to or higher than 45 million" according to Article 33 DSA.
3See for example Art. 34(c) which covers “any actual or foreseeable negative effects on civic discourse and electoral processes." This if further discussed in Recital 84 “[...] When assessing the systemic risks identified in this Regulation, those providers should also focus on the information that is not illegal, but contributes to the systemic risks identified in this Regulation. Such providers should therefore pay particular attention on how their services are used to disseminate or amplify misleading or deceptive content, including disinformation."
illegal content (Arts. 34–35, Rec. 80 and 87). Alternative approaches, such as community-based moderation, face challenges of scale, time constraints, increasingly convincing AI-generated material, linguistic diversity, evolving semiotics, and heterogeneous content types.
In response, platforms have increasingly explored AI-based support for moderation and reporting work-flows. Yet, AI assistance can amplify systemic risks rather than mitigate them, and human collaboration with decision-support systems is error-prone. A dominant approach to address these challenges has combined decision support with explainable AI (XAI), aiming to foster appropriate reliance through post-hoc rationales. This paradigm assumes relatively reliable recommendations and treats explanations as justificatory aids. In legal classification tasks, however, users must deliberate among competing interpretations before committing to a category, and a single suggested label may anchor decisions even when incorrect.
A growing body of work shows that appropriate reliance remains difficult to achieve—even with XAI—and that ex-planations can increase overreliance via illusory explanatory depth, particularly in low-friction or conversational settings or when explanations are hard to verify.
To address these limitations, Miller proposed evaluative AI: decision support that provides evidence for and against multiple options rather than prescribing a recommendation. This hypothesis-driven approach surfaces structured arguments, statutory elements, and evidentiary considerations without prescribing an outcome, thereby supporting user deliberation and mitigating over- and under-reliance. Evaluative AI is particularly suited to legal reporting contexts, where responsibility rests with the user and overconfident automation may undermine agency and legal reasoning.
In this work, we examine how different architectures of AI assistance influence provision-level accuracy, the quality of users’ substantiated explanations, and robustness under erroneous AI output in illegal-content reporting. We do not frame AI assistance as a mechanism for automated compliance or enforcement. Instead, we focus on user–AI interaction at the reporting interface and its implications for users’ ability to meet the procedural requirements of Article 16 DSA relevant to all intermediary services. References to VLOP/VLOSE’s systemic-risk provisions are limited to downstream effects of misclassification at the reporting stage, rather than claims of direct compliance or risk-mitigation effects (Arts. 34–35).
To that extent, our study is guided by the following overarching research question—“How do different AI assistance architectures in illegal-content reporting interfaces shape user decision quality, explanation quality, and robustness to erroneous AI output in Article 16 DSA reporting tasks?” Based on this question and informed by existing literature, we derive six hypotheses that henceforth structure our empirical analysis. As described below, H1–H3 examine provision-level accuracy under correct and erroneous AI assistance. H4 examines whether assistance architecture improves the quality of users’ own substantiated explanations. H5 tests whether different AI error conditions differentially degrade performance and whether EvalAI attenuates these losses relative to XAI.
H6 examines downstream risk-relevant outcomes of misclassification at the reporting stage, namely over-removal and misrouting.
• H1: Under AI error conditions, participants using EvalAI will achieve higher provision-level accuracy than participants using XAI, indicating lower overreliance on incorrect AI outputs.
• H2: Under AI error conditions, participants using EvalAI will select legally closer incorrect provisions than participants using XAI (smaller misclassification distance), but this reduction will be attenuated for Out-of-Scope errors relative to Near-Miss and Overbreadth errors.
• H3: When AI assistance is correct, EvalAI and XAI will yield comparable provision-level accuracy, but decisions under XAI will be faster than under EvalAI.
• H4: Participants using EvalAI will produce higher-quality substantiated explanations than those in Baseline and XAI, as reflected in stronger legal element coverage and reasoning depth.
• H5: AI error conditions (Near-Miss, Overbreadth, Out-of-Scope) will differ in the magnitude of their impact on reporting performance, and EvalAI will differentially attenuate these error-specific performance drops relative to XAI.
• H6: Relative to XAI, EvalAI will better align with downstream DSA-relevant reporting goals by reducing over-removal under Overbreadth errors and reducing misrouting under Out-of-Scope errors.
Hypotheses. We investigate these hypotheses through a controlled experiment (N = 450) using a custom reporting interface modeled on existing social-media workflows. Participants evaluated a simulated feed con-taining illegal content grounded in German criminal law,4 selected mandatory legal categories, and submitted reports under one of three assistance conditions. The primary outcomes in our study capture provision-level legal accuracy; while secondary outcomes assess explanation quality, decision time, and trust-related measures. This paper contributes to FAccT scholarship on content moderation and platform governance in three ways:
introducing an experimental paradigm for AI-assisted content reporting under regulatory constraints;
providing empirical evidence on how assistance conditions affect accuracy, reasoning quality, decision time, and trust under correct and erroneous AI assistance; and
deriving design and policy implications for DSA-aligned reporting interfaces.
2 Related Work
Prior research examines how generative AI tools support human tasks across domains such as design, finance, creative writing, and content moderation. Across these settings, effectiveness depends not only on model quality, but also on how assistance is presented, how easily outputs can be verified, and whether the interface supports appropriate reliance rather than passive acceptance.
Content Moderation and Reporting. User reporting plays a central role in content moderation ecosystems, yet the scale and complexity of online platforms exceed what purely human moderation can sustain. As a result, platforms rely on automated moderation systems that classify user-generated content and trigger governance outcomes. Hybrid approaches that combine automation with user input show promise, but also raise concerns about overreliance on AI.
Content moderation is often described as operating across two phases: ex ante moderation before publication, typically relying on automated filtering, and ex post moderation after publication, where user reporting plays a key role. Reporting interfaces, however, are often opaque and misaligned with user expectations and mental models. Users are reluctant to engage deeply with reporting, though they value nuanced options, proactive platform action, and personal moderation settings. User characteristics further shape flagging behavior. While accessible platform help centers can aid policy understanding, most users lack the legal knowledge to file substantiated reports, echoing broader findings on the inaccessibility of rules like the terms of service.
Legal Design Challenges. The DSA highlights legal design (Art. 25) and bans features that “deceive or manipulate [...] or otherwise materially distort or impair" user decision-making as exemplified in the Commission’s first DSA-based fine against X for the “use of the ‘blue checkmark’ for ‘verified accounts’ [which] deceives users.". The DSA also requires “user-friendly" reporting mechanisms (Art. 16). Yet, which design solutions meet these standards remains unclear. As regulation reshapes design practice, interdisciplinary approaches become necessary. Poorly implemented legal UX can take the form of dark patterns, raising usability, legal, 4We selected this context as a member-state example where illegal-content reporting was regulated prior to the DSA.
Supporting Illegal Content Reporting under the Digital Services Act with LLMs FAccT ’26, June 25–28, 2026, Montreal, QC, Canada and ethical concerns under which reporting mechanisms as Interface Interference can be subsumed, driven by Language Inaccessibility and Complex Language. AI-supported reporting. Large Language Models (LLMs) as a form of AI, increasingly shape moderation systems, including detecting harmful content across languages and generating “Statements of Reason" under the DSA. LLMs are also applied in legal practice and research to simplify statutes and provide advice, but they remain prone to hallucinations and misclassifications, raising ethical concerns.
Salimzadeh et al. show that complexity and uncertainty shape user performance in AI-supported tasks; in legal content reporting, this makes meaningful contestability especially important because interfaces must enable users to understand, challenge, and revise automated reasoning.
A central challenge in AI-supported reporting is not only whether AI assistance is present, but how it is structured and how susceptible users become to inevitable output errors. Conventional explainability often follows a recommend-and-justify logic, foregrounding a single predicted option and rationale. Such designs may reduce friction, but can also anchor users to erroneous outputs when the reasoning is difficult to verify. By contrast, more deliberative interfaces may reduce harmful overreliance, though often at a cost in effort and usability. Taken together, this literature suggests that assistance architectures foregrounding a single recommendation may be especially vulnerable to harmful reliance under AI error, whereas evaluative architectures that present evidence for and against alternatives may better align with human decision processes and support correction and recovery.
This expectation informs H1 and the premise of H2, which test whether Evaluative AI improves provision-level accuracy and reduces misclassification distance relative to Conventional XAI when the AI advice is wrong.
Recent work proposes more argumentative and contestable forms of AI support. User judgments are also sensitive to how alternatives are presented under model error, highlighting the importance of choice structure and choice independence. Together, these lines of work point toward evaluative forms of AI assistance that support deliberation across alternatives rather than merely justifying one recommendation. This informs H3, which examines whether the two AI-assisted conditions remain comparable in accuracy when the AI advice is correct, while differing in decision time.
Because evaluative interfaces expose users to structured supporting and opposing considerations, they may support richer user justifications as well as better selections, motivating H4. Prior work on automation bias, error tolerance, and contestability also suggests that plausible, overbroad, and legally irrelevant errors may shape reliance differently and carry distinct downstream consequences. This motivates H5, which tests whether AI error conditions differ in their performance impact and whether Evaluative AI attenuates those penalties relative to Conventional XAI. It also motivates H6, which examines whether these differences extend to downstream risk-relevant outcomes such as disproportionate intervention under Overbreadth errors and routing failures under Out-of-Scope errors (described in detail in the following section).
Together, these strands of work motivate our comparison of manual reporting, conventional XAI, and evaluative AI under systematically controlled legal error conditions in a user-facing illegal-content reporting task.
3 Experimental Design and Procedure
We conducted a controlled experiment with a between-subjects factor (assistance condition) and a micro within-subject manipulation (correct AI assistance vs. AI error condition) applied only in the AI-assisted arms. Participants were blind to their condition assignment.
Assistance conditions. Participants were randomly assigned to one of three conditions:
• Baseline: Manual reporting without AI assistance (control).
• EvalAI: Evaluative AI, presenting structured pros and cons for multiple candidate legal provisions, inspired by Miller.
• XAI: Conventional XAI, presenting a single suggested legal provision with a brief rationale.
Each participant evaluated a 12-post feed 1: 10 benign filler posts (identical across participants) and two illegal posts drawn from four legal case categories (§ 86a symbols [L6], § 111 incitement [L10], § 130 hate speech [L15], drug sales [L27]; Appendix I.3). The two illegal posts were A/B variants of the same case. In AI-assisted conditions, one illegal post included correct AI assistance and the other included an injected AI error condition. Error order was counterbalanced.
AI assistance architectures. Figure 1 illustrates the reporting interface in the two AI-assisted conditions. In all conditions, participants saw the same candidate legal provisions. What varied across conditions was the assistance architecture. In XAI, the interface pre-highlighted one AI-suggested provision and displayed a short rationale (recommend-and-justify). In EvalAI, no provision was preselected; instead, participants could inspect structured supporting and opposing arguments for candidate provisions (deliberation-oriented). In Baseline, no AI-generated reasoning was shown.
AI error conditions. To capture distinct human–AI risk profiles, we introduced three randomized AI error conditions that meaningfully operationalize errors in our context:
• Near-Miss (NM): A semantically close but incorrect provision (e.g., § 111 instead of § 130), representing a setting in which the AI output is plausible and may increase automation bias risk.
• Overbreadth (OB): A related but overly broad provision (e.g., § 166 instead of § 130), representing a setting that may encourage over-removal or overly expansive legal interpretation.
• Out-of-Scope (OS): A legally irrelevant provision, representing a setting in which the AI output misroutes the report away from the appropriate legal frame.
These conditions are not intended to encode a universal severity ranking. Rather, they operationalize dis-tinct reporting risks at the human–AI interface: plausible misguidance (Near-Miss), overbroad categorization (Overbreadth), and legally irrelevant routing (Out-of-Scope).
Condition manifestation. In both AI-assisted conditions, the injected error pointed to the same incorrect legal provision for a given case; only the presentation architecture differed. In XAI, the incorrect provision was highlighted and an explanation was provided. In EvalAI the same incorrect provision received comparatively stronger supporting arguments, while competing categories were comparatively downplayed; outside the targeted mappings, categories were often represented with weaker, minimal, or placeholder pro/con content rather than fully elaborated arguments. In Baseline, no AI assistance was provided and no error manipulation was applied. Representative mappings are listed in Appendix Table 3.
Procedure. The study consisted of a pre-survey, the reporting task, and a post-survey (Appendix D–E; Fig. 2). Pre-survey. Participants completed measures of need for cognition, AI literacy (MAILS), trust in automation, and legal knowledge, including attention checks (Appendix D.10–D.11). Reporting task. Participants reviewed the 12-post feed under their assigned condition. Reporting required selecting a legal category (with AI support, where applicable), completing an Explanation of Reasons (statutory element marking and textual evidence; Art. 16(a) DSA), choosing a moderation action, and indicating proportionality. To preserve internal validity, progression was locked and cross-post comparison was prevented. Post-survey. Participants completed measures of perceived usefulness and ease of use (TAM), sense of agency, and workload (NASA–TLX short).
Participants exposed to AI errors were debriefed.
Cases. We modeled the reporting interface on Facebook’s illegal-content workflow using German statutory categories (Fig. 7). Legal cases were derived from German case law, translated into fictionalized scenarios using
ChatGPT (Appendix I.2), and cross-checked by two authors. Four cases were selected to span easy and hard classification challenges (See, Appendix Table 5).
Variables. The independent variables were assistance condition (between-subjects), AI trial type (correct-assistance vs. error trial; within-subjects), and AI error condition (NM, OB, OS) for error trials only.
Primary outcomes were (a) categorization accuracy (binary match to ground truth) and (b) misclassification distance (0 = correct; 1 = near-miss; 2 = related-but-wrong; 3 = out-of-scope). Secondary measures included decision time, Explanation-of-Reasons coverage, and explanation quality (1–5 rubric; two independent coders; Appendix G). Additional measures captured moderation actions, false positives on benign posts, interaction logs, self-reported confidence, trust, workload, and demographic covariates.
Participants. We recruited Prolific participants located in Germany (fluent in English; approval rate ≥ 90%). Compensation was set to a fixed £1.60 base payment (10–15 min) in addition to a bonus up to £2.00 (i.e., a maximum of £3.60). Participants could earn bonus payments by correctly categorizing malicious posts and by providing high-quality explanations of reasons, assessed based on coverage of key legal elements and proportionality considerations. We recruited 489 participants in total, and after exclusions due to technical issues or false-positive reports on benign filler posts, N =450 were retained, with 150 participants per condition (EvalAI, XAI, and Baseline; Appendix Table 4). The final sample comprised 293 men, 152 women, and 5 non-binary participants, with mean age M =31.32 years (SD =8.7, range 18–71).
4 Methodology
We ran a controlled experiment following Ozanne et al. using a custom web app (SocialNet) that mimics a social-media reporting workflow. We evaluate six hypotheses using models matched to outcome type: logistic models for binary outcomes, ordered logit models for misclassification distance, linear models for decision time and coder-mean explanation ratings, and paired difference-in-differences specifications for error-induced performance changes. Full model equations and particulars, robustness checks, and supplementary analyses are reported in the appendix. Our goal is to test whether different assistance conditions shape reporting performance and downstream risk-relevant outcomes in DSA-aligned reporting tasks (Table 1).
H1: Provision accuracy under AI error. Among AI-assisted participants (Evaluative AI vs. Conventional XAI), we estimate the effect of Evaluative AI on provision-level accuracy on illegal-content error trials (one error trial per participant by design). The outcome of interest here is binary provision accuracy. We fit a logistic model with AI error condition and case fixed effects, reporting robust standard errors, odds ratios, 95% confidence intervals, and two-sided p -values. Robustness checks add pre-task covariates and exclude participants who reported benign filler posts.
H2: Misclassification distance under AI error. On illegal-content error trials in the AI-assisted conditions, H2a tests whether Evaluative AI reduces misclassification distance relative to Conventional XAI, and H2b tests whether this effect varies across AI error conditions, with particular attention to Out-of-Scope errors. The outcome is an ordinal misclassification distance score (0= correct, 3= farthest misclassification). We estimate proportional-odds ordered logit models with participant clustered robust standard errors. Robustness checks add pre-task covariates, exclude participants who reported benign filler posts, and compare conditions using non-parametric tests.
H3: No-error trials (correct AI assistance). We analyze the no-error trial in the two AI-assisted conditions only (one no-error trial per participant; N =300, 150 per condition). H3a tests provision-level accuracy using a binomial GLM with robust standard errors and a TOST equivalence test on the difference in proportions (δ = 0.05). H3b tests decision time using OLS on log-transformed response times with robust standard errors, with a one-tailed Mann–Whitney U test as a non-parametric check. Robustness checks add case fixed effects where available.
H4: Explanation-of-Reasons quality. We analyze illegal-content reports with completed coder ratings across all three assistance conditions (Baseline, XAI, EvalAI). Each report was independently rated by two trained annotators. We examine four coder-rated Explanation-of-Reasons dimensions: element coverage, proportionality reasoning, reasoning depth, and perceived overall quality. For each dimension, the dependent variable is the mean of the two coder ratings. We estimate linear models with (AI) assistance condition as a categorical predictor, participant-clustered robust standard errors, and pairwise contrasts across conditions. Participants could contribute up to two rated reports, which motivates clustering at the participant level. Inter-rater reliability was acceptable to high across dimensions (Appendix), supporting aggregation.
H5: AI error condition impact and attenuation. For AI-assisted participants only, we construct a paired difference-in-differences panel in which each participant contributes one illegal-content no-error trial and one illegal-content error trial. The participant’s AI error condition (Near-Miss, Overbreadth, or Out-of-Scope) is assigned at the pair level. We examine two classes of outcomes: provision accuracy and the four coder-rated Explanation-of-Reasons dimensions. Provision accuracy is analyzed using a paired difference-in-differences logistic model with participant-clustered robust standard errors. Explanation outcomes are analyzed primarily using participant-level first-difference OLS models on coder-mean ratings, with ordered-logit sensitivity analyses reported in the appendix.
H5a tests whether the error penalty differs by AI error condition; H5b tests whether Evaluative AI attenuates these error-specific performance drops relative to Conventional XAI.
H6: DSA-relevant downstream risk outcomes. H6 evaluates whether, relative to Conventional XAI, Evalu-ative AI yields user reporting behavior more consistent with downstream risk-relevant outcomes in DSA-aligned reporting. Analyses are restricted to AI-assisted error trials and estimated separately by AI error condition. For Overbreadth errors, we model over-removal as a binary indicator of whether the participant selected a severe enforcement action. For Out-of-Scope errors, we model misrouting as a binary indicator of whether the participant failed to route the case as out-of-scope. Each outcome is analyzed with a binomial GLM with a logit link and participant-clustered standard errors. When variation is insufficient for model estimation, we report descriptive statistics only.
5 Results
We first report provision-level outcomes under AI error (H1–H2), then comparisons under correct AI assistance (H3), followed by Explanation-of-Reasons outcomes (H4–H5) and downstream risk-relevant outcomes (H6).
H1: Provision accuracy under AI error Evaluative AI substantially improved provision-level accuracy under AI error. Participants in the Evaluative arm were correct on 74.0% of error trials (n = 150), compared to 46.0% in the Conventional XAI arm (n = 150), a difference of 28 percentage points. In a logit GLM with heteroskedasticity-robust (HC1) standard errors, including AI error condition and case fixed effects, Evaluative assistance was associated with an approximately ten-fold increase in odds of selecting the correct legal provision relative to Conventional XAI (OR = 10.96, 95% CI [4.35, 27.65], p <.001). Results were numerically stable across estimation choices: a covariate-adjusted model including pre-task measures (MAILS, NCS-6, legal knowledge, Propensity to Trust Automation) yielded a comparable effect (OR = 11.10, 95% CI [4.52, 27.20], p <.001).
The Evaluative advantage was present across all error conditions (Fig. 3). Independent of assistance condition, Out-of-Scope errors were associated with higher baseline accuracy than Near-Miss errors (OR = 10.34, 95% CI [3.77, 28.41], p <.001), while Overbreadth errors showed a smaller but still positive shift relative to Near-Miss (OR = 2.71, 95% CI [1.09, 6.72], p =.032). Moderation analyses indicated that the Evaluative advantage was attenuated for Out-of-Scope errors relative to Near-Miss errors (Evaluative×Out-of-Scope: OR = 0.07, 95% CI [0.02, 0.27], p <.001), while no reliable attenuation was observed for Overbreadth errors (Evaluative×Overbreadth: OR = 0.40, 95% CI [0.11, 1.46], p =.165). Case fixed effects were not statistically significant overall; full model tables and balance checks are reported in the Appendix.
Conclusion. H1 is supported: when the AI makes an error, Evaluative AI substantially improves provision accuracy, increasing correctness by 28 percentage points and yielding an order-of-magnitude increase in the odds of selecting the correct legal provision relative to Conventional XAI for the reference AI error condition (Near-Miss).
H2: Misclassification distance under AI error Evaluative AI reduced misclassification distance relative to Conventional XAI on AI-error trials. In the primary ordered-logit specification (without case fixed effects), the Evaluative coefficient was negative and statistically significant (ˆβ = −1.53, SE = 0.40, z = −3.88, p <.001), indicating a shift in probability mass toward smaller (less severe) distances. Interpreted as a proportional-odds effect, this corresponds to OR = 0.22 with a 95% confidence interval of [0.10, 0.47].5 Figure 4 visualizes the corresponding distributions: compared to Conventional XAI, Evaluative AI concentrates more probability mass on distances 0–1 and less on distances 2–3. Error condition moderated the effect.
Relative to Near-Miss errors as the reference AI error condition, Out-of-Scope errors were associated with smaller distances overall (ˆγOS = −1.34, p =.001). The Evaluative advantage was significantly attenuated for Out-of-Scope errors (Evaluative×Out-of-Scope: ˆφEval×OS = 1.62, p =.007). By contrast, there was no evidence in the interaction tests that Evaluative’s 5Threshold (cutpoint) parameters are reported in the Appendix but are not substantively interpreted.
improvement differed between Overbreadth and Near-Miss errors (Evaluative×Overbreadth: ˆφEval×OB = −0.07, p =.898). Restricting the sample to Near-Miss and Overbreadth errors (H2b specification), the Evaluative effect remained strong at the Near-Miss baseline (ˆβ = −1.61, SE = 0.40, z = −4.02, p <.001; OR = 0.20, 95% CI [0.09, 0.44]), while the Evaluative×Overbreadth interaction again was not statistically significant (ˆφEval×OB = −0.06, p =.919), indicating comparable improvements for Near-Miss and Overbreadth errors. Results were robust across alternative specifications. Adding covariates (MAILS, NCS-6, legal knowledge, and Propensity to Trust Automation), including case fixed effects as a design-control robustness check, or excluding benign reporters did not alter conclusions.
Non-parametric Mann–Whitney tests provided convergent evidence, showing significantly smaller misclassification distances under Evaluative AI overall (U = 8128, one-tailed p <.001), as well as within Near-Miss (U = 734, p <.001) and Overbreadth errors (U = 841.5, p =.0003).
Conclusion. H2a is supported: when the AI is wrong, Evaluative AI yields smaller misclassification distances than Conventional XAI. H2b receives partial support: the Evaluative benefit is attenuated for Out-of-Scope errors, while differences between Near-Miss and Overbreadth errors are not detectable in the interaction tests.
H3: No-error trials H3a: Accuracy. When the AI was correct, provision accuracy was high in both conditions and did not differ reliably between assistance forms. Descriptively, accuracy was 86.7% under Conventional XAI and 82.0% under Evaluative AI. In the binomial GLM with HC1 heteroskedasticity-robust standard errors (one observation per participant) and no case fixed effects, the Evaluative vs. Conventional XAI contrast was not statistically significant (OR = 0.65, 95% CI [0.35, 1.22], p =.192). Results were unchanged when adding case fixed effects as a design-control robustness check. A TOST equivalence test with margin δ = ±0.05 did not support equivalence. While the upper bound test was satisfied (p =.011), the lower bound test failed (p =.468), indicating that equivalence could not be established under the specified margin. H3b: Decision time.
Decision time analyses showed a consistent speed advantage for Conventional XAI. In the OLS model on log-transformed decision times with robust (HC1) standard errors, Evaluative AI was slower than Conventional XAI (coef = 0.13, SE = 0.07, p =.064), indicating a directionally consistent delay that did not reach conventional significance. Results were substantively unchanged when adding case fixed effects as a design-control robustness check. The non-parametric Mann–Whitney U test on raw decision times corroborated this directional effect: decisions were significantly faster under Conventional XAI than Evaluative AI (U = 10012, one-tailed p =.0498). Median decision times were 203,690 ms for Conventional XAI and 235,769 ms for Evaluative AI, corresponding to a substantial practical difference.
Conclusion. H3b is supported: when the AI assistance is correct, XAI enables faster decisions than EvalAI. H3a remains inconclusive: although the accuracy difference is not statistically significant, equivalence between Evaluative and Conventional XAI is not established under the prespecified margin.
H4: Explanation-of-Reasons (EoR) quality We tested whether Evaluative AI improves the quality of users’ Explanation-of-Reasons (EoR) submissions relative to Baseline reporting and Conventional XAI across four coder-rated dimensions: element coverage, proportionality reasoning, reasoning depth, and perceived overall quality. Across all four dimensions, we found no reliable differences between assistance conditions. In linear regression models on coder-mean ratings with participant-clustered standard errors, coefficients comparing Evaluative AI to Conventional XAI were small and statistically indistinguishable from zero (all |t | < 0.70, all p >.48). Baseline reporting likewise did not significantly differ from Conventional XAI on any dimension.
For element coverage, the largest estimated contrast was between Evaluative AI and Baseline reporting (ˆβ = 0.14, SE = 0.11, t = 1.24, p =.216), but this difference was not statistically significant. No evidence of differences was observed between Evaluative AI and Conventional XAI (ˆβ = 0.007, p =.953). Figure 5b illustrates the distribution of perceived overall quality ratings. Distributions for Evaluative AI and Conventional XAI largely overlap across rating levels, with no systematic shifts consistent with improved explanation quality.
Conclusion. H4 is not supported. Although Evaluative AI improves decision accuracy and error handling (H1–H2), it does not reliably improve coder-rated explanation quality relative to Baseline reporting or XAI.
H5: AI error condition impact and attenuation Accuracy. In the paired difference-in-differences logistic model, AI error conditions substantially reduced provision accuracy relative to correct AI trials, and the magnitude of this error penalty varied across AI error conditions. Relative to Near-Miss as the reference AI error condition, the accuracy decline was significantly smaller for Out-of-Scope errors (error trial×OS > 0, p <.001) and smaller for Overbreadth errors (error trial×OB > 0, p =.025). Near-Miss errors produced the largest accuracy drops. Both Overbreadth and Out-of-Scope errors were significantly less damaging than Near-Miss errors, with Out-of-Scope errors showing the smallest declines. Differences between Overbreadth and Out-of-Scope were marginal. Wald contrasts confirmed that error penalties differed by AI error condition (OS vs. Near-Miss: p <.001; OS vs.
Overbreadth: p =.054). Attenuation in accuracy. Evaluative AI significantly reduced the baseline error penalty relative to Conventional XAI (Evaluative AI×error trial > 0, p <.001), indicating overall attenuation of error-induced accuracy losses. However, this attenuation was moderated by AI error condition. The three-way interaction for Out-of-Scope errors was negative and statistically significant (Evaluative AI×error trial×OS < 0, p <.001), indicating that the additional attenuation provided by Evaluative AI was smaller for Out-of-Scope errors than for Near-Miss errors. No reliable moderation was observed for Overbreadth errors (Evaluative AI×error trial×OB, p =.232).
Taken together, Evaluative AI most strongly buffered the accuracy penalty associated with Near-Miss errors, while incremental gains were reduced for AI error conditions that were already less harmful under Conventional XAI. EoR quality (coder ratings; first-difference models). We next examined whether AI error conditions affected explanation quality differently. Using participant-level first-difference linear models on continuous coder-mean ratings, error trials were not associated with statistically reliable declines in EoR quality across any dimension. In particular, Out-of-Scope errors did not produce larger reductions than Near-Miss or Overbreadth errors in element coverage, proportionality reasoning, reasoning depth, or perceived overall quality (all p >.39). Accordingly, H5a is not supported for EoR quality. Attenuation in EoR quality.
Evaluative AI did not significantly attenuate error-related changes in explanation quality on any EoR dimension. Interaction terms between Evaluative AI and Out-of-Scope errors were small and non-significant across all measures (all p >.39). Thus, H5b is not supported for EoR quality. Robustness. Sensitivity analyses using proportional-odds (ordered logit) models on discretized coder-mean ratings yielded substantively same conclusions, with no sign reversals or emergent significant effects (Appendix).
Conclusion. H5a and H5b are supported for accuracy but not for explanation quality. Evaluative AI selectively attenuates error-induced accuracy losses—most strongly for Near-Miss errors—while providing limited additional benefit for error conditions that are already less harmful under Conventional XAI. Across AI error conditions, however, Evaluative AI does not reliably improve the structure or completeness of user-provided explanations under AI error, indicating that its primary benefit lies in supporting correct decision outcomes rather than downstream explanation quality.
H6: DSA-relevant downstream risk outcomes Overbreadth errors (H6(i)). The Overbreadth panel comprised 102 error trials from 102 participants (54 Conventional XAI; 48 Evaluative AI). Severe enforcement actions were frequent in both conditions, with higher raw over-removal rates under Evaluative AI (91.7%) than Conventional XAI (81.5%). A binomial GLM with participant-clustered standard errors indicated that the Evaluative AI condition was associated with higher odds of selecting a removal-type action relative to Conventional XAI (log-odds β = 0.92,
SE= 0.64; OR= 2.50, 95% CI [0.72, 8.68]), though this difference was not statistically significant (p = 0.149). Thus, we find no evidence that Evaluative AI reduces disproportionate enforcement under Overbreadth errors. Out-of-Scope errors (H6(ii)). The Out-of-Scope panel included 94 error trials from 94 participants (50 Evaluative AI; 44 Conventional XAI). Misrouting was near-universal: all Evaluative AI trials (100%) and nearly all Conventional XAI trials (95.5%) failed to select the out-of-scope routing option. Because the outcome exhibited insufficient variation, a regression model could not be meaningfully estimated. We therefore report descriptive statistics only and treat H6(ii) as not identifiable in this dataset.
Conclusion. Overall, H6(i) is not supported: Evaluative AI did not reduce over-removal under Overbreadth errors and showed descriptively higher, though non-significant, rates of severe enforcement. H6(ii) was not identifiable due to near-zero variance in out-of-scope routing behavior, indicating a measurement or interface limitation rather than a treatment effect. A summary of the hypotheses can be found in Appendix A.
6 Discussion
This study examined how different AI assistance architectures shape user reporting behavior in content moderation workflows, focusing on accuracy under AI error, explanation quality, and implications for DSA-relevant reporting objectives. The results reveal a consistent pattern: Evaluative AI improves decision outcomes when the AI is wrong, but does not reliably improve the quality of user-provided explanations. This distinction is consequential for governance-critical settings, where correct decisions and higher quality user provided justifications do not necessarily coincide.
Reducing overreliance under AI error. Under error conditions, Evaluative AI substantially increased provision accuracy (H1) and reduced misclassification distance (H2), indicating less harmful reliance on incorrect model suggestions. These effects are consistent with evaluative assistance functioning as a deliberation scaffold: by surfacing competing considerations rather than a single recommended category, the interface makes disagreement explicit and discourages one-shot acceptance of erroneous outputs. Even when users did not fully correct an AI error, evaluative support helped avoid the most severe misclassifications, keeping users’ selections closer to legally relevant provisions. For moderation pipelines, this can improve the quality of information entering downstream processes under uncertainty.
Selective attenuation and targeted deployment. The paired DiD analyses show that Evaluative AI does not uniformly mitigate all error-related failures. Benefits were strongest for Near-Miss errors, attenuated for Out-of-Scope errors, and not detectably different from Near-Miss for Overbreadth errors. This pattern argues against a one-size-fits-all assistance policy. Instead, the results suggest that evaluative support may be most useful in high-uncertainty or high-risk situations—for example, when model confidence is low, category entropy is high, or user choices conflict with AI suggestions. Such targeted use could preserve robustness benefits while limiting added time and cognitive load, a key consideration given that evaluative interfaces slowed decisions when the AI was correct (H3b) and full human–AI review is not scalable in high-volume moderation contexts.
Why explanation quality did not improve. Despite accuracy gains, Evaluative AI did not reliably improve coder-rated Explanation-of-Reasons (EoR) quality (H4–H5). This dissociation suggests a mechanism mismatch: evaluative assistance primarily supports choice revision (selecting a better provision) rather than argument construction (articulating richer justifications). In our design, the reporting interface already required users to engage with statutory elements and evidence fields, which may have elevated baseline explanation quality across conditions and compressed variance. As prior work shows, providing AI explanations or guidance alone does not reliably improve reasoning quality or calibration.
Improving EoRs likely requires assistance that directly targets articulation—such as gap-highlighting, prompts for missing elements, or retrieval-backed evidence cues—without generating legal arguments on the user’s behalf.
Implications for DSA-oriented design. The DSA requires user-friendly reporting mechanisms and suf-ficiently substantiated notices, but does not specify how interfaces should support users in meeting these standards. Our findings indicate that Evaluative AI can improve the reliability of user reports under AI error by reducing overreliance and severe misclassification, thereby improving the quality of information entering moderation pipelines at the reporting stage. However, improved classification accuracy should not be conflated with legally sufficient explanations: Evaluative AI supports better decisions, but not necessarily better justifications. It is therefore best understood as a performance-enhancing design intervention under uncertainty, not as a mechanism that enforces regulatory compliance.
Downstream risks and platform costs. Downstream risk proxies (H6) highlight important limits. Under Overbreadth errors, over-removal rates were high in both conditions and descriptively higher under Evaluative AI, with no statistically reliable reduction in disproportionate enforcement. For Out-of-Scope errors, misrouting was near-universal, leaving insufficient variation to identify treatment effects. These results underscore that accuracy gains do not automatically translate into proportionate enforcement and suggest that interface defaults can dominate user choices. More broadly, lowering reporting friction could increase the volume of actionable notices, raising moderation costs and governance burdens.
Limitations and future work. This study relied on a custom reporting interface with a curated set of legal provisions, which may have reduced real-world ambiguity and increased baseline performance relative to deployment settings. The structured explanation fields, while useful for measurement, may scaffold reasoning in ways that differ from deployed platform environments. Moreover, the experiment captures short-term behavior and cannot assess learning, strategic adaptation, or adversarial misuse over time. Future work should evaluate evaluative assistance in platform-integrated settings, link user reports to moderator outcomes, and explore adap-tive or constrained writing support that better aligns gains in decision accuracy with meaningful improvements in explanation quality and downstream governance outcomes.
7 Conclusion
This study demonstrates that Evaluative AI improves user accuracy under AI error in illegal content reporting by reducing harmful overreliance relative to conventional XAI. Whereas conventional XAI interfaces encourage one-shot acceptance of incorrect suggestions and increase susceptibility to biases favorable to machine outputs, Evaluative AI supports more cautious and contestable decision-making under uncertainty. At the same time, we find no reliable improvements in the quality of user-provided explanations, indicating that gains in decision accuracy do not automatically translate into more substantiated or higher-quality reasoning.
These findings underscore that reporting interfaces are an important design space under the Digital Services Act and can improve performance under the right conditions. Evaluative AI is best understood as a deliberation scaffold that improves decision outcomes in error-prone settings, rather than as a mechanism that guarantees legally sufficient explanations or independently satisfies regulatory requirements. Important trade-offs remain: evaluative assistance increases deliberation time and may raise the volume of actionable reports, with implications for platform costs and downstream governance burden.
Future work should examine longitudinal effects on trust, learning, and strategic behavior; test generalizability across jurisdictions and platform contexts; and explore complementary support modalities—such as adaptive or constrained assistance—that directly target explanation articulation. Advancing such designs will be critical for balancing accuracy, user agency, and regulatory accountability in human–AI collaboration.